CPU inference can run a model that fits available memory
A local LLM can generate text without a discrete GPU when the runtime supports CPU inference and the complete workload fits system memory. Start with a modest task: rewrite a paragraph, extract a field from a short note or explain a small function. Whether the result is useful depends on answer quality and measured waiting time on your machine, not simply whether the model loads.
This workflow uses a directly installed local Ollama runtime. It is separate from gpuOS's connected GPU service. The local AI guide explains the broader choice of runtime, interface and data path; the Ollama and LM Studio comparison covers interface differences. No model tag establishes a universal CPU speed or minimum RAM requirement.
Budget RAM for weights, context and the rest of the machine
Scroll horizontally to see every column.
| Memory component | What changes it | Useful check |
|---|---|---|
| Model weights | Model size and exact quantized artifact | Record the downloaded variant, not just its family name. |
| Context and KV cache | Prompt length, output allowance and concurrent requests | Test your largest actual input with explicit context settings. |
| Runtime allocations | Inference backend and processing buffers | Observe process and system memory while generating. |
| Other applications | Editor, browser and background workloads | Keep enough free memory for the surrounding workflow. |
Download size is not total runtime memory. The quantization guide explains weight precision, while the context and KV cache guide explains prompt-dependent allocations. A smaller quantized model may fit where a larger variant does not, but test whether it still follows your required instructions.
Measure system pressure while the representative prompt runs. Excessive swapping can make a model that technically fits too slow to use. Begin with one request and a short context, then increase input size deliberately. Do not interpret gpuOS catalog VRAM estimates as CPU RAM guarantees or treat a long advertised context window as a practical default for your computer.
Download one explicit local model before testing
ollama --version
ollama pull qwen3:0.6b
ollama listInstall the runtime using its supported instructions before running these commands. The Ollama Qwen3 0.6B listing verifies this exact example tag. It is a small candidate for a first functional check, not a recommendation that it will solve every task. Keep the listed model identifier and runtime version with your results.
Select another available local model if this candidate misses required behavior. Change model size and precision separately so you can understand which change affected quality or memory. Review the model's license before integrating its output into your application. Keep the first prompt synthetic while verifying that requests stay on the intended machine.
Confirm the execution device instead of inferring it from the URL
# Ollama must already be installed, running, and have this model downloaded.
# Run on a machine without a usable accelerator, then inspect ollama ps.
curl --noproxy '*' --fail-with-body http://127.0.0.1:11434/api/generate \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3:0.6b",
"prompt": "Explain RAM and storage in two short sentences.",
"stream": false,
"think": false,
"options": {"num_ctx": 2048, "num_predict": 128, "temperature": 0}
}' > cpu-response.json
ollama psOn a machine without a usable accelerator, inspect the loaded model after the request. The Ollama FAQ defines ollama ps: 100% CPU identifies system-memory loading, 100% GPU identifies GPU loading and a mixed value identifies partial offloading. Keep that observation with your sample. Loopback routing alone does not prove CPU execution.
If you need a deliberate CPU comparison on a machine with an accelerator, use a runtime configuration that explicitly disables offloading and verify its logs. Current llama.cpp server options include --device none and --no-op-offload. Check the installed version's help before adapting the alternative below; it expects an already downloaded compatible GGUF file.
llama-server --model ./models/model.gguf --device none --no-op-offload \
--ctx-size 2048 --host 127.0.0.1 --port 8080This alternative does not use Ollama's /api/generate endpoint. Follow its server documentation for requests and inspect startup allocations. It is a separate local-runtime option, not a statement that gpuOS enrolls every CPU or accelerator supported by upstream software.
Separate model loading, prompt work and output generation
Save the answer as well as its timing fields. The Ollama usage reference defines token counts and durations in nanoseconds. Loading, prompt evaluation and output generation describe different costs; dividing generated tokens by generation time does not measure the entire application's latency. A non-streaming response also cannot show when the first token reached the client.
import json
from pathlib import Path
sample = json.loads(Path("cpu-response.json").read_text())
if not isinstance(sample, dict) or sample.get("error") or sample.get("done") is not True:
raise SystemExit("Expected a completed local generation response.")
fields = ("eval_count", "eval_duration", "load_duration", "prompt_eval_duration", "total_duration")
if any(type(sample.get(field)) is not int or sample[field] < 0 for field in fields):
raise SystemExit("Expected nonnegative integer timing and token fields.")
if sample["eval_count"] == 0 or sample["eval_duration"] == 0:
raise SystemExit("This sample has no measurable generated output.")
print("Generated tokens:", sample["eval_count"])
print("Generation tokens/s:", format(sample["eval_count"] * 1_000_000_000 / sample["eval_duration"], ".2f"))
for field in ("load_duration", "prompt_eval_duration", "total_duration"):
print(field + ":", format(sample[field] / 1_000_000_000, ".3f"), "s")
if sample.get("done_reason") == "length":
print("Generation reached its limit; inspect whether the answer is complete.")Run python3 read-cpu-timing.py beside the saved response. There are no sample throughput claims in this calculation. Compare repeated warm requests and a cold start, then repeat with your normal background applications. For a longer prompt, inspect both memory and prompt evaluation time. The benchmark guide explains streaming latency and repeatable comparisons.
Keep the CPU workflow only when it passes your acceptance check
- Use a small prompt set with expected fields, facts or code behavior, including a case the model should answer as unknown.
- Record exact model, precision, context, actual token counts, CPU placement and other active workloads.
- Check the slowest representative input, not just an easy warm request.
- Reduce context or choose another model when memory pressure or latency exceeds your application's budget.
- Compare a supported GPU deployment or a cloud route when the CPU configuration fails required quality or timing.
gpuOS runs Ollama on the GPU machine you connect, with prompts and outputs passing through its hosted gateway. Local versus cloud AI helps compare that route with a direct local runtime. cpuOS supplies separate bounded Python and Node execution jobs for trusted team code; it does not serve language-model inference or turn its CPU job worker into this local LLM runtime.