gpuos

Hardware · 5 min read · updated Oct 7, 2026

Run a local LLM without a GPU: a CPU inference workflow

Run a local language model on a CPU. Choose a small quantized model, budget RAM and context, verify the execution device and measure your own latency.

On this page

CPU inference can run a model that fits available memory

A local LLM can generate text without a discrete GPU when the runtime supports CPU inference and the complete workload fits system memory. Start with a modest task: rewrite a paragraph, extract a field from a short note or explain a small function. Whether the result is useful depends on answer quality and measured waiting time on your machine, not simply whether the model loads.

This workflow uses a directly installed local Ollama runtime. It is separate from gpuOS's connected GPU service. The local AI guide explains the broader choice of runtime, interface and data path; the Ollama and LM Studio comparison covers interface differences. No model tag establishes a universal CPU speed or minimum RAM requirement.

Budget RAM for weights, context and the rest of the machine

Scroll horizontally to see every column.

Memory componentWhat changes itUseful check
Model weightsModel size and exact quantized artifactRecord the downloaded variant, not just its family name.
Context and KV cachePrompt length, output allowance and concurrent requestsTest your largest actual input with explicit context settings.
Runtime allocationsInference backend and processing buffersObserve process and system memory while generating.
Other applicationsEditor, browser and background workloadsKeep enough free memory for the surrounding workflow.

Download size is not total runtime memory. The quantization guide explains weight precision, while the context and KV cache guide explains prompt-dependent allocations. A smaller quantized model may fit where a larger variant does not, but test whether it still follows your required instructions.

Measure system pressure while the representative prompt runs. Excessive swapping can make a model that technically fits too slow to use. Begin with one request and a short context, then increase input size deliberately. Do not interpret gpuOS catalog VRAM estimates as CPU RAM guarantees or treat a long advertised context window as a practical default for your computer.

Download one explicit local model before testing

Connected preparation: inspect the runtime and download the example
ollama --version
ollama pull qwen3:0.6b
ollama list

Install the runtime using its supported instructions before running these commands. The Ollama Qwen3 0.6B listing verifies this exact example tag. It is a small candidate for a first functional check, not a recommendation that it will solve every task. Keep the listed model identifier and runtime version with your results.

Select another available local model if this candidate misses required behavior. Change model size and precision separately so you can understand which change affected quality or memory. Review the model's license before integrating its output into your application. Keep the first prompt synthetic while verifying that requests stay on the intended machine.

Confirm the execution device instead of inferring it from the URL

A bounded non-streaming request to the local Ollama server
# Ollama must already be installed, running, and have this model downloaded.
# Run on a machine without a usable accelerator, then inspect ollama ps.
curl --noproxy '*' --fail-with-body http://127.0.0.1:11434/api/generate \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3:0.6b",
    "prompt": "Explain RAM and storage in two short sentences.",
    "stream": false,
    "think": false,
    "options": {"num_ctx": 2048, "num_predict": 128, "temperature": 0}
  }' > cpu-response.json
ollama ps

On a machine without a usable accelerator, inspect the loaded model after the request. The Ollama FAQ defines ollama ps: 100% CPU identifies system-memory loading, 100% GPU identifies GPU loading and a mixed value identifies partial offloading. Keep that observation with your sample. Loopback routing alone does not prove CPU execution.

If you need a deliberate CPU comparison on a machine with an accelerator, use a runtime configuration that explicitly disables offloading and verify its logs. Current llama.cpp server options include --device none and --no-op-offload. Check the installed version's help before adapting the alternative below; it expects an already downloaded compatible GGUF file.

Alternative llama.cpp setup with offloading disabled; supply your own GGUF
llama-server --model ./models/model.gguf --device none --no-op-offload \
  --ctx-size 2048 --host 127.0.0.1 --port 8080

This alternative does not use Ollama's /api/generate endpoint. Follow its server documentation for requests and inspect startup allocations. It is a separate local-runtime option, not a statement that gpuOS enrolls every CPU or accelerator supported by upstream software.

Separate model loading, prompt work and output generation

Save the answer as well as its timing fields. The Ollama usage reference defines token counts and durations in nanoseconds. Loading, prompt evaluation and output generation describe different costs; dividing generated tokens by generation time does not measure the entire application's latency. A non-streaming response also cannot show when the first token reached the client.

read-cpu-timing.py: analyze cpu-response.json without packages
import json
from pathlib import Path

sample = json.loads(Path("cpu-response.json").read_text())
if not isinstance(sample, dict) or sample.get("error") or sample.get("done") is not True:
    raise SystemExit("Expected a completed local generation response.")
fields = ("eval_count", "eval_duration", "load_duration", "prompt_eval_duration", "total_duration")
if any(type(sample.get(field)) is not int or sample[field] < 0 for field in fields):
    raise SystemExit("Expected nonnegative integer timing and token fields.")
if sample["eval_count"] == 0 or sample["eval_duration"] == 0:
    raise SystemExit("This sample has no measurable generated output.")
print("Generated tokens:", sample["eval_count"])
print("Generation tokens/s:", format(sample["eval_count"] * 1_000_000_000 / sample["eval_duration"], ".2f"))
for field in ("load_duration", "prompt_eval_duration", "total_duration"):
    print(field + ":", format(sample[field] / 1_000_000_000, ".3f"), "s")
if sample.get("done_reason") == "length":
    print("Generation reached its limit; inspect whether the answer is complete.")

Run python3 read-cpu-timing.py beside the saved response. There are no sample throughput claims in this calculation. Compare repeated warm requests and a cold start, then repeat with your normal background applications. For a longer prompt, inspect both memory and prompt evaluation time. The benchmark guide explains streaming latency and repeatable comparisons.

Keep the CPU workflow only when it passes your acceptance check

  • Use a small prompt set with expected fields, facts or code behavior, including a case the model should answer as unknown.
  • Record exact model, precision, context, actual token counts, CPU placement and other active workloads.
  • Check the slowest representative input, not just an easy warm request.
  • Reduce context or choose another model when memory pressure or latency exceeds your application's budget.
  • Compare a supported GPU deployment or a cloud route when the CPU configuration fails required quality or timing.

gpuOS runs Ollama on the GPU machine you connect, with prompts and outputs passing through its hosted gateway. Local versus cloud AI helps compare that route with a direct local runtime. cpuOS supplies separate bounded Python and Node execution jobs for trusted team code; it does not serve language-model inference or turn its CPU job worker into this local LLM runtime.

Questions

Can I run a local language model without a dedicated GPU?
Yes, with a CPU-capable runtime and enough available system memory for weights, context and runtime allocations. Test a small model and your actual prompts; loading successfully does not guarantee useful quality or latency.
How do I verify that Ollama actually used the CPU?
Inspect ollama ps while the model is loaded. The documented PROCESSOR column distinguishes CPU, GPU and mixed loading. A localhost API address alone does not identify the execution device.
Does cpuOS provide the CPU inference endpoint in this guide?
No. This guide uses a separate directly installed model runtime. cpuOS executes short trusted Python and Node jobs on your Docker worker, while gpuOS connects GPU inference through a hosted gateway.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.