gpuos

Hardware · 7 min read · updated Oct 9, 2026

Ollama not using your GPU: diagnose CPU fallback and partial offload

Check Ollama model placement, NVIDIA visibility, container access and memory pressure. Diagnose CPU fallback with scoped commands and server logs.

On this page

1. Identify the server that receives your request

Start with the request route: application, API endpoint, Ollama server and inference machine. Inspect the machine running that server. A laptop's GPU monitor cannot explain inference on a remote node, and a local Ollama CLI can be pointed at a different server from your application. The Ollama API guide distinguishes direct Ollama requests from the hosted gpuOS gateway.

  • Record the application endpoint, inference host, installed Ollama version and exact local model tag. A gpuOS catalog id and an Ollama tag can differ.
  • Identify the deployment: native desktop app, Linux service, manually started server or an existing container. Inspect that deployment's configuration privately, rather than printing a full environment or process command line.
  • Confirm the request uses a local model. A cloud-hosted model does not run on your GPU, even when the request passes through a local Ollama instance.
  • For a remote server, run the following check on that host through your existing administrative connection. The explicit loopback address targets its local Ollama listener; it does not expose a port.
Inspect a native Ollama listener on the inference host
OLLAMA_HOST=http://127.0.0.1:11434 ollama --version
OLLAMA_HOST=http://127.0.0.1:11434 ollama ps

These commands assume the standard local port. If your deployment uses another port, substitute its known local address. A connection failure is an endpoint or service problem to resolve before investigating GPU placement. Record client and server version differences if the CLI reports them.

2. Read model placement before judging utilization

Inspect ollama ps while an existing application request is running or shortly afterward, before the model unloads. Its PROCESSOR column describes loaded placement, not a live percentage of GPU computation. Ollama documents the distinction in its GPU placement FAQ.

Scroll horizontally to see every column.

ObservationMeaningNext check
100% GPUThe model is loaded entirely on the GPUObserve a real request and measure warm latency
100% CPUThe model is loaded in system memoryCheck GPU discovery, deployment access and intentional CPU settings
A CPU/GPU percentage splitThe model is partially loaded on eachCheck model size, allocated context and available memory
No model rowsNo model is currently loaded on this serverCorrelate the request endpoint, model tag and observation time

A read-only alternative is Ollama's running-models endpoint. Its result includes the model identity and memory information. An empty models array is not evidence that GPU discovery failed: the model may have unloaded, the request may use another server, or no local inference has started.

Read loaded models from the same local server
curl --noproxy '*' --fail --silent --show-error --max-time 10 \
  http://127.0.0.1:11434/api/ps

3. On NVIDIA hosts, separate driver visibility from inference

Run this read-only query on the NVIDIA inference host, then repeat it while your existing application is generating an answer. It reports every visible GPU so a request on another device is not mistaken for inactivity on the first card. NVIDIA documents these query fields.

Inspect NVIDIA driver, memory and sampled GPU activity
nvidia-smi --query-gpu=index,name,driver_version,memory.total,memory.used,utilization.gpu --format=csv
  • If the command cannot communicate with the driver or reports no device, investigate host-level visibility before changing the model. A missing command alone does not diagnose the hardware.
  • A successful result proves the driver exposes a GPU to this host command. It does not prove that the Ollama service, container or selected model can use it.
  • Memory can remain allocated while an idle model performs no inference. A low utilization sample between requests is compatible with GPU placement.
  • Other processes can also occupy GPU memory or generate utilization. Correlate the sample with this model's placement, request timing and server logs; utilization alone cannot identify the workload.

Some devices or environments return N/A for unsupported metrics, including GPU utilization on MIG-enabled devices. An unavailable metric is not a zero reading. See the nvidia-smi reference for reporting limits.

Do not use one utilization snapshot as a speed benchmark. After confirming placement, compare cold loading and warm generation separately using the local LLM benchmark method.

4. Follow the branch for your actual GPU and deployment

Scroll horizontally to see every column.

DeploymentWhat to inspectReference
Native NVIDIA on Linux or WindowsExact GPU support, driver requirements and discovery errors in the server logOllama hardware support
NVIDIA in a Linux container or Windows WSL2Container GPU access and NVIDIA Container Toolkit configuration, separately from host visibilityOllama Docker deployment
Native Ollama on an Apple GPUMetal backend and model placement; nvidia-smi is not the diagnostic for Apple hardwareOllama Metal support
Ollama in Docker Desktop on macOSThe documented absence of GPU acceleration in this deploymentOllama Docker FAQ
AMD on Linux or WindowsThe supported GPU/OS matrix and matching driver/backend stack; on Linux, device access and permissionsOllama AMD support

For an existing container named ollama, inspect placement inside it with the command below. Substitute the known container name if yours differs. This selects its loopback listener rather than assuming that a working host driver is sufficient.

Read placement inside an existing Ollama container
docker exec -e OLLAMA_HOST=http://127.0.0.1:11434 ollama ollama ps

Review only the relevant server settings in your deployment configuration. A GPU-selection variable that hides every device or an explicit CPU library override can explain CPU inference. The server's configuration matters; changing a variable in a client shell does not reconfigure an already running service. Use the hardware reference for backend-specific selectors and the troubleshooting reference for discovery errors.

GPU support in upstream Ollama does not establish support for that hardware in the gpuOS node installer. Check the gpuOS quickstart for its current deployment requirements.

5. For partial offload, inspect the whole memory budget

A CPU/GPU split is evidence of partial placement, not proof of a broken driver. Model weights, allocated context, runtime buffers, concurrent requests and other loaded models share the available capacity. A model download size or a card's advertised memory cannot settle this check. Use the VRAM guide to shortlist models and the context and KV cache guide to plan the remaining budget.

  • Record the exact model variant and quantization, along with the effective context shown by your Ollama version. Current context settings are more useful than an assumed default.
  • Observe the system with one representative request, then compare it with the intended concurrency. Record which other models and applications are resident at each observation.
  • If the complete workload cannot fit, test a smaller model or lower-memory variant, or a shorter context that still accommodates the task. Change one factor at a time and recheck task quality.
  • Repeat the check at the largest context and concurrency your application intends to admit. A short isolated prompt cannot validate production capacity.

Ollama explains that larger contexts require more memory and that parallel requests increase context allocation. No fixed number of users or context tokens fits every model and GPU. If CPU execution meets your task's latency requirements, the local LLM without a GPU guide provides a separate evaluation path.

6. Review a small log window on the inference machine

Correlate the request time with discovery, backend selection, model loading and memory errors. Ollama's troubleshooting guide documents the platform-specific log locations. Choose only the command matching your deployment; do not collect every host log.

Linux with the standard Ollama systemd service
journalctl -u ollama --since "15 minutes ago" --lines 80 --no-pager
Existing Docker container named ollama
docker logs --since 15m --tail 80 ollama
Native macOS Ollama app
tail -n 80 ~/.ollama/logs/server.log

For a manually started server, inspect its existing terminal. On Windows, open the current server.log in %LOCALAPPDATA%\Ollama. Log wording varies by release, so record the actual error and its timestamp rather than expecting a sample message verbatim.

A discovery failure after Linux suspend/resume or after a container previously worked is a different symptom from a model that never fit in memory. Follow the matching official troubleshooting branch. Driver reloads, container-runtime changes and restarts affect running work, so arrange maintenance before applying a fix.

Review logs locally. Before sharing an excerpt, remove tokens, authorization headers, private endpoints, prompts, outputs and identifying paths. A short log window can still contain sensitive values; it is not automatically safe to publish.

7. Verify the fix with placement and a repeatable request

Keep a small before/after record: inference host, deployment type, runtime version, local model tag, context, concurrency, placement and the error observed. After the selected maintenance or workload change, repeat the same non-sensitive application request and inspect the same server. A useful result explains which observation changed.

  1. Confirm the intended model appears in ollama ps during the request and record whether placement meets your deployment goal.
  2. Check that the relevant discovery or loading error has cleared, then measure warm response time and task correctness separately.
  3. Retest the intended context and concurrency before sending normal traffic. Keep a failing minimal example if the problem remains.

For gpuOS, a connected node or successful gateway health response does not prove GPU placement or that the selected model meets a latency target. Inspect Ollama on the connected inference machine and validate an actual model request. The gpuOS quickstart covers API routing and errors; the benchmark guide covers performance measurements.

If upstream GPU discovery still fails, use Ollama's issue tracker with the version, hardware, deployment type, reproducible symptom and redacted error excerpt. These references were reviewed on 9 October 2026; recheck the hardware matrix when changing runtime or driver versions.

Questions

Why does Ollama use the CPU when nvidia-smi sees my GPU?
Host driver visibility and access by the Ollama process are separate checks. Inspect the actual service or container, GPU-selection settings, model placement and discovery logs. Available memory also matters for the exact model, context and workload.
Is partial CPU/GPU offload always an error?
It means the model is loaded partly in system memory and partly on the GPU. Check the full memory budget and measure whether the resulting latency suits the task. A smaller variant or shorter context can be worth testing, but neither guarantees acceptable quality or performance.
Does low GPU utilization mean Ollama has fallen back to CPU?
No. A loaded GPU model can be idle between requests, and a sample can miss a short burst of computation. Check placement while a real request runs, then correlate utilization, memory and server logs. Activity from other processes is another possible explanation.
Does a healthy gpuOS gateway prove inference is using my GPU?
No. Gateway reachability and inference placement are different observations. Inspect Ollama on the connected inference host, send a request to the intended deployed model, and measure its response under the context and concurrency your application needs.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.