1. Identify the server that receives your request
Start with the request route: application, API endpoint, Ollama server and inference machine. Inspect the machine running that server. A laptop's GPU monitor cannot explain inference on a remote node, and a local Ollama CLI can be pointed at a different server from your application. The Ollama API guide distinguishes direct Ollama requests from the hosted gpuOS gateway.
- Record the application endpoint, inference host, installed Ollama version and exact local model tag. A gpuOS catalog id and an Ollama tag can differ.
- Identify the deployment: native desktop app, Linux service, manually started server or an existing container. Inspect that deployment's configuration privately, rather than printing a full environment or process command line.
- Confirm the request uses a local model. A cloud-hosted model does not run on your GPU, even when the request passes through a local Ollama instance.
- For a remote server, run the following check on that host through your existing administrative connection. The explicit loopback address targets its local Ollama listener; it does not expose a port.
OLLAMA_HOST=http://127.0.0.1:11434 ollama --version
OLLAMA_HOST=http://127.0.0.1:11434 ollama psThese commands assume the standard local port. If your deployment uses another port, substitute its known local address. A connection failure is an endpoint or service problem to resolve before investigating GPU placement. Record client and server version differences if the CLI reports them.
2. Read model placement before judging utilization
Inspect ollama ps while an existing application request is running or shortly afterward, before the model unloads. Its PROCESSOR column describes loaded placement, not a live percentage of GPU computation. Ollama documents the distinction in its GPU placement FAQ.
Scroll horizontally to see every column.
| Observation | Meaning | Next check |
|---|---|---|
| 100% GPU | The model is loaded entirely on the GPU | Observe a real request and measure warm latency |
| 100% CPU | The model is loaded in system memory | Check GPU discovery, deployment access and intentional CPU settings |
| A CPU/GPU percentage split | The model is partially loaded on each | Check model size, allocated context and available memory |
| No model rows | No model is currently loaded on this server | Correlate the request endpoint, model tag and observation time |
A read-only alternative is Ollama's running-models endpoint. Its result includes the model identity and memory information. An empty models array is not evidence that GPU discovery failed: the model may have unloaded, the request may use another server, or no local inference has started.
curl --noproxy '*' --fail --silent --show-error --max-time 10 \
http://127.0.0.1:11434/api/ps3. On NVIDIA hosts, separate driver visibility from inference
Run this read-only query on the NVIDIA inference host, then repeat it while your existing application is generating an answer. It reports every visible GPU so a request on another device is not mistaken for inactivity on the first card. NVIDIA documents these query fields.
nvidia-smi --query-gpu=index,name,driver_version,memory.total,memory.used,utilization.gpu --format=csv- If the command cannot communicate with the driver or reports no device, investigate host-level visibility before changing the model. A missing command alone does not diagnose the hardware.
- A successful result proves the driver exposes a GPU to this host command. It does not prove that the Ollama service, container or selected model can use it.
- Memory can remain allocated while an idle model performs no inference. A low utilization sample between requests is compatible with GPU placement.
- Other processes can also occupy GPU memory or generate utilization. Correlate the sample with this model's placement, request timing and server logs; utilization alone cannot identify the workload.
Some devices or environments return N/A for unsupported metrics, including GPU utilization on MIG-enabled devices. An unavailable metric is not a zero reading. See the nvidia-smi reference for reporting limits.
Do not use one utilization snapshot as a speed benchmark. After confirming placement, compare cold loading and warm generation separately using the local LLM benchmark method.
4. Follow the branch for your actual GPU and deployment
Scroll horizontally to see every column.
| Deployment | What to inspect | Reference |
|---|---|---|
| Native NVIDIA on Linux or Windows | Exact GPU support, driver requirements and discovery errors in the server log | Ollama hardware support |
| NVIDIA in a Linux container or Windows WSL2 | Container GPU access and NVIDIA Container Toolkit configuration, separately from host visibility | Ollama Docker deployment |
| Native Ollama on an Apple GPU | Metal backend and model placement; nvidia-smi is not the diagnostic for Apple hardware | Ollama Metal support |
| Ollama in Docker Desktop on macOS | The documented absence of GPU acceleration in this deployment | Ollama Docker FAQ |
| AMD on Linux or Windows | The supported GPU/OS matrix and matching driver/backend stack; on Linux, device access and permissions | Ollama AMD support |
For an existing container named ollama, inspect placement inside it with the command below. Substitute the known container name if yours differs. This selects its loopback listener rather than assuming that a working host driver is sufficient.
docker exec -e OLLAMA_HOST=http://127.0.0.1:11434 ollama ollama psReview only the relevant server settings in your deployment configuration. A GPU-selection variable that hides every device or an explicit CPU library override can explain CPU inference. The server's configuration matters; changing a variable in a client shell does not reconfigure an already running service. Use the hardware reference for backend-specific selectors and the troubleshooting reference for discovery errors.
GPU support in upstream Ollama does not establish support for that hardware in the gpuOS node installer. Check the gpuOS quickstart for its current deployment requirements.
5. For partial offload, inspect the whole memory budget
A CPU/GPU split is evidence of partial placement, not proof of a broken driver. Model weights, allocated context, runtime buffers, concurrent requests and other loaded models share the available capacity. A model download size or a card's advertised memory cannot settle this check. Use the VRAM guide to shortlist models and the context and KV cache guide to plan the remaining budget.
- Record the exact model variant and quantization, along with the effective context shown by your Ollama version. Current context settings are more useful than an assumed default.
- Observe the system with one representative request, then compare it with the intended concurrency. Record which other models and applications are resident at each observation.
- If the complete workload cannot fit, test a smaller model or lower-memory variant, or a shorter context that still accommodates the task. Change one factor at a time and recheck task quality.
- Repeat the check at the largest context and concurrency your application intends to admit. A short isolated prompt cannot validate production capacity.
Ollama explains that larger contexts require more memory and that parallel requests increase context allocation. No fixed number of users or context tokens fits every model and GPU. If CPU execution meets your task's latency requirements, the local LLM without a GPU guide provides a separate evaluation path.
6. Review a small log window on the inference machine
Correlate the request time with discovery, backend selection, model loading and memory errors. Ollama's troubleshooting guide documents the platform-specific log locations. Choose only the command matching your deployment; do not collect every host log.
journalctl -u ollama --since "15 minutes ago" --lines 80 --no-pagerdocker logs --since 15m --tail 80 ollamatail -n 80 ~/.ollama/logs/server.logFor a manually started server, inspect its existing terminal. On Windows, open the current server.log in %LOCALAPPDATA%\Ollama. Log wording varies by release, so record the actual error and its timestamp rather than expecting a sample message verbatim.
A discovery failure after Linux suspend/resume or after a container previously worked is a different symptom from a model that never fit in memory. Follow the matching official troubleshooting branch. Driver reloads, container-runtime changes and restarts affect running work, so arrange maintenance before applying a fix.
Review logs locally. Before sharing an excerpt, remove tokens, authorization headers, private endpoints, prompts, outputs and identifying paths. A short log window can still contain sensitive values; it is not automatically safe to publish.
7. Verify the fix with placement and a repeatable request
Keep a small before/after record: inference host, deployment type, runtime version, local model tag, context, concurrency, placement and the error observed. After the selected maintenance or workload change, repeat the same non-sensitive application request and inspect the same server. A useful result explains which observation changed.
- Confirm the intended model appears in ollama ps during the request and record whether placement meets your deployment goal.
- Check that the relevant discovery or loading error has cleared, then measure warm response time and task correctness separately.
- Retest the intended context and concurrency before sending normal traffic. Keep a failing minimal example if the problem remains.
For gpuOS, a connected node or successful gateway health response does not prove GPU placement or that the selected model meets a latency target. Inspect Ollama on the connected inference machine and validate an actual model request. The gpuOS quickstart covers API routing and errors; the benchmark guide covers performance measurements.
If upstream GPU discovery still fails, use Ollama's issue tracker with the version, hardware, deployment type, reproducible symptom and redacted error excerpt. These references were reviewed on 9 October 2026; recheck the hardware matrix when changing runtime or driver versions.