An advertised context window is a limit, not a memory budget
A model's context window describes how much tokenized information it can use in one generation. Your configured runtime context can be smaller than the model's advertised maximum. Increasing that runtime setting needs more memory; Ollama documents the setting and inspection commands in its context length guide.
Choose a context budget from the application you are serving. A short classification endpoint, document assistant and coding agent have different input sizes. Treat the advertised maximum as a capability to evaluate, rather than a setting to enable automatically on every deployment. Check whether the longest requests still meet your latency and answer-quality targets.
Budget system prompts, history, retrieval and the answer
Count the entire request, including the system instruction, previous turns, retrieved text and any tool descriptions or results. Reserve space for the answer as well. Tokenization depends on the model and input language, so a character count or page count is only a rough proxy. Use actual request usage or a matching tokenizer when available.
Scroll horizontally to see every column.
| Input component | Question to ask before sending it |
|---|---|
| System instructions | Does each repeated rule affect the current task? |
| Conversation history | Which earlier turns are still needed to answer correctly? |
| RAG passages | Which retrieved sections support this question? |
| Tool data | Can the result be narrowed to the fields the model needs? |
| Output allowance | How much space does a complete useful answer need? |
For example, an application can define separate input and output limits, reject oversized uploads early and show users which document sections were included. If you summarize old turns, preserve identifiers, decisions and unresolved questions. Test the summary process itself, since reducing text is useful only if it keeps the facts needed for the next answer.
The KV cache holds attention state alongside model weights
Autoregressive models reuse attention key and value tensors from earlier tokens through a KV cache. That cache is distinct from the model's weights. Its memory depends on the architecture, token capacity, cache precision and active sequences. Models with sliding or chunked attention can behave differently from full-attention models. Hugging Face documents these cache strategies.
This distinction explains why a small model download can still fail on a long request. Reducing weight precision changes one part of the memory budget; it does not automatically change the precision or capacity of the cache. Avoid applying a single per-token memory constant across unrelated model families without checking their architecture and runtime.
When a workload exceeds memory, change one factor at a time: shorten inputs, lower the configured context, reduce active requests or choose a smaller model. Record the result of each change. That gives you a capacity boundary for this deployment instead of a guess based only on parameter count.
One long conversation and four parallel requests are different tests
Ollama's parallel-request allocation scales with configured context length and parallelism. Its concurrency and KV cache FAQ describes the relevant server settings, queuing and cache precision options. For a shared service, test simultaneous requests explicitly; a successful single-user run does not establish team capacity.
- Start with one request and record peak GPU memory, total latency and input/output token counts.
- Repeat with the busiest simultaneous workload you expect, using different prompts so shared prefixes do not hide its cost.
- Keep output limits fixed across the comparison, then add a test with the longest allowed answer.
- Record queued requests and failures alongside successful responses. Waiting in a queue still affects the user's experience.
gpuOS key quotas and request limits help control API access. They do not reserve memory or guarantee a latency target. Decide how your application handles busy periods: a bounded queue, an explicit retry message or fewer concurrent jobs. Choose that behavior from measured service capacity.
Inspect Ollama locally and validate the application route
ollama ps
nvidia-smiUse ollama ps to inspect loaded models, processor placement and allocated context where your installed version reports it. Use nvidia-smi on NVIDIA hardware to observe device memory during a representative request. Also record runtime versions, other GPU processes and whether the model was already loaded.
These commands run on your GPU machine. gpuOS runs inference through Ollama on that node, while the hosted gateway receives prompts and outputs. Test both the local engine and your application's gpuOS endpoint when diagnosing latency, because the application path adds routing and network time. Configure runtime settings on the node using the supported deployment workflow, then verify their effect rather than assuming an API client field changed them.
Use retrieval to send relevant evidence within a known limit
A larger window is not a substitute for selecting useful passages. For a document assistant, retrieve relevant sections, keep source labels and ask questions that require a citation. Evaluate missing evidence and conflicting passages, not only a question whose answer appears at the beginning of the prompt. The private RAG guide covers the retrieval pipeline.
Set an input budget for retrieval, a separate answer allowance and an explicit fallback when the evidence is insufficient. Save these limits with your deployment configuration. Repeat the benchmark after increasing document size or concurrency; a context change can alter the service your users experience even when the weights stay identical.