gpuos

Hardware · 4 min read · updated Oct 7, 2026

LLM context window and KV cache: plan GPU memory for real traffic

Plan LLM context length, KV cache and concurrency on your GPU. Budget chat history, RAG passages and outputs, then validate Ollama memory and latency.

On this page

An advertised context window is a limit, not a memory budget

A model's context window describes how much tokenized information it can use in one generation. Your configured runtime context can be smaller than the model's advertised maximum. Increasing that runtime setting needs more memory; Ollama documents the setting and inspection commands in its context length guide.

Choose a context budget from the application you are serving. A short classification endpoint, document assistant and coding agent have different input sizes. Treat the advertised maximum as a capability to evaluate, rather than a setting to enable automatically on every deployment. Check whether the longest requests still meet your latency and answer-quality targets.

Budget system prompts, history, retrieval and the answer

Count the entire request, including the system instruction, previous turns, retrieved text and any tool descriptions or results. Reserve space for the answer as well. Tokenization depends on the model and input language, so a character count or page count is only a rough proxy. Use actual request usage or a matching tokenizer when available.

Scroll horizontally to see every column.

Input componentQuestion to ask before sending it
System instructionsDoes each repeated rule affect the current task?
Conversation historyWhich earlier turns are still needed to answer correctly?
RAG passagesWhich retrieved sections support this question?
Tool dataCan the result be narrowed to the fields the model needs?
Output allowanceHow much space does a complete useful answer need?

For example, an application can define separate input and output limits, reject oversized uploads early and show users which document sections were included. If you summarize old turns, preserve identifiers, decisions and unresolved questions. Test the summary process itself, since reducing text is useful only if it keeps the facts needed for the next answer.

The KV cache holds attention state alongside model weights

Autoregressive models reuse attention key and value tensors from earlier tokens through a KV cache. That cache is distinct from the model's weights. Its memory depends on the architecture, token capacity, cache precision and active sequences. Models with sliding or chunked attention can behave differently from full-attention models. Hugging Face documents these cache strategies.

This distinction explains why a small model download can still fail on a long request. Reducing weight precision changes one part of the memory budget; it does not automatically change the precision or capacity of the cache. Avoid applying a single per-token memory constant across unrelated model families without checking their architecture and runtime.

When a workload exceeds memory, change one factor at a time: shorten inputs, lower the configured context, reduce active requests or choose a smaller model. Record the result of each change. That gives you a capacity boundary for this deployment instead of a guess based only on parameter count.

One long conversation and four parallel requests are different tests

Ollama's parallel-request allocation scales with configured context length and parallelism. Its concurrency and KV cache FAQ describes the relevant server settings, queuing and cache precision options. For a shared service, test simultaneous requests explicitly; a successful single-user run does not establish team capacity.

  • Start with one request and record peak GPU memory, total latency and input/output token counts.
  • Repeat with the busiest simultaneous workload you expect, using different prompts so shared prefixes do not hide its cost.
  • Keep output limits fixed across the comparison, then add a test with the longest allowed answer.
  • Record queued requests and failures alongside successful responses. Waiting in a queue still affects the user's experience.

gpuOS key quotas and request limits help control API access. They do not reserve memory or guarantee a latency target. Decide how your application handles busy periods: a bounded queue, an explicit retry message or fewer concurrent jobs. Choose that behavior from measured service capacity.

Inspect Ollama locally and validate the application route

Inspect loaded models and NVIDIA GPU memory on the node
ollama ps
nvidia-smi

Use ollama ps to inspect loaded models, processor placement and allocated context where your installed version reports it. Use nvidia-smi on NVIDIA hardware to observe device memory during a representative request. Also record runtime versions, other GPU processes and whether the model was already loaded.

These commands run on your GPU machine. gpuOS runs inference through Ollama on that node, while the hosted gateway receives prompts and outputs. Test both the local engine and your application's gpuOS endpoint when diagnosing latency, because the application path adds routing and network time. Configure runtime settings on the node using the supported deployment workflow, then verify their effect rather than assuming an API client field changed them.

Use retrieval to send relevant evidence within a known limit

A larger window is not a substitute for selecting useful passages. For a document assistant, retrieve relevant sections, keep source labels and ask questions that require a citation. Evaluate missing evidence and conflicting passages, not only a question whose answer appears at the beginning of the prompt. The private RAG guide covers the retrieval pipeline.

Set an input budget for retrieval, a separate answer allowance and an explicit fallback when the evidence is insufficient. Save these limits with your deployment configuration. Repeat the benchmark after increasing document size or concurrency; a context change can alter the service your users experience even when the weights stay identical.

Questions

Why does a model fit at 4K context but fail at a longer context?
Weights are only part of loaded memory. A larger context can require more KV cache and runtime allocations. Inspect memory on the node and test the same model with the intended context and concurrent requests.
Does Q4 weight quantization also make the KV cache four-bit?
No. Weight quantization and cache precision are separate settings. Check your installed Ollama version and backend support before changing cache precision, then evaluate response quality and memory again.
Should I send every retrieved document into the context window?
Select passages that support the current question and reserve room for the answer. Test whether the assistant finds and cites the right evidence under realistic document sizes, rather than filling the window because space is available.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.