Define which part of inference you are measuring
A useful benchmark answers an application question: how quickly a chat starts, how long a summary takes or how many jobs finish while people share the GPU. Record latency and throughput separately. The NVIDIA inference metrics reference distinguishes first-token latency, request latency and token throughput.
Scroll horizontally to see every column.
| Metric | Measure | Why it matters |
|---|---|---|
| Time to first token (TTFT) | Request start to the first generated token received | How soon streamed output begins |
| Total request latency | Request start to completed response | How long the full task takes |
| Generation tokens/s | Generated tokens divided by generation time | Speed while the model produces its answer |
| Aggregate output throughput | All completed output tokens divided by benchmark wall time | Capacity at a stated concurrency |
Label the measurement boundary. Engine timing, gateway timing and client wall time describe different parts of the path. A tokens-per-second number without a boundary, input length and concurrency cannot establish how your application will behave.
Keep the model and workload fixed for a fair comparison
- Record the GPU, driver, Ollama version, model artifact and quantization, along with other processes using the device.
- Keep the prompt, chat template, configured context, output limit and generation settings the same for each variant being compared.
- Record actual input and output token counts; a requested maximum does not guarantee a response of that length.
- Separate cold starts from warm requests and state how the model was warmed before collecting results.
- Use representative short and long requests, then test simultaneous traffic separately.
Include both repeated and varied prompts if your application has both. Reusing an identical prefix can change cache behavior, so describe that condition instead of presenting a cached run as every request's performance. Also save the answers: a faster configuration is useful only if it still completes the task correctly.
Collect an Ollama response with native timing fields
On a machine with Ollama running and qwen3:8b already pulled, the command below makes one non-streaming local request and saves its JSON response. Run a warm-up request first if you are measuring warm inference. This is an engine-side sample, not a gpuOS gateway request or a complete benchmark suite.
curl --fail-with-body http://localhost:11434/api/generate \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3:8b",
"prompt": "Explain the difference between RAM and VRAM in three sentences.",
"stream": false,
"options": {"temperature": 0, "num_predict": 128, "num_ctx": 4096}
}' > benchmark.jsonOllama's generate endpoint reference defines eval_count, eval_duration, prompt_eval_duration, load_duration and total_duration. Durations use nanoseconds. Keep those engine fields separate from any client stopwatch measurements.
python3 - <<'PY'
import json
from pathlib import Path
sample = json.loads(Path("benchmark.json").read_text())
if sample.get("error"):
raise SystemExit(sample["error"])
tokens = sample["eval_count"]
generation_seconds = sample["eval_duration"] / 1_000_000_000
if generation_seconds <= 0:
raise SystemExit("No positive generation duration in this sample")
print(f"Generated tokens: {tokens}")
print(f"Generation tokens/s: {tokens / generation_seconds:.2f}")
for field in ("load_duration", "prompt_eval_duration", "total_duration"):
print(f"{field}: {sample[field] / 1_000_000_000:.3f} s")
PYThis calculation reports whatever your machine measured; it contains no example performance claim. Save each run separately with its prompt and configuration. Missing timing fields should stop analysis rather than become zero-valued ‘fast’ results.
Measure TTFT with streaming and a client clock
A non-streaming response cannot reveal when the first token arrived. For a streaming test, start a monotonic client timer before sending the request, record the first non-empty generated token and stop at the completed response. Initial role-only or empty events do not establish TTFT. Define whether reasoning output or answer output counts as ‘first’ for your application.
The first network event may contain several tokens, and event boundaries are not necessarily token boundaries. Count tokens from the model's reported usage or a matching tokenizer, rather than counting streaming chunks. For comparisons through gpuOS, use the same client location and network conditions so you do not confuse a different route with a different model speed.
Measure repeated requests and then increase concurrency
One sample helps diagnose a request, but it does not describe a service. Collect enough repeated samples to show variation, preserve errors and report the number of requests. Show a median and a slower-tail percentile when the sample supports it. Avoid claiming reliable p99 behavior from a handful of measurements.
Increase concurrency in controlled steps while keeping prompt sizes and output limits fixed. Record completed requests, failures, TTFT and total latency at each step. Aggregate throughput can improve while each user waits longer. Choose the operating point from the latency and error rate your application accepts, rather than the largest throughput number alone.
For a larger test, NVIDIA GenAI-Perf documents workload controls and aggregated metrics. Check tool and endpoint compatibility before using its configuration against your deployment.
Publish a result another person can reproduce
Keep a short report with hardware, exact model and precision, runtime version, context, workload token counts, warm-up method, concurrency, measurement boundary and failed requests. Include quality checks next to timing results. This makes a comparison useful after a driver change or a new model download.
gpuOS runs Ollama on your GPU, and its hosted gateway receives prompts and outputs while recording request usage. Use node-side observations to diagnose the engine and application-side timing to assess the full route. The context guide and quantization guide identify the capacity choices to revisit when a benchmark no longer meets your target.