Choose for the workload you need to serve
Start with Ollama when you want a short path from downloading a model to calling it on one machine. Evaluate vLLM when a dedicated GPU service needs concurrent requests, explicit serving configuration or parallelism across GPUs. Both can serve a model through an OpenAI-compatible API; that common client interface does not make their memory use, scheduler or deployment interchangeable.
There is no engine-wide tokens-per-second ranking on this page. Hardware, weight format, prompt length, output length and simultaneous requests change the result. The examples below establish working endpoints; the test protocol helps you choose using your own workload.
gpuOS currently runs inference with Ollama. This guide compares the engines independently; it does not imply that a gpuOS node can switch to vLLM.
Ollama vs vLLM at a glance
Scroll horizontally to see every column.
| Decision | Ollama | vLLM |
|---|---|---|
| First local model | Pull a library tag, then run or call it | Install a compatible runtime, then serve a model repository |
| Client API | Native /api routes and selected OpenAI /v1 routes | OpenAI-compatible routes plus serving-specific endpoints |
| Memory and parallel requests | Configure context, loaded models and request parallelism | Paged KV-cache management and explicit serving limits |
| Weights | Check the exact tag and quantization in the Ollama library | Check the model architecture, dtype and supported quantization |
| Multi-GPU plan | Check how the selected model is placed on your GPUs | Evaluate the documented tensor, pipeline or data parallel setup |
| Team access | Local API has no workspace keys; add an access layer | A server key is not a complete multi-user access layer |
Ollama supports concurrent processing when memory allows it. vLLM stores attention cache in blocks, which is part of its serving design. Neither fact proves which engine is faster for your request mix. See Ollama concurrency configuration, vLLM paged attention and vLLM parallelism options.
Launch each engine before comparing it
Install each runtime from its official instructions and check your supported GPU and driver combination. Run one engine at a time on the GPU. These commands use Qwen3 8B, but the default Ollama download and the Hugging Face weights are different formats: this is a setup check, not a controlled performance comparison.
# With Ollama installed and its local service running:
ollama --version
ollama pull qwen3:8b
curl --fail-with-body http://127.0.0.1:11434/api/chat \
-H "Content-Type: application/json" \
-d '{"model":"qwen3:8b","stream":false,"think":false,"messages":[{"role":"user","content":"Reply with one short sentence."}]}'
ollama ps
# Before using this GPU for vLLM:
ollama stop qwen3:8b# With vLLM installed in its own compatible environment:
vllm --version
vllm serve Qwen/Qwen3-8B \
--host 127.0.0.1 --port 8000 --max-model-len 8192
# Wait for the server to finish loading before sending a request.The unquantized vLLM example can require substantially more VRAM than the gpuOS catalog's Q4_K_M estimate for Qwen3 8B. Do not use that estimate as a requirement for these different weights. Install references: Ollama, vLLM GPU runtime and vLLM serve arguments.
Use the same API shape, with the correct model name
# vLLM server from the previous section:
curl --fail-with-body http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen3-8B","stream":false,"max_tokens":256,"messages":[{"role":"user","content":"Explain an HTTP status code in one sentence."}]}'
# Ollama's OpenAI-compatible route:
curl --fail-with-body http://127.0.0.1:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"qwen3:8b","stream":false,"max_tokens":256,"messages":[{"role":"user","content":"Explain an HTTP status code in one sentence."}]}'For a non-streaming response, inspect choices[0].message.content, finish_reason and usage. A thinking model may spend its output budget on reasoning, leaving little or no visible answer. Check its documented reasoning controls before treating an empty answer as an HTTP failure.
Validate every application feature you depend on: chat template, tools, structured output, embeddings and streaming. Compatibility with chat completions does not establish compatibility with every OpenAI operation. Route references: Ollama compatibility and vLLM online serving. The Ollama API guide explains native and compatible payloads separately.
Benchmark without changing the experiment halfway through
- Record GPU model and count, VRAM, driver, engine version, exact model revision, quantization, tokenizer, context limit and launch settings.
- Use the same saved prompts, chat template where possible, sampling settings and output budget. Check answer quality alongside speed. If the weight formats differ, label the result as a comparison of two deployment configurations.
- Measure a cold first request separately. Warm the model with a disposable request before measuring normal serving, and record whether repeated prefixes are cached.
- Test one, four and eight simultaneous requests only if the configurations can fit them. Include short prompts and representative long inputs, not only repeated one-sentence requests.
- Report completed requests, failures, actual generated tokens, full response latency and peak memory. For a streaming app, additionally measure time to first visible content with an SSE-aware client.
- Repeat each run and save raw results. Compare latency distributions and failure rates rather than selecting the fastest single response.
No results are published here because this protocol has not been run on a stated GPU. The probe below measures completed-response HTTP latency, not time to first token or sustained production capacity.
Save an initial HTTP probe for both endpoints
Save the following as probe.py; it uses Python 3's standard library. Start at concurrency one, run a warm-up separately, then collect each run into a JSONL file. Replace its short synthetic prompt with a saved workload before making a capacity decision.
import concurrent.futures
import json
import os
import time
import urllib.error
from urllib.parse import urlsplit
from urllib.request import HTTPRedirectHandler, Request, build_opener
base = os.environ.get("LLM_BASE_URL", "http://127.0.0.1:11434/v1").rstrip("/")
model = os.environ.get("LLM_MODEL", "qwen3:8b")
key = os.environ.get("LLM_API_KEY", "")
workers = int(os.environ.get("CONCURRENCY", "1"))
origin = urlsplit(base)
if not origin.hostname or origin.username or origin.password or origin.query or origin.fragment:
raise ValueError("LLM_BASE_URL must use a trusted server origin and API path")
if origin.scheme != "https" and not (
origin.scheme == "http" and origin.hostname in {"localhost", "127.0.0.1", "::1"}
):
raise ValueError("Use HTTPS for remote endpoints; HTTP is allowed only on loopback")
class NoRedirect(HTTPRedirectHandler):
def redirect_request(self, req, fp, code, msg, headers, newurl):
return None
opener = build_opener(NoRedirect())
headers = {"Content-Type": "application/json"}
if key:
headers["Authorization"] = "Bearer " + key
def run_one(number):
body = {
"model": model, "stream": False, "max_tokens": 128,
"temperature": 0,
"messages": [{"role": "user", "content":
f"Give one practical example of task {number}: testing an HTTP API."}],
}
request = Request(base + "/chat/completions",
data=json.dumps(body).encode(), headers=headers)
started = time.perf_counter()
try:
with opener.open(request, timeout=120) as response:
data = json.load(response)
content = data["choices"][0]["message"].get("content") or ""
return {"id": number, "ok": True,
"seconds": time.perf_counter() - started,
"usage": data.get("usage", {}),
"has_content": bool(content)}
except (urllib.error.URLError, TimeoutError, ValueError, KeyError, IndexError):
return {"id": number, "ok": False,
"seconds": time.perf_counter() - started}
started = time.perf_counter()
with concurrent.futures.ThreadPoolExecutor(max_workers=workers) as executor:
results = list(executor.map(run_one, range(20)))
for result in results:
print(json.dumps(result))
print(json.dumps({"batch_seconds": time.perf_counter() - started,
"requests": len(results), "failures": sum(not r["ok"] for r in results)}))LLM_BASE_URL=http://127.0.0.1:11434/v1 \
LLM_MODEL=qwen3:8b CONCURRENCY=1 python3 probe.py > ollama-c1.jsonl
LLM_BASE_URL=http://127.0.0.1:8000/v1 \
LLM_MODEL=Qwen/Qwen3-8B CONCURRENCY=1 python3 probe.py > vllm-c1.jsonl
# Repeat with CONCURRENCY=4, then 8, if memory permits.Each request records seconds, usage, has_content and ok; the final line records elapsed batch time and failures. An HTTP success with has_content: false needs model-level investigation. When output-token counts exist, sum them and divide by batch time for this run's aggregate rate. Do not compare a single request's decode rate with that aggregate rate.
For an authenticated endpoint, set LLM_API_KEY in the client environment. The probe refuses redirects so the Bearer key stays on the configured origin, and requires HTTPS for remote servers. Keep the key out of result files and prompts.
Ollama's native API also reports server-side timing in nanoseconds. eval_count / (eval_duration / 1e9) describes generation rate for that response; it excludes parts of end-to-end latency. See Ollama usage metrics.
Choose an operating model as well as an engine
Scroll horizontally to see every column.
| Observed problem | What to check first |
|---|---|
| Slow first request | Weight loading and warm-up, separately from normal serving |
| Good single-user latency, poor team latency | Queueing, concurrency, input lengths and memory limits |
| Out-of-memory at a longer prompt | Context and KV-cache headroom, not just weight size |
| Chat fails despite a successful model load | Supported architecture, chat template and request fields |
| A remote client cannot connect | Bind address, private network or tunnel, and the access layer |
Keep the local-only servers above private. For a shared service, provide TLS, authentication, request limits, logging policy and monitoring across all exposed routes. vLLM's --api-key authenticates selected route prefixes; it does not secure every endpoint on that server. See the vLLM security guide.
With gpuOS, your GPU node uses Ollama and connects outbound to the hosted gateway. Applications use workspace keys and supported OpenAI-compatible routes. Prompts and outputs traverse that gateway; quotas bound API access, not GPU performance. Follow the gpuOS quickstart for that deployment, or the Ollama team access comparison for the boundary between engine and access layer.