gpuos

Comparisons · 7 min read · updated Oct 7, 2026

Ollama vs vLLM: which inference engine should you choose?

Compare Ollama and vLLM for local models and shared GPU serving. Launch both, check API differences and benchmark your workload without misleading rankings.

On this page

Choose for the workload you need to serve

Start with Ollama when you want a short path from downloading a model to calling it on one machine. Evaluate vLLM when a dedicated GPU service needs concurrent requests, explicit serving configuration or parallelism across GPUs. Both can serve a model through an OpenAI-compatible API; that common client interface does not make their memory use, scheduler or deployment interchangeable.

There is no engine-wide tokens-per-second ranking on this page. Hardware, weight format, prompt length, output length and simultaneous requests change the result. The examples below establish working endpoints; the test protocol helps you choose using your own workload.

gpuOS currently runs inference with Ollama. This guide compares the engines independently; it does not imply that a gpuOS node can switch to vLLM.

Ollama vs vLLM at a glance

Scroll horizontally to see every column.

DecisionOllamavLLM
First local modelPull a library tag, then run or call itInstall a compatible runtime, then serve a model repository
Client APINative /api routes and selected OpenAI /v1 routesOpenAI-compatible routes plus serving-specific endpoints
Memory and parallel requestsConfigure context, loaded models and request parallelismPaged KV-cache management and explicit serving limits
WeightsCheck the exact tag and quantization in the Ollama libraryCheck the model architecture, dtype and supported quantization
Multi-GPU planCheck how the selected model is placed on your GPUsEvaluate the documented tensor, pipeline or data parallel setup
Team accessLocal API has no workspace keys; add an access layerA server key is not a complete multi-user access layer

Ollama supports concurrent processing when memory allows it. vLLM stores attention cache in blocks, which is part of its serving design. Neither fact proves which engine is faster for your request mix. See Ollama concurrency configuration, vLLM paged attention and vLLM parallelism options.

Launch each engine before comparing it

Install each runtime from its official instructions and check your supported GPU and driver combination. Run one engine at a time on the GPU. These commands use Qwen3 8B, but the default Ollama download and the Hugging Face weights are different formats: this is a setup check, not a controlled performance comparison.

Ollama: download, request and check GPU placement
# With Ollama installed and its local service running:
ollama --version
ollama pull qwen3:8b
curl --fail-with-body http://127.0.0.1:11434/api/chat \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3:8b","stream":false,"think":false,"messages":[{"role":"user","content":"Reply with one short sentence."}]}'
ollama ps
# Before using this GPU for vLLM:
ollama stop qwen3:8b
vLLM: start a local-only inference server
# With vLLM installed in its own compatible environment:
vllm --version
vllm serve Qwen/Qwen3-8B \
  --host 127.0.0.1 --port 8000 --max-model-len 8192
# Wait for the server to finish loading before sending a request.

The unquantized vLLM example can require substantially more VRAM than the gpuOS catalog's Q4_K_M estimate for Qwen3 8B. Do not use that estimate as a requirement for these different weights. Install references: Ollama, vLLM GPU runtime and vLLM serve arguments.

Use the same API shape, with the correct model name

Two OpenAI-compatible chat requests
# vLLM server from the previous section:
curl --fail-with-body http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen3-8B","stream":false,"max_tokens":256,"messages":[{"role":"user","content":"Explain an HTTP status code in one sentence."}]}'

# Ollama's OpenAI-compatible route:
curl --fail-with-body http://127.0.0.1:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3:8b","stream":false,"max_tokens":256,"messages":[{"role":"user","content":"Explain an HTTP status code in one sentence."}]}'

For a non-streaming response, inspect choices[0].message.content, finish_reason and usage. A thinking model may spend its output budget on reasoning, leaving little or no visible answer. Check its documented reasoning controls before treating an empty answer as an HTTP failure.

Validate every application feature you depend on: chat template, tools, structured output, embeddings and streaming. Compatibility with chat completions does not establish compatibility with every OpenAI operation. Route references: Ollama compatibility and vLLM online serving. The Ollama API guide explains native and compatible payloads separately.

Benchmark without changing the experiment halfway through

  1. Record GPU model and count, VRAM, driver, engine version, exact model revision, quantization, tokenizer, context limit and launch settings.
  2. Use the same saved prompts, chat template where possible, sampling settings and output budget. Check answer quality alongside speed. If the weight formats differ, label the result as a comparison of two deployment configurations.
  3. Measure a cold first request separately. Warm the model with a disposable request before measuring normal serving, and record whether repeated prefixes are cached.
  4. Test one, four and eight simultaneous requests only if the configurations can fit them. Include short prompts and representative long inputs, not only repeated one-sentence requests.
  5. Report completed requests, failures, actual generated tokens, full response latency and peak memory. For a streaming app, additionally measure time to first visible content with an SSE-aware client.
  6. Repeat each run and save raw results. Compare latency distributions and failure rates rather than selecting the fastest single response.

No results are published here because this protocol has not been run on a stated GPU. The probe below measures completed-response HTTP latency, not time to first token or sustained production capacity.

Save an initial HTTP probe for both endpoints

Save the following as probe.py; it uses Python 3's standard library. Start at concurrency one, run a warm-up separately, then collect each run into a JSONL file. Replace its short synthetic prompt with a saved workload before making a capacity decision.

Python: non-streaming HTTP probe
import concurrent.futures
import json
import os
import time
import urllib.error
from urllib.parse import urlsplit
from urllib.request import HTTPRedirectHandler, Request, build_opener

base = os.environ.get("LLM_BASE_URL", "http://127.0.0.1:11434/v1").rstrip("/")
model = os.environ.get("LLM_MODEL", "qwen3:8b")
key = os.environ.get("LLM_API_KEY", "")
workers = int(os.environ.get("CONCURRENCY", "1"))
origin = urlsplit(base)
if not origin.hostname or origin.username or origin.password or origin.query or origin.fragment:
    raise ValueError("LLM_BASE_URL must use a trusted server origin and API path")
if origin.scheme != "https" and not (
    origin.scheme == "http" and origin.hostname in {"localhost", "127.0.0.1", "::1"}
):
    raise ValueError("Use HTTPS for remote endpoints; HTTP is allowed only on loopback")

class NoRedirect(HTTPRedirectHandler):
    def redirect_request(self, req, fp, code, msg, headers, newurl):
        return None

opener = build_opener(NoRedirect())
headers = {"Content-Type": "application/json"}
if key:
    headers["Authorization"] = "Bearer " + key

def run_one(number):
    body = {
        "model": model, "stream": False, "max_tokens": 128,
        "temperature": 0,
        "messages": [{"role": "user", "content":
            f"Give one practical example of task {number}: testing an HTTP API."}],
    }
    request = Request(base + "/chat/completions",
        data=json.dumps(body).encode(), headers=headers)
    started = time.perf_counter()
    try:
        with opener.open(request, timeout=120) as response:
            data = json.load(response)
        content = data["choices"][0]["message"].get("content") or ""
        return {"id": number, "ok": True,
            "seconds": time.perf_counter() - started,
            "usage": data.get("usage", {}),
            "has_content": bool(content)}
    except (urllib.error.URLError, TimeoutError, ValueError, KeyError, IndexError):
        return {"id": number, "ok": False,
            "seconds": time.perf_counter() - started}

started = time.perf_counter()
with concurrent.futures.ThreadPoolExecutor(max_workers=workers) as executor:
    results = list(executor.map(run_one, range(20)))
for result in results:
    print(json.dumps(result))
print(json.dumps({"batch_seconds": time.perf_counter() - started,
    "requests": len(results), "failures": sum(not r["ok"] for r in results)}))
Record each engine and concurrency level separately
LLM_BASE_URL=http://127.0.0.1:11434/v1 \
LLM_MODEL=qwen3:8b CONCURRENCY=1 python3 probe.py > ollama-c1.jsonl

LLM_BASE_URL=http://127.0.0.1:8000/v1 \
LLM_MODEL=Qwen/Qwen3-8B CONCURRENCY=1 python3 probe.py > vllm-c1.jsonl
# Repeat with CONCURRENCY=4, then 8, if memory permits.

Each request records seconds, usage, has_content and ok; the final line records elapsed batch time and failures. An HTTP success with has_content: false needs model-level investigation. When output-token counts exist, sum them and divide by batch time for this run's aggregate rate. Do not compare a single request's decode rate with that aggregate rate.

For an authenticated endpoint, set LLM_API_KEY in the client environment. The probe refuses redirects so the Bearer key stays on the configured origin, and requires HTTPS for remote servers. Keep the key out of result files and prompts.

Ollama's native API also reports server-side timing in nanoseconds. eval_count / (eval_duration / 1e9) describes generation rate for that response; it excludes parts of end-to-end latency. See Ollama usage metrics.

Choose an operating model as well as an engine

Scroll horizontally to see every column.

Observed problemWhat to check first
Slow first requestWeight loading and warm-up, separately from normal serving
Good single-user latency, poor team latencyQueueing, concurrency, input lengths and memory limits
Out-of-memory at a longer promptContext and KV-cache headroom, not just weight size
Chat fails despite a successful model loadSupported architecture, chat template and request fields
A remote client cannot connectBind address, private network or tunnel, and the access layer

Keep the local-only servers above private. For a shared service, provide TLS, authentication, request limits, logging policy and monitoring across all exposed routes. vLLM's --api-key authenticates selected route prefixes; it does not secure every endpoint on that server. See the vLLM security guide.

With gpuOS, your GPU node uses Ollama and connects outbound to the hosted gateway. Applications use workspace keys and supported OpenAI-compatible routes. Prompts and outputs traverse that gateway; quotas bound API access, not GPU performance. Follow the gpuOS quickstart for that deployment, or the Ollama team access comparison for the boundary between engine and access layer.

Questions

Is vLLM always faster than Ollama?
No universal ranking is established here. Measure the same hardware, representative prompts, output budgets and simultaneous requests. If you use different weight formats, compare the full configurations and include answer quality.
Can Ollama serve several requests at once?
Yes, when memory permits. Context length, loaded models and the configured request parallelism affect memory use and queueing. Test your intended concurrency rather than assuming a single-request VRAM estimate covers it.
Can I replace Ollama with vLLM by changing the base URL?
For a basic supported chat-completions client, the URL and model identifier may be enough. You must also check request fields, reasoning behavior, tools, structured output and streaming before switching an application.
Does gpuOS support vLLM nodes?
The current gpuOS catalog and node deployment use Ollama. The vLLM launch example is a separate local deployment, not a gpuOS engine option.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.