gpuos

Getting started · 6 min read · updated Oct 7, 2026

Ollama API: native endpoints, OpenAI compatibility and access

Call Ollama with curl, Python and Node. Compare native and OpenAI routes, parse streaming, request embeddings and connect through the gpuOS gateway.

On this page

Pick the API family before writing the client

A local Ollama server exposes native /api/* endpoints and selected OpenAI-compatible /v1/* endpoints. Use the native API when you need Ollama's runtime options or timing fields. Use compatible chat completions when your application already speaks that contract. Start with http://127.0.0.1:11434; append the endpoint path once.

Scroll horizontally to see every column.

OperationNative OllamaOpenAI-compatible Ollama
ConversationPOST /api/chat; messagesPOST /v1/chat/completions; messages
Text completionPOST /api/generate; promptPOST /v1/completions; prompt
EmbeddingsPOST /api/embed; inputPOST /v1/embeddings; input
List local modelsGET /api/tagsGET /v1/models
Read a full chat answermessage.contentchoices[0].message.content
Read a chat streamNewline-delimited JSONServer-sent events with data payloads

The routes are related, but payloads and response parsers differ. Follow the reference for the family you selected: native chat, native text generation, model listing and OpenAI compatibility.

Make a native chat request with curl

With Ollama installed and its service running, pull the exact tag, check that it appears in the model list, then send a short request. The following example disables streaming so the answer is one JSON object.

curl: a complete native response
ollama pull qwen3:8b
curl --fail-with-body http://127.0.0.1:11434/api/tags
curl --fail-with-body http://127.0.0.1:11434/api/chat \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3:8b",
    "messages": [{"role":"user","content":"Explain a GPU in one sentence."}],
    "stream": false,
    "think": false,
    "options": {"num_ctx":8192,"num_predict":128}
  }'

Inspect message.content, done and done_reason; the generated wording is not fixed. num_ctx selects the runtime context budget and num_predict limits generated tokens for this native request. These are Ollama options, not portable fields to copy into a compatible request. think: false is selected here for Qwen3 so the short demonstration does not spend its budget on thinking output.

Use ollama ps after the request to inspect memory placement. Longer contexts need additional memory even when the weights fit. See Ollama context and GPU checks and the VRAM calculator.

Parse native streaming one JSON line at a time

Save this as chat.py and run python3 chat.py after the download. It requires no Python package. Native chat streams successive JSON objects; read each line independently rather than calling json.load on the entire stream. See the Ollama streaming format.

Python: native NDJSON streaming
import json
import urllib.error
import urllib.request

payload = {
    "model": "qwen3:8b",
    "messages": [{"role": "user", "content": "Explain a GPU in one sentence."}],
    "stream": True,
    "think": False,
    "options": {"num_ctx": 8192, "num_predict": 128},
}
request = urllib.request.Request(
    "http://127.0.0.1:11434/api/chat",
    data=json.dumps(payload).encode(),
    headers={"Content-Type": "application/json"},
)
completed = False
try:
    with urllib.request.urlopen(request, timeout=120) as response:
        for line in response:
            if not line.strip():
                continue
            chunk = json.loads(line)
            if "error" in chunk:
                raise RuntimeError(chunk["error"])
            print(chunk.get("message", {}).get("content", ""), end="", flush=True)
            if chunk.get("done"):
                completed = True
                print()
                print(json.dumps({
                    "done_reason": chunk.get("done_reason"),
                    "output_tokens": chunk.get("eval_count"),
                    "generation_ns": chunk.get("eval_duration"),
                }))
                break
    if not completed:
        raise RuntimeError("Stream ended without a final done chunk")
except urllib.error.HTTPError as error:
    raise SystemExit(f"Ollama HTTP {error.code}") from error

Expect answer fragments followed by a JSON summary containing done_reason, output-token count and generation duration when the server supplies them. The text can differ on every run. An error chunk or a connection ending before done: true is a failed or incomplete response, even if the initial HTTP status was successful. Ollama error handling documents this distinction.

This parser is for native /api/chat. Do not feed it a /v1/chat/completions stream: that stream uses SSE framing and may include a terminal data: [DONE] event.

Call the compatible chat API from curl or Node

curl: the compatible chat-completions route
curl --fail-with-body http://127.0.0.1:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3:8b","stream":false,"reasoning_effort":"none","max_tokens":128,"messages":[{"role":"user","content":"Explain a GPU in one sentence."}]}'
Node: compatible chat without an SDK
// Save as chat.mjs and run with a current Node.js version.
const response = await fetch("http://127.0.0.1:11434/v1/chat/completions", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    model: "qwen3:8b", stream: false, max_tokens: 128,
    reasoning_effort: "none",
    messages: [{ role: "user", content: "Explain a GPU in one sentence." }],
  }),
  signal: AbortSignal.timeout(120_000),
})
if (!response.ok) throw new Error("Ollama HTTP " + response.status)
const data = await response.json()
if (!data.choices?.length) throw new Error("Missing chat choice")
console.log(data.choices[0].message.content)

To use the OpenAI SDK instead, set its base URL to http://127.0.0.1:11434/v1 and provide the placeholder key its constructor requires. Ollama's local server ignores that key; it does not establish authentication. For SDK setup, see our Python integration and Node integration, using the local URL and tag above rather than a gpuOS workspace configuration.

Ollama documents a subset of the OpenAI API. Supported routes and reasoning controls depend on the installed version and model. Check Ollama compatibility before relying on tools, response formats or Responses API behavior; those features are not automatically forwarded by another gateway.

Request embeddings with an embedding model

Two embedding routes, different response envelopes
ollama pull bge-m3
curl --fail-with-body http://127.0.0.1:11434/api/embed \
  -H "Content-Type: application/json" \
  -d '{"model":"bge-m3","input":["A GPU runs model inference.","A CPU executes application code."],"truncate":false}'

curl --fail-with-body http://127.0.0.1:11434/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model":"bge-m3","input":["A GPU runs model inference.","A CPU executes application code."]}'

The native route returns an embeddings array containing one vector per input. The compatible route uses data items with an embedding field. Check vector count, order and dimensions before writing to an index; do not mix vectors from different models in the same search space. Native truncate: false rejects inputs beyond the context budget instead of silently shortening them. See the native embed reference.

Use the same model and preprocessing when indexing documents and embedding queries. Changing models requires rebuilding the index and checking retrieval quality. For a complete retrieval workflow, see private RAG on your own GPU.

Keep local tags and gpuOS model ids separate

Scroll horizontally to see every column.

ModelDirect Ollama requestgpuOS gateway request
Qwen3 8Bqwen3:8bqwen3-8b
gpt-oss 20Bgpt-oss:20bgpt-oss-20b
BGE-M3bge-m3bge-m3

Ollama takes the installed model tag. gpuOS takes a stable catalog id and routes it to the deployed model's Ollama tag. A colon-to-hyphen replacement is not a general conversion rule: check the model page and workspace deployment. The same text happens to work for BGE-M3.

Changing only a URL while keeping the wrong model name is a common migration error. Check the model catalog before calling a workspace, and the local /api/tags response before calling Ollama directly.

Choose a private tunnel or an authenticated gateway

For a single developer connecting to an existing private GPU host, an SSH local-forward keeps the Ollama listener on the remote loopback interface. With Ollama running there, replace the host below and leave the SSH session open; then the local examples above use the tunnel.

Remote access through an SSH local-forward
ssh -N -L 127.0.0.1:11434:127.0.0.1:11434 user@your-gpu-host

Ensure local port 11434 is free before opening that tunnel. Shared applications need an access layer with TLS and authentication, plus a deliberate request and logging policy. Setting OLLAMA_HOST=0.0.0.0 changes reachability; it does not create user authentication. Ollama's default bind is loopback. Ollama networking reference.

Use the gpuOS gateway with a deployed catalog id
# With a connected node, deployed model and workspace API key:
curl --fail-with-body https://gpuos.si/v1/chat/completions \
  -H "Authorization: Bearer $GPUOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-8b","stream":false,"messages":[{"role":"user","content":"Explain a GPU in one sentence."}]}'

gpuOS uses an outbound node connection and workspace API keys for its supported chat completions, completions and embeddings routes. Applications call the hosted gateway, which is part of the prompt and output data path. The examples for Ollama's native /api/* endpoints are local examples, not gpuOS gateway endpoints. Follow the quickstart and self-hosted API guide for workspace setup.

Diagnose the failing layer

Scroll horizontally to see every column.

SymptomFirst check
Connection refusedServer or tunnel is running; address and port match
404 or model not foundEndpoint family, installed local tag or deployed catalog id
JSON parsing fails on a streamNDJSON for native routes; SSE for compatible routes
Stream stops without a final resultTreat as incomplete; inspect server and proxy errors
Slow first responseModel load time, then memory placement with ollama ps
Memory error after a longer promptContext, simultaneous requests and other loaded models
gpuOS 401 / 429 / 503Workspace key / quota or rate limit / node and model availability

Capture the status code, endpoint family, model identifier and whether the response completed. Avoid logging secrets or full prompts while diagnosing transport errors. Retry transient failures with a bounded policy; a retry cannot repair a wrong model id, invalid request or insufficient memory. Native Ollama and the gpuOS gateway have different error boundaries, so consult the Ollama errors reference and gpuOS quickstart for the server you called.

Questions

What is the default Ollama API URL?
The local server normally listens at http://127.0.0.1:11434. Native chat uses /api/chat; a compatible chat-completions client uses the /v1 base URL and the /chat/completions route.
Does the local Ollama API need an API key?
A direct local request does not need a workspace key. The OpenAI SDK may require a placeholder key, which local Ollama ignores. A hosted gpuOS request requires a real workspace key; the two configurations are distinct.
Why does json.load fail when I call /api/chat?
Native chat streams by default, so the body can contain multiple newline-delimited JSON objects. Set stream to false for one JSON response, or parse each line and require the final done chunk.
Can I call Ollama's native /api/chat through gpuOS?
The current gpuOS public gateway exposes supported OpenAI-compatible chat completions, completions and embeddings. Use its /v1 routes and catalog ids. The native /api examples in this guide call Ollama directly.
Can I use a chat model to build an embedding index?
Choose a model whose deployed endpoint supports embeddings. Use the same embedding model for documents and queries, and validate the returned vector dimensions and retrieval quality before building the index.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.