Why self-host an OpenAI-compatible API
Most AI (or SI, super intelligence) code today is written against the OpenAI API. If your own GPU speaks the same API, every SDK, framework and tool you already use keeps working: you change a base URL and a key, not your code.
Self-hosting gives you three things an external API cannot: prompts and documents never leave infrastructure you control, the marginal cost of a token is close to zero once the GPU is paid for, and no provider can change or retire the model under you.
Three ways to do it
| Approach | What you get | What you still build |
|---|---|---|
| Ollama or vLLM on a public port | An OpenAI-compatible endpoint on one machine | TLS, authentication, per-user keys, quotas, metering, firewall rules |
| A gateway like LiteLLM in front | Keys and routing | GPU nodes, model deployment, VRAM planning, the engine itself |
| gpuos | Agent, model deployment, gateway, keys, quotas, metering | Nothing beyond the install command |
Exposing Ollama directly is fine for a single developer. It has no notion of users or keys, so as soon as a team or an application depends on it you end up rebuilding a gateway.
Step by step with gpuos
- Create a free workspace. The Community plan covers one node with one GPU.
- In Nodes, click Add node and run the install command on your GPU machine. It installs Ollama and the
gpuos-agentservice; the agent only makes outbound HTTPS connections. - In Models, deploy a model. The VRAM gauge shows whether it fits before anything downloads.
- In API keys, create a key, optionally with a monthly token quota and a requests-per-minute limit.
- Call
https://gpuos.si/v1with any OpenAI client.
curl https://gpuos.si/v1/chat/completions \
-H "Authorization: Bearer $GPUOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "qwen3-32b", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'What production use needs
- One key per app or teammate, so you can revoke one without breaking the others.
- Quotas and rate limits per key: a monthly token cap and a requests-per-minute limit stop one runaway script from saturating the GPU.
- Metering: prompt and completion tokens, latency and status per key, model and node.
- OpenAI error format: 401, 403, 404, 429 and 503 errors that SDKs already know how to handle.
- No inbound port: the agent connects out, so the GPU box stays behind its firewall.
gpuos covers all five. The quickstart lists every error code and its meaning.