gpuos

Getting started · 2 min read · updated Oct 6, 2026

How to self-host an OpenAI-compatible API on your own GPU

Run open LLMs like Qwen3 and GLM-4 on your own GPU and expose them through an OpenAI-compatible API with keys, quotas and metering, without opening a port.

Why self-host an OpenAI-compatible API

Most AI (or SI, super intelligence) code today is written against the OpenAI API. If your own GPU speaks the same API, every SDK, framework and tool you already use keeps working: you change a base URL and a key, not your code.

Self-hosting gives you three things an external API cannot: prompts and documents never leave infrastructure you control, the marginal cost of a token is close to zero once the GPU is paid for, and no provider can change or retire the model under you.

Three ways to do it

ApproachWhat you getWhat you still build
Ollama or vLLM on a public portAn OpenAI-compatible endpoint on one machineTLS, authentication, per-user keys, quotas, metering, firewall rules
A gateway like LiteLLM in frontKeys and routingGPU nodes, model deployment, VRAM planning, the engine itself
gpuosAgent, model deployment, gateway, keys, quotas, meteringNothing beyond the install command

Exposing Ollama directly is fine for a single developer. It has no notion of users or keys, so as soon as a team or an application depends on it you end up rebuilding a gateway.

Step by step with gpuos

  1. Create a free workspace. The Community plan covers one node with one GPU.
  2. In Nodes, click Add node and run the install command on your GPU machine. It installs Ollama and the gpuos-agent service; the agent only makes outbound HTTPS connections.
  3. In Models, deploy a model. The VRAM gauge shows whether it fits before anything downloads.
  4. In API keys, create a key, optionally with a monthly token quota and a requests-per-minute limit.
  5. Call https://gpuos.si/v1 with any OpenAI client.
First request
curl https://gpuos.si/v1/chat/completions \
  -H "Authorization: Bearer $GPUOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen3-32b", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'

What production use needs

  • One key per app or teammate, so you can revoke one without breaking the others.
  • Quotas and rate limits per key: a monthly token cap and a requests-per-minute limit stop one runaway script from saturating the GPU.
  • Metering: prompt and completion tokens, latency and status per key, model and node.
  • OpenAI error format: 401, 403, 404, 429 and 503 errors that SDKs already know how to handle.
  • No inbound port: the agent connects out, so the GPU box stays behind its firewall.

gpuos covers all five. The quickstart lists every error code and its meaning.

Questions

Is a self-hosted model as good as GPT-class APIs?
For many internal tasks, yes: 32B models such as Qwen3 32B and GLM-4 32B handle summaries, extraction, RAG answers and code assistance well. The largest frontier models still lead on the hardest reasoning tasks.
What hardware do I need?
An NVIDIA GPU with 8 GB of VRAM runs 8B models; 24 GB runs 32B models at 4-bit. See the VRAM guide for every model.
Does my code have to change?
Only the base URL, the API key and the model name.

Related

Run it on your own GPU

Free for one GPU. Connect a machine in one command and call your models through one OpenAI-compatible API.