Docs · 15 minutes
An OpenAI-compatible API on your own GPU
This guide takes a GPU server, a homelab box or a rented machine to a private LLM endpoint your code can call with the official OpenAI SDK. You need a Linux server with an NVIDIA GPU, a Mac with Apple Silicon or a Windows PC with an NVIDIA GPU, and administrator access.
- 1
Create a workspace
Sign up with email, GitHub or Google. Your first workspace is created on the free Community plan: one node, one GPU, two members and the base model catalog.
- 2
Connect your GPU machine
In Nodes, click Add node and give it a name. gpuos shows a one-line install command with a token for that node. Run it on the machine:
Linux and macOS (Terminal) curl -fsSL https://gpuos.si/install.sh | sudo sh -s -- --token gpuos_node_…Windows (PowerShell as administrator) $env:GPUOS_TOKEN="gpuos_node_…"; irm https://gpuos.si/install.ps1 | iexThe dashboard shows the command for your system. It installs Ollama if it is missing and runs
gpuos-agentin the background: a systemd service on Linux, a launchd service on macOS, a scheduled task on Windows. Within seconds the node shows up as online with its GPUs and memory. On a Mac, the GPU budget is about two thirds to three quarters of the unified memory, and the agent keeps the Mac awake while it runs. - 3
Deploy a model
Open Models and pick one from the catalog. Before you confirm, the VRAM gauge shows how much memory the model takes on that node, for example “6.5 GB of 24 GB” for Qwen3 8B. The agent downloads the weights and the model turns Ready.
- 4
Create an API key
In API keys, create a key per app or teammate. You can cap it with a monthly token quota and a requests-per-minute limit. The full key is shown once; store it as
GPUOS_API_KEY. - 5
Call it like OpenAI
Point any OpenAI client at
https://gpuos.si/v1. Chat completions, completions, embeddings and streaming are supported.GET /v1/modelslists the models deployed in your workspace.Python from openai import OpenAI client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…") stream = client.chat.completions.create( model="qwen3-8b", messages=[{"role": "user", "content": "Summarize our refund policy."}], stream=True, ) for chunk in stream: print(chunk.choices[0].delta.content or "", end="")Node.js import OpenAI from "openai" const client = new OpenAI({ baseURL: "https://gpuos.si/v1", apiKey: process.env.GPUOS_API_KEY }) const reply = await client.chat.completions.create({ model: "qwen3-8b", messages: [{ role: "user", content: "Summarize our refund policy." }], }) console.log(reply.choices[0].message.content)curl curl https://gpuos.si/v1/chat/completions \ -H "Authorization: Bearer $GPUOS_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "qwen3-8b", "messages": [{"role": "user", "content": "Hello"}]}'Embeddings (Python) from openai import OpenAI client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…") result = client.embeddings.create(model="bge-m3", input=["first text", "second text"]) print(len(result.data[0].embedding))
Errors
Errors use the OpenAI format ({"error": {"message", "type", "code"}}), so SDKs raise their usual exceptions.
| Status | Code | Meaning |
|---|---|---|
| 401 | invalid_api_key | The key is missing, wrong or revoked. |
| 403 | plan_required | The model is in the Pro catalog and the workspace is on Community. |
| 404 | model_not_found | The model id is not in the gpuos catalog. |
| 429 | rate_limit_exceeded | The key's requests-per-minute limit is reached. Retry after the retry-after header. |
| 429 | insufficient_quota | The key used its monthly token quota. |
| 503 | model_not_deployed | No node in the workspace has this model deployed. |
| 503 | node_unavailable | The model is deployed, but no node serving it is online. |
Questions
- Do I need to open a port on my GPU server?
- No. The gpuos agent only makes outbound HTTPS connections: it sends a heartbeat every 10 seconds and holds a long poll to receive requests. Your server can stay behind a firewall or NAT.
- Which machines are supported?
- Linux with systemd (Ubuntu 22.04 or newer recommended) on x86_64 or ARM64 with an NVIDIA GPU; Macs with Apple Silicon (M1 or newer), which use their unified memory; and 64-bit Windows 10 or 11, ideally with an NVIDIA GPU. Without a GPU, models still run on the CPU, slowly.
- Does it work with LangChain, LlamaIndex, Continue or Open WebUI?
- Yes. Anything that accepts an OpenAI base URL and API key works: set the base URL to https://gpuos.si/v1 and use your gpuos key and a model id from the catalog. The integrations pages have copy-paste setups for each tool.
- How are tokens counted?
- Prompt and completion tokens come from the engine's usage report for every request, streamed or not, and are shown per key, model and node on the Usage page.
Your first token in 15 minutes
Free for one GPU, no card needed. Upgrade to Pro for unlimited nodes and the full model catalog.