gpuos

Docs · 15 minutes

An OpenAI-compatible API on your own GPU

This guide takes a GPU server, a homelab box or a rented machine to a private LLM endpoint your code can call with the official OpenAI SDK. You need a Linux server with an NVIDIA GPU, a Mac with Apple Silicon or a Windows PC with an NVIDIA GPU, and administrator access.

  1. 1

    Create a workspace

    Sign up with email, GitHub or Google. Your first workspace is created on the free Community plan: one node, one GPU, two members and the base model catalog.

  2. 2

    Connect your GPU machine

    In Nodes, click Add node and give it a name. gpuos shows a one-line install command with a token for that node. Run it on the machine:

    Linux and macOS (Terminal)
    curl -fsSL https://gpuos.si/install.sh | sudo sh -s -- --token gpuos_node_…
    Windows (PowerShell as administrator)
    $env:GPUOS_TOKEN="gpuos_node_…"; irm https://gpuos.si/install.ps1 | iex

    The dashboard shows the command for your system. It installs Ollama if it is missing and runs gpuos-agent in the background: a systemd service on Linux, a launchd service on macOS, a scheduled task on Windows. Within seconds the node shows up as online with its GPUs and memory. On a Mac, the GPU budget is about two thirds to three quarters of the unified memory, and the agent keeps the Mac awake while it runs.

  3. 3

    Deploy a model

    Open Models and pick one from the catalog. Before you confirm, the VRAM gauge shows how much memory the model takes on that node, for example “6.5 GB of 24 GB” for Qwen3 8B. The agent downloads the weights and the model turns Ready.

  4. 4

    Create an API key

    In API keys, create a key per app or teammate. You can cap it with a monthly token quota and a requests-per-minute limit. The full key is shown once; store it as GPUOS_API_KEY.

  5. 5

    Call it like OpenAI

    Point any OpenAI client at https://gpuos.si/v1. Chat completions, completions, embeddings and streaming are supported. GET /v1/models lists the models deployed in your workspace.

    Python
    from openai import OpenAI
    
    client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…")
    stream = client.chat.completions.create(
        model="qwen3-8b",
        messages=[{"role": "user", "content": "Summarize our refund policy."}],
        stream=True,
    )
    for chunk in stream:
        print(chunk.choices[0].delta.content or "", end="")
    Node.js
    import OpenAI from "openai"
    
    const client = new OpenAI({ baseURL: "https://gpuos.si/v1", apiKey: process.env.GPUOS_API_KEY })
    const reply = await client.chat.completions.create({
      model: "qwen3-8b",
      messages: [{ role: "user", content: "Summarize our refund policy." }],
    })
    console.log(reply.choices[0].message.content)
    curl
    curl https://gpuos.si/v1/chat/completions \
      -H "Authorization: Bearer $GPUOS_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model": "qwen3-8b", "messages": [{"role": "user", "content": "Hello"}]}'
    Embeddings (Python)
    from openai import OpenAI
    
    client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…")
    result = client.embeddings.create(model="bge-m3", input=["first text", "second text"])
    print(len(result.data[0].embedding))

Errors

Errors use the OpenAI format ({"error": {"message", "type", "code"}}), so SDKs raise their usual exceptions.

StatusCodeMeaning
401invalid_api_keyThe key is missing, wrong or revoked.
403plan_requiredThe model is in the Pro catalog and the workspace is on Community.
404model_not_foundThe model id is not in the gpuos catalog.
429rate_limit_exceededThe key's requests-per-minute limit is reached. Retry after the retry-after header.
429insufficient_quotaThe key used its monthly token quota.
503model_not_deployedNo node in the workspace has this model deployed.
503node_unavailableThe model is deployed, but no node serving it is online.

Questions

Do I need to open a port on my GPU server?
No. The gpuos agent only makes outbound HTTPS connections: it sends a heartbeat every 10 seconds and holds a long poll to receive requests. Your server can stay behind a firewall or NAT.
Which machines are supported?
Linux with systemd (Ubuntu 22.04 or newer recommended) on x86_64 or ARM64 with an NVIDIA GPU; Macs with Apple Silicon (M1 or newer), which use their unified memory; and 64-bit Windows 10 or 11, ideally with an NVIDIA GPU. Without a GPU, models still run on the CPU, slowly.
Does it work with LangChain, LlamaIndex, Continue or Open WebUI?
Yes. Anything that accepts an OpenAI base URL and API key works: set the base URL to https://gpuos.si/v1 and use your gpuos key and a model id from the catalog. The integrations pages have copy-paste setups for each tool.
How are tokens counted?
Prompt and completion tokens come from the engine's usage report for every request, streamed or not, and are shown per key, model and node on the Usage page.

Your first token in 15 minutes

Free for one GPU, no card needed. Upgrade to Pro for unlimited nodes and the full model catalog.