# gpuos

> gpuos.si turns your NVIDIA GPUs into a private super intelligence (SI) cloud. Deploy open models like Qwen3, GLM-4 and gpt-oss in one click and serve them behind one OpenAI-compatible API with keys, quotas and usage metering. Your data stays on your hardware, in the EU.

- Positioning: The GPU OS for super intelligence. SI stands for super intelligence, the name the US government adopted for AI in September 2026; .si is also Slovenia's country domain. gpuos runs SI (AI) models on hardware you control.
- What it is: a control plane, an agent and an OpenAI-compatible gateway that turn your own NVIDIA GPUs into a private AI cloud.
- API: OpenAI-compatible at https://gpuos.si/v1 (GET /models, POST /chat/completions, /completions, /embeddings, with streaming). Works with the official OpenAI SDKs, LangChain and LlamaIndex by changing the base URL and key.
- Agent: gpuos-agent (Rust) runs on Linux with systemd, reports GPUs through nvidia-smi, manages models in Ollama and only makes outbound HTTPS connections, so no inbound port is needed.
- Install: `curl -fsSL https://gpuos.si/install.sh | sudo sh -s -- --token gpuos_node_…` with a token from the dashboard.
- Governance: API keys per app or user, monthly token quotas, requests-per-minute limits, usage per key, model and node, workspaces with owner/admin/member roles.
- Hosting: built and operated in the EU (France).

## Docs

- [Quickstart: an OpenAI-compatible API on your own GPU](https://gpuos.si/docs/quickstart): install the agent, deploy a model, create a key, call it from the OpenAI SDK.
- [Model catalog](https://gpuos.si/models): every supported model with VRAM, license and GPU fit.
- [Super intelligence (SI) on your own GPUs](https://gpuos.si/super-intelligence): what SI and .si mean, and how to run private SI models.

## Guides

- [How to self-host an OpenAI-compatible API on your own GPU](https://gpuos.si/guides/self-host-openai-compatible-api): Run open LLMs like Qwen3 and GLM-4 on your own GPU and expose them through an OpenAI-compatible API with keys, quotas and metering, without opening a port.
- [How much VRAM do you need to run an LLM?](https://gpuos.si/guides/how-much-vram-to-run-an-llm): VRAM needed for Qwen3, GLM-4, gpt-oss, DeepSeek R1 distill, Mistral Small, Gemma 3 and more, the rule of thumb behind it, and which GPUs fit each model.
- [Run Qwen3 32B on a Hetzner GPU server (GEX45)](https://gpuos.si/guides/run-qwen3-on-hetzner-gex45): Set up a Hetzner GEX45 with an RTX PRO 4000 Blackwell 24 GB, install the NVIDIA driver and gpuos, and serve Qwen3 32B through an OpenAI-compatible API.
- [Private RAG on your own GPU: embeddings and chat, self-hosted](https://gpuos.si/guides/private-rag-on-your-own-gpu): Build retrieval-augmented generation that never sends documents to a third party: BGE-M3 embeddings and Qwen3 32B answers on your GPU, called through LangChain.
- [GDPR and LLMs: keeping AI and SI data in the EU](https://gpuos.si/guides/gdpr-compliant-llm-eu): What changes for GDPR when you run open models on your own GPUs instead of a US API, what gpuos stores, and a checklist for your records of processing.
- [An Ollama alternative for teams: Ollama with gpuOS](https://gpuos.si/guides/ollama-vs-gpuos): Ollama is the easiest way to run an LLM on one machine. gpuos builds on it for teams: API keys, quotas, metering, several nodes and no open port.

## Integrations

- [OpenAI Python SDK](https://gpuos.si/integrations/openai-python): Use the official openai package with your own GPUs by changing the base URL and the key.
- [OpenAI Node.js SDK](https://gpuos.si/integrations/openai-node): Call self-hosted models from TypeScript or JavaScript with the official openai package.
- [LangChain](https://gpuos.si/integrations/langchain): ChatOpenAI and OpenAIEmbeddings work with gpuos models for chains, agents and RAG.
- [LlamaIndex](https://gpuos.si/integrations/llamaindex): Use gpuos models in LlamaIndex with the OpenAILike LLM and an OpenAI-compatible embedding model.
- [Vercel AI SDK](https://gpuos.si/integrations/vercel-ai-sdk): Stream gpuos models into Next.js and React apps with the AI SDK's OpenAI-compatible provider.
- [Continue](https://gpuos.si/integrations/continue): Use self-hosted coding models in VS Code and JetBrains through Continue.
- [Open WebUI](https://gpuos.si/integrations/open-webui): Give your team a ChatGPT-style interface on top of models running on your GPUs.
- [n8n](https://gpuos.si/integrations/n8n): Run n8n AI workflows and agents on your own GPUs with an OpenAI credential pointing at gpuos.
- [Aider](https://gpuos.si/integrations/aider): Pair-program in your terminal with coding models running on your GPUs.
- [LiteLLM](https://gpuos.si/integrations/litellm): Add your own GPUs to a LiteLLM proxy next to cloud providers, with gpuos as an OpenAI-compatible backend.

## Models

- [Qwen3 32B](https://gpuos.si/models/qwen3-32b): Chat, 32B dense, about 22 GB VRAM at Q4_K_M, Apache 2.0. API model id `qwen3-32b`.
- [GLM-4 32B 0414](https://gpuos.si/models/glm-4-32b): Chat, 32B dense, about 20.5 GB VRAM at Q4_K_M, MIT. API model id `glm-4-32b`.
- [Qwen3 30B A3B](https://gpuos.si/models/qwen3-30b-a3b): Fast MoE, 30B MoE (3B active), about 20 GB VRAM at Q4_K_M, Apache 2.0. API model id `qwen3-30b-a3b`.
- [gpt-oss 20B](https://gpuos.si/models/gpt-oss-20b): Agentic coding, 21B MoE, about 14 GB VRAM at MXFP4, Apache 2.0. API model id `gpt-oss-20b`.
- [DeepSeek R1 Distill Qwen 32B](https://gpuos.si/models/deepseek-r1-distill-qwen-32b): Reasoning, 32B dense, about 22 GB VRAM at Q4_K_M, MIT. API model id `deepseek-r1-distill-qwen-32b`.
- [GLM-Z1 32B 0414](https://gpuos.si/models/glm-z1-32b): Reasoning, 32B dense, about 20.5 GB VRAM at Q4_K_M, MIT. API model id `glm-z1-32b`.
- [Mistral Small 3.2 24B](https://gpuos.si/models/mistral-small-24b): Chat, 24B dense, about 17 GB VRAM at Q4_K_M, Apache 2.0. API model id `mistral-small-24b`.
- [Qwen3 8B](https://gpuos.si/models/qwen3-8b): Small & fast, 8B dense, about 6.5 GB VRAM at Q4_K_M, Apache 2.0. API model id `qwen3-8b`.
- [GLM-4 9B](https://gpuos.si/models/glm-4-9b): Small & fast, 9B dense, about 6.5 GB VRAM at Q4_K_M, MIT. API model id `glm-4-9b`.
- [Qwen3-VL 8B](https://gpuos.si/models/qwen3-vl-8b): Vision, 8B dense, about 7.5 GB VRAM at Q4_K_M, Apache 2.0. API model id `qwen3-vl-8b`.
- [Gemma 3 27B](https://gpuos.si/models/gemma-3-27b): Chat, 27B dense, about 18 GB VRAM at Q4_K_M, Gemma Terms of Use. API model id `gemma-3-27b`.
- [BGE-M3](https://gpuos.si/models/bge-m3): Embeddings, 568M, about 1.5 GB VRAM at FP16, MIT. API model id `bge-m3`.

## Pricing

- Community: free. 1 node, 1 GPU, 2 members, base model catalog.
- Pro: $29 per GPU per month, or $290 per GPU per year. Unlimited nodes and members, full model catalog, usage metering.
- Team add-on: +$99 per workspace per month (SAML, roles, audit log, advanced routing policies).
- gpuOS Cloud: +$49 per month for a control plane hosted by gpuos.
- Prices in USD, excluding VAT.

## Sibling product: cpuOS

gpuOS and cpuOS are two halves of one operating system for super intelligence: models think on GPUs (gpuOS), agents act on CPUs (cpuOS). cpuOS gives every AI agent task its own Firecracker microVM sandbox in under a second, to run the code a model wrote, drive a headless browser or test a repository. Hosted in the EU or self-hosted, billed per second, free while paused. Agents built on the gpuOS API use cpuOS for their tool calls.

- [cpuOS home](https://cpuos.si): secure sandboxes for AI agents.
- [cpuOS guides](https://cpuos.si/guides): running LLM-generated code safely, Firecracker vs Docker, building a code interpreter with open models.
- [cpuOS integrations](https://cpuos.si/integrations): LangChain, Vercel AI SDK, OpenAI Agents SDK, MCP and more.
- [cpuOS llms.txt](https://cpuos.si/llms.txt)

## Optional

- [Full text for LLMs](https://gpuos.si/llms-full.txt)
- [Home page](https://gpuos.si)
