gpuos

Agentic coding · Free plan

Run gpt-oss 20B on your own GPU

OpenAI's open-weight 21B MoE model, shipped natively in MXFP4. Built for agentic coding and tool use, with a long 128K context, and it fits in about 14 GB.

With gpuos, gpt-oss 20B runs on Ollama at MXFP4 on your machine and is served as gpt-oss-20b through one OpenAI-compatible endpoint, with API keys, quotas and usage metering.

Parameters
21B MoE
Quantization
MXFP4
VRAM (4–8K ctx)
≈ 14 GB
Context window
128K tokens
License
Apache 2.0
Engine
Ollama
Ollama tag
gpt-oss:20b
API
/v1/chat/completions

What gpt-oss 20B is good at

Which GPUs can run gpt-oss 20B?

Catalog estimate at MXFP4, not a measured benchmark. Chat estimates assume a short context; embedding memory depends on input and batch size. More in how much VRAM an LLM needs.

Check this estimate against your GPU with the VRAM calculator

Local setup and workload validation

Install a GPU-compatible Ollama runtime, confirm the driver with nvidia-smi, then pull and inspect the exact tag. Download size is not the same as runtime VRAM.

On your GPU machine
ollama pull gpt-oss:20b
ollama show gpt-oss:20b
ollama run gpt-oss:20b "Reply with a short greeting"
ollama ps

The 14 GB estimate uses native MXFP4, not Q4_K_M. All resident expert weights matter for memory. A 128K model context window does not imply a 128K request fits within that estimate; validate tool calls and context with your runtime.

For a reproducible check, record the GPU, driver, Ollama version, model tag and quantization. Test one short input first, then the actual context and batch/concurrency you need. Separate model loading from warm inference and record peak memory and errors. No throughput benchmark is claimed on this page.

Runtime source: Ollama model card. Tags can change; inspect the pulled model before comparing results.

Call gpt-oss 20B with the OpenAI SDK

Same SDKs, same request format. Only the base URL, the key and the model name change. Setup for LangChain, Continue, Open WebUI and more is in integrations.

Python
from openai import OpenAI

client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…")
stream = client.chat.completions.create(
    model="gpt-oss-20b",
    messages=[{"role": "user", "content": "Summarize our refund policy."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
Node.js
import OpenAI from "openai"

const client = new OpenAI({ baseURL: "https://gpuos.si/v1", apiKey: process.env.GPUOS_API_KEY })
const reply = await client.chat.completions.create({
  model: "gpt-oss-20b",
  messages: [{ role: "user", content: "Summarize our refund policy." }],
})
console.log(reply.choices[0].message.content)
curl
curl https://gpuos.si/v1/chat/completions \
  -H "Authorization: Bearer $GPUOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-oss-20b", "messages": [{"role": "user", "content": "Hello"}]}'

Questions

How much VRAM does gpt-oss 20B need?
About 14 GB at MXFP4 with a short context (4–8K tokens). The smallest common GPU that fits it is the RTX 4070 Ti Super 16 GB. Longer contexts and more concurrent requests need extra headroom for the KV cache.
Can I use gpt-oss 20B commercially?
Yes. gpt-oss 20B is released under the Apache 2.0 license, which allows commercial use.
Is gpt-oss 20B compatible with the OpenAI API?
Yes. Through gpuos, gpt-oss 20B is served at /v1/chat/completions with the model id "gpt-oss-20b", so the official OpenAI SDKs, LangChain and LlamaIndex work by changing the base URL and the API key.
Which gpuos plan includes gpt-oss 20B?
gpt-oss 20B is in the base catalog, available on the free Community plan (1 node, 1 GPU).

Related models

Plan your gpt-oss 20B deployment with gpuOS

Join early access for onboarding, or read the quickstart to evaluate the Community workflow on your own GPU.