Small & fast · Free plan
Run Qwen3 8B on your own GPU
A small, fast Qwen3 model that runs on 8 GB cards. A good default for classification, extraction, routing and short answers at high volume.
With gpuos, Qwen3 8B runs on Ollama at Q4_K_M on your machine and is served as qwen3-8b through one OpenAI-compatible endpoint, with API keys, quotas and usage metering.
- Parameters
- 8B dense
- Quantization
- Q4_K_M
- VRAM (4–8K ctx)
- ≈ 6.5 GB
- Context window
- 32K tokens
- License
- Apache 2.0
- Engine
- Ollama
- Ollama tag
- qwen3:8b
- API
- /v1/chat/completions
What Qwen3 8B is good at
- Classification and extraction
- High-volume tasks
- Small GPUs
Which GPUs can run Qwen3 8B?
Catalog estimate at Q4_K_M, not a measured benchmark. Chat estimates assume a short context; embedding memory depends on input and batch size. More in how much VRAM an LLM needs.
Check this estimate against your GPU with the VRAM calculator
- RTX 4060 Ti 8 GBFits
- RTX 4070 Ti Super 16 GBFits
- RTX 4000 Ada 20 GBFits
- RTX 3090 / 4090 24 GBFits
- RTX PRO 4000 Blackwell 24 GB (Hetzner GEX45)Fits
- NVIDIA L4 24 GBFits
- RTX 5090 32 GBFits
- A100 / H100 80 GBFits
- RTX PRO 6000 Blackwell 96 GB (Hetzner GEX131)Fits
Local setup and workload validation
Install a GPU-compatible Ollama runtime, confirm the driver with nvidia-smi, then pull and inspect the exact tag. Download size is not the same as runtime VRAM.
ollama pull qwen3:8b
ollama show qwen3:8b
ollama run qwen3:8b "Reply with a short greeting"
ollama psThe 6.5 GB Q4_K_M estimate is for short context. An 8 GB GPU leaves little room for larger context or concurrent requests. Thinking mode can produce substantially more output than a short-answer task.
For a reproducible check, record the GPU, driver, Ollama version, model tag and quantization. Test one short input first, then the actual context and batch/concurrency you need. Separate model loading from warm inference and record peak memory and errors. No throughput benchmark is claimed on this page.
Runtime source: Ollama model card. Tags can change; inspect the pulled model before comparing results.
Call Qwen3 8B with the OpenAI SDK
Same SDKs, same request format. Only the base URL, the key and the model name change. Setup for LangChain, Continue, Open WebUI and more is in integrations.
from openai import OpenAI
client = OpenAI(base_url="https://gpuos.si/v1", api_key="gpuos_key_…")
stream = client.chat.completions.create(
model="qwen3-8b",
messages=[{"role": "user", "content": "Summarize our refund policy."}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")import OpenAI from "openai"
const client = new OpenAI({ baseURL: "https://gpuos.si/v1", apiKey: process.env.GPUOS_API_KEY })
const reply = await client.chat.completions.create({
model: "qwen3-8b",
messages: [{ role: "user", content: "Summarize our refund policy." }],
})
console.log(reply.choices[0].message.content)curl https://gpuos.si/v1/chat/completions \
-H "Authorization: Bearer $GPUOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "qwen3-8b", "messages": [{"role": "user", "content": "Hello"}]}'Questions
- How much VRAM does Qwen3 8B need?
- About 6.5 GB at Q4_K_M with a short context (4–8K tokens). The smallest common GPU that fits it is the RTX 4060 Ti 8 GB. Longer contexts and more concurrent requests need extra headroom for the KV cache.
- Can I use Qwen3 8B commercially?
- Yes. Qwen3 8B is released under the Apache 2.0 license, which allows commercial use.
- Is Qwen3 8B compatible with the OpenAI API?
- Yes. Through gpuos, Qwen3 8B is served at /v1/chat/completions with the model id "qwen3-8b", so the official OpenAI SDKs, LangChain and LlamaIndex work by changing the base URL and the API key.
- Which gpuos plan includes Qwen3 8B?
- Qwen3 8B is in the base catalog, available on the free Community plan (1 node, 1 GPU).
Related models
Plan your Qwen3 8B deployment with gpuOS
Join early access for onboarding, or read the quickstart to evaluate the Community workflow on your own GPU.