The rule of thumb
At 4-bit quantization (Q4_K_M), a model needs about 0.6 GB of VRAM per billion parameters, plus 1 to 2 GB for the KV cache and runtime overhead at a short context of 4 to 8K tokens.
- 8B model: about 5 to 7 GB
- 24B model: about 15 to 17 GB
- 32B model: about 20 to 22 GB
- 70B model: about 42 to 45 GB, beyond a single 24 GB card
Mixture-of-experts models need memory for all their weights, not only the active ones: Qwen3 30B A3B activates 3B parameters per token but still needs about 20 GB.
VRAM per model
| Model | Job | Quant | VRAM | Smallest GPU that fits |
|---|---|---|---|---|
| BGE-M3 | Embeddings | FP16 | 1.5 GB | RTX 4060 Ti 8 GB |
| Qwen3 8B | Small & fast | Q4_K_M | 6.5 GB | RTX 4060 Ti 8 GB |
| GLM-4 9B | Small & fast | Q4_K_M | 6.5 GB | RTX 4060 Ti 8 GB |
| Qwen3-VL 8B | Vision | Q4_K_M | 7.5 GB | RTX 4060 Ti 8 GB |
| gpt-oss 20B | Agentic coding | MXFP4 | 14 GB | RTX 4070 Ti Super 16 GB |
| Mistral Small 3.2 24B | Chat | Q4_K_M | 17 GB | RTX 4000 Ada 20 GB |
| Gemma 3 27B | Chat | Q4_K_M | 18 GB | RTX 4000 Ada 20 GB |
| Qwen3 30B A3B | Fast MoE | Q4_K_M | 20 GB | RTX 3090 / 4090 24 GB |
| GLM-4 32B 0414 | Chat | Q4_K_M | 20.5 GB | RTX 3090 / 4090 24 GB |
| GLM-Z1 32B 0414 | Reasoning | Q4_K_M | 20.5 GB | RTX 3090 / 4090 24 GB |
| Qwen3 32B | Chat | Q4_K_M | 22 GB | RTX 3090 / 4090 24 GB |
| DeepSeek R1 Distill Qwen 32B | Reasoning | Q4_K_M | 22 GB | RTX 3090 / 4090 24 GB |
Figures are for short contexts. Long prompts and many concurrent requests grow the KV cache, so keep 10 to 20% headroom.
Context length and concurrency
The KV cache grows linearly with the number of tokens in flight. A 32B model that fits in 22 GB at 4K tokens can run out of memory at 32K tokens, or with eight long requests at once. If you need long contexts on a 24 GB card, prefer a 24B or MoE model, or a smaller quantization.
Picking a GPU
- 8 GB (RTX 4060 Ti): 8 to 9B models, embeddings, vision at 8B.
- 16 GB (RTX 4070 Ti Super): gpt-oss 20B, Mistral Small 24B with care.
- 24 GB (RTX 3090 / 4090, L4, RTX PRO 4000 Blackwell in Hetzner's GEX45): every 32B model in the catalog.
- 96 GB (RTX PRO 6000 Blackwell): 70B models and long contexts.
gpuos shows a VRAM gauge with your node's real memory before every deployment, so you see the fit before anything downloads.