gpuos

Hardware · 2 min read · updated Oct 6, 2026

How much VRAM do you need to run an LLM?

VRAM needed for Qwen3, GLM-4, gpt-oss, DeepSeek R1 distill, Mistral Small, Gemma 3 and more, the rule of thumb behind it, and which GPUs fit each model.

The rule of thumb

At 4-bit quantization (Q4_K_M), a model needs about 0.6 GB of VRAM per billion parameters, plus 1 to 2 GB for the KV cache and runtime overhead at a short context of 4 to 8K tokens.

  • 8B model: about 5 to 7 GB
  • 24B model: about 15 to 17 GB
  • 32B model: about 20 to 22 GB
  • 70B model: about 42 to 45 GB, beyond a single 24 GB card

Mixture-of-experts models need memory for all their weights, not only the active ones: Qwen3 30B A3B activates 3B parameters per token but still needs about 20 GB.

VRAM per model

ModelJobQuantVRAMSmallest GPU that fits
BGE-M3EmbeddingsFP161.5 GBRTX 4060 Ti 8 GB
Qwen3 8BSmall & fastQ4_K_M6.5 GBRTX 4060 Ti 8 GB
GLM-4 9BSmall & fastQ4_K_M6.5 GBRTX 4060 Ti 8 GB
Qwen3-VL 8BVisionQ4_K_M7.5 GBRTX 4060 Ti 8 GB
gpt-oss 20BAgentic codingMXFP414 GBRTX 4070 Ti Super 16 GB
Mistral Small 3.2 24BChatQ4_K_M17 GBRTX 4000 Ada 20 GB
Gemma 3 27BChatQ4_K_M18 GBRTX 4000 Ada 20 GB
Qwen3 30B A3BFast MoEQ4_K_M20 GBRTX 3090 / 4090 24 GB
GLM-4 32B 0414ChatQ4_K_M20.5 GBRTX 3090 / 4090 24 GB
GLM-Z1 32B 0414ReasoningQ4_K_M20.5 GBRTX 3090 / 4090 24 GB
Qwen3 32BChatQ4_K_M22 GBRTX 3090 / 4090 24 GB
DeepSeek R1 Distill Qwen 32BReasoningQ4_K_M22 GBRTX 3090 / 4090 24 GB

Figures are for short contexts. Long prompts and many concurrent requests grow the KV cache, so keep 10 to 20% headroom.

Context length and concurrency

The KV cache grows linearly with the number of tokens in flight. A 32B model that fits in 22 GB at 4K tokens can run out of memory at 32K tokens, or with eight long requests at once. If you need long contexts on a 24 GB card, prefer a 24B or MoE model, or a smaller quantization.

Picking a GPU

  • 8 GB (RTX 4060 Ti): 8 to 9B models, embeddings, vision at 8B.
  • 16 GB (RTX 4070 Ti Super): gpt-oss 20B, Mistral Small 24B with care.
  • 24 GB (RTX 3090 / 4090, L4, RTX PRO 4000 Blackwell in Hetzner's GEX45): every 32B model in the catalog.
  • 96 GB (RTX PRO 6000 Blackwell): 70B models and long contexts.

gpuos shows a VRAM gauge with your node's real memory before every deployment, so you see the fit before anything downloads.

Questions

Can I run a 32B model on a 24 GB GPU?
Yes, at 4-bit (Q4_K_M): Qwen3 32B needs about 22 GB and GLM-4 32B about 20.5 GB at short context.
Can I run two models on one GPU?
Yes if their combined VRAM fits. Otherwise Ollama swaps them in and out, and the first request after a swap is slower.
Does the CPU or system RAM matter?
Much less. Models that do not fit in VRAM can spill to system RAM, but generation becomes many times slower.

Related

Run it on your own GPU

Free for one GPU. Connect a machine in one command and call your models through one OpenAI-compatible API.