Free hardware planning tool
LLM VRAM calculator: which models fit your GPU?
Choose an open model and a GPU to compare their memory requirements. Qwen3-8B is estimated at 6.5 GB in Q4_K_M, Qwen3-32B at 22 GB, and gpt-oss-20b at 14 GB in MXFP4. No account required.
Check a model against your GPU
Fits according to the estimate
- Estimated requirement
- ≈ 6.5 GB
- Available VRAM
- 24 GB
- Remaining VRAM
- 17.5 GB
Catalog estimates, not measured benchmarks. Chat estimates assume a short 4–8K context and one request. Embedding memory depends on input length and batch size. Longer contexts, concurrent requests and multiple loaded models need additional memory; 5% headroom is a screening rule, not a guarantee.
How the estimate works
We use the gpuOS catalog’s per-model memory estimates for the listed quantization. These already account for runtime overhead at short context; we do not add another generic KV-cache estimate. A model passes the fit check when its estimate uses at most 95% of your GPU’s VRAM.
MoE models need memory for their total resident weights, not just the experts active per token. MXFP4 is not interchangeable with GGUF Q4_K_M. BGE-M3 is an embedding model, so this tool does not estimate a generation KV cache for it.
This is a capacity check for a single model on a single GPU. It does not predict tokens per second or memory at the model’s maximum context window. Confirm the exact runtime, context and workload on your hardware before production use.
VRAM estimates by model
| Model | Quantization | Estimated VRAM |
|---|---|---|
| Qwen3 32B | Q4_K_M | ≈ 22 GB |
| GLM-4 32B 0414 | Q4_K_M | ≈ 20.5 GB |
| Qwen3 30B A3B | Q4_K_M | ≈ 20 GB |
| gpt-oss 20B | MXFP4 | ≈ 14 GB |
| DeepSeek R1 Distill Qwen 32B | Q4_K_M | ≈ 22 GB |
| GLM-Z1 32B 0414 | Q4_K_M | ≈ 20.5 GB |
| Mistral Small 3.2 24B | Q4_K_M | ≈ 17 GB |
| Qwen3 8B | Q4_K_M | ≈ 6.5 GB |
| GLM-4 9B | Q4_K_M | ≈ 6.5 GB |
| Qwen3-VL 8B | Q4_K_M | ≈ 7.5 GB |
| Gemma 3 27B | Q4_K_M | ≈ 18 GB |
| BGE-M3 | FP16 | ≈ 1.5 GB |
Read the VRAM planning guide for context and concurrency limits, or follow the Qwen3 deployment guide for a 24 GB GPU.