gpuos

Free hardware planning tool

LLM VRAM calculator: which models fit your GPU?

Choose an open model and a GPU to compare their memory requirements. Qwen3-8B is estimated at 6.5 GB in Q4_K_M, Qwen3-32B at 22 GB, and gpt-oss-20b at 14 GB in MXFP4. No account required.

Check a model against your GPU

Fits according to the estimate

Estimated requirement
≈ 6.5 GB
Available VRAM
24 GB
Remaining VRAM
17.5 GB

Catalog estimates, not measured benchmarks. Chat estimates assume a short 4–8K context and one request. Embedding memory depends on input length and batch size. Longer contexts, concurrent requests and multiple loaded models need additional memory; 5% headroom is a screening rule, not a guarantee.

How the estimate works

We use the gpuOS catalog’s per-model memory estimates for the listed quantization. These already account for runtime overhead at short context; we do not add another generic KV-cache estimate. A model passes the fit check when its estimate uses at most 95% of your GPU’s VRAM.

MoE models need memory for their total resident weights, not just the experts active per token. MXFP4 is not interchangeable with GGUF Q4_K_M. BGE-M3 is an embedding model, so this tool does not estimate a generation KV cache for it.

This is a capacity check for a single model on a single GPU. It does not predict tokens per second or memory at the model’s maximum context window. Confirm the exact runtime, context and workload on your hardware before production use.

VRAM estimates by model

ModelQuantizationEstimated VRAM
Qwen3 32BQ4_K_M≈ 22 GB
GLM-4 32B 0414Q4_K_M≈ 20.5 GB
Qwen3 30B A3BQ4_K_M≈ 20 GB
gpt-oss 20BMXFP4≈ 14 GB
DeepSeek R1 Distill Qwen 32BQ4_K_M≈ 22 GB
GLM-Z1 32B 0414Q4_K_M≈ 20.5 GB
Mistral Small 3.2 24BQ4_K_M≈ 17 GB
Qwen3 8BQ4_K_M≈ 6.5 GB
GLM-4 9BQ4_K_M≈ 6.5 GB
Qwen3-VL 8BQ4_K_M≈ 7.5 GB
Gemma 3 27BQ4_K_M≈ 18 GB
BGE-M3FP16≈ 1.5 GB

Read the VRAM planning guide for context and concurrency limits, or follow the Qwen3 deployment guide for a 24 GB GPU.