gpuos

Hardware · 6 min read · updated Oct 7, 2026

Which LLMs can run on a 24 GB GPU?

Choose LLMs for a 24 GB GPU by task, quantization and memory headroom. Compare catalog estimates and verify context, concurrency and offloading.

On this page

Choose the workload before the largest model

A 24 GB GPU can be a useful starting point for quantized chat, coding and reasoning models, plus smaller vision and embedding models. In the gpuOS catalog, Qwen3 32B and Qwen3-Coder 30B A3B clear the short-context screening check when loaded alone. That does not establish their maximum usable context or serving capacity.

For an interactive assistant, begin with a smaller baseline and compare its answers with a larger candidate. For RAG, budget retrieval and generation separately. For a coding tool, test complete repository tasks, not a short greeting. Memory fit narrows the candidates; your evaluation chooses the model.

Every memory value below comes from the current gpuOS catalog. These are estimates, not measured results on an RTX 3090, RTX 4090 or NVIDIA L4. The cards have different hardware characteristics even when their memory capacity matches.

A shortlist by task, with memory headroom

Scroll horizontally to see every column.

ModelEvaluation useQuantizationEstimated VRAM24 GB minus estimateWhat to check
Qwen3 8BExtraction and a first chat baselineQ4_K_M6.5 GB17.5 GBLeave room to test longer prompts and other workloads.
Qwen3 14BChat with more memory headroom than 32BQ4_K_M11.5 GB12.5 GBEvaluate quality before moving to a larger model.
Qwen3 32BSingle-model chat evaluationQ4_K_M22 GB2 GBShort context first; little room for another model.
Qwen2.5-Coder 14BCode completion and editing evaluationQ4_K_M12 GB12 GBRun your own repository tasks and tests.
Qwen3-Coder 30B A3BLarger coding-model evaluationQ4_K_M22 GB2 GBResident MoE weights still consume memory.
gpt-oss 20BTool-use, code and reasoning evaluationMXFP414 GB10 GBKeep the native MXFP4 format and verify tool behavior in your runtime.
DeepSeek R1 Distill Qwen 32BReasoning evaluationQ4_K_M22 GB2 GBCheck answer correctness and reasoning output length.
Qwen3-VL 8BImage and document understandingQ4_K_M7.5 GB16.5 GBTest real image sizes and the actual input route.
BGE-M3Dense retrieval and RAG indexingFP161.5 GB22.5 GBEmbedding batches need their own memory test.
Qwen3 Embedding 0.6BInstruction-aware retrieval evaluationQ8_02 GB22 GBUse the same embedding configuration for the index and queries.

The estimate assumes one short-context request, generally 4–8K tokens for text generation. The remaining-memory column is simple subtraction, not usable extra context or promised capacity. Image inputs, embedding batches, other GPU processes and the runtime can change the peak.

Browse the full model catalog for other candidates or use the VRAM calculator for a different GPU size. Check the model's license before putting it into a customer-facing application.

Quantization, MoE and context are separate decisions

Keep the exact model tag and quantization in your test record. A result for Q4_K_M does not describe the same model at FP16 or another quantization. Test quality after any weight-format change using the same tasks and acceptance criteria.

Qwen3-Coder 30B A3B is estimated at 22 GB in this catalog. Its active expert count does not remove the other experts from the resident weight budget. gpt-oss 20B uses MXFP4 and is estimated at 14 GB; do not substitute a generic Q4 estimate for its native format.

A published context window is an architectural limit, not a 24 GB capacity guarantee. The catalog's gpt-oss 120B estimate is 80 GB, so it is not a single-24-GB-GPU candidate here. CPU offloading and multi-GPU deployment are different configurations and need separate evaluation.

Can chat and embeddings stay loaded together?

Add the estimates if you want two models resident at the same time. The table uses the same 5% screening margin as the calculator. Passing this arithmetic is only a first check; it does not model simultaneous cache allocation or batching.

Scroll horizontally to see every column.

PairCatalog estimatesSum24 GB screening result
Qwen3 8B + BGE-M36.5 GB + 1.5 GB8 GBClears the screening margin; validate both workloads together.
Qwen2.5-Coder 14B + Qwen3 Embedding 0.6B12 GB + 2 GB14 GBClears the screening margin; validate both workloads together.
Qwen3 32B + BGE-M322 GB + 1.5 GB23.5 GBFits the arithmetic only; misses the 5% screening margin.
Qwen3-Coder 30B A3B + Qwen3-VL 8B22 GB + 7.5 GB29.5 GBExceeds capacity; do not plan simultaneous residency.

Alternatively, index documents in a separate batch and let the embedding model unload before generation. That reduces overlap but introduces model-loading work when queries need fresh embeddings. Measure this tradeoff against your latency requirement.

Ollama can queue requests and unload idle models when the next model needs memory. A deployment entry is not proof that every model is simultaneously loaded. Inspect the runtime during overlapping requests. Ollama memory and concurrency behavior.

Verify memory and GPU offloading on your machine

On a machine with Ollama already installed and running, start with one model and a bounded context. The following request uses Ollama's native local route so its context option is explicit; it is not a request to the gpuOS gateway.

Local short-context memory check
ollama --version
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv
ollama pull qwen3:8b
curl --fail-with-body http://127.0.0.1:11434/api/chat \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3:8b","stream":false,"messages":[{"role":"user","content":"Return a short checklist for reviewing an API key rotation."}],"options":{"num_ctx":4096,"num_predict":128},"keep_alive":"5m"}'
ollama ps
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv

Check the PROCESSOR column in ollama ps for GPU residency and, when present in your version, the CONTEXT column. Record the model digest/tag and runtime version. Watch nvidia-smi during requests to capture the peak; a reading after completion alone can miss a transient allocation.

If the model uses a CPU/GPU split, label the result as offloaded. A successful response does not prove full GPU residency. Increase context deliberately and repeat the check, since runtime defaults can vary by version and available memory. Ollama context configuration.

Test the workload that will actually run

  1. Prepare representative inputs: short, typical and near-limit prompts; actual images for vision; realistic repository files for code. Use labeled expected outcomes where possible.
  2. Record the configuration: GPU, driver, Ollama version, model digest, quantization, context, parallelism and competing GPU processes.
  3. Separate cold and warm requests: include model-load time in the cold result. Record errors and latency distributions rather than one favorable request.
  4. Increase one variable at a time: context first, then overlap two requests, then add a second model. Track peak memory, queueing and offloading at each step.
  5. Check answers as well as speed: extraction accuracy, code tests, citation correctness or image-reading errors. Keep the smallest candidate that passes your requirements.

If the larger model misses your latency or memory target, shorten the supplied context, retrieve fewer passages, serialize work, evaluate a smaller model or move to larger hardware. A higher API rate limit cannot increase the GPU's capacity.

Serve the selected model to an application

For a local prototype, keep Ollama on the local machine. For a shared gpuOS deployment, follow the quickstart, connect the GPU node, deploy the catalog model and create a server-side API key. Use catalog IDs with gpuOS and Ollama tags with the local engine.

gpuOS routes requests through its hosted gateway to Ollama on your connected node. Prompts and results pass through that gateway, so this is not an air-gapped data path. Review that flow before moving sensitive material into a shared application.

Choose the next recipe for your workload: OpenAI Python, Continue, Aider or the embeddings API tutorial.

Questions

What is the best LLM for an RTX 4090 or RTX 3090?
There is no task-independent winner. Both provide a 24 GB memory budget, but memory fit does not establish speed or answer quality. Start with a smaller baseline, compare a larger candidate on your workload and measure on the actual card.
Can a 32B model fit in 24 GB?
Qwen3 32B is estimated at 22 GB at Q4_K_M for a short-context request. That clears the catalog screening margin on 24 GB, but long context, parallel requests and other resident models need a separate memory test.
Can I use the model's maximum context on a 24 GB GPU?
The published context window does not guarantee that its full allocation fits on your card. Set the context explicitly, measure peak memory and inspect GPU offloading before increasing it.
Does selecting several models in gpuOS keep all of them loaded?
Deployment availability and runtime residency are different. Ollama loads models on demand and can evict idle models. Check overlapping requests and model-loading latency if your application depends on two models being available together.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.