Choose the workload before the largest model
A 24 GB GPU can be a useful starting point for quantized chat, coding and reasoning models, plus smaller vision and embedding models. In the gpuOS catalog, Qwen3 32B and Qwen3-Coder 30B A3B clear the short-context screening check when loaded alone. That does not establish their maximum usable context or serving capacity.
For an interactive assistant, begin with a smaller baseline and compare its answers with a larger candidate. For RAG, budget retrieval and generation separately. For a coding tool, test complete repository tasks, not a short greeting. Memory fit narrows the candidates; your evaluation chooses the model.
Every memory value below comes from the current gpuOS catalog. These are estimates, not measured results on an RTX 3090, RTX 4090 or NVIDIA L4. The cards have different hardware characteristics even when their memory capacity matches.
A shortlist by task, with memory headroom
Scroll horizontally to see every column.
| Model | Evaluation use | Quantization | Estimated VRAM | 24 GB minus estimate | What to check |
|---|---|---|---|---|---|
| Qwen3 8B | Extraction and a first chat baseline | Q4_K_M | 6.5 GB | 17.5 GB | Leave room to test longer prompts and other workloads. |
| Qwen3 14B | Chat with more memory headroom than 32B | Q4_K_M | 11.5 GB | 12.5 GB | Evaluate quality before moving to a larger model. |
| Qwen3 32B | Single-model chat evaluation | Q4_K_M | 22 GB | 2 GB | Short context first; little room for another model. |
| Qwen2.5-Coder 14B | Code completion and editing evaluation | Q4_K_M | 12 GB | 12 GB | Run your own repository tasks and tests. |
| Qwen3-Coder 30B A3B | Larger coding-model evaluation | Q4_K_M | 22 GB | 2 GB | Resident MoE weights still consume memory. |
| gpt-oss 20B | Tool-use, code and reasoning evaluation | MXFP4 | 14 GB | 10 GB | Keep the native MXFP4 format and verify tool behavior in your runtime. |
| DeepSeek R1 Distill Qwen 32B | Reasoning evaluation | Q4_K_M | 22 GB | 2 GB | Check answer correctness and reasoning output length. |
| Qwen3-VL 8B | Image and document understanding | Q4_K_M | 7.5 GB | 16.5 GB | Test real image sizes and the actual input route. |
| BGE-M3 | Dense retrieval and RAG indexing | FP16 | 1.5 GB | 22.5 GB | Embedding batches need their own memory test. |
| Qwen3 Embedding 0.6B | Instruction-aware retrieval evaluation | Q8_0 | 2 GB | 22 GB | Use the same embedding configuration for the index and queries. |
The estimate assumes one short-context request, generally 4–8K tokens for text generation. The remaining-memory column is simple subtraction, not usable extra context or promised capacity. Image inputs, embedding batches, other GPU processes and the runtime can change the peak.
Browse the full model catalog for other candidates or use the VRAM calculator for a different GPU size. Check the model's license before putting it into a customer-facing application.
Quantization, MoE and context are separate decisions
Keep the exact model tag and quantization in your test record. A result for Q4_K_M does not describe the same model at FP16 or another quantization. Test quality after any weight-format change using the same tasks and acceptance criteria.
Qwen3-Coder 30B A3B is estimated at 22 GB in this catalog. Its active expert count does not remove the other experts from the resident weight budget. gpt-oss 20B uses MXFP4 and is estimated at 14 GB; do not substitute a generic Q4 estimate for its native format.
A published context window is an architectural limit, not a 24 GB capacity guarantee. The catalog's gpt-oss 120B estimate is 80 GB, so it is not a single-24-GB-GPU candidate here. CPU offloading and multi-GPU deployment are different configurations and need separate evaluation.
Can chat and embeddings stay loaded together?
Add the estimates if you want two models resident at the same time. The table uses the same 5% screening margin as the calculator. Passing this arithmetic is only a first check; it does not model simultaneous cache allocation or batching.
Scroll horizontally to see every column.
| Pair | Catalog estimates | Sum | 24 GB screening result |
|---|---|---|---|
| Qwen3 8B + BGE-M3 | 6.5 GB + 1.5 GB | 8 GB | Clears the screening margin; validate both workloads together. |
| Qwen2.5-Coder 14B + Qwen3 Embedding 0.6B | 12 GB + 2 GB | 14 GB | Clears the screening margin; validate both workloads together. |
| Qwen3 32B + BGE-M3 | 22 GB + 1.5 GB | 23.5 GB | Fits the arithmetic only; misses the 5% screening margin. |
| Qwen3-Coder 30B A3B + Qwen3-VL 8B | 22 GB + 7.5 GB | 29.5 GB | Exceeds capacity; do not plan simultaneous residency. |
Alternatively, index documents in a separate batch and let the embedding model unload before generation. That reduces overlap but introduces model-loading work when queries need fresh embeddings. Measure this tradeoff against your latency requirement.
Ollama can queue requests and unload idle models when the next model needs memory. A deployment entry is not proof that every model is simultaneously loaded. Inspect the runtime during overlapping requests. Ollama memory and concurrency behavior.
Verify memory and GPU offloading on your machine
On a machine with Ollama already installed and running, start with one model and a bounded context. The following request uses Ollama's native local route so its context option is explicit; it is not a request to the gpuOS gateway.
ollama --version
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv
ollama pull qwen3:8b
curl --fail-with-body http://127.0.0.1:11434/api/chat \
-H "Content-Type: application/json" \
-d '{"model":"qwen3:8b","stream":false,"messages":[{"role":"user","content":"Return a short checklist for reviewing an API key rotation."}],"options":{"num_ctx":4096,"num_predict":128},"keep_alive":"5m"}'
ollama ps
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csvCheck the PROCESSOR column in ollama ps for GPU residency and, when present in your version, the CONTEXT column. Record the model digest/tag and runtime version. Watch nvidia-smi during requests to capture the peak; a reading after completion alone can miss a transient allocation.
If the model uses a CPU/GPU split, label the result as offloaded. A successful response does not prove full GPU residency. Increase context deliberately and repeat the check, since runtime defaults can vary by version and available memory. Ollama context configuration.
Test the workload that will actually run
- Prepare representative inputs: short, typical and near-limit prompts; actual images for vision; realistic repository files for code. Use labeled expected outcomes where possible.
- Record the configuration: GPU, driver, Ollama version, model digest, quantization, context, parallelism and competing GPU processes.
- Separate cold and warm requests: include model-load time in the cold result. Record errors and latency distributions rather than one favorable request.
- Increase one variable at a time: context first, then overlap two requests, then add a second model. Track peak memory, queueing and offloading at each step.
- Check answers as well as speed: extraction accuracy, code tests, citation correctness or image-reading errors. Keep the smallest candidate that passes your requirements.
If the larger model misses your latency or memory target, shorten the supplied context, retrieve fewer passages, serialize work, evaluate a smaller model or move to larger hardware. A higher API rate limit cannot increase the GPU's capacity.
Serve the selected model to an application
For a local prototype, keep Ollama on the local machine. For a shared gpuOS deployment, follow the quickstart, connect the GPU node, deploy the catalog model and create a server-side API key. Use catalog IDs with gpuOS and Ollama tags with the local engine.
gpuOS routes requests through its hosted gateway to Ollama on your connected node. Prompts and results pass through that gateway, so this is not an air-gapped data path. Review that flow before moving sensitive material into a shared application.
Choose the next recipe for your workload: OpenAI Python, Continue, Aider or the embeddings API tutorial.