gpuos

Hardware · 4 min read · updated Oct 7, 2026

LLM quantization explained: choose Q4, Q8, FP16 or BF16

Choose a quantized LLM for your GPU. Compare Q4, Q8, FP16 and BF16, estimate weight memory, and test answer quality with real context lengths.

On this page

What changes when you quantize model weights

Quantization represents model weights with fewer bits. It reduces storage and memory use, with a possible change in answer quality. A quantized model still needs memory for its context, inference buffers and runtime. Download size alone cannot tell you whether a deployment will fit your GPU. Hugging Face explains the precision and accuracy tradeoff.

Compare variants of the same base model before changing model families. If you switch both the model and its precision, you cannot tell which change caused a different answer. Record the original repository, exact variant, prompt template and inference settings alongside your evaluation results.

Q4, Q8, FP16 and BF16 refer to different choices

Scroll horizontally to see every column.

FormatWhat the label describesWhat to check
Q4_K_MA GGUF quantization recipe using roughly four-bit weights, with mixed tensor precisionExact artifact size and quality on your prompts
Q8_0An eight-bit GGUF weight quantization formatWhether the larger variant leaves enough context memory
FP16 / F1616-bit floating-point weightsHardware support and total runtime memory
BF16A different 16-bit floating-point representationModel availability and runtime support

FP16 and BF16 use the same number of bits per weight but have different numerical ranges and precision. Q8_0 is not an FP8 format. Likewise, a GGUF file is a container, not a guarantee that every tensor has four-bit precision. The llama.cpp quantization reference describes recipes, mixed tensors and quantization options.

Use the full suffix when sharing a deployment configuration. Saying only ‘Q4’ hides differences between recipes and publishers. In gpuOS, choose the exact catalog variant, then verify the deployed model rather than assuming an upstream model name identifies one unique artifact.

Calculate weight bytes, then add the rest of memory

The basic arithmetic is parameter count multiplied by bits per weight, divided by eight. For a hypothetical model with exactly eight billion parameters, the following values describe idealized weight storage in decimal GB. They exclude quantization metadata, higher-precision tensors, KV cache and runtime allocations. They are calculation examples, not measured VRAM requirements for an 8B catalog model.

Scroll horizontally to see every column.

Ideal precisionCalculationWeight bytes only
4 bits8,000,000,000 × 4 ÷ 84 GB
8 bits8,000,000,000 × 8 ÷ 88 GB
16 bits8,000,000,000 × 16 ÷ 816 GB

Keep units consistent when comparing a file listing with a memory monitor: one decimal GB is one billion bytes, while one GiB is 1,073,741,824 bytes. The gpuOS VRAM calculator uses catalog estimates to help shortlist variants. Confirm the result under your intended workload before buying hardware or assigning traffic.

Evaluate the errors that matter to your application

Build a small evaluation set before choosing the smallest download. Include ordinary requests and the cases that cause expensive mistakes: extracting a number from a table, keeping a required JSON field, citing the right passage, following a language instruction or selecting the right tool. Write expected outcomes so another teammate can repeat the check.

  • Compare the same prompts on two precisions of the same base model, keeping context and generation settings fixed.
  • Score task success separately from writing style. A fluent answer can contain the wrong identifier or omit a required condition.
  • Include long documents and multilingual inputs if they appear in production, rather than evaluating only short English chat.
  • Inspect failures directly. Record which inputs fail and whether a larger variant actually fixes them.

Choose acceptance criteria around the product: valid fields, correct citations, useful code or accurate classification. A model that passes your checks at Q4 can be a sensible deployment; one that fails a required behavior needs a different variant, model or application design. No quantization label guarantees task accuracy.

Leave room for context and concurrent requests

A precision change is useful only when the complete workload fits. Test the maximum prompt size you intend to accept, the output allowance and simultaneous requests. Do not fill the GPU with weights and treat the remaining memory as an unlimited chat history budget. The context and KV cache guide explains the separate capacity check.

gpuOS runs Ollama on the GPU machine you connect. On that machine, inspect ollama ps and your GPU memory monitor while requests run. A catalog estimate is a planning input; loaded memory, CPU offloading and latency are deployment observations. Save both in your notes so teammates can reproduce the decision.

Make the final choice with a repeatable comparison

  1. Shortlist two available variants that fit your estimated hardware budget.
  2. Run the task evaluation, including the longest representative inputs.
  3. Measure warm latency and output throughput using the local LLM benchmark method.
  4. Test the expected request concurrency and keep the results, exact artifact and runtime version together.
  5. Deploy the variant that meets your quality and service targets with memory headroom, then repeat after material model or runtime changes.

Questions

Does Q4 mean an 8B model needs exactly 4 GB of VRAM?
No. That arithmetic covers idealized four-bit weight storage only. Quantization metadata, mixed-precision tensors, KV cache and runtime allocations add memory. Validate the catalog estimate on the GPU with your real context and concurrency.
Is Q8 always better than Q4 for a local LLM?
A larger precision variant can preserve behavior that a smaller one loses, but its value depends on the task and available memory. Compare the same model on representative prompts, then measure whether the quality difference matters to your application.
Can I compare FP16 and BF16 by download size?
Both are 16-bit floating-point formats, so ideal weight storage is similar. Their numerical representation differs. Check the actual model artifact, inference support and total loaded memory rather than choosing from the label alone.

Related

Run it on your own GPU

Connect your GPU, deploy a catalog model and test the hosted API on a representative request. The quickstart explains the setup and data flow.