All models › Guides › GGUF Quantizations Explained: Q4_K_M, IQ3, Q8_0 — Which to Pick?
GGUF Quantizations Explained: Q4_K_M, IQ3, Q8_0 — Which to Pick?
Quantization shrinks model weights from 16-bit floats to fewer bits. Less memory, slightly worse outputs. Here is the decoder ring.
The naming scheme
- Q4 = ~4 bits per weight (4.5-5 in practice — some tensors stay 16-bit). K = k-quant (block-wise scales, better than legacy Q4_0/Q4_1). M/L/S/XL = medium/large/small/extra-large variant — which tensors get more bits.
- IQ3_XXS / IQ2_M = importance-matrix quants: calibrated on real text, better quality per bit than plain Q3/Q2, slightly slower to load.
- UD-* = unsloth's dynamic quants: they keep attention/embeddings at higher precision automatically.
- Q8_0 ≈ visually lossless. F16/BF16 = the original model, no compression.
Memory cost per quant
| Model size | Q2_K | Q3_K_M | IQ4_XS | Q4_K_M | Q5_K_M | Q6_K | Q8_0 |
|---|---|---|---|---|---|---|---|
| 8B | 5.2 | 5.8 | 6.2 | 6.9 | 7.8 | 8.8 | 10.9 |
| 14B | 8.0 | 9.0 | 9.8 | 10.9 | 12.5 | 14.2 | 17.9 |
| 32B | 16.5 | 18.7 | 20.4 | 23.1 | 26.6 | 30.5 | 38.9 |
| 70B | 34.2 | 39.0 | 42.9 | 48.7 | 56.4 | 65.0 | 83.3 |
Rule-of-thumb: params × bits/8 + ~10% tensor overhead + ~1.5GB KV cache at 8K context. Exact numbers per model (measured GGUF files) are on each model page.
Which one do I pick?
- Default: Q4_K_M (or IQ4_XS / UD-Q4_K_XL). ~2-3% quality loss vs F16 on most benchmarks — the community consensus sweet spot.
- VRAM headroom? Q5_K_M/Q6_K — diminishing returns but free quality.
- VRAM tight? IQ3_XXS/Q3_K_M — noticeable degradation on reasoning tasks; only if it means running at all.
- Below Q3 — chat still works, math/code/logic suffer visibly. Last resort.
Every ModelFit model page lists the measured file size of each quant actually published for that model — pick the largest one that fits your card.
Doesn't fit your machine? Rent a GPU by the hour instead — a 24GB RTX 4090 starts around $0.30-0.50/h (Vast.ai, RunPod), or use the model via a hosted API. Buying instead? Which GPU should I buy? →
More guides
How Much VRAM Do You Need to Run a Local LLM? (7B, 13B, 70B)VRAM needed to run 7B, 13B, 30B and 70B LLMs locally at every quantiza…GGUF Quantizations Explained: Q4_K_M, IQ3, Q8_0 — Which to Pick?What Q4_K_M, Q5_K_M, IQ3_XXS, Q8_0 and F16 actually mean, how much qua…Ollama vs llama.cpp vs LM Studio: Which Local LLM Runtime?Honest comparison of Ollama, llama.cpp and LM Studio for running local…RTX 3090 vs RTX 4090 for Local LLMs: Real Numbers3090 vs 4090 for LLM inference: same 24GB VRAM, 936 vs 1008 GB/s bandw…How to Run a 70B LLM on 24GB VRAM (Honest Options)Can you run Llama 70B or DeepSeek on one 24GB GPU? Three real options:…