Can I run this LLM? Real file sizes, live from Hugging Face.
All models › Guides › GGUF Quantizations Explained: Q4_K_M, IQ3, Q8_0 — Which to Pick?

GGUF Quantizations Explained: Q4_K_M, IQ3, Q8_0 — Which to Pick?

Quantization shrinks model weights from 16-bit floats to fewer bits. Less memory, slightly worse outputs. Here is the decoder ring.

The naming scheme

Memory cost per quant

Model sizeQ2_KQ3_K_MIQ4_XSQ4_K_MQ5_K_MQ6_KQ8_0
8B5.25.86.26.97.88.810.9
14B8.09.09.810.912.514.217.9
32B16.518.720.423.126.630.538.9
70B34.239.042.948.756.465.083.3

Rule-of-thumb: params × bits/8 + ~10% tensor overhead + ~1.5GB KV cache at 8K context. Exact numbers per model (measured GGUF files) are on each model page.

Which one do I pick?

  1. Default: Q4_K_M (or IQ4_XS / UD-Q4_K_XL). ~2-3% quality loss vs F16 on most benchmarks — the community consensus sweet spot.
  2. VRAM headroom? Q5_K_M/Q6_K — diminishing returns but free quality.
  3. VRAM tight? IQ3_XXS/Q3_K_M — noticeable degradation on reasoning tasks; only if it means running at all.
  4. Below Q3 — chat still works, math/code/logic suffer visibly. Last resort.

Every ModelFit model page lists the measured file size of each quant actually published for that model — pick the largest one that fits your card.

Doesn't fit your machine? Rent a GPU by the hour instead — a 24GB RTX 4090 starts around $0.30-0.50/h (Vast.ai, RunPod), or use the model via a hosted API. Buying instead? Which GPU should I buy? →

More guides

How Much VRAM Do You Need to Run a Local LLM? (7B, 13B, 70B)VRAM needed to run 7B, 13B, 30B and 70B LLMs locally at every quantiza…GGUF Quantizations Explained: Q4_K_M, IQ3, Q8_0 — Which to Pick?What Q4_K_M, Q5_K_M, IQ3_XXS, Q8_0 and F16 actually mean, how much qua…Ollama vs llama.cpp vs LM Studio: Which Local LLM Runtime?Honest comparison of Ollama, llama.cpp and LM Studio for running local…RTX 3090 vs RTX 4090 for Local LLMs: Real Numbers3090 vs 4090 for LLM inference: same 24GB VRAM, 936 vs 1008 GB/s bandw…How to Run a 70B LLM on 24GB VRAM (Honest Options)Can you run Llama 70B or DeepSeek on one 24GB GPU? Three real options:…
Share this page: 𝕏 Post Reddit Hacker News Telegram WhatsApp More…