How to Run a 70B LLM on 24GB VRAM (Honest Options)
A 70B dense model at Q4_K_M measures ~40-43GB. It does not fit 24GB. Here is what actually works, in order of sanity.
Option 1 — run a MoE model instead (best)
Modern MoE models give near-70B quality at 70B-total / 3B-active scale. A 30B-A3B MoE at Q4 measures ~18-20GB and fits one 24GB card — and decodes fast because only ~3B params are read per token. Check measured sizes: models that fit 24GB.
Option 2 — CPU+GPU offload (slow but real)
llama.cpp can split layers: hot layers in VRAM, the rest in system RAM (--n-gpu-layers 20 + as many as fit). Rule: offloaded layers run at RAM bandwidth (~50GB/s vs ~1000GB/s VRAM). A 70B with 60% offloaded decodes at ~2-4 tok/s — usable for batch jobs, painful for chat. Needs 48-64GB system RAM.
Option 3 — extreme quants (quality cost)
IQ2_XXS/Q2_K shrink 70B to ~24-27GB — it might squeeze in with offload, but at Q2 quality reasoning degrades badly. Almost always a worse deal than option 1.
Option 4 — rent, don't buy
If you need a true dense 70B occasionally: an A100/H100 hour on a marketplace costs less than a coffee. For daily use, two used 3090s (48GB) run 70B Q4 natively for ~$1400.
Measured examples
Exact file sizes per quant for every 70B-class model we track — see each model page from the 48GB list and 64GB list.