All models › Guides › How Much VRAM Do You Need to Run a Local LLM? (7B, 13B, 70B)
How Much VRAM Do You Need to Run a Local LLM? (7B, 13B, 70B)
The short answer: params × bits ÷ 8, plus KV cache, plus ~10% overhead. But the exact file sizes differ between repos — always check the measured size of the GGUF you actually download. Here is the full picture.
VRAM by model size and quantization
| Model size | Q2_K | Q3_K_M | IQ4_XS | Q4_K_M | Q5_K_M | Q6_K | Q8_0 |
|---|---|---|---|---|---|---|---|
| 7B | 4.8 | 5.3 | 5.6 | 6.2 | 7.0 | 7.9 | 9.7 |
| 8B | 5.2 | 5.8 | 6.2 | 6.9 | 7.8 | 8.8 | 10.9 |
| 13B | 7.6 | 8.5 | 9.2 | 10.3 | 11.7 | 13.3 | 16.7 |
| 14B | 8.0 | 9.0 | 9.8 | 10.9 | 12.5 | 14.2 | 17.9 |
| 27B | 14.1 | 16.0 | 17.5 | 19.7 | 22.7 | 26.0 | 33.1 |
| 32B | 16.5 | 18.7 | 20.4 | 23.1 | 26.6 | 30.5 | 38.9 |
| 70B | 34.2 | 39.0 | 42.9 | 48.7 | 56.4 | 65.0 | 83.3 |
| 235B | 111.4 | 127.5 | 140.4 | 159.8 | 185.7 | 214.8 | 276.2 |
Rule-of-thumb: params × bits/8 + ~10% tensor overhead + ~1.5GB KV cache at 8K context. Exact numbers per model (measured GGUF files) are on each model page.
The three costs people forget
- KV cache grows with context. 8K context on a 27B model costs ~1-4GB; 128K context can cost more than the weights themselves. Use the context selector on any model page to see it.
- Runtime overhead. CUDA context, compute buffers and fragmentation eat ~8-15% — that's why a 23.5GB model does not fit a 24GB card.
- MoE models keep every expert in RAM/VRAM. A 30B-A3B MoE with 3B active params still needs the full 30B's bytes resident — it is only faster than dense 30B, not smaller.
What fits what (rule of thumb)
- 8GB — 7-8B models at Q4_K_M, 12-14B at Q2/Q3. Ranked list →
- 12GB — 8B at Q8, 14B at Q4-Q5. Ranked list →
- 16GB — 14B at Q6-Q8, 27-32B at Q3-Q4. Ranked list →
- 24GB — 27-32B at Q4-Q5, 70B MoE (A3B) at Q4. Ranked list →
- 48GB+ — 70B dense at Q4-Q5. Ranked list →
Not enough VRAM? Three escape hatches
CPU+GPU offload (llama.cpp --n-gpu-layers): splits layers between VRAM and system RAM — works, but every offloaded layer pays PCIe speed. Smaller quant: Q4_K_M → IQ3/Q2 saves 30-50% memory at a real quality cost. Rent a GPU: a cloud RTX 4090 costs ~$0.30-0.50/h when you need the big model just occasionally.
Doesn't fit your machine? Rent a GPU by the hour instead — a 24GB RTX 4090 starts around $0.30-0.50/h (Vast.ai, RunPod), or use the model via a hosted API. Buying instead? Which GPU should I buy? →
More guides
How Much VRAM Do You Need to Run a Local LLM? (7B, 13B, 70B)VRAM needed to run 7B, 13B, 30B and 70B LLMs locally at every quantiza…GGUF Quantizations Explained: Q4_K_M, IQ3, Q8_0 — Which to Pick?What Q4_K_M, Q5_K_M, IQ3_XXS, Q8_0 and F16 actually mean, how much qua…Ollama vs llama.cpp vs LM Studio: Which Local LLM Runtime?Honest comparison of Ollama, llama.cpp and LM Studio for running local…RTX 3090 vs RTX 4090 for Local LLMs: Real Numbers3090 vs 4090 for LLM inference: same 24GB VRAM, 936 vs 1008 GB/s bandw…How to Run a 70B LLM on 24GB VRAM (Honest Options)Can you run Llama 70B or DeepSeek on one 24GB GPU? Three real options:…