Can I run this LLM? Real file sizes, live from Hugging Face.
All models › Guides › How Much VRAM Do You Need to Run a Local LLM? (7B, 13B, 70B)

How Much VRAM Do You Need to Run a Local LLM? (7B, 13B, 70B)

The short answer: params × bits ÷ 8, plus KV cache, plus ~10% overhead. But the exact file sizes differ between repos — always check the measured size of the GGUF you actually download. Here is the full picture.

VRAM by model size and quantization

Model sizeQ2_KQ3_K_MIQ4_XSQ4_K_MQ5_K_MQ6_KQ8_0
7B4.85.35.66.27.07.99.7
8B5.25.86.26.97.88.810.9
13B7.68.59.210.311.713.316.7
14B8.09.09.810.912.514.217.9
27B14.116.017.519.722.726.033.1
32B16.518.720.423.126.630.538.9
70B34.239.042.948.756.465.083.3
235B111.4127.5140.4159.8185.7214.8276.2

Rule-of-thumb: params × bits/8 + ~10% tensor overhead + ~1.5GB KV cache at 8K context. Exact numbers per model (measured GGUF files) are on each model page.

The three costs people forget

  1. KV cache grows with context. 8K context on a 27B model costs ~1-4GB; 128K context can cost more than the weights themselves. Use the context selector on any model page to see it.
  2. Runtime overhead. CUDA context, compute buffers and fragmentation eat ~8-15% — that's why a 23.5GB model does not fit a 24GB card.
  3. MoE models keep every expert in RAM/VRAM. A 30B-A3B MoE with 3B active params still needs the full 30B's bytes resident — it is only faster than dense 30B, not smaller.

What fits what (rule of thumb)

Not enough VRAM? Three escape hatches

CPU+GPU offload (llama.cpp --n-gpu-layers): splits layers between VRAM and system RAM — works, but every offloaded layer pays PCIe speed. Smaller quant: Q4_K_M → IQ3/Q2 saves 30-50% memory at a real quality cost. Rent a GPU: a cloud RTX 4090 costs ~$0.30-0.50/h when you need the big model just occasionally.

Doesn't fit your machine? Rent a GPU by the hour instead — a 24GB RTX 4090 starts around $0.30-0.50/h (Vast.ai, RunPod), or use the model via a hosted API. Buying instead? Which GPU should I buy? →

More guides

How Much VRAM Do You Need to Run a Local LLM? (7B, 13B, 70B)VRAM needed to run 7B, 13B, 30B and 70B LLMs locally at every quantiza…GGUF Quantizations Explained: Q4_K_M, IQ3, Q8_0 — Which to Pick?What Q4_K_M, Q5_K_M, IQ3_XXS, Q8_0 and F16 actually mean, how much qua…Ollama vs llama.cpp vs LM Studio: Which Local LLM Runtime?Honest comparison of Ollama, llama.cpp and LM Studio for running local…RTX 3090 vs RTX 4090 for Local LLMs: Real Numbers3090 vs 4090 for LLM inference: same 24GB VRAM, 936 vs 1008 GB/s bandw…How to Run a 70B LLM on 24GB VRAM (Honest Options)Can you run Llama 70B or DeepSeek on one 24GB GPU? Three real options:…
Share this page: 𝕏 Post Reddit Hacker News Telegram WhatsApp More…