Which GPU should I buy to run LLMs locally? (2026)
Pick any model — we show the cheapest hardware that runs it, from measured GGUF sizes at the recommended quant. Prices are street/used-market approximations (USD), refreshed monthly.
Cheapest hardware by memory tier (static table)
| Hardware | Usable VRAM | Street price | Runs (recommended quant, 8K) |
|---|---|---|---|
| Used RTX 3070 8GB | 7 GB | ~$220 | LFM2.5-2.6B, cohere-transcribe-03-2026 |
| New RTX 4060 8GB | 7 GB | ~$280 | LFM2.5-2.6B, cohere-transcribe-03-2026 |
| Used RTX 3060 12GB | 11 GB | ~$250 | Ternary-Bonsai-2-27B, Gemma-4-E4B-Uncensored-HauhauCS-Aggressive |
| Used RTX 3080 12GB | 11 GB | ~$380 | Ternary-Bonsai-2-27B, Gemma-4-E4B-Uncensored-HauhauCS-Aggressive |
| Used RTX 4070 12GB | 11 GB | ~$430 | Ternary-Bonsai-2-27B, Gemma-4-E4B-Uncensored-HauhauCS-Aggressive |
| New RTX 4060 Ti 16GB | 15 GB | ~$420 | Ternary-Bonsai-2-27B, gemma-4-E4B-it |
| New RTX 5080 16GB | 15 GB | ~$1000 | Ternary-Bonsai-2-27B, gemma-4-E4B-it |
| Used RTX 3090 24GB | 22 GB | ~$650 | Qwen3.8-27B, Ternary-Bonsai-2-27B |
| Used RTX 4090 24GB | 22 GB | ~$1600 | Qwen3.8-27B, Ternary-Bonsai-2-27B |
| New RTX 5090 32GB | 29 GB | ~$2500 | Qwen3.8-27B, Ternary-Bonsai-2-27B |
| MacBook Air 16GB (M-series) | 11 GB | ~$1000 | Ternary-Bonsai-2-27B, Gemma-4-E4B-Uncensored-HauhauCS-Aggressive |
| MacBook Pro 24GB | 17 GB | ~$1600 | Ternary-Bonsai-2-27B, gemma-4-E4B-it |
| MacBook Pro Max 64GB | 44 GB | ~$3200 | Qwen3.8-27B, Ternary-Bonsai-2-27B |
| Mac Studio Ultra 256GB | 177 GB | ~$6000 | Qwen3.8-27B, Ternary-Bonsai-2-27B |
Usable = physical × 0.92 (fragmentation headroom); Macs × 0.75 (macOS wired limit). Two cheap 3090s (48GB pooled, ~$1300) beat one 4090 for models needing 25-48GB if your board has 2× x8 slots.
Doesn't fit your machine? Rent a GPU by the hour instead — a 24GB RTX 4090 starts around $0.30-0.50/h (Vast.ai, RunPod), or use the model via a hosted API. Buying instead? Which GPU should I buy? →