Can I run this LLM? Real file sizes, live from Hugging Face.
All models › gemma-3-12b-it-qat

gemma-3-12b-it-qat requirements — can you run it?

~11.8B parameters · 26 quantizations measured · updated 2025-05-12 · source: unsloth/gemma-3-12b-it-qat-GGUF

Every quantization: real file size vs what you actually need

QuantFile (measured)VRAM/RAM needed (8K ctx)Runs on (first fit)
UD-IQ1_S3.1 GB7.3 GBRTX 3060 12GB
UD-IQ1_M3.2 GB7.5 GBRTX 3060 12GB
UD-IQ2_XXS3.6 GB7.8 GBRTX 3060 12GB
UD-IQ2_M4.4 GB8.6 GBRTX 3060 12GB
Q2_K4.8 GB9.0 GBRTX 3060 12GB
Q2_K_L4.8 GB9.0 GBRTX 3060 12GB
UD-IQ3_XXS4.8 GB9.1 GBRTX 3060 12GB
UD-Q2_K_XL4.9 GB9.1 GBRTX 3060 12GB
Q3_K_S5.5 GB9.7 GBRTX 3060 12GB
Q3_K_M6.0 GB10.2 GBRTX 3060 12GB
UD-Q3_K_XL6.1 GB10.4 GBRTX 3060 12GB
IQ4_XS6.6 GB10.8 GBRTX 3060 12GB
IQ4_NL6.9 GB11.1 GBRTX 4060 Ti 16GB
Q4_06.9 GB11.1 GBRTX 4060 Ti 16GB
Q4_K_S6.9 GB11.2 GBRTX 4060 Ti 16GB
Q4_K_M7.3 GB11.5 GBRTX 4060 Ti 16GB
UD-Q4_K_XL7.4 GB11.7 GBRTX 4060 Ti 16GB
Q4_17.6 GB11.8 GBRTX 4060 Ti 16GB
Q5_K_S8.2 GB12.5 GBRTX 4060 Ti 16GB
Q5_K_M8.4 GB12.7 GBRTX 4060 Ti 16GB
UD-Q5_K_XL8.5 GB12.7 GBRTX 4060 Ti 16GB
Q6_K9.7 GB13.9 GBRTX 4060 Ti 16GB
UD-Q6_K_XL10.6 GB14.8 GBRTX 3090 24GB
Q8_012.5 GB16.7 GBRTX 3090 24GB
UD-Q8_K_XL14.4 GB18.8 GBRTX 3090 24GB
BF1623.5 GB28.6 GBRTX 5090 32GB

KV cache computed exactly from the model config (GQA formula, 8K context). MoE models keep all experts in memory — total size counts, not just active params.

Can I run gemma-3-12b-it-qat on my GPU? (UD-IQ2_XXS recommended quant)

HardwareMemoryVerdict
RTX 3060 12GB12 GBRUNS
RTX 4060 Ti 16GB16 GBRUNS
RTX 3090 24GB24 GBRUNS
RTX 4090 24GB24 GBRUNS
RTX 5090 32GB32 GBRUNS
RTX PRO 6000 96GB96 GBRUNS
Mac 16GB unified16 GBRUNS
Mac 32GB unified32 GBRUNS
Mac 64GB unified64 GBRUNS
Mac 128GB unified128 GBRUNS
Mac 256GB unified256 GBRUNS
Mac 512GB unified512 GBRUNS
32GB system RAM (CPU)32 GBRUNS
64GB system RAM (CPU)64 GBRUNS
128GB system RAM (CPU)128 GBRUNS
256GB system RAM (CPU)256 GBRUNS
1TB server RAM (CPU)1024 GBRUNS
2TB server RAM (CPU)2048 GBRUNS

How much VRAM does gemma-3-12b-it-qat need?

The honest answer: 7.8 GB at the UD-IQ2_XXS quant with 8K context — that is the measured file size (3.6 GB) plus KV cache and runtime overhead. The smallest published build needs 7.3 GB; the lossless one needs 28.6 GB.

Can I run it with CPU offload?

Yes, if combined RAM+VRAM ≥ file size — but generation speed is limited by memory bandwidth. Expect roughly: DDR5 dual channel ~10-30 tok/s for small MoE actives, single digits for big ones, 0.1-1 tok/s when experts page from disk. CPU-only is fine for batch jobs, painful for chat.

Why do other sites show different numbers?

Most pages compute weights from a formula (params × bits/8). We use the actual file sizes published in GGUF repos, which include the embedding table, unquantized tensors, and container overhead — that is why our numbers can differ from a naive calculation.

FAQ

How much VRAM does gemma-3-12b-it-qat need?

The measured answer: 7.8 GB at the UD-IQ2_XXS quant with 8K context. The smallest published build needs 7.3 GB; the lossless (F16/BF16) build needs 28.6 GB. These are real GGUF file sizes from Hugging Face, not formula estimates.

Can I run gemma-3-12b-it-qat on a 24GB GPU (RTX 3090/4090)?

Yes — at UD-IQ2_XXS (7.8 GB). Quants that fit 24GB: UD-IQ1_S, UD-IQ1_M, UD-IQ2_XXS, UD-IQ2_M, Q2_K, Q2_K_L, UD-IQ3_XXS, UD-Q2_K_XL, Q3_K_S, Q3_K_M, UD-Q3_K_XL, IQ4_XS, IQ4_NL, Q4_0, Q4_K_S, Q4_K_M, UD-Q4_K_XL, Q4_1, Q5_K_S, Q5_K_M, UD-Q5_K_XL, Q6_K, UD-Q6_K_XL, Q8_0, UD-Q8_K_XL.

Can I run gemma-3-12b-it-qat on CPU with system RAM?

Yes, if your usable RAM is at least the file size. At UD-IQ2_XXS you need ~7.8 GB of RAM. CPU generation is memory-bandwidth-bound: expect single-digit tok/s for large MoE models, 10-30 tok/s for small active params, and below 1 tok/s when experts page from disk.

Why do VRAM numbers for gemma-3-12b-it-qat differ between sites?

Most sites compute weights from a formula (params x bits/8). ModelFit uses the actual GGUF file sizes published on Hugging Face, which include the embedding table, unquantized tensors, and container overhead — so our numbers reflect what you really download and load.

More models

Qwen3-Coder-30B-A3B-Instructmeasured requirementsQwen3.8-27Bmeasured requirementsOrnith-1.5-9Bmeasured requirementsOrnith-1.5-35B-A3Bmeasured requirementsOrnith-1.0-9Bmeasured requirementsHuihui-Qwen3.8-27B-abliteratedmeasured requirementsQwen3.8-27B-Uncensoredmeasured requirementsQwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTPmeasured requirements