Yes — Gemma 3 12B runs on CPU only 32GB
Gemma 3 12B (12B) needs 13GB VRAM at Q8 vs your 32GB VRAM. Estimated ~3–6 tok/s (RAM offload).
Multimodal 12B. Q4 ~8 GB, Q5 ~10 GB. Excellent on 12 GB cards.
| quant | weights | + KV 8k | fits 0GB? |
|---|---|---|---|
| Q4 | 8.1GB | ~8.7GB | × tight/no |
| Q5 | 10GB | ~10.6GB | × tight/no |
| Q8 | 13GB | ~13.6GB | × tight/no |
| FP16 | 24GB | ~24.6GB | × tight/no |
File size ≠ runtime. KV grows with context — halve to 4k/2k if long chats OOM. Leave 1–2GB free.
Can Gemma 3 12B run on CPU only 32GB?
Gemma 3 12B (12B) needs 13GB VRAM at Q8 vs your 32GB VRAM. Estimated ~3–6 tok/s (RAM offload). Model context 128k. Compare with all models or re-check with your exact specs on the bench.
Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.
Start Q8. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
About ~3–6 tok/s (RAM offload) when 100% on GPU. Any CPU offload drops speed 5–20× — check ollama ps.
Gemma — check terms before commercial use. Filter commercial-safe on the bench.