Yes — Gemma 3 27B runs on Mac Studio 128GB
Gemma 3 27B (27B) needs 54GB VRAM at FP16 vs your 128GB unified. Estimated ~8–11 tok/s.
Google’s dense 27B with vision. Q4 ~17 GB — same class as Qwen3.8-27B.
| quant | weights | + KV 8k | fits 128GB? |
|---|---|---|---|
| Q4 | 17GB | ~18GB | ✓ yes |
| Q5 | 20GB | ~21GB | ✓ yes |
| Q8 | 29GB | ~30GB | ✓ yes |
| FP16 | 54GB | ~55GB | ✓ yes |
File size ≠ runtime. KV grows with context — halve to 4k/2k if long chats OOM. Leave 1–2GB free.
Can Gemma 3 27B run on Mac Studio 128GB?
Gemma 3 27B (27B) needs 54GB VRAM at FP16 vs your 128GB unified. Estimated ~8–11 tok/s. Model context 128k. Compare with all models or re-check with your exact specs on the bench.
Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.
Start FP16. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
About ~8–11 tok/s when 100% on GPU. Any CPU offload drops speed 5–20× — check ollama ps.
Gemma — check terms before commercial use. Filter commercial-safe on the bench.