Yes — Llama 4 Scout runs on Mac Studio 128GB
Llama 4 Scout (109B MoE) needs 68GB VRAM at Q5 vs your 128GB unified. Estimated ~36–49 tok/s.
Meta’s natively multimodal MoE. Q4 is ~55 GB — 64 GB unified or 2× 24 GB + RAM.
| quant | weights | + KV 8k | fits 128GB? |
|---|---|---|---|
| Q3 | 42GB | ~42.6GB | ✓ yes |
| Q4 | 55GB | ~55.6GB | ✓ yes |
| Q5 | 68GB | ~68.6GB | ✓ yes |
| Q8 | 109GB | ~109.6GB | ✓ yes |
| FP16 | 218GB | ~218.6GB | × tight/no |
File size ≠ runtime. KV grows with context — halve to 4k/2k if long chats OOM. Leave 1–2GB free.
Can Llama 4 Scout run on Mac Studio 128GB?
Llama 4 Scout (109B MoE) needs 68GB VRAM at Q5 vs your 128GB unified. Estimated ~36–49 tok/s. Model context 10240k. Compare with all models or re-check with your exact specs on the bench.
Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.
Start Q5. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
About ~36–49 tok/s when 100% on GPU. Any CPU offload drops speed 5–20× — check ollama ps.
Llama 4 — check terms before commercial use. Filter commercial-safe on the bench.