Yes — Llama 3.3 70B runs on 2x RTX 3090 (48GB)
Llama 3.3 70B (70B) needs 40GB VRAM at Q4 vs your 24GB VRAM. Estimated ~26–35 tok/s.
Still the 70B workhorse. Q4 wants ~40 GB (2× 24 GB or 48 GB unified).
| quant | weights | + KV 8k | fits 48GB? |
|---|---|---|---|
| Q3 | 32GB | ~33.4GB | ✓ yes |
| Q4 | 40GB | ~41.4GB | ✓ yes |
| Q5 | 50GB | ~51.4GB | × tight/no |
| Q8 | 75GB | ~76.4GB | × tight/no |
| FP16 | 140GB | ~141.4GB | × tight/no |
File size ≠ runtime. KV grows with context — halve to 4k/2k if long chats OOM. Leave 1–2GB free.
Can Llama 3.3 70B run on 2x RTX 3090 (48GB)?
Llama 3.3 70B (70B) needs 40GB VRAM at Q4 vs your 24GB VRAM. Estimated ~26–35 tok/s. Model context 128k. Compare with all models or re-check with your exact specs on the bench.
Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.
Start Q4. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
About ~26–35 tok/s when 100% on GPU. Any CPU offload drops speed 5–20× — check ollama ps.
Llama 3.3 — check terms before commercial use. Filter commercial-safe on the bench.