Yes — Devstral Small runs on 2x RTX 3090 (48GB)
Devstral Small (24B) needs 25GB VRAM at Q8 vs your 24GB VRAM. Estimated ~40–55 tok/s.
Mistral’s coding specialist. Same 24B footprint, tuned for repos and tools.
| quant | weights | + KV 8k | fits 48GB? |
|---|---|---|---|
| Q4 | 14GB | ~15GB | ✓ yes |
| Q5 | 17GB | ~18GB | ✓ yes |
| Q8 | 25GB | ~26GB | ✓ yes |
| FP16 | 48GB | ~49GB | × tight/no |
File size ≠ runtime. KV grows with context — halve to 4k/2k if long chats OOM. Leave 1–2GB free.
Can Devstral Small run on 2x RTX 3090 (48GB)?
Devstral Small (24B) needs 25GB VRAM at Q8 vs your 24GB VRAM. Estimated ~40–55 tok/s. Model context 128k. Compare with all models or re-check with your exact specs on the bench.
Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.
Start Q8. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
About ~40–55 tok/s when 100% on GPU. Any CPU offload drops speed 5–20× — check ollama ps.
Yes — Apache-2.0 allows commercial use.