Yes — Nemotron 3 Nano 30B-A3B runs on RTX 5090 32GB
Nemotron 3 Nano 30B-A3B (30B-A3B) needs 22GB VRAM at Q5 vs your 32GB VRAM. Estimated ~508+ tok/s (blazing).
3B active MoE. Fast on consumer GPUs, official GGUF from ggml-org.
| quant | weights | + KV 8k | fits 32GB? |
|---|---|---|---|
| Q4 | 18GB | ~18.2GB | ✓ yes |
| Q5 | 22GB | ~22.2GB | ✓ yes |
| Q8 | 32GB | ~32.2GB | × tight/no |
| FP16 | 60GB | ~60.2GB | × tight/no |
File size ≠ runtime. KV grows with context — halve to 4k/2k if long chats OOM. Leave 1–2GB free.
Can Nemotron 3 Nano 30B-A3B run on RTX 5090 32GB?
Nemotron 3 Nano 30B-A3B (30B-A3B) needs 22GB VRAM at Q5 vs your 32GB VRAM. Estimated ~508+ tok/s (blazing). Model context 128k. Compare with all models or re-check with your exact specs on the bench.
Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.
Start Q5. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
About ~508+ tok/s (blazing) when 100% on GPU. Any CPU offload drops speed 5–20× — check ollama ps.
NVIDIA — check terms before commercial use. Filter commercial-safe on the bench.