No — Qwen3.8-Flash-Next won't fit RTX 5090 32GB at Q4
Qwen3.8-Flash-Next needs 48GB+ and exceeds RTX 5090 32GB. Try a smaller quant or model.
180B hybrid (6B active + n-gram memory), 256k multimodal with thinking + tools. UD-Q4 ≈ 100 GB — 128 GB unified or 2×24 GB + RAM.
Can Qwen3.8-Flash-Next run on RTX 5090 32GB?
Qwen3.8-Flash-Next needs 48GB+ and exceeds RTX 5090 32GB. Try a smaller quant or model. Model context 256k. Compare with all models or re-check with your exact specs on the bench.
Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.
Start Q4. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
Slower than GPU — this combo needs offload. Pick a smaller model for interactive chat.
Qwen-Community — check terms before commercial use. Filter commercial-safe on the bench.