No — DeepSeek-V4 Flash won't fit RTX 4090 24GB at Q4
DeepSeek-V4 Flash needs 76GB+ and exceeds RTX 4090 24GB. Try a smaller quant or model.
284B MoE with 13B active and 1M context. Even aggressive UD-IQ2 needs ~75 GB combined — 128 GB unified or multi-GPU + fat RAM.
Can DeepSeek-V4 Flash run on RTX 4090 24GB?
DeepSeek-V4 Flash needs 76GB+ and exceeds RTX 4090 24GB. Try a smaller quant or model. Model context 1024k. Compare with all models or re-check with your exact specs on the bench.
Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.
Start Q4. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
Slower than GPU — this combo needs offload. Pick a smaller model for interactive chat.
Yes — MIT allows commercial use.