No — Qwen2.5-Coder 32B won't fit RTX 3060 12GB at Q4
Qwen2.5-Coder 32B needs 16GB+ and exceeds RTX 3060 12GB. Try a smaller quant or model.
Still one of the best local coding models. Q4 ~20 GB.
Can Qwen2.5-Coder 32B run on RTX 3060 12GB?
Qwen2.5-Coder 32B needs 16GB+ and exceeds RTX 3060 12GB. Try a smaller quant or model. Model context 32k. Compare with all models or re-check with your exact specs on the bench.
Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.
Start Q4. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
Slower than GPU — this combo needs offload. Pick a smaller model for interactive chat.
Yes — Apache-2.0 allows commercial use.