Yes — InternVL3 8B runs on RTX 4090 24GB
InternVL3 8B (8B) needs 16GB VRAM at FP16 vs your 24GB VRAM. Estimated ~36–49 tok/s.
Strong open vision-language at 8B. Q4 ~6 GB. Screenshot / document QA.
| quant | weights | + KV 8k | fits 24GB? |
|---|---|---|---|
| Q4 | 6GB | ~6.4GB | ✓ yes |
| Q5 | 7.5GB | ~7.9GB | ✓ yes |
| Q8 | 10GB | ~10.4GB | ✓ yes |
| FP16 | 16GB | ~16.4GB | ✓ yes |
File size ≠ runtime. KV grows with context — halve to 4k/2k if long chats OOM. Leave 1–2GB free.
Can InternVL3 8B run on RTX 4090 24GB?
InternVL3 8B (8B) needs 16GB VRAM at FP16 vs your 24GB VRAM. Estimated ~36–49 tok/s. Model context 32k. Compare with all models or re-check with your exact specs on the bench.
Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.
Start FP16. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
About ~36–49 tok/s when 100% on GPU. Any CPU offload drops speed 5–20× — check ollama ps.
Yes — Apache-2.0 allows commercial use.