can i run · seo

Yes — Llama 3.3 70B runs on 2x RTX 3090 (48GB)

Llama 3.3 70B (70B) needs 40GB VRAM at Q4 vs your 24GB VRAM. Estimated ~26–35 tok/s.

Still the 70B workhorse. Q4 wants ~40 GB (2× 24 GB or 48 GB unified).

best fit on this rig · weights + KV + headroom
Q4 · 40GB weights+1.4GB KV @ 8kfits GPU~26–35 tok/s
quantweights+ KV 8kfits 48GB?
Q332GB~33.4GB✓ yes
Q440GB~41.4GB✓ yes
Q550GB~51.4GB× tight/no
Q875GB~76.4GB× tight/no
FP16140GB~141.4GB× tight/no

File size ≠ runtime. KV grows with context — halve to 4k/2k if long chats OOM. Leave 1–2GB free.

ollama run llama3.3:70b
Run it: 1) install Ollama · 2) copy command · 3) verify with ollama ps (want 100% GPU).
faq · straight answers

Can Llama 3.3 70B run on 2x RTX 3090 (48GB)?

Llama 3.3 70B (70B) needs 40GB VRAM at Q4 vs your 24GB VRAM. Estimated ~26–35 tok/s. Model context 128k. Compare with all models or re-check with your exact specs on the bench.

What if it OOMs?

Unload with ollama stop, cap layers --num-gpu 28, cut context to 2k. See the OOM fixer on the homepage.

Q4 or Q8 on this rig?

Start Q4. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.

How fast will it be?

About ~26–35 tok/s when 100% on GPU. Any CPU offload drops speed 5–20× — check ollama ps.

Is Llama 3.3 commercial-safe?

Llama 3.3 — check terms before commercial use. Filter commercial-safe on the bench.

Aarav — RTX 4090Qwen3.8 27B Q4 · 2m agoPriya — M4 16GBPhi-4 Mini Q8 · 4m agoRohan — RTX 3060 12GBGemma 3 12B Q5 · 6m agoSneha — 2x RTX 3090Llama 3.3 70B Q4 · 9m agoVikram — 8GB laptopQwen3 8B Q4 · 11m agoAnanya — M4 Max 64GBDeepSeek-R1 32B Q4 · 14m agoKabir — RTX 4070 12GBMistral Small 3.2 Q4 · 18m agoIsha — Mac Studio 128GBGLM-5.3 Flash IQ2 · 22m agoAarav — RTX 4090Qwen3.8 27B Q4 · 2m agoPriya — M4 16GBPhi-4 Mini Q8 · 4m agoRohan — RTX 3060 12GBGemma 3 12B Q5 · 6m agoSneha — 2x RTX 3090Llama 3.3 70B Q4 · 9m agoVikram — 8GB laptopQwen3 8B Q4 · 11m agoAnanya — M4 Max 64GBDeepSeek-R1 32B Q4 · 14m agoKabir — RTX 4070 12GBMistral Small 3.2 Q4 · 18m agoIsha — Mac Studio 128GBGLM-5.3 Flash IQ2 · 22m agoAarav — RTX 4090Qwen3.8 27B Q4 · 2m agoPriya — M4 16GBPhi-4 Mini Q8 · 4m agoRohan — RTX 3060 12GBGemma 3 12B Q5 · 6m agoSneha — 2x RTX 3090Llama 3.3 70B Q4 · 9m agoVikram — 8GB laptopQwen3 8B Q4 · 11m agoAnanya — M4 Max 64GBDeepSeek-R1 32B Q4 · 14m agoKabir — RTX 4070 12GBMistral Small 3.2 Q4 · 18m agoIsha — Mac Studio 128GBGLM-5.3 Flash IQ2 · 22m ago
live telemetry

real-time global visitors

Coarse location inferred from browser timezone without fingerprinting or tracking cookies. Every blinking node represents an active hardware builder browsing the open-weight registry.

growingearly builders · be earlylive0 active
  • recent pings · timezone-derived0 nodes
  • listening for hardware builders…
Live Mesh
[@] You[!] New (<20s)[+] Active Node[│] Radar Sweep
Nodes: 0