catalog

every model we match against

Curated open-weight catalog with GGUF sizes. Live Hugging Face noise lives on drops. Filter by VRAM ceiling to see what a card can hold at Q4.

38 models
Z.ai

GLM-5.3

The big GLM. 2-bit still wants ~245 GB combined memory. Homelab / Mac Studio Ultra territory.

params
753B MoE
license
GLM-5.3
Q4-ish
420 GB
context
1024k
tasks
chat · code · reason · agent
Z.ai

GLM-5.3 Flash

Frontier-class multimodal MoE. 1-bit Unsloth GGUF ~93 GB; 3-bit wants ~120 GB unified. Not a 24 GB card model.

params
321B MoE
license
MIT
Q4-ish
200 GB
context
1024k
tasks
chat · code · reason · vision · agent
Qwen

Qwen3.8-Flash-Next

180B hybrid (6B active + n-gram memory), 256k multimodal with thinking + tools. UD-Q4 ≈ 100 GB — 128 GB unified or 2×24 GB + RAM.

params
MoE 180B · 6B active
license
Qwen-Community
Q4-ish
100 GB
context
256k
tasks
chat · code · reason · vision · agent
DeepSeek

DeepSeek-V4 Flash

284B MoE with 13B active and 1M context. Even aggressive UD-IQ2 needs ~75 GB combined — 128 GB unified or multi-GPU + fat RAM.

params
MoE 284B · 13B active
license
MIT
Q4-ish
160 GB
context
1024k
tasks
chat · code · reason · agent
Qwen

Qwen3.8 27B

Best all-rounder you can actually run. Vision + reasoning + 256k context, Apache-2.0, Unsloth Q4 around 17 GB.

huggingfaceGGUF downloadollama run qwen3.8:27b
params
27.8B dense
license
Apache-2.0
Q4-ish
17 GB
context
256k
tasks
chat · code · reason · vision · agent
Black Forest Labs

FLUX.1 dev

Higher-fidelity FLUX. Wants 16–24 GB. Non-commercial weights.

params
12B
license
Non-commercial
Q4-ish
10 GB
context
n/a
tasks
image
Meta

Llama 4 Scout

Meta’s natively multimodal MoE. Q4 is ~55 GB — 64 GB unified or 2× 24 GB + RAM.

huggingfaceGGUF downloadollama run llama4:scout
params
109B MoE
license
Llama 4
Q4-ish
55 GB
context
10240k
tasks
chat · code · vision
Black Forest Labs

FLUX.1 schnell

Apache image model. 8–12 GB VRAM at fp8/nf4, 24 GB comfortable at fp16.

params
12B
license
Apache-2.0
Q4-ish
8 GB
context
n/a
tasks
image
DeepSeek

DeepSeek-R1 Distill 32B

The distill that still feels like R1. Q4 on a 24 GB card is the intended home.

huggingfaceGGUF downloadollama run deepseek-r1:32b
params
32B
license
MIT
Q4-ish
20 GB
context
128k
tasks
reason · code · chat
Mistral

Devstral Small

Mistral’s coding specialist. Same 24B footprint, tuned for repos and tools.

huggingfaceGGUF downloadollama run devstral
params
24B
license
Apache-2.0
Q4-ish
14 GB
context
128k
tasks
code · agent
Qwen

Qwen3 32B

Dense 32B for 24 GB cards at Q4. Serious coding and long-horizon agent work.

huggingfaceGGUF downloadollama run qwen3:32b
params
32B
license
Apache-2.0
Q4-ish
20 GB
context
128k
tasks
chat · code · reason · agent
Mistral

Mistral Small 3.2

Apache-2.0 24B with vision. Q4 ~14 GB — sweet spot for 16 GB cards.

huggingfaceGGUF downloadollama run mistral-small3.2
params
24B
license
Apache-2.0
Q4-ish
14 GB
context
128k
tasks
chat · code · vision · agent
OpenAI

Whisper large-v3

Speech-to-text. Q8 ~3 GB VRAM, CPU-ok. Pair with any chat model.

huggingfaceGGUF downloadwhisper.cpp or faster-whisper large-v3
params
1.5B
license
Apache-2.0
Q4-ish
1.5 GB
context
n/a
tasks
audio
Google

Gemma 3 27B

Google’s dense 27B with vision. Q4 ~17 GB — same class as Qwen3.8-27B.

huggingfaceGGUF downloadollama run gemma3:27b
params
27B
license
Gemma
Q4-ish
17 GB
context
128k
tasks
chat · vision · code · reason
Qwen

Qwen2.5-Coder 32B

Still one of the best local coding models. Q4 ~20 GB.

huggingfaceGGUF downloadollama run qwen2.5-coder:32b
params
32B
license
Apache-2.0
Q4-ish
20 GB
context
32k
tasks
code · agent
Meta

Llama 3.3 70B

Still the 70B workhorse. Q4 wants ~40 GB (2× 24 GB or 48 GB unified).

huggingfaceGGUF downloadollama run llama3.3:70b
params
70B
license
Llama 3.3
Q4-ish
40 GB
context
128k
tasks
chat · code · reason
NVIDIA

Nemotron 3 Nano 30B-A3B

3B active MoE. Fast on consumer GPUs, official GGUF from ggml-org.

params
30B-A3B
license
NVIDIA
Q4-ish
18 GB
context
128k
tasks
chat · code · agent
Qwen

Qwen3.5 9B

Unified vision-language 9B. The 8 GB card champion — multilingual, sharp at code, Q4 sits around 5.7 GB with cache room.

huggingfaceGGUF downloadollama run qwen3.5:9b
params
9B
license
Apache-2.0
Q4-ish
5.7 GB
context
128k
tasks
chat · code · reason · vision
Google

Gemma 3 12B

Multimodal 12B. Q4 ~8 GB, Q5 ~10 GB. Excellent on 12 GB cards.

huggingfaceGGUF downloadollama run gemma3:12b
params
12B
license
Gemma
Q4-ish
8.1 GB
context
128k
tasks
chat · vision · code
OpenGVLab

InternVL3 8B

Strong open vision-language at 8B. Q4 ~6 GB. Screenshot / document QA.

params
8B
license
Apache-2.0
Q4-ish
6 GB
context
32k
tasks
vision · chat
DeepSeek

DeepSeek-R1 Distill 8B

Distilled reasoner. Long traces, strong math, fits any 8 GB card at Q4.

huggingfaceGGUF downloadollama run deepseek-r1:8b
params
8B
license
MIT
Q4-ish
5.2 GB
context
32k
tasks
reason · code · chat
Nomic

Nomic Embed Text v1.5

Local RAG embeddings. CPU is enough. Matryoshka dims.

huggingfaceollama pull nomic-embed-text
params
137M
license
Apache-2.0
Q4-ish
0.3 GB
context
8k
tasks
embed
OpenBMB

MiniCPM-V 4.5

OCR and screenshot specialist. Fits 8 GB at Q4.

huggingfaceGGUF downloadollama run minicpm-v
params
8B
license
Apache-2.0
Q4-ish
5.8 GB
context
32k
tasks
vision · chat
Microsoft

Phi-4 14B

Small model, dense knowledge. Q4 ~9 GB, Q5 ~11 GB. Punches above 14B on STEM.

params
14B
license
MIT
Q4-ish
9 GB
context
16k
tasks
chat · code · reason
Qwen

Qwen3 8B

Thinking-mode 8B. Fast, cheap, still beats most 7B-class models at reasoning.

huggingfaceGGUF downloadollama run qwen3:8b
params
8B
license
Apache-2.0
Q4-ish
5.2 GB
context
128k
tasks
chat · code · reason
hexgrad

Kokoro 82M

Tiny high-quality TTS. Runs on CPU. The local voice stack default.

params
82M
license
Apache-2.0
Q4-ish
0.5 GB
context
n/a
tasks
audio
Stability

Stable Diffusion 3.5 Medium

Runs on 8 GB. Open-ish weights, good default if FLUX is too heavy.

params
2.5B
license
Stability Community
Q4-ish
6 GB
context
n/a
tasks
image
Cohere

Command R7B

RAG-native 7B with tools. Non-commercial license. Great retrieval chat.

huggingfaceGGUF downloadollama run command-r7b
params
8B
license
CC-BY-NC
Q4-ish
5 GB
context
128k
tasks
chat · agent
IBM

Granite 3.3 8B

Apache, long context, enterprise-clean. Solid 8 GB citizen.

huggingfaceGGUF downloadollama run granite3.3:8b
params
8B
license
Apache-2.0
Q4-ish
5.1 GB
context
128k
tasks
chat · code · reason
Qwen

Qwen2.5 7B

Battle-tested 7B. Still a great 6–8 GB default if you want something boring and stable.

huggingfaceGGUF downloadollama run qwen2.5:7b
params
7B
license
Apache-2.0
Q4-ish
4.7 GB
context
32k
tasks
chat · code
Ai2

OLMo 2 13B

Fully open (data + code + weights). Q4 ~8 GB. The research-honest pick.

huggingfaceGGUF downloadollama run olmo2:13b
params
13B
license
Apache-2.0
Q4-ish
8 GB
context
4k
tasks
chat
OpenBMB

MiniCPM5 2B

Sept-2026 edge refresh: 2.5B, 128k context, tool-calling. CPU-first; Q4 ≈ 1.6 GB.

params
2.5B
license
Apache-2.0
Q4-ish
1.6 GB
context
128k
tasks
chat · code · agent
01.AI

Yi 1.5 9B

Bilingual EN/ZH 9B. Q4 ~6 GB. Still a good 8–12 GB pick.

huggingfaceGGUF downloadollama run yi:9b
params
9B
license
Apache-2.0
Q4-ish
5.8 GB
context
4k
tasks
chat · code
Qwen

Qwen3 4B

Laptop / 6 GB card model. Surprisingly useful for chat and light coding.

huggingfaceGGUF downloadollama run qwen3:4b
params
4B
license
Apache-2.0
Q4-ish
2.6 GB
context
32k
tasks
chat · code
Google

Gemma 3 4B

Tiny multimodal. Phone-class, also great on iGPU laptops.

huggingfaceGGUF downloadollama run gemma3:4b
params
4B
license
Gemma
Q4-ish
2.8 GB
context
128k
tasks
chat · vision
Microsoft

Phi-4 Mini

CPU-friendly 3.8B with 128k context. Fine default for 8 GB RAM machines.

huggingfaceGGUF downloadollama run phi4-mini
params
3.8B
license
MIT
Q4-ish
2.5 GB
context
128k
tasks
chat · code
Hugging Face

SmolLM3 3B

On-device 3B from HF. Dual-mode thinking, tiny footprint.

huggingfaceGGUF downloadollama run smollm3
params
3B
license
Apache-2.0
Q4-ish
2 GB
context
64k
tasks
chat · code
Meta

Llama 3.2 3B

Edge / CPU-only. Fine for summarization and short chat on 8 GB laptops.

huggingfaceGGUF downloadollama run llama3.2:3b
params
3B
license
Llama 3.2
Q4-ish
2 GB
context
128k
tasks
chat

every model · vram requirements · quant size · ollama command

model VRAM requirements table — run locally.

Each row lists params, the Q4 VRAM requirement in GB, the smallest quant that runs locally, and the copy-paste Ollama command with a GGUF download link. Compare rows to decide which model is better for your rig. Sizes exclude KV cache — leave ~10% headroom.

modelparamsVRAM requirements (Q4)runs locally fromollama command
Qwen3.8 27B27.8B dense17 GB17 GB (Q4)ollama run qwen3.8:27b
Qwen3.5 9B9B5.7 GB5.7 GB (Q4)ollama run qwen3.5:9b
Qwen3 8B8B5.2 GB5.2 GB (Q4)ollama run qwen3:8b
Qwen3 32B32B20 GB16 GB (Q3)ollama run qwen3:32b
Qwen3 4B4B2.6 GB2.6 GB (Q4)ollama run qwen3:4b
GLM-5.3 Flash321B MoE200 GB93 GB (IQ2)llama-cli -hf unsloth/GLM-5.3-Flash-GGUF:UD-IQ2_M
GLM-5.3753B MoE420 GB245 GB (IQ2)llama-cli -hf unsloth/GLM-5.3-GGUF:UD-IQ2_M
DeepSeek-R1 Distill 8B8B5.2 GB5.2 GB (Q4)ollama run deepseek-r1:8b
DeepSeek-R1 Distill 32B32B20 GB16 GB (Q3)ollama run deepseek-r1:32b
DeepSeek-V4 FlashMoE 284B · 13B active160 GB76 GB (IQ2)llama-cli -hf unsloth/DeepSeek-V4-Flash-GGUF:UD-IQ2_M
Llama 4 Scout109B MoE55 GB42 GB (Q3)ollama run llama4:scout
Llama 3.3 70B70B40 GB32 GB (Q3)ollama run llama3.3:70b
Llama 3.2 3B3B2 GB2 GB (Q4)ollama run llama3.2:3b
Gemma 3 12B12B8.1 GB8.1 GB (Q4)ollama run gemma3:12b
Gemma 3 27B27B17 GB17 GB (Q4)ollama run gemma3:27b
Gemma 3 4B4B2.8 GB2.8 GB (Q4)ollama run gemma3:4b
Phi-4 14B14B9 GB9 GB (Q4)ollama run phi4
Phi-4 Mini3.8B2.5 GB2.5 GB (Q4)ollama run phi4-mini
Mistral Small 3.224B14 GB14 GB (Q4)ollama run mistral-small3.2
Devstral Small24B14 GB14 GB (Q4)ollama run devstral
Qwen2.5-Coder 32B32B20 GB16 GB (Q3)ollama run qwen2.5-coder:32b
Qwen3.8-Flash-NextMoE 180B · 6B active100 GB48 GB (IQ2)llama-cli -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
Nemotron 3 Nano 30B-A3B30B-A3B18 GB18 GB (Q4)llama-cli -hf ggml-org/NVIDIA-Nemotron-3-Nano-30B-A3B-GGUF:Q4_K_M
OLMo 2 13B13B8 GB8 GB (Q4)ollama run olmo2:13b
SmolLM3 3B3B2 GB2 GB (Q4)ollama run smollm3
FLUX.1 schnell12B8 GB8 GB (Q4)Not an LLM — run via ComfyUI / diffusers
FLUX.1 dev12B10 GB10 GB (Q4)see GGUF download
Stable Diffusion 3.5 Medium2.5B6 GB6 GB (Q4)see GGUF download
Whisper large-v31.5B1.5 GB1.5 GB (Q4)whisper.cpp or faster-whisper large-v3
Kokoro 82M82M0.5 GB0.5 GB (FP16)see GGUF download
Nomic Embed Text v1.5137M0.3 GB0.3 GB (Q8)ollama pull nomic-embed-text
InternVL3 8B8B6 GB6 GB (Q4)llama-cli -hf unsloth/InternVL3-8B-GGUF:Q4_K_M --mmproj auto
MiniCPM-V 4.58B5.8 GB5.8 GB (Q4)ollama run minicpm-v
MiniCPM5 2B2.5B1.6 GB1.6 GB (Q4)llama-cli -hf openbmb/MiniCPM5-2B-GGUF:Q4_K_M
Qwen2.5 7B7B4.7 GB4.7 GB (Q4)ollama run qwen2.5:7b
Granite 3.3 8B8B5.1 GB5.1 GB (Q4)ollama run granite3.3:8b
Command R7B8B5 GB5 GB (Q4)ollama run command-r7b
Yi 1.5 9B9B5.8 GB5.8 GB (Q4)ollama run yi:9b
Aarav — RTX 4090Qwen3.8 27B Q4 · 2m agoPriya — M4 16GBPhi-4 Mini Q8 · 4m agoRohan — RTX 3060 12GBGemma 3 12B Q5 · 6m agoSneha — 2x RTX 3090Llama 3.3 70B Q4 · 9m agoVikram — 8GB laptopQwen3 8B Q4 · 11m agoAnanya — M4 Max 64GBDeepSeek-R1 32B Q4 · 14m agoKabir — RTX 4070 12GBMistral Small 3.2 Q4 · 18m agoIsha — Mac Studio 128GBGLM-5.3 Flash IQ2 · 22m agoAarav — RTX 4090Qwen3.8 27B Q4 · 2m agoPriya — M4 16GBPhi-4 Mini Q8 · 4m agoRohan — RTX 3060 12GBGemma 3 12B Q5 · 6m agoSneha — 2x RTX 3090Llama 3.3 70B Q4 · 9m agoVikram — 8GB laptopQwen3 8B Q4 · 11m agoAnanya — M4 Max 64GBDeepSeek-R1 32B Q4 · 14m agoKabir — RTX 4070 12GBMistral Small 3.2 Q4 · 18m agoIsha — Mac Studio 128GBGLM-5.3 Flash IQ2 · 22m agoAarav — RTX 4090Qwen3.8 27B Q4 · 2m agoPriya — M4 16GBPhi-4 Mini Q8 · 4m agoRohan — RTX 3060 12GBGemma 3 12B Q5 · 6m agoSneha — 2x RTX 3090Llama 3.3 70B Q4 · 9m agoVikram — 8GB laptopQwen3 8B Q4 · 11m agoAnanya — M4 Max 64GBDeepSeek-R1 32B Q4 · 14m agoKabir — RTX 4070 12GBMistral Small 3.2 Q4 · 18m agoIsha — Mac Studio 128GBGLM-5.3 Flash IQ2 · 22m ago
live telemetry

real-time global visitors

Coarse location inferred from browser timezone without fingerprinting or tracking cookies. Every blinking node represents an active hardware builder browsing the open-weight registry.

growingearly builders · be earlylive0 active
  • recent pings · timezone-derived0 nodes
  • listening for hardware builders…
Live Mesh
[@] You[!] New (<20s)[+] Active Node[│] Radar Sweep
Nodes: 0