can i run · seo
Yes — MiniCPM5 2B runs on Mac M4 16GB MiniCPM5 2B (2.5B) needs 5GB VRAM at FP16 vs your 16GB unified. Estimated ~11–14 tok/s.
Sept-2026 edge refresh: 2.5B, 128k context, tool-calling. CPU-first; Q4 ≈ 1.6 GB.
best fit on this rig · weights + KV + headroom
FP16 · 5GB weights +0.2GB KV @ 8k fits GPU ~11–14 tok/s
quant weights + KV 8k fits 16GB? Q4 1.6GB ~1.8GB ✓ yes Q5 1.9GB ~2.1GB ✓ yes Q8 2.7GB ~2.9GB ✓ yes FP16 5GB ~5.2GB ✓ yes
File size ≠ runtime. KV grows with context — halve to 4k/2k if long chats OOM. Leave 1–2GB free.
llama-cli -hf openbmb/MiniCPM5-2B-GGUF:Q4_K_M
Run it: 1) install Ollama · 2) copy command · 3) verify with ollama ps (want 100% GPU).
faq · straight answers
Can MiniCPM5 2B run on Mac M4 16GB? MiniCPM5 2B (2.5B) needs 5GB VRAM at FP16 vs your 16GB unified. Estimated ~11–14 tok/s. Model context 128k. Compare with all models or re-check with your exact specs on the bench .
What if it OOMs?
Unload with ollama stop , cap layers --num-gpu 28 , cut context to 2k. See the OOM fixer on the homepage.
Q4 or Q8 on this rig?
Start FP16. Bigger model at Q4 beats smaller at Q8. Only go Q8 if 2GB+ headroom remains.
How fast will it be?
About ~11–14 tok/s when 100% on GPU. Any CPU offload drops speed 5–20× — check ollama ps.
Is Apache-2.0 commercial-safe?
Yes — Apache-2.0 allows commercial use.
MiniCPM5 2B on other rigs
live 18 checks/hr · 3 active
Aarav — RTX 4090 → Qwen3.8 27B Q4 · 2m ago ◆ Priya — M4 16GB → Phi-4 Mini Q8 · 4m ago ◆ Rohan — RTX 3060 12GB → Gemma 3 12B Q5 · 6m ago ◆ Sneha — 2x RTX 3090 → Llama 3.3 70B Q4 · 9m ago ◆ Vikram — 8GB laptop → Qwen3 8B Q4 · 11m ago ◆ Ananya — M4 Max 64GB → DeepSeek-R1 32B Q4 · 14m ago ◆ Kabir — RTX 4070 12GB → Mistral Small 3.2 Q4 · 18m ago ◆ Isha — Mac Studio 128GB → GLM-5.3 Flash IQ2 · 22m ago ◆ Aarav — RTX 4090 → Qwen3.8 27B Q4 · 2m ago ◆ Priya — M4 16GB → Phi-4 Mini Q8 · 4m ago ◆ Rohan — RTX 3060 12GB → Gemma 3 12B Q5 · 6m ago ◆ Sneha — 2x RTX 3090 → Llama 3.3 70B Q4 · 9m ago ◆ Vikram — 8GB laptop → Qwen3 8B Q4 · 11m ago ◆ Ananya — M4 Max 64GB → DeepSeek-R1 32B Q4 · 14m ago ◆ Kabir — RTX 4070 12GB → Mistral Small 3.2 Q4 · 18m ago ◆ Isha — Mac Studio 128GB → GLM-5.3 Flash IQ2 · 22m ago ◆ Aarav — RTX 4090 → Qwen3.8 27B Q4 · 2m ago ◆ Priya — M4 16GB → Phi-4 Mini Q8 · 4m ago ◆ Rohan — RTX 3060 12GB → Gemma 3 12B Q5 · 6m ago ◆ Sneha — 2x RTX 3090 → Llama 3.3 70B Q4 · 9m ago ◆ Vikram — 8GB laptop → Qwen3 8B Q4 · 11m ago ◆ Ananya — M4 Max 64GB → DeepSeek-R1 32B Q4 · 14m ago ◆ Kabir — RTX 4070 12GB → Mistral Small 3.2 Q4 · 18m ago ◆ Isha — Mac Studio 128GB → GLM-5.3 Flash IQ2 · 22m ago ◆
social proof · timezone-derived · no cookies
live telemetry
real-time global visitors Coarse location inferred from browser timezone without fingerprinting or tracking cookies. Every blinking node represents an active hardware builder browsing the open-weight registry.
growing early builders · be early live 0 active
recent pings · timezone-derived 0 nodes listening for hardware builders… [1] 🗺️ ASCII World Radar [2] ⚡ Node Cluster Matrix
Live Mesh
┌────────────────────────────────────────────────────────────────┐
│· -180° · · · · -90° · · · · · · 0° · · · · · · +90° · · +180° ·│
│ __..---..__ .---. _..----.._ _.._ │
│ .' · NA · '. / · \ .' · EU · '. .'·AP │
│ / *US/CA* · \| GL |· / *UK/DE/FR* \ / *JP/CN│
│| [W-COAST] · || · · · | · | [CENTRAL] · | | *IN/KR│
│ \ [E-COAST]· / \ · / · \ [MEDIT] · / \ *SG/AU│
│ '. · · · .' '---' '. · · · .' '. · │
│ '-·_____·-' · · '-·____·-' '--· │
│ · · · · · · · · · · · · · · · · · · · · · · · · · · · · · · │
│ _.._ _.._ · · · _.._ │
│ .' · '. .' · '. .' · '. │
│ / LATAM \ / AFRICA \ / ANZ \ │
│ | *BR/MX/AR|· | *ZA/KE/NG| | *AU/NZ* | │
│ \ · · / \ · · / | · · | │
│ '. · .' '. · .' \ · · / │
│ '..' '..' '.____.' │
│[LAT -60°] · · · · · · · · · · · · · · · · · · · · · ·[LAT +60°]│
│ >> ACTIVE NODE RADAR · TELEMETRY MESH // SYNCED VIA TIMEZONE << │
└────────────────────────────────────────────────────────────────┘ [@] You [!] New (<20s) [+] Active Node [│] Radar Sweep
Nodes: 0