This Tiny Coding Model Shouldn’t Be This Good (MiniCPM5)
MiniCPM5-2B is a tiny 2B-parameter local LLM that claims to beat Qwen3.5-4B on SWE-bench Verified, but the real story gets much more interesting once you actually run it. In this video, I test MiniCPM5-2B locally with MLX and llama.cpp, look at its coding, reasoning, long-context, and tool-calling performance, and break down where it genuinely competes with models twice its size. 🔗 Relevant Links OpenBMB - https://www.openbmb.cn/model/minicpm5-2b HuggingFace - https://huggingface.co/spaces/openbmb/MiniCPM5-2B-Demo ❤️ More about us Radically better observability stack: https://betterstack.com/ Written tutorials: https://betterstack.com/community/ Example projects: https://github.com/BetterStackHQ 📱 Socials Twitter: https://twitter.com/betterstackhq Instagram: https://www.instagram.com/betterstackhq/ TikTok: https://www.tiktok.com/@betterstack LinkedIn: https://www.linkedin.com/company/betterstack 📌 Chapters: 00:00 MiniCPM5-2B vs Qwen3.5-4B 00:32 Why 2B Models Usually Fail at Agent Work 01:20 How MiniCPM5-2B Was Trained for Agents 01:56 Running MiniCPM5-2B Locally with MLX 02:52 Tool Calling with llama.cpp 03:15 MiniCPM5-2B Architecture Explained 04:00 MiniCPM5-2B Benchmark Results 05:00 The 92% Runaway Repetition Problem 05:24 The One Setting That Fixes It 06:15 MiniCPM5-2B Limitations and Fine Print 06:50 Should You Use MiniCPM5-2B?