This Tiny Coding Model Shouldn’t Be This Good (MiniCPM5)

Better Stack
22,369 views September 15, 2026

MiniCPM5-2B is a tiny 2B-parameter local LLM that claims to beat Qwen3.5-4B on SWE-bench Verified, but the real story gets much more interesting once you actually run it. In this video, I test MiniCPM5-2B locally with MLX and llama.cpp, look at its coding, reasoning, long-context, and tool-calling performance, and break down where it genuinely competes with models twice its size. 🔗 Relevant Links OpenBMB - https://www.openbmb.cn/model/minicpm5-2b HuggingFace - https://huggingface.co/spaces/openbmb/MiniCPM5-2B-Demo ❤️ More about us Radically better observability stack: https://betterstack.com/ Written tutorials: https://betterstack.com/community/ Example projects: https://github.com/BetterStackHQ 📱 Socials Twitter: https://twitter.com/betterstackhq Instagram: https://www.instagram.com/betterstackhq/ TikTok: https://www.tiktok.com/@betterstack LinkedIn: https://www.linkedin.com/company/betterstack 📌 Chapters: 00:00 MiniCPM5-2B vs Qwen3.5-4B 00:32 Why 2B Models Usually Fail at Agent Work 01:20 How MiniCPM5-2B Was Trained for Agents 01:56 Running MiniCPM5-2B Locally with MLX 02:52 Tool Calling with llama.cpp 03:15 MiniCPM5-2B Architecture Explained 04:00 MiniCPM5-2B Benchmark Results 05:00 The 92% Runaway Repetition Problem 05:24 The One Setting That Fixes It 06:15 MiniCPM5-2B Limitations and Fine Print 06:50 Should You Use MiniCPM5-2B?

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close