OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device
MarkTechPost Sana Hassan
OpenBMB dropped MiniCPM5-2B, a 2.52B model built to run on device. It scores 53.9 across 34 tests and is strongest on tools, code, and long context.
Based on reporting by MarkTechPost, Sana Hassan — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenBMB has released MiniCPM5-2B, the second checkpoint in its MiniCPM5 line and the follow-up to MiniCPM5-1B. It’s a dense causal language model with 2,516,756,480 parameters, 1,981,982,720 of them outside the embeddings, and it uses 42 layers with grouped-query attention. The context window is native, not bolted on: 131,072 tokens.
The practical pitch is simple. This is a standard LlamaForCausalLM model, so it can be loaded by mainstream engines without custom kernels or a model-code fork. OpenBMB also ships the weights under Apache 2.0, and says they run through vLLM, SGLang, Transformers, llama.cpp, Ollama, LM Studio, MLX and FlagOS.
On the benchmark sheet, MiniCPM5-2B averages 53.9 across 34 rows. OpenBMB compares it with same-size models like LFM2.5-2.6B, Qwen3.5-2B and Gemma-4-E2B-it, while also listing bigger references such as Qwen3.5-4B, granite-4.2-3B, Nemotron-3-Nano-4B, Gemma-4-E4B-it and LFM2.5-8B-A1B. The best baseline in that table is Qwen3.5-4B at 51.1.
The model’s shape is pretty clear from the scores. It does well on code reasoning, with 69.1 on LiveCodeBench v6 and 46.4 on SWE-bench Verified, and it shows its biggest margin on tool use, including 97.1 on τ²-Bench Telecom and 66.6 on BFCL v4. It also leads on NoLiMa with 68.1, but it slips behind on some broader knowledge tests, including 70.8 on MMLU-Pro and 8.9 on Humanity’s Last Exam.
OpenBMB says the training stack runs from SFT to RL to on-policy distillation. The post-training phase starts with 400B tokens of deep-thinking SFT, then uses specialized RL teachers for math, code, agentic tasks and writing, and finishes with OPD, which merges 16 RL experts into one shipped model. The company says the open data release includes Ultra-FineWeb, Ultra-FineWeb-L3, UltraX, UltraData-Code, UltraData-Math, UltraData-SFT-2605, UltraData-SFT-Agent-2609 with 500K agent samples, and UltraData-RL-2609 with more than 80K RL samples. It also published intermediate checkpoints, so the Base, Midtrain and SFT-only stages can be checked directly.
My take — AI-written commentary, not fact-checked reporting
This is the sort of release that actually deserves attention: open weights, open data, and intermediate checkpoints, not just a flashy average score. MiniCPM5-2B looks tuned for work that needs tools and context, which is more useful than another model bragging about general trivia it will forget by next Tuesday. The industry could use fewer victory laps and more receipts like this.
Read more about this at: MarkTechPost
Related stories
Someone Fine-Tuned OpenBMB’s MiniCPM5-1B on Claude Fable 5 Traces to Ship a 657MB Local Thinking Model
MarkTechPost · 2 months ago ·
30
Liquid AI Releases LFM2.5-2.6B: An On-Device Agentic Model With 128K Context, Tool Calling, And Open Weights
MarkTechPost · 1 month ago ·
45
Data Machina #250
Substack · 2 years ago ·
32