“Impressive level of openness”: Xiaomi goes way beyond the usual open-weight playbook with MiMo-V2.6
The New Stack Paul Sawers ● Covered by 6 sources
Xiaomi released MiMo-V2.6, a big open model with a 1M-token window. The surprise is how much of the training and RL setup it says it will open up.
Based on reporting by The New Stack, Paul Sawers — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Xiaomi has put out MiMo-V2.6, a new open model that reaches into frontier-model territory on paper and, more unusually, opens up a large chunk of the machinery behind it. The flagship MiMo-V2.6-Pro is a trillion-parameter model with 42 billion parameters active at a time, can handle text, images, audio and video, and uses a one-million-token context window. Xiaomi says it performs at frontier level across coding, agentic tasks, cybersecurity, multimodal work and research. Artificial Analysis puts it first among the 114 large open-weight models it tracks, with an Intelligence Index score of 46.
The more interesting part is how Xiaomi trained it, and how much of that it showed in public. The company livestreamed reinforcement-learning runs through a dashboard for five days starting on September 15, letting people watch metrics from production RL in real time. By the end, the dashboard showed $854,044 in RL costs for MiMo-V2.6-Flash and $2,620,670 for Pro, or about $3.5 million combined. That figure covers only the RL stage. Xiaomi has not said what pretraining cost.
That kind of public bill is still rare. Last year, MiniMax said the RL phase of its 456-billion-parameter MiniMax-M1 cost $534,700 in GPU rental, and DeepSeek put the RL training of its 671-billion-parameter R1 at $294,000. Xiaomi’s setup is larger, and its RL run was aimed at longer, agentic tasks. Fuli Luo, who leads Xiaomi’s MiMo team after previously working at DeepSeek, said the team spent almost six months pushing RL by scaling compute, environments and grading resources, and said more details would be open-sourced over the coming weeks.
Xiaomi also released the model weights under the MIT license, along with a technical report and a 9-billion-parameter Qwen-based model meant as a starting point for more agentic RL research. The company says it has fully open-sourced more than 7,000 task environments, plus an end-to-end training framework and lightweight agent harnesses. At least for now, only the three model releases are visible on Hugging Face, with the larger set of environments and tools not yet surfaced there. Even so, people in the research community are already focusing less on the benchmark score than on the package around it. The model is one thing. The training playground may be the bigger story.
My take — AI-written commentary, not fact-checked reporting
This is the sort of openness that actually matters, not the usual “trust us, the weights are public” routine. If Xiaomi really ships the environments and RL tooling, that’s more useful to researchers than another victory lap over a leaderboard. Open AI talk is cheap; open training machinery is where the grown-up arguments start.
Read more about this at: The New Stack
Related stories
The Next Top Open Model, Google Voice Agents, DeepSeek Shrinks Caches: Plus a letter on open weight cybersecurity capabilities
The Batch ·
45
Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
MarkTechPost · 1 month ago ·
5
The Unbearable Cheapness of Open Weight Models
jamesoclaire.com · 3 months ago ·
15