Xiaomi’s LLM-Core team has published a technical report on MiMo-V2.6, an omni-modal model family whose gains the company attributes to scaling reinforcement learning compute rather than pre-training.
The report describes scaling along three axes. Throughput: asynchronous training that consumes 1,568 samples and roughly 2.7 to 3.7 billion tokens per step, at context lengths of up to a million tokens. Environments: code, general, visual and cyber tasks under a mixture of agent harnesses. Grading: groupwise agentic grading, meant to give more accurate reward signals on long tasks and to push the model towards shorter answers. Two details stand out for anyone running similar training. To hold the run stable, the team freezes the mixture-of-experts router, and it describes a multi-layer defence against reward hacking rather than a single filter.
What is actually open matters here. Xiaomi says it is releasing the training dynamics, the reinforcement learning environments and the framework, and weights are published on Hugging Face under an MIT licence in Pro and Flash variants. The weight files put Pro at about 1.02 trillion parameters in total and Flash at about 311 billion, although for a mixture-of-experts model the active parameter count is what governs serving cost, and the report’s abstract makes no headline benchmark claim at all.
Source: arXiv:2610.11959, announced 9 October 2026.
