On September 22, Xiaomi released and open-sourced the MiMo-V2.6 family: the omnimodal flagship MiMo-V2.6-Pro, the efficient inference model MiMo-V2.6-Flash, and an UltraSpeed mode that pushes output up to 20x. The team also released model weights, the technical report, a distilled model, and supporting RL research resources, claiming this training run is "possibly one of the largest single RL training runs by an open-source model team." According to media reports, the large-scale reinforcement learning run cost roughly $3.47 million.

[1][2]

This is not an ordinary model drop. Xiaomi explicitly anchors the generation to the RSI (recursive self-improvement) path: in the official framing, "scaling RL compute on verifiable, complex tasks, so the model can continuously expand its capability frontier through exploration and feedback." In plain terms — instead of spending budget on larger pretraining corpora, pour money into reinforcement learning and let the model iterate on real agent tasks. The open-sourcing goes beyond weights: 7k+ high-quality RL task environments covering software engineering, vulnerability reproduction, knowledge work, and web design and development, with the training recipe released alongside.

The results bought the open-source camp a top spot. Xiaomi states MiMo-V2.6-Pro ranks first among open-weight models on the comprehensive intelligence index (AA Index) and is the highest-ranked Chinese model on the leaderboard — above Kimi K3 and GLM-5.3. The flagship is natively omnimodal (text, image, video, audio in one model), built on a trillion-parameter-scale sparse MoE architecture, with a distilled Qwen-9B variant and the UltraSpeed tier. The narrative significance: the "open vs closed gap" story is shifting from "one generation behind" to "one tier behind," and the tier is being closed by RL compute, not data scale.

Do the math. $3.47 million for one RL run is pocket change for a closed-source giant and a heavy bet for an open-source team — what it buys is not parameters but the position of "number one open-weight model on the intelligence index." The most interesting thing to watch is not Xiaomi itself but the signal: when the open-source camp starts stacking real money into single RL runs, the competitive axis between open and closed shifts from "can you build it" to "dare you burn compute on training." Open-sourcing the RL recipe and task environments also lets the community reproduce the full path to "open number one" for the first time. Note the attribution: the ranking and the "possibly largest" claim are official and media-stated; flagship scores rely on third-party aggregates like the AA Index; the $3.47 million figure is media-reported. Until independently reproduced, all of these should be treated as claims rather than settled facts.

The bigger context is Xiaomi's unusual position in this race. Unlike most open-weight publishers — research institutes or model startups — Xiaomi ships hardware: phones, laptops, and a desktop client that MiMo-V2.6 now feeds. The UltraSpeed tier and the MiMoDesktop release are not accessories; they are the deployment story that justifies burning $3.47 million on a single RL run. A model that is strong on agentic tasks and fast enough for local hardware can sit inside Xiaomi's own devices, turning a training cost into a product moat instead of a research trophy. The same logic explains why the open-source package includes distillation targets and RL task environments: every developer who builds on MiMo-V2.6 extends Xiaomi's ecosystem reach beyond its own hardware. Meanwhile, the Chinese open-source field is visibly consolidating around RL as the differentiator — Kimi K3 and GLM-5.3 traded the same currency — which means the next frontier-model fights will be fought less over pretraining data and more over who can run bigger, better-verifiable RL loops. Whether Xiaomi can sustain this cadence, and whether the AA Index ranking survives independent reproduction, are the two open questions worth tracking.

[1][2]
Early-morning GPU server room, an engineer's back before rows of racks, right hand holding a deployment checklist, a monitor showing an RL reward curve climbing slowly, rack indicator lights blinking in rows, cold light reflected on the floor, dust lit by screen glow
Burning compute on training, AI-generated illustration, not a news photo