Xiaomi’s MiMo team did not merely drop another model card on September 22, 2026. It open-sourced the MiMo-V2.6 series—native multimodal Pro and Flash checkpoints, plus Distill-Qwen-9B and a pile of reinforcement-learning research resources—and attached the kind of training receipt most labs keep in a vault.
According to Xiaomi’s own release (and TechNode’s write-up), the two flagship runs completed about 30 RL steps each in under six days, chewing through roughly 750,000 trajectories per model. Sticker shock, itemized: about $2.62 million for Pro and $850,000 for Flash—call it $3.47 million if you like round numbers and mild existential dread.

Benchmarks with caveats (as all good benches should)
Xiaomi says MiMo-V2.6-Pro scored 46 on the Artificial Analysis Intelligence Index, edging Kimi K3 (44) and GLM-5.3 (45) in the company’s cited comparison and landing as the top open-weight entry on that table. Closed-source leaders still sit nearer 53 on the same index story—so this is a loud open-source milestone, not a claim that the frontier keys just changed hands.
On agent-style benches, Xiaomi claims Pro is competitive with Claude Opus 5 and GPT-5.6 Sol on several evaluations, while Flash comprehensively beats the prior MiMo-V2.5-Pro. Treat vendor leaderboards as progress markers: useful, braggable, and still waiting for independent rematches.

What’s in the box
Both Pro and Flash are pitched as fully multimodal—text plus richer perception—with expanded 3D spatial reasoning, computer-use chops, and a new desktop client for people who prefer apps over raw API curl. API pricing, Xiaomi stresses, stays aligned with the V2.5 generation: more intelligence, same sticker on the meter.
Parameter-count gossip has been floating around community write-ups; a vLLM recipe for MiMo-V2.6-Pro-RL describes the Pro MoE as roughly 1.02 trillion total parameters with about 42 billion active per token. That is a serving story as much as a brag—expect serious GPU math if you self-host the full Pro stack.
Weights and accompanying materials are on Hugging Face under Xiaomi’s MiMo collection (MIT-licensed open weights per Xiaomi’s open-source framing), alongside RL environments, training-framework pieces, and the Distill-Qwen-9B companion for researchers who want a smaller on-ramp into the same agentic RL playground.
Why the RL live-blog matters
Scaling reinforcement learning is usually a closed-door sport: batch sizes, grader compute, reward-hacking defenses, frozen MoE routers, and the occasional “whoops, the dataset was cursed for three hours.” Xiaomi’s public materials lean into that grind—larger asynchronous batches, multi-task harnesses spanning code/general/visual/cyber, and groupwise grading meant to keep long-horizon agents honest.
For the open-source community, the interesting artifact is not only the score. It is the combination of MIT-weight availability, documented RL cost, and tooling meant to let outsiders reproduce pieces of the loop. Whether those reproductions hit the same DeepSWE-style jumps Xiaomi advertises is the next chapter.
Open resources, not just weights
Beyond the Pro and Flash checkpoints, Xiaomi is pushing the Distill-Qwen-9B companion and a grab bag of RL research materials: thousands of task environments spanning software engineering, vulnerability reproduction, knowledge work, and web design; an end-to-end training framework built around open components; and lightweight harness pieces meant so labs can remix prompts, tools, and context managers instead of cloning one brittle stack. That is the difference between “here is a GGUF, good luck” and “here is a playground.”
Xiaomi also says the V2.6 generation expands from vibe coding into something closer to a “vibe world”—3D game scaffolding, Blender-style modeling assists, embodied-sim control loops, and computer-use agents that click through GUIs with visual feedback. Marketing language? Sure. But it maps to the multimodal pitch: models that are supposed to see, act, and iterate, not only autocomplete your README.
Geeknewz take
MiMo-V2.6 will not end the closed-vs-open argument overnight. A 46 versus ~53 gap is still a gap. What it does do is raise the open-weight floor while publishing a training bill most competitors would redact. If you care about agents that can see, click, and ship code without a mystery invoice, this drop is mandatory reading—and mandatory downloading.
Sources: TechNode; primary details from Xiaomi MiMo.
