Xiaomi releases MiMo-V2.6 weights after large-scale reinforcement learning on agent tasks
Xiaomi released MIT-licensed MiMo-V2.6-Pro-RL and Flash-RL checkpoints, with a technical report describing mixed-task reinforcement learning across coding, general work, vision, and cybersecurity. Xiaomi reports that its six-day training run raised Pro's DeepSWE v1.1 score from 58.4 to 72.6; Artificial Analysis separately measured an Intelligence Index score of 46, placing Pro among the leading open-weight models.
Why it made the cut: Strong independently measured capability combined with downloadable weights gives builders a consequential alternative for running and adapting long-horizon agents. The training gains remain vendor-reported, benchmark scores do not guarantee operational reliability, and Xiaomi's “self-improvement” framing describes a designed reinforcement-learning process rather than demonstrated autonomous recursive improvement. Xiaomi also announced training environments and RL tooling; their complete release was not independently verified here.
Technical report · Official announcement · Pro model weights · All released checkpoints · Independent evaluation (Artificial Analysis)
Link to this post