Research worth keeping up with

AI Research Newsletter.

Consequential AI research and releases, with the context that makes them matter.

A curated digest of work from across the field. Prepared with AI assistance and primary-source links; these are summaries of others’ research.

Daily · Up to three updatesSignal over volume

Reinforcement learning

← All updates · 1 update

Xiaomi releases MiMo-V2.6 weights after large-scale reinforcement learning on agent tasks

Xiaomi released MIT-licensed MiMo-V2.6-Pro-RL and Flash-RL checkpoints, with a technical report describing mixed-task reinforcement learning across coding, general work, vision, and cybersecurity. Xiaomi reports that its six-day training run raised Pro's DeepSWE v1.1 score from 58.4 to 72.6; Artificial Analysis separately measured an Intelligence Index score of 46, placing Pro among the leading open-weight models.

Why it made the cut: Strong independently measured capability combined with downloadable weights gives builders a consequential alternative for running and adapting long-horizon agents. The training gains remain vendor-reported, benchmark scores do not guarantee operational reliability, and Xiaomi's “self-improvement” framing describes a designed reinforcement-learning process rather than demonstrated autonomous recursive improvement. Xiaomi also announced training environments and RL tooling; their complete release was not independently verified here.

Technical report · Official announcement · Pro model weights · All released checkpoints · Independent evaluation (Artificial Analysis)