ME AI message, noting Beating AI breaking news: after the release of MiMo-V2.6, Xiaomi MiMo负责人罗福莉 further explained this round of reinforcement learning. She said she had previously been involved in DeepSeek R1, and that the research innovations and engineering challenges behind MiMo-V2.6, in her view, have already surpassed R1. She also explained why MiMo uses both MixRL and MOPD at the same time. MixRL mixes verifiable tasks—such as code, general agents, vision, and network security—into the same round of reinforcement learning for joint training. MOPD, on the other hand, handles ultra-long, hard-to-verify, or more subjective tasks: it trains them separately first, and then integrates those capabilities back into the main model. For tasks like games and 3D, the runtime is too long and it’s also hard to automatically judge right or wrong. If these tasks are put in the same RL round as code-related tasks, training will be noticeably slowed. MiMo therefore trains these kinds of tasks separately, and then uses MOPD to fold the learned capabilities back into the main model. (Source: ME)
