Luo Fuli, who leads the development of Xiaomi's MiMo large-scale model, announced that since the open-source release of MiMo-v2.5 in April, his team has invested nearly half a year in exploring the expansion of reinforcement learning boundaries. Presently, MiMo-V2.6 is undergoing an intermediate phase of reinforcement learning training. The team has successfully integrated three new functionalities: enhanced computational capabilities, enriched environmental and tool sets, and Grader Compute for performance evaluation. Details regarding these advancements will be progressively open-sourced over the next few weeks.
Furthermore, Luo Fuli shared a live-stream link showcasing the model training process. The live-stream indicated that MiMo-V2.6 is on the verge of official launch, with the total training expenditure surpassing US$1.25 million, underscoring Xiaomi's commitment to pushing the frontiers of AI technology.
