Xiaomi Live-Streams MiMo-V2.6 Training Process; Luo Fuli Reveals Half-Year Dedication to Reinforcement Learning Research
2 day ago / Read about 0 minute
Author:小编   

Luo Fuli, who leads the development of Xiaomi's MiMo large-scale model, announced that since the open-source release of MiMo-v2.5 in April, his team has invested nearly half a year in exploring the expansion of reinforcement learning boundaries. Presently, MiMo-V2.6 is undergoing an intermediate phase of reinforcement learning training. The team has successfully integrated three new functionalities: enhanced computational capabilities, enriched environmental and tool sets, and Grader Compute for performance evaluation. Details regarding these advancements will be progressively open-sourced over the next few weeks.

Furthermore, Luo Fuli shared a live-stream link showcasing the model training process. The live-stream indicated that MiMo-V2.6 is on the verge of official launch, with the total training expenditure surpassing US$1.25 million, underscoring Xiaomi's commitment to pushing the frontiers of AI technology.