On the evening of June 8, Xiaomi's MiMo technology team officially rolled out the Xiaomi MiMo-V2.5-Pro-UltraSpeed mode. Leveraging full-link engineering optimization, this mode, for the first time, attains an inference speed surpassing 1000 tokens per second on general-purpose GPUs, eliminating the necessity for custom chips while preserving the model's original capabilities intact. This significant achievement breaks through the industry's long-standing bottleneck, where achieving high speed, low power consumption, and compatibility with general-purpose GPUs simultaneously seemed unattainable. From now until June 23, this innovative mode will be accessible on a limited-time, application-only basis. Interested users are invited to apply for API access to experience its cutting-edge performance firsthand. Previously, Xiaomi AI has made a series of remarkable strides in enhancing model capabilities, reducing inference costs, and boosting efficiency. On April 23, MiMo-V2.5-Pro clinched the top spot in both the Comprehensive Intelligence Index and Agent Index among open-source models in a globally recognized evaluation. On May 27, the pricing for series model APIs was slashed by up to 99%, accompanied by adjustments to the billing system. And on June 8, the UltraSpeed mode set a new benchmark for inference speed in trillion-parameter models, further solidifying Xiaomi's position at the forefront of AI innovation.
