Huawei has successfully integrated an on-device Mixture of Experts (MoE) full-modal large model, boasting a total of 30 billion parameters (with 2 billion actively utilized), into its latest-generation triple-folding smartphone, the Mate XT 2. This cutting-edge model operates on the Kirin 9050 Pro chip, with its computational prowess bolstered by the Da Vinci architecture NPU. Initially unveiled in June of this year, this model has undergone specific optimizations tailored for Kirin chips. Leveraging advanced techniques such as expert prediction, it significantly reduces memory consumption while boosting inference throughput. The implementation is slated for rollout on Kirin chips and smartphones in the autumn season.
