On September 22, at the Snapdragon Summit 2026, Qualcomm revealed its partnership with StepFun, AI Infinite, and FORESEE. Together, they have successfully accomplished the on-device adaptation and inference optimization of the StepFun StepEdge-Omni 30B-MoE model, all powered by the sixth-generation Snapdragon 8 Elite mobile platform. This model boasts a staggering 30 billion parameters and leverages a Mixture of Experts (MoE) architecture. This innovative approach ensures that only specific modules are activated according to task demands, thereby significantly cutting down on computational and memory requirements. Thanks to meticulous optimization efforts, the partners have managed to slash the model's operational memory needs by over 50%. This breakthrough enables the model to run smoothly on smartphones, eliminating the necessity of relying on cloud services for multiple tasks. Test data indicates that the model's pre-filling throughput performance soars above 330 Tokens/s, while its decoding throughput performance exceeds 28 Tokens/s.
