During the 2025 The Bund Summit, Ant Group and Renmin University of China came together to present a groundbreaking achievement in the industry: the debut of 'LLaDA-MoE,' the first native Mixture-of-Experts (MoE) architecture diffusion language model. Built entirely from the ground up using approximately 20 terabytes of data, this model exhibits remarkable scalability and stability when subjected to industrial-scale, large-scale training scenarios. It surpasses the capabilities of its predecessors, LLaDA1.0/1.5 and Dream-7B, by a significant margin. Notably, 'LLaDA-MoE' achieves performance on par with autoregressive models, all while dramatically accelerating inference speed by several fold. This breakthrough challenges the long-held notion that 'language models must adhere to an autoregressive framework.' It excels in multi-task environments, demonstrating performance equivalent to a 3-billion-parameter dense model, yet it does so by activating only 1.4 billion parameters. The collaborative team from Ant Group and Renmin University tackled core technical hurdles over a three-month period, resulting in an average performance boost of 8.4% across multiple benchmark tests. In the near future, the model will be made fully open-source. This release will encompass not just the model weights but also a proprietary inference framework and an inference engine meticulously optimized for the parallel characteristics of dLLM. Simultaneously, the relevant code and technical reports will be made available to the wider community. Looking ahead, Ant Group remains committed to investing in AGI (Artificial General Intelligence) research and development, leveraging dLLM as a foundation to drive further technological advancements.
