The InclusionAI team, a part of Ant Group, has recently open-sourced the LLaDA-Image, a 6B-parameter text-to-image model. This model employs a comprehensive diffusion architecture. During the pre-training and mid-training stages, it builds visual priors using over 90% pure image samples. Subsequently, in the SFT phase, text-image pairs are introduced to ensure alignment with language. LLaDA-Image tops the list among open-source models on both the Chinese and English leaderboards of Qwen-Image-Bench, with its overall score nestled between that of GPT Image 1 and Imagen 4.0 Ultra. The model is versatile, supporting both text-to-image generation and instruction-based image editing. Through TwinFlow distillation technology, it streamlines the inference process, reducing steps from 50 to just 2-4, all while preserving high-quality image generation. At present, the model weights, training code, and a comprehensive guide have all been made publicly available.
