On September 20th, reports emerged indicating that the Qwen large model team has officially open-sourced the Qwen-Image-2.1 image model. This model strikes an optimal balance between generation effectiveness, inference efficiency, and operational costs. Qwen-Image-2.1 seamlessly integrates text-to-image generation and image editing capabilities into a unified platform. Its visual generation component is notably compact, featuring only 7 billion parameters, while inherently supporting the generation and editing of transparent images.
The upgraded model showcases four primary attributes:
Firstly, it embodies a compact yet robust design, offering outstanding cost-efficiency. This is achieved through a streamlined model architecture and inference optimizations, ensuring a harmonious equilibrium between generation quality and computational expenses.
Secondly, it natively accommodates transparency and merges creation with modification. This enables the generation of both standard and transparent images based on user prompts, while also facilitating transparent layer editing and image matting operations.
Thirdly, it provides extensive editing capabilities and preserves intricate details across multiple domains. The model supports up to 10 reference images, enhancing local editing precision. It maintains high fidelity for portraits and product imagery, catering to a broad spectrum of editing tasks.
Fourthly, it delivers lifelike textures and refined typography. By refining text layout, human figure lighting, and detail representation, the model ensures that generated results are not only accurate but also aesthetically captivating.
