Alibaba’s Qwen has released Qwen-Image-2.1, an open-weight image-generation and editing model with a 7 billion parameter visual component. Qwen announces in an official blog post that the model combines text-to-image generation, image editing, and native transparency support in one system.
The release targets a growing demand for image models that can do more than create a new picture from a text prompt. Qwen-Image-2.1 can also work with existing images, alter selected areas, extract subjects from photographs, and create images with transparent backgrounds. The company says the model is designed to balance output quality, speed, and computing cost.
Qwen describes its model as compact compared with many high-end image systems. Its visual generation component uses 32 Single-Stream DiT layers and contains 7 billion parameters. The company says its internal Qwen-Image-Bench comparison places the model competitively against both open and closed models, though the post does not provide enough independent evidence to verify that claim.
Transparency becomes part of the main model
The most notable addition is native support for transparency. Rather than requiring a dedicated background-removal tool, Qwen-Image-2.1 can generate images with an alpha channel directly from a prompt. An alpha channel stores which parts of an image are opaque and which are transparent.
This could be useful for teams creating product assets, social graphics, presentation materials, or design elements that need to be placed over different backgrounds. The model can also edit an existing transparent image without replacing its background. According to Qwen, it can change details such as a person’s expression or text in a transparent layer.
For regular photographs, the model can isolate a requested subject and return it as an RGBA image. RGBA files contain red, green, blue, and transparency information. This process can make it easier to reuse a person or object in a new composition.
Up to 10 references and local edits
Qwen-Image-2.1 accepts up to 10 reference images in a single request. Qwen presents this as a way to combine separate assets, such as individual portraits in a group photo or clothing items in a virtual try-on image. Other examples include placing several furniture products in a room design.
The model also offers several ways to direct edits to a specific area:
- Colored circles can identify several regions with different instructions.
- Painted annotations can indicate where new content should appear.
- A separate mask can preserve the original image while defining the editable area.
Qwen says the model has improved its ability to preserve identity in portrait edits and maintain product details such as labels, textures, and shapes. These are important capabilities for marketing and e-commerce work, where an edited image must still represent a person or product accurately. As with other generative image systems, users should review output closely for visual errors, altered text, or inaccurate details.
The company also highlights panorama creation, infographic expansion, and storyboard generation as supported tasks. It says the model can build a wider scene from a selfie, turn a product or model photograph into an information-rich graphic, or develop character reference images into sequential storyboards.
For performance, Qwen says it uses a mixed-granularity attention system and KV cache reuse. In practical terms, this means the model can retain static input information, including reference images and editing instructions, rather than recalculating it at every generation step. The company says this reduces memory use and speeds up workflows involving several images.
Qwen-Image-2.1 is released as an open-weight model, allowing developers and organizations to run and adapt it within the terms of its distribution. Its combination of transparency, multi-reference editing, and localized changes positions it as a tool for production-oriented visual workflows rather than prompt-only image generation.
Stay up to date
AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox: