Alibaba has released Qwen-Image-3.0, an image generator designed for practical applications such as newspaper layouts, complex infographics, and other information-dense visual content. The model processes inputs up to 4,500 tokens and renders text as small as ten pixels, mathematical formulas, and twelve languages in a legible way in a single pass. It is currently available only through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon. The model is meant to handle practical work such as newspaper layouts, storyboards, and exam sheets, not just produce attractive images.

Qwen-Image-3.0 accepts prompts of up to 4,500 tokens, giving the model enough room to create dense layouts in one pass rather than assemble them from several images. One demo packs nine separate infographics into a 3 x 3 grid, each with its own text, formulas, and illustrations. The panels cover topics ranging from safe following distances near tunnels to the detachment speed of a projectile from a rotating cylinder. Other panels explain the liver fluke life cycle, right-sided chest pain, and Sylow theorems for groups of order 72. The 3x3 grid shows how text-heavy and formula-rich content from engineering, philosophy, physics, medicine, math, finance, and cell biology can be laid out in a single coherent image.

According to the Qwen team, the model can produce legible text as small as ten pixels. Its examples include a whale shark infographic packed with text and a full page from a fictional algebraic geometry paper. The paper contains multi-line LaTeX equations with subscripts, superscripts, braces, fractions, sums, and products. Other demos show a simulated newspaper page and red handwritten comments that resemble notes from a teacher. Qwen-Image-3.0 also aims for photographic detail in portraits and objects, including visible pores, skin texture, and individual strands of hair.

Source: thedecoder