QwenCloud has published a product page for Qwen-Image-3.0-Pro, an image-generation model the company positions as a layout tool rather than a picture generator. According to the vendor, the model accepts prompts up to 4.5k tokens and can assemble dense, multi-panel compositions, images nested inside images, in a single generation pass. QwenCloud lists newspapers, storyboards, restaurant menus and exam papers as example outputs.

The company also claims the model renders text as small as 10 pixels legibly, along with fine physical details such as pores and individual hair strands. QwenCloud says Qwen-Image-3.0-Pro natively supports 12 languages and can mimic interfaces like web pages, games and livestreams.

None of these claims come with independent benchmarks or sample images on the page reviewed. Legible embedded text has been a persistent failure mode for diffusion models, so a vendor asserting reliable 10px rendering plus complex single-pass layout composition is claiming to fix a problem most competitors still struggle with.

Published by QwenCloud on August 6, 2026.