アリババ「Qwen-Image-3.0」発表、12言語対応の画像生成モデル
アリババのQwen-Image-3.0は、単なる画像生成から「実用性」を追求し、複雑なレイアウトの生成や微細なテキスト・詳細描写を実現する新世代モデルとして登場した。
キーポイント
Rich Content(豊富なコンテンツ)
最大4.5kトークンの入力に対応し、新聞やストーリーボード、試験用紙など複雑なレイアウトを単一の生成処理で正確に再現できる。
Authentic Details(本物の詳細)
10pxの微小テキストも正確にレンダリングし、肌の毛穴や髪の毛 strands などの微細なディテールを写実的に描写する能力を持つ。
Deep Knowledge(深い知識)
12言語をネイティブにサポートし、ウェブページやゲーム、ライブ配信などの主流インターフェースをシミュレートできる。
実用性の追求
「見栄え」だけでなく「使えるツール」としての価値を重視し、画像生成を本格的な生産性ツールへと進化させた。
重要な引用
Qwen-Image-3.0 is not just pursuing “good-looking” — it is pursuing “useful”, making image generation a truly deployable productivity tool.
the entire image above was generated by Qwen-Image-3.0 in a single pass, rather than being stitched together from multiple images.
影響分析・編集コメントを表示
影響分析
この発表は、AI画像生成技術が単なるクリエイティブツールから、複雑なドキュメント作成や教育・医療資料の生成など実務業務を担う生産性ツールへと進化することを示しています。特に大規模なコンテキスト処理と微細な描写能力の向上は、業界標準を再定義する重要な転換点となり、今後のAI応用領域を大幅に拡大させるでしょう。
編集コメント
アリババが公開したQwen-Image-3.0は、画像生成AIの「実用性」への転換点を象徴するモデルです。複雑なレイアウトや微細なテキスト描写が可能になったことで、クリエイティブ領域だけでなく、ドキュメント作成や教育資料など実務現場での活用がさらに加速すると予想されます。
imageQWEN CHAT
Qwen-Image シリーズの第 3 世代となる基盤画像生成モデル「Qwen-Image-3.0」をリリースします。Qwen-Image-1.0 のキーワードが「精度」であり、Qwen-Image-2.0 が「精度」「多様性」「完全性」「美しさ」「真正性」だったのに対し、Qwen-Image-3.0 の核心は一言で表せます。「リアル(実)」です。
この「リアル」は、以下の 3 つの側面で体現されています。
- 豊富なコンテンツ: 最大 4,500 トークンの入力に対応し、新聞やストーリーボード、試験用紙といった複雑なレイアウトを難なく生成します。
- 本物のディテール: 10px の小さな文字も正確に描画でき、毛穴や髪の毛一本に至るまで、微細なレベルで生々しく再現します。
- 深い知識: 12 か国語のネイティブなレンダリングをサポートし、Web ページやゲーム、ライブ配信など主流のインターフェースをシミュレーション。豊富な世界知識に基づいて生成を行います。
一言で言えば、Qwen-Image-3.0 は単に「美しい」画像を作るだけでなく、「実用的」であることを目指しています。これにより、画像生成は真に現場で活用できる生産性ツールへと進化しました。
豊富なコンテンツ
まずは Qwen-Image-3.0 が生成した以下の画像をご覧ください。

必ず JSON 形式で返してください。translation フィールドのみ。他のフィールドは一切追加しないこと — 余計なフィールドを書こうとして本文がトークン上限で打ち切られる事故を防ぐため:
{"translation": "リライト全文"}
ご覧の通り、Qwen-Image-3.0 は数式スライドを正確に描画できます。空間関係や数学記号、定理の説明といった豊富な視覚情報を適切に配置し、相対位置も正しく保たれた、情報量の多い画像として出力されています。
しかし、これが Qwen-Image-3.0 の真の強みではありません。実はこれはその能力の「1/9」に過ぎません。なぜなら、この画像は Qwen-Image-3.0 が生成した複雑な 3×3 グリッドの単一セルにすぎないからです。元の画像をご覧ください。

その通り、上記の画像全体は Qwen-Image-3.0 によって一度の生成処理で描かれました。複数の画像を拼接したものではありません。この画像が難しい理由は、各セルが複雑なインフォグラフィックである点にあります。この 3×3 グリッド全体を正確に記述するには、なんと 3,700 トークンが必要です。
これら 3,700 トークンは、トンネル安全に関する漫画、空間幾何の授業、「出師表」の文体分析、物理学的な投射運動、生物学の寄生虫解説、右胸痛の医療図解、群論のサイロウ定理、銀行内部統制管理のインフォグラフィック、そして細胞 DNA 構造の比較という 9 つの異なるテーマを完全に描写しなければなりません。各セルには正確な中国語と英語のテキスト、数式、グラフ、キャラクターなどが含まれています。
それでも Qwen-Image-3.0 はこれを難なくこなします。その理由は、許容する指示の長さを 4,500 トークンに引き上げたからです。これにより、極めて複雑で情報密度の高い視覚レイアウトを理解し、描画することが可能になりました。
これは前述した「リッチコンテンツ」の特徴の一つです。コンテンツは水平方向へ拡張できます。この水平展開は、モデルが持つセマンティックな並置(意味の対比)と空間制御の強さを示しています。つまり、単一の画像内で複数の概念を秩序立てて配置し、互いに干渉させずに描画できる能力のことです。
水平方向への拡張に加え、「深さ」もリッチコンテンツのもう一つの重要な特徴です。
水平展開が「1 つのキャンバス上に並列に配置できる要素はいくつまでか」というテストである一方、深さはモデルのセマンティックな分解(意味の分解)と論理的なネスト(入れ子構造)の能力を問うものです。単一の画像内で、複数のインターフェースを階層ごとに描画できるかどうかを試すのです。
以下の例では、1 つの指示で外側から内側へと、VSCode のプログラミングインターフェース → Qwen Chat のチャット画面 → WeChat のメッセージ画面 → ドリップコーヒーのポスターというように表示しています。各レイヤーはそれぞれの UI が持つ本物のスタイルとディテールを保持しており、「絵の中の絵の中の絵」といった視覚的な奥行きを生み出します。

上記の 2 つの例は、「リッチコンテンツ」が水平方向と垂直方向の両方の次元において何を意味するかを物語っています。
本物のディテール
「リッチコンテンツ」が「どれだけ描くか」という問いに答えるなら、「オーセンティック・ディテール」は「いかに細かく描くか」という問いに応えるものです。Qwen-Image-3.0 は、微細なレベルの描写精度において新たな高みに到達しました。10px の小さな文字も明確に読み取れるようになり、毛穴や髪の毛一本一本まで繊細に表現され、肌の質感は写真のようなリアリティを備えています。
まずは、小さな文字の正確な描画から見ていきましょう。
以下はサメザメに関する知識インフォグラフィックです。大量のテキストとイラストが含まれていますが、Qwen-Image-3.0 はすべての領域を正確にレンダリングできます。

学術論文は、小さな文字の描画能力に対する究極のストレステストです。密度の高い LaTeX 数式、添字や上付き文字、ギリシャ文字、定理番号など、一文字たりとも誤りが許されません。

このモデルは、代数幾何学の分野における学術論文の 1 ページ全体をレンダリングします。複雑な数式の導出が複数行にわたって含まれていますが、上付き文字や添字、中括弧、分数バー、そして多行整列といった LaTeX の組版要素もすべて正確に表現されています。小さなフォントサイズであっても、優れた可読性を維持しています。
Qwen-Image-3.0 は、リアルな紙媒体上の細かな文字も生成できます。以下の例は同モデルが作成した新聞ですが、密度の高いテキストを正確に再現するだけでなく、新聞らしい質感や雰囲気まで忠実にシミュレートしています。

編集作業においても、同様に細かな文字の生成が可能です。例えば以下のケースでは、モデルがリアルなスタイルのアノテーションを出力しています。
image
image
モデルは、本ページに赤い手書きのアノテーションを重ねています。下線、波線、丸印、矢印、そして短いコメントなど、高校生の授業ノートのような自然で流れるような筆跡を完璧に再現し、そのスタイルを忠実に模倣しています。
文字やレイアウトの精巧な再現に加え、「Authentic Details(本物の詳細)」という点は、質感の描写においても際立っています。以下は、極めて繊細な質感を捉えたポートレート写真の例です。
image
image
ポートレートに限らず、モデルは他の物体の繊細な質感も描写することが可能です。
画像編集タスクにおいても、Qwen-Image-3.0 は細部まで豊かなディテールを持つ画像を生成できます。
損傷や欠落がある伝統的な絵画に対しては、元の芸術様式や筆致を忠実に維持しながら、欠けた部分を補完して修復することが可能です。
このモデルは、元の作品の筆致に忠実にエッジを復元し、墨の濃淡や羽の質感、構図のバランスを保ちつつ、カビの斑点や損傷の跡を除去して修復を行います。
深い知識
「豊かなコンテンツ」は「どれほど複雑な描画が可能か」、「本物のような詳細」は「どれほどリアルに描けるか」、そして「深い知識」は「どの程度の広範囲を描画できるか」を示します。Qwen-Image-3.0 は、12 の言語、複数のフォント、100 以上の芸術様式、多彩な UI インターフェースの描画能力を備えており、これらはすべてモデルが持つ世界知識への深い理解に基づいています。
以下の 3 つの例では、それぞれ日本語、韓国語、スペイン語を正確に描画しています。

image
image
複数の言語の正確な描画に加え、モデルは豊富な世界知識も備えています。特に、さまざまなリアルな UI インターフェースを生成することが可能です。


また、このモデルの強力な世界知識を活用してインフォグラフィックを作成することも可能です。以下の例では、実際の画像を基に複雑なインフォグラフィックを生成しています。
image
image
元の昆虫写真の主要な被写体を保ちつつ、分類情報や形態学的注釈、拡大した詳細ビュー、スケールバーといった専門的な要素を追加。これにより、学術論文にそのまま使用できる研究用図版が完成します。
モデルが既に持つ世界知識に加え、インターネットへの接続機能によって最新の情報を取得することも可能です。例えば、「7 月 21 日の杭州の天気予報画像を作成して」と指示すれば、その場で生成されます。
image.png)
また、特定の IP(知的財産)キャラクターを見つけて、それらを基に画像を創作することもできます。例えば、「斉白石とゴッホが Qwen-Image-3.0 をライブ配信で紹介している様子」を生成させることも可能です。

結論
Qwen-Image-1.0 が「精度」を追求し、2.0 で「精度、多様性、完全性、美しさ、そして真実味」を実現したのに対し、3.0 ではついに「リアルさ」へと進化しました。私たちが一貫して目指してきたのは、画像生成技術を単に「使えるもの」「見栄えの良いもの」から、「実際に役立つもの」へと昇華させることです。
「豊富なコンテンツ」「真実味のある詳細描写」「深い知識」という 3 つのコア機能を備えた Qwen-Image-3.0 は、新聞の PDF データ処理や短編ドラマのストーリーボード作成、複雑な UI インターフェースの設計など、高付加価値な業務シーンにおいて大きな飛躍を遂げました。画像生成モデルの能力がさらに向上するにつれ、デザイン、コンテンツ制作、教育、EC 分野など、より多くの領域で真の生産性価値を発揮していくと確信しています。
今回のアップデートにおける主要なポイントについては以上です。Qwen-Image-3.0 をぜひご活用ください!
原文を表示

We are launching Qwen-Image-3.0, the third-generation foundational image generation model in the Qwen-Image series. If the keyword for Qwen-Image-1.0 was “Precision”, and the keywords for Qwen-Image-2.0 were “Precision, Variety, Completeness, Beauty, and Authenticity”, then the core of Qwen-Image-3.0 comes down to a single word — “Real” (实).
This “Real” is embodied across three dimensions:
- Rich Content: Supports up to 4.5k token input, effortlessly generating complex layouts such as newspapers, storyboards, and exam papers.
- Authentic Details: Supports precise rendering of text as small as 10px, vividly reproducing details like pores and hair strands with lifelike, micro-level depiction.
- Deep Knowledge: Supports native rendering of 12 languages, simulates mainstream interfaces such as web pages, games, and livestreams, and draws on rich world knowledge.
In a word, Qwen-Image-3.0 is not just pursuing “good-looking” — it is pursuing “useful”, making image generation a truly deployable productivity tool.
Rich Content
Let’s start with the image below, generated by Qwen-Image-3.0:

As you can see, Qwen-Image-3.0 can accurately render a math slide, including spatial relationships, mathematical symbols, theorem descriptions, and other rich visual content. This content is laid out reasonably, with proper relative positioning, and looks rich in information.
However, this is not the true strength of Qwen-Image-3.0. In fact, this is only “1/9” of its real capability, because this image is actually just one cell of a complex 3×3 grid generated by Qwen-Image-3.0. Let’s look at the original image:

That’s right — the entire image above was generated by Qwen-Image-3.0 in a single pass, rather than being stitched together from multiple images. The difficulty of this image lies in the fact that each cell is a complex infographic; to precisely describe the full 3×3 grid takes a full 3.7k tokens.
These 3.7k tokens must fully depict a tunnel safety comic, a spatial geometry lesson, a stylistic analysis of “Chu Shi Biao” (Memorial on Dispatching the Troops), physics projectile motion, a biology parasitology explainer, a medical diagram of right-side chest pain, the Sylow theorems of group theory, a bank internal-control management infographic, and a cell DNA structure comparison — each cell containing precise Chinese and English text, formulas, charts, cartoon characters, and more.
And yet this remains effortless for Qwen-Image-3.0 — because Qwen-Image-3.0 raises the acceptable instruction length to 4.5k tokens, which means the model can understand and render extremely complex, information-dense visual layouts.
This is an important characteristic of the “Rich Content” we mentioned: content can expand horizontally. Horizontal expansion reflects the model’s strength in semantic juxtaposition and spatial control — the ability to lay out multiple concepts in an orderly fashion within a single image and render them without mutual interference.
Beyond horizontal expansion, depth is another important characteristic of Rich Content.
Horizontal expansion tests “how many parallel elements can be placed on a single canvas,” while depth tests the model’s semantic deconstruction and logical nesting — whether it can render multiple nested interfaces layer by layer within a single image. The following example uses a single instruction to display, from outer to inner: a VSCode programming interface → a Qwen Chat interface → a Wechat interface → a pour-over coffee poster. Each layer preserves the authentic style and details of its respective UI, forming a “picture-in-picture-in-picture” visual depth.

The two examples above illustrate the meaning of “Rich Content” along both the horizontal and vertical dimensions.
Authentic Details
If “Rich Content” addresses the question of “how much to draw,” then “Authentic Details” addresses the question of “how finely to draw.” Qwen-Image-3.0 reaches a new height in the rendering precision of micro-level details: 10px small text is clearly legible, pores and hair strands are rendered in fine detail, and skin texture approaches photographic realism. Let’s start with the precise rendering of small text.
Below is a knowledge infographic about whale sharks, containing a large amount of text and illustrations. Qwen-Image-3.0 is able to accurately render every region.

Academic papers are the ultimate stress test for small-text rendering — dense LaTeX formulas, subscripts and superscripts, Greek letters, and theorem numbering, where not a single symbol can go wrong.

The model renders a full page of an academic paper in the field of algebraic geometry, including multiple lines of complex formula derivations. LaTeX typesetting elements such as superscripts, subscripts, curly braces, fraction bars, and multi-line alignment are all accurately presented, maintaining excellent readability even at small font sizes.
Qwen-Image-3.0 can also generate fine text on realistic paper. The example below is a newspaper generated by Qwen-Image-3.0, in which the model not only accurately generates dense text but also simulates the authentic look of a newspaper.

In editing tasks, we can also generate fine text. For example, in the case below, the model produces annotations with a realistic style.


The model overlays realistic red handwritten annotations onto the book page — underlines, wavy lines, circles, arrows, and short comments — with natural, fluent handwriting that perfectly simulates the style of a high school student’s class notes.
Beyond the fine reproduction of text and layout, “Authentic Details” also stands out in texture depiction. Below are two portrait photography examples in which the model captures extremely delicate textures.


Beyond portraits, the model can also depict the delicate textures of other objects.
.png)
.png)
.png)

In editing tasks, we can also generate images with rich details.




Given a damaged or incomplete traditional painting, the model can restore the missing parts while faithfully maintaining the original artistic style and brushwork.


The model completes the restoration of the eagle-combat painting with brushwork consistent with the original, preserving the ink-wash gradients, feather texture, and compositional balance while removing mold spots and signs of damage.
Deep Knowledge
“Rich Content” answers “how complex can it draw,” “Authentic Details” answers “how lifelike can it draw,” and “Deep Knowledge” answers “how broadly can it draw.” Qwen-Image-3.0 possesses rendering capabilities covering 12 languages, multiple fonts, 100+ artistic styles, and a variety of UI interfaces — all backed by the model’s deep understanding of world knowledge.
In the three examples below, the model accurately renders Japanese, Korean, and Spanish respectively.



Beyond accurate rendering of multiple languages, the model also possesses rich world knowledge. In particular, the model can generate various realistic UI interfaces.


We can also leverage the model’s powerful world knowledge to create infographics. In the example below, we generate a complex infographic based on a real image.


While preserving the main subject of the original insect photograph, the model adds professional elements such as taxonomic information, morphological annotations, magnified detail views, and a scale bar, producing a research figure ready for direct use in academic publication.
In addition to the world knowledge the model already possesses, it can also connect to the internet to retrieve the latest world knowledge. For example, we can ask the model to generate a weather forecast image for Hangzhou on July 21.
After editing.png)
The model can also find specific IP figures and create based on them. For instance, we can generate an image of Qi Baishi and Van Gogh introducing Qwen-Image-3.0 in a livestream room.

Conclusion
From the “Precision” of Qwen-Image-1.0, to the “Precision, Variety, Completeness, Beauty, and Authenticity” of Qwen-Image-2.0, and now to the “Real” of Qwen-Image-3.0 — the goal we have always pursued is to move image generation from “usable” to “practical,” and from “good-looking” to “useful.”
Supported by its three core features — “Rich Content, Authentic Details, and Deep Knowledge” — Qwen-Image-3.0 achieves significant breakthroughs in high-value productivity scenarios such as newspaper PDFs, short-drama storyboards, and complex UI interfaces. We believe that as the capabilities of image generation models continue to improve, they will unlock genuine productivity value in even more fields, including design, content creation, education, and e-commerce.
That concludes the main highlights of this update. We hope you enjoy using Qwen-Image-3.0!
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み