OpenRouter、画像生成モデルの能力を評価するベンチマーク「Visual Image Benchmarks」を公開
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
OpenRouter Blog
OpenRouter は、画像生成モデルの能力を視覚的に評価するための「Visual Image Benchmarks」を発表し、7 つのカテゴリーに分類された難易度の高いプロンプトを用いて 39 のモデルを比較可能にするツールを提供した。
AI深層分析を開く2026年8月23日 02:51
AI深層分析
キーポイント
画像モデル評価の課題と新ツールの登場
従来の LLM ベンチマークやアリーナスコアでは捉えきれない画像生成モデルの詳細な能力を、視覚的な出力サンプルで即座に比較できる「Visual Image Benchmarks」が OpenRouter によって公開された。
7 つのカテゴリーによる境界テスト
不可能なシーンの描写、数え上げ、テキスト生成、空間関係、否定表現、編集処理、一貫性の維持という 7 つの難問カテゴリーが設定され、各モデルの能力差を明確に区別する設計となっている。
価格と生成時間の可視化
ベンチマーク結果はグリッド形式で表示され、モデル名とともにコストと生成時間でのソートが可能となり、ユーザーは要件に応じて最適なモデルを迅速に選定できる。
動画・音声評価への拡張計画
画像モデルの評価ツールは現在段階にあるが、OpenRouter は将来的に動画や音声モデルの品質評価にもこのアプローチを適用し、マルチモーダルな評価基盤へと拡大する意向を示している。
画像モデルの能力境界を検証する7つの課題カテゴリ
不可能なシーン、数え上げ、テキスト、空間関係、否定、編集、一貫性の7つのカテゴリーに分類された課題が用意されている。
重要な引用
Unlike choosing a text model, where we have a vast array of LLM benchmarks, picking an image model can feel arbitrary.
We've selected a set of challenging prompts designed to differentiate the capabilities of models, and show every result in a grid with sorting for both price and generation time.
It's similarly challenging to evaluate video and audio models to understand their quality and capabilities.
Each challenge is designed to differentiate the capabilities of image models.
編集コメントを表示
編集コメント
画像生成モデルの品質評価において、主観的なアリーナスコアの限界を克服する具体的なアプローチが示された点は非常に価値がある。特に「不可能なシーン」や「否定表現」といった難問カテゴリーは、実務でのモデル選定における盲点を突く重要な指標となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
テキストモデルを選ぶ際、LLM ベンチマークが多数存在するため比較は容易ですが、画像生成モデルを選ぶ際は、その選定基準が曖昧に感じられることがあります。出力サンプルは往々にして「見栄えのするもの」に厳選されたものであり、LLM を評価者として用いる手法では、人間なら即座に気づくような細部まで捉えきれていないのが現状です。
アレーナ(Arena)スコアも有用ですが、これはあくまで「ユーザーがどちらを好むか」を測る指標であり、「モデルが実際に何ができるか」を評価するものではありません。
そこで本日、Visual Image Benchmarks を公開します。これにより、OpenRouter が提供するすべての画像生成モデル(2026 年 8 月時点で 39 モデル)の能力を素早く評価することが可能になります。
ここでは、各モデルの能力差を明確に示すよう設計された難易度の高いプロンプトセットを用意しました。すべての結果をグリッド形式で表示し、価格や生成時間でのソートも可能です。
画像モデルの能力限界を試すためのプロンプト
各課題は、画像生成モデルの能力差を明確に区別することを目的としています。当初は以下の 7 つのカテゴリーに分類しています。
- 非現実的なシーン。 リムまで満たされたワイングラスや、閉じられた傘などです。これらはトレーニングデータにおいて「通常の状態」でしか出現しないため、モデルが学習しにくいケースです。
- 数え上げ。 指が 3 本、特定の枚数のカードやサイコロの目など、正確な数を要求するものです。
- テキスト。 ポスターに 1 行の長い文字列を正確に描画させること、あるいは同一フレーム内で複数の言語を正しく表示させることです。
- 空間関係。 物体の重なり(オクルージョン)や鏡像反射など、奥行きや位置関係を正しく表現する能力です。
- 否定。 縞模様のないゼブラや、広告がないタイムズスクエアなど、「ないもの」を描画させる課題です。
- 編集。 リファレンス画像に基づき、最小限の変更を加える、オブジェクトを削除する、人物を削除するなど、画像編集の能力を試すものです。
一貫性。製品や 4 つの参照対象を新しいシーンで安定して保持できるか。
プロンプトは、視覚的に即座に評価できるよう設計されています。例えば、「ワイングラスを完全に満たす」という指示をモデルが正確に実行できるかどうかを確認できます。

今後の展開:視覚評価と他のモダリティも追加予定
動画や音声モデルの評価は同様に困難です。その品質や能力を理解するためには、画像モデルと同様に多様なモダリティに対応し、新たに追加されるすべての画像モデルに合わせて常に最新の状態を維持していく必要があります。
ぜひ、画像ベンチマークをチェックしてください。次回のモデル評価で挑戦できるプロンプトをお持ちの方は、Discord の #feedback チャンネルまで共有ください。
OpenRouter API と Chat を活用した画像生成
これらのベンチマークを通じてモデルを評価した後、画像生成 API や Chat を使って、ご自身のプロンプトで実際に試してみてください。ご自身のコンテンツ上でどのように動作するかを確認できます。
原文を表示
Unlike choosing a text model, where we have a vast array of LLM benchmarks, picking an image model can feel arbitrary. Output samples tend to be curated eye candy and LLM-as-a-judge evals can’t yet capture the details a human would notice instantly. While arena scores help, they evaluate which outputs people prefer rather than what a model can actually do.
Today we’re launching Visual Image Benchmarks to help you quickly evaluate the capabilities of all the image models we offer (39 as of August ‘26). We’ve selected a set of challenging prompts designed to differentiate the capabilities of models, and show every result in a grid with sorting for both price and generation time.
Prompts designed to test the boundaries of image model capabilities
Each challenge is designed to differentiate the capabilities of image models. We’ve initially grouped them into seven families:
- Improbable scenes. A wine glass filled level with the rim, umbrellas that are closed. Training data is full of the ordinary version of both.
- Counting. Three fingers, specific numbers of cards and dice.
- Text. One long exact string on a poster, and several languages in the same frame.
- Spatial relations. Occlusion and mirror reflections.
- Negation. A zebra with no stripes, a Times Square with no advertising.
- Editing. Minimal diffs, object removal, person removal, all from a reference image.
- Consistency. Holding a product or four reference subjects steady across a new scene.
The prompts are written so you can instantly evaluate them visually. For example, can a model follow the instruction to fully fill a wine glass?

More visual evals and additional modalities coming soon
It’s similarly challenging to evaluate video and audio models to understand their quality and capabilities. We intend to expand this tool out across modalities as well as keep it up to date with all the new image models we add.
Check out our image benchmarks today! If you have any prompts that could challenge the next round of models, share it with us in #feedback on our Discord.
Generating images via the OpenRouter API and Chat
Once you’ve evaluated models via these benchmarks, try them on your own prompts through the image generation API or Chat to see how they perform on your own content.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み