多モーダルモデルの中間視覚状態利用を検証する評価フレームワーク「See2Think」を公開
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Hugging Face Daily Papers
See2Think は、マルチモーダルモデルが推論過程で中間視覚状態を本当に利用しているかを診断する新たな評価フレームワークであり、既存ベンチマークの限界を克服してモデルの視覚推論能力とレンダリングの忠実性を厳密に評価する。
AI深層分析を開く2026年8月4日 15:22
AI深層分析
キーポイント
See2ThinkBench の導入
12 のタスクカテゴリにわたる 1,200 のオープンエンドな視覚依存問題を収録した新しいベンチマークセットを公開する。
VAoT によるプロセス分析
テキスト思考、視覚行動、レンダリング状態、推論の連鎖を記録する Visual Action-of-Thought (VAoT) を導入し、中間状態の生成と利用を詳細に診断する。
視覚推論の依存性とボトルネック
モデルや環境によって視覚推論の結果が強く依存しており、忠実なレンダリングが最大のボトルネックであることが示された。
フィードバック介入による検証
タスク関連のノイズのあるフィードバックを与えた際、モデルは視覚状態に行動を依存し、精度が 10 ポイント以上低下することが確認された。
重要な引用
Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used.
Faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains.
編集コメントを表示
編集コメント
既存のベンチマークが抱える「最終回答偏重」の問題を解決する画期的なアプローチであり、モデルのブラックボックス化された推論プロセスを可視化する重要な一歩となる。今後のマルチモーダルモデル開発において、中間状態の品質評価が標準的な工程として組み込まれる可能性が高い。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
マルチモーダル大規模言語モデルは、推論過程でスケッチや注釈、ツール、中間画像を頻繁に活用していますが、それらが本当にこれらの視覚状態に依存しているのかはまだ不明です。既存のベンチマークは、カバー範囲が狭いタスクコレクションや部分的にテキストだけで解けるサンプルによる制限、そして最終回答のみを重視し、中間視覚状態がどのように生成・描画され・利用されるかを診断しない評価手法によって限界があります。
そこで私たちは、See2ThinkBench と Visual Action-of-Thought (VAoT) からなる統合的な評価フレームワーク「See2Think」を導入しました。See2ThinkBench には、2D 構造化、3D シーン、現実世界の推論にまたがる 12 のタスクカテゴリで構成される 1,200 のオープンエンドかつ視覚依存型の問題が含まれています。VAoT は、4 つの制御された推論設定下でのテキスト思考、視覚的アクション、描画状態、およびその後の推論を記録します。
代表的なプロプライエタリおよびオープンソースのマルチモーダルモデルを評価した結果、視覚的推論はモデルと環境に強く依存しており、どのタスクでも一貫して優位な設定は見つかりませんでした。プロセス分析では、モデルが関連する視覚操作を選択する傾向がある一方、忠実な描画が最も明確なボトルネックであることが示されました。また、高いフィードバックの受容率が必ずしも精度向上につながるとは限りません。
タスクに関連するノイズのあるフィードバック下では、モデルは視覚状態に行動面で依存することが確認され、制御された介入により精度が 10 パーセントポイント以上低下しました。
原文を表示
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み