インタラクティブ・ベンチマーク
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
ArXiv cs.AI
研究者らは、従来のベンチマークが飽和・主観性・汎化性の問題を抱えると指摘し、モデルの能動的情報獲得能力を評価する「インタラクティブ・ベンチマーク」を提案した。この枠組みは予算制約下での対話的推論能力を測定する。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
arXiv:2603.04737v1 発表タイプ: 新規
要約: 標準的なベンチマークは、飽和、主観性、および一般化性能の低さにより、信頼性が低下しつつあります。モデルの知能を適切に評価するには、モデルが能動的に情報を獲得する能力を測定することが重要であると我々は考えます。本論文では、予算制約下でのインタラクティブなプロセスにおいてモデルの推論能力を評価する、統一的な評価パラダイム「Interactive Benchmarks(インタラクティブ・ベンチマーク)」を提案します。この枠組みを、二つの設定で具体化します。一つは「Interactive Proofs(インタラクティブ証明)」であり、モデルが審判(判定者)と対話し、論理学や数学における客観的な真実や解答を推論します。もう一つは「Interactive Games(インタラクティブゲーム)」であり、モデルが長期的効用を最大化するために戦略的に推論します。実験結果から、インタラクティブ・ベンチマークはモデルの知能を堅牢かつ忠実に評価できること、また、インタラクティブなシナリオにおいては依然として大幅な改善の余地があることが示されました。プロジェクトページ: https://github.com/interactivebench/interactivebench
原文を表示
arXiv:2603.04737v1 Announce Type: new
Abstract: Standard benchmarks have become increasingly unreliable due to saturation, subjectivity, and poor generalization. We argue that evaluating model's ability to acquire information actively is important to assess model's intelligence. We propose Interactive Benchmarks, a unified evaluation paradigm that assesses model's reasoning ability in an interactive process under budget constraints. We instantiate this framework across two settings: Interactive Proofs, where models interact with a judge to deduce objective truths or answers in logic and mathematics; and Interactive Games, where models reason strategically to maximize long-horizon utilities. Our results show that interactive benchmarks provide a robust and faithful assessment of model intelligence, revealing that there is still substantial room to improve in interactive scenarios. Project page: https://github.com/interactivebench/interactivebench
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み