並行世界における検索エージェントの評価
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
ArXiv cs.AI
研究者らが、LLMに統合された検索エージェントの評価における課題(高品質なベンチマーク構築の困難さと静的ベンチマークの陳腐化)を指摘し、新たな評価手法の必要性を論じている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
arXiv:2603.04751v1 Announce Type: new
アブストラクト: Web検索ツールの統合により、LLMがオープンワールド、リアルタイム、ロングテールの問題に対処する能力は大幅に拡張されました。しかし、これらの検索エージェントの評価には、多大な困難が伴います。第一に、高品質な深層検索ベンチマークの構築は非常に高コストであり、未検証の合成データは信頼性の低い情報源に起因する問題を抱えがちです。第二に、静的ベンチマークは動的陳腐化の問題に直面します。インターネット情報が変化するにつれ、深い調査を必要とする複雑なクエリは、そのトピックの認知度上昇により単純な検索タスクへと変質し、グランドトゥルースも時間の経過とともに陳腐化します。第三に、帰属の曖昧さが評価を妨げます。エージェントの性能は、実際の検索・推論能力というより、そのパラメトリックメモリに大きく依存してしまうためです。最後に、特定の商用検索エンジンへの依存は、再現性を損なう変動要因となります。これらの課題に対処するため、本研究ではパラレルワールドにおける検索エージェント評価のための新たなフレームワーク「Mind-ParaWorld」を提案します。具体的には、MPWは実世界のエンティティ名をサンプリングし、モデルの知識カットオフを超えた将来のシナリオと質問を生成します。続いて、ParaWorld Law Modelが、分割不可能なアトミックファクトの集合と、各質問に対する独自のグランドトゥルースを構築します。評価時には、エージェントは実世界の結果を検索する代わりに、これらの不変のアトミックファクトに基づいてSERPを動的に生成するParaWorld Engine Modelと対話します。我々は、19のドメインにまたがる1,608インスタンスから成るインタラクティブなベンチマーク「MPW-Bench」を公開します。3種類の評価設定による実験結果から、検索エージェントは完全な情報が与えられた場合の証拠統合には優れる一方で、その性能は、未知の検索環境における証拠収集とカバレッジの不足、さらに、信頼性の低い証拠十分性判断と「いつ停止するか」の決定というボトルネックによって制限されていることが明らかになりました。
原文を表示
arXiv:2603.04751v1 Announce Type: new
Abstract: Integrating web search tools has significantly extended the capability of LLMs to address open-world, real-time, and long-tail problems. However, evaluating these Search Agents presents formidable challenges. First, constructing high-quality deep search benchmarks is prohibitively expensive, while unverified synthetic data often suffers from unreliable sources. Second, static benchmarks face dynamic obsolescence: as internet information evolves, complex queries requiring deep research often degrade into simple retrieval tasks due to increased popularity, and ground truths become outdated due to temporal shifts. Third, attribution ambiguity confounds evaluation, as an agent's performance is often dominated by its parametric memory rather than its actual search and reasoning capabilities. Finally, reliance on specific commercial search engines introduces variability that hampers reproducibility. To address these issues, we propose a novel framework, Mind-ParaWorld, for evaluating Search Agents in a Parallel World. Specifically, MPW samples real-world entity names to synthesize future scenarios and questions situated beyond the model's knowledge cutoff. A ParaWorld Law Model then constructs a set of indivisible Atomic Facts and a unique ground-truth for each question. During evaluation, instead of retrieving real-world results, the agent interacts with a ParaWorld Engine Model that dynamically generates SERPs grounded in these inviolable Atomic Facts. We release MPW-Bench, an interactive benchmark spanning 19 domains with 1,608 instances. Experiments across three evaluation settings show that, while search agents are strong at evidence synthesis given complete information, their performance is limited not only by evidence collection and coverage in unfamiliar search environments, but also by unreliable evidence sufficiency judgment and when-to-stop decisions-bottlenecks.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み