ミストラル、文書検索に「ナビゲーション型ループ」を導入し精度向上
本文の状態
日本語全文を表示中
詳細モードで約16分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
Mistral は単発のドキュメント検索から、複雑な文書内をナビゲートして情報を検証する「Agentic Search」へ移行し、FinanceBench で正答率を約3倍に向上させる成果を発表した。
AI深層分析を開く2026年8月22日 22:00
AI深層分析
キーポイント
検索手法の転換と機能強化
Mistral は従来のワンショット検索から、情報を発見・検証するマルチステップのループ構造を持つ「Agentic Search」へ移行し、複雑な文書内でのナビゲーションを可能にした。
ベンチマークにおける精度向上
FinanceBench において財務書類の正答率が 26.7% から 86%(約3倍)に、OfficeQA Pro では表を含む多文書質問で +45.6 ポイント改善する成果を報告している。
既存インフラとの統合とツール
Mistral Search Toolkit や Studio、Vibe に組み込まれており、search、open、navigate、read、grep の 5 つのツールを用いて既存の検索インデックスを拡張して利用する。
データセキュリティとコスト削減
クラウドまたはオンプレミスでのアイソレーション境界を越えずに機密領域データを扱えるほか、標的型ナビゲーションによりレイテンシとトークン使用量を削減する。
OfficeQA Pro ベンチマークでの大幅な性能向上
表形式が多く複数の文書にまたわる質問に対して、従来の手法から +45.6 ポイントの精度向上を達成した。
重要な引用
Mistral Agentic Search delivers more accurate search results while reducing turns, token use, and latency against FinanceBench and OfficeQA Pro benchmarks.
Agentic Search is the retrieval layer that enables AI systems to navigate, read, and verify information inside even the most complex documents.
Agentic Search delivers to 3x correctness on financial filings, from 26.7% to 86%, based on FinanceBench.
The limitation is more pronounced on dense, complex data and documents.
編集コメントを表示
編集コメント
単発の検索結果に依存する従来の手法から、AI が自ら文書内を探索・検証する能動的なアプローチへの転換は、実務レベルでの AI 活用における信頼性向上に寄与する。特に財務や法務といった高精度が求められる領域での適用可能性が高く、開発者が直面する課題解決の新たな選択肢を提供している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
思考
要約
Mistral Agentic Search は、FinanceBench および OfficeQA Pro のベンチマークにおいて、検索結果の精度を向上させつつ、必要なターン数やトークン使用量、レイテンシを削減します。Agentic Search は、AI システムが最も複雑な文書内でも情報を探索し、読み込み、検証することを可能にする検索層です。Mistral Search Toolkit やライブラリーを通じて利用可能です。

Mistral Agentic Search は、モデルが組織内で最も複雑なデータや文書を検索・探索できるようにすることで、企業が AI システムからより良い結果を得られるよう支援します。Agentic Search は、どこに保存されているデータソース間を横断して情報を発見し、確認し、検証するための多段階の検索ループを導入しました。この機能は Mistral Search Toolkit を通じて利用でき、Studio と Vibe の両方に組み込まれた ライブラリー でも使用可能です。
機密性の高いドメイン固有データへの対応。Mistral のポータブルでオープンなツールにより、クラウド上やオンプレミス環境における分離境界を越えることなく、データから価値を引き出すことが可能になります。
検索結果の改善。モデルは、取得されたチャンクを超えて、長大で密度の高い文書内や複数のソース間を探索・移動できるようになります。
既存のインデックスへのアクセス。Agentic Search は、search、open、navigate、read、grep の 5 つのツールを活用し、既存の検索インデックスを基盤に構築されています。
精度の向上。Agentic Search は、FinanceBench ベンチマークに基づく財務報告書において正答率を 26.7%から 86%へと引き上げ、3 倍の正確性を達成します。また、表が多く含まれる文書を対象とした OfficeQA Pro ベンチマークにおける複数ドキュメントへの質問では、6.3%から 51.9%へ 45.6 ポイントもの大幅な向上を記録しています。
遅延の低減とトークン使用量の削減
標的型ナビゲーションにより、エージェント検索は p90 レイテンシを最大 39.6% 短縮できます。重複する検索回数を減らすことで、トークン消費量を最大 3 分の 1 に抑えることが可能です。
データが競争優位性を生む
競争優位性は、長年にわたる実運用で築かれたものです。それは「あなたのデータ」「あなたのプロセス」、そして「あなたのドメイン知識」に支えられています。独自知識は成功の鍵であると同時に機密情報でもあり、隔離境界や分割デプロイメント、セルフホスト型プラットフォームといった壁の向こう側に存在します。
この独自知識は、財務報告書、法務契約書、社内リソース、政府文書などに蓄積されています。これらは長く密度の高いドキュメントであり、従来の検索手法では効果的にナビゲートすることができません。
継続的に学習し改善していくエージェントは、競争優位性を積み上げる手助けをしますが、セキュリティ上の理由からこれらのエージェントが機密データや独自知識にアクセスできないケースが多々あります。AI から真のインパクトを引き出すには、最先端の推論能力と、最も機微な資料へ安全にアクセスできる検索ツールを組み合わせたアプローチが必要です。
従来の RAG の限界
従来のワンショット型 RAG(Retrieval-Augmented Generation)は、固定されたテキストチャンク群を検索し、モデルが単一のパスで回答を生成する仕組みです。これが機能するのは、答えが上位結果のいずれかに明確に含まれている場合に限られます。しかし、長いレポート内の情報をたどる必要がある場合や、参照先を追跡する場合、複数の文書を比較検討する場合、あるいは根拠となる証拠を検証しなければならない場合には、この手法は力を発揮できません。
この限界は、密度が高く複雑なデータやドキュメントにおいてより顕著に現れます。質問への回答に必要な情報が複数の文書に散在していたり、特定の表や脚注、条項の中に埋もれていたりするケースがあるからです。ワンショット RAG に基づく検索が、最先端 AI の真価を十分に活用できず、信頼性の高い回答を提供できないのには、主に以下の3つの理由があります:
推論を伴わない検索の限界:モデルは初期段階で選択されたチャンクのみから回答を生成する必要があります。情報が不完全だったり関連性が低かったりする場合でも、モデルは「別の文書が必要だ」「別のセクションを確認すべきだ」「より多くの文脈が必要だ」と判断して検索を修正することができません。この制約が、モデルの推論能力による影響を著しく制限しています。
チャンク単位の限界:重要なデータは複雑なマルチモーダルドキュメントに格納されていることが多々あります。「第3四半期の企業の有効税率はいくらか?」と問われた際、インデックスは正しい文書を見つけられるかもしれませんが、その文書を「開く」「表までナビゲートする」「周囲の文脈を読み取る」「回答を検証する」といった操作を行うことはできません。
反復処理の欠如:多くの質問には、正解にたどり着くために複数の検索パスが必要です。モデルは検索条件を精緻化したり、有望な文書を精査したり、参照先を追跡したり、複数のソースを比較したり、既に見た情報を追跡したり、最初の結果が不十分な場合は新たな経路を試したりする必要があります。ワンショット RAG(Retrieval-Augmented Generation)では、これらの次のステップを実行する方法がありません。
エージェント型検索の活用
1953 年の各暦月の報告値のみを厳密に使用した場合、米国国防および関連活動への支出合計額(名目価格ベースで百万ドル単位)はいくらになりますか?
経路 3:ツール呼び出し(検索×2 → 読取)
search("national defense expenditures monthly 1953") → 月別報文書(年次途中まで)
search("…1953 November December 1954 to date") → treasury_bulletin_1954_02.pdf p.15(表 3、1953 年の全 12 ヶ月分)を表面化
「read(treasury_bulletin_1954_02.pdf, p.15)」を実行すると、1953年の月次値をまとめた完全な表 3 が取得されます。
表 3(単位:百万ドル)
| 1月 | 2月 | 3月 | 4月 | 5月 | 6月 | 7月 | 8月 | 9月 | 10月 | 11月 | 12月 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 3,632 | 3,501 | 3,789 | 3,891 | 3,746 | 4,056 | 3,890 | 3,519 | 3,787 | 3,647 | 3,540 | 3,465 |
合計 44,463。
エージェント型検索の仕組み
Mistral Search Toolkit は、クラウド上またはオンプレミス環境にある重要な複雑なデータをインデックス化するためのオープンモジュールを提供します。エージェント型検索はこのインデックスを基盤とし、モデルにファイルシステム操作に似た 5 つのツールを与えます。
「search」は既存のインデックスを活用して、コーパス全体から関連文書を検索します。
「open」は特定の文書を開きます。
「navigate」はその中にあるページ、セクション、または領域へ移動します。
readは、指定された場所のコンテンツを取得します。grepは、開いているドキュメント内でパターンを検索します。
モデルは初期の上位 k 件結果のみを回答に利用するのではなく、検索で見つけた情報を精査し、検索条件を調整し、関連文書を開いて特定のセクションへ移動し、回答を作成する前にソース資料を読み込むことができます。インデックスが有望な情報源を特定し、アジェンティック・サーチ(Agentic Search)がその中やそれら全体で何を精査すべきかを決定します。
One-shot RAG

エージェント型検索

これらのツールは、ファインチューニングやモデル固有のトレーニングを必要としません。モデルが推論やツールの利用においてより熟練するにつれて、情報検索の精度も向上し、インフラストラクチャの変更は不要です。これは重要な特性であり、検索品質はチャンキング戦略によって制限されるのではなく、モデルの能力に応じてスケールします。
エージェント型検索が適しているケース
- 長文ドキュメント。回答が特定のページや、表、条項、図、脚注に含まれる可能性のある提出書類、契約書、マニュアル、技術仕様書、報告書など。
- 複数ソースにまたがる質問。回答に至るまでに複数のドキュメントから証拠を見つけ、比較し、整合性を取る必要がある調査。
- 検証が必須の回答。財務数値、法的条項、規制参照、運用データなど、回答を安定した特定のドキュメント位置として参照できるケース。
- 表や構造化されたドキュメント。行、列、ページ上の位置、あるいは周囲の文脈に意味が依存する財務諸表、政府記録、スキャンされた PDF など。単なる記述テキストだけでは不十分な場合です。
インデックス型検索(インデックス・リトリーバル)が適しているケース
- 直接参照。回答が最初に取得されるチャンクのいずれかに含まれる可能性が高い、短く整理されたドキュメント。
- 大量の検索。推論やナビゲーションを行わずに、関連する記述を返す必要があるキーワード検索またはセマンティック検索。
- 単純で予測可能な質問。回答のソースと場所が事前にわかっているユースケース。追加の取得ステップを行っても結果は改善しない場合です。
こうした検索タスクには、ワンショット型 RAG(Retrieval-Augmented Generation)で十分なケースが多いです。ただし、モデルが初期結果を超えてソース資料を調査する必要がある質問に対しては、エージェント型検索を追加する必要があります。どちらの場合でも、適切に設定されたインデックスが基盤として機能し続けます。
より関連性の高い結果を、より高速に
エージェント型検索のベンチマークでは、業界標準の評価指標を用い、Mistral Search Toolkit スタック(デフォルトのチャンク分割、デフォルトのランキング、チューニングなし)をそのまま使用しました。これらの数値は「下限」を示すものであり、「上限」ではありません。つまり、ユースケースに合わせたチューニングを行うことで、さらに結果の質を向上させることができます。
このベンチマークでは、Mistral Search Toolkit を用いて 2 つのモデルを検証しました。小規模モデルであるMistral Medium 3.5(MM 3.5)と、大規模モデルであるZ.ai GLM-5.2(GLM-5.2)です。
ベンチマーク結果は明確で、エージェント型ループが実質的な品質向上をもたらす一方、ナビゲーション機能によって精度が高まり、無駄なトークン数やターン数、レイテンシが削減されました。第一者モデルと第三者モデルの両方で同様のパフォーマンスパターンが観測されたことから、エージェント型検索は特定のモデルに依存しない(モデル非依存)ことが示唆され、新しいモデルが登場すれば検索品質も向上すると考えられます。
FinanceBench: 368 件の SEC 提出書類、150 の質問
FinanceBench(Islam et al., 2023)は、368 件の SEC 提出書類(10-K / 10-Q / 8-K)を対象に財務関連の質問応答を評価するベンチマークです。各書類は平均約 147 ページで、合計約 53,900 ページに及びます。表が多く含まれる長文の財務ドキュメントが対象となります。回答の評価は、人間によるラベルと較正された LLM ジャッジによって行われます。
financebench_-one-shot-rag-vs-agentic-search-glm-1.svg)
検索結果:
検索に特化したエージェント型ループが、品質向上の最大の要因となります。ワンショット RAG から検索専用ループへ移行することで、MM 3.5 では精度が +47.3 ポイント、GLM-5.2 では +52.6 ポイント向上し、両モデルで約 3 倍の改善が見られました。モデルが反復的に検索できるため、最初の結果が弱くても挽回でき、クエリを洗練させ、インデックスを実用的なツールとして活用できます。
ナビゲーション機能も精度向上に寄与します。オープン、ナビゲート、リーディング、グリップの機能を追加すると、さらに精度が向上します(MM 3.5 で +8.7 ポイント、GLM-5.2 で +6.7 ポイント)。これは、複雑な文書において、広範囲な検索を繰り返すよりも、ターゲットを絞った掘り下げ検索の方が優れていることを意味しています。
より優れた検索ツールにより、トークン数とパフォーマンス効率も向上します。ナビゲーション機能を備えた完全なループは、検索専用ループと比較して、より少ないトークンで多くの質問に正解するようになります(MM 3.5 でトークン使用量が -23.9%、GLM-5.2 で -33.7%)。検索ツールは追加のオーバーヘッドではなく、無駄な検索リトライを精密なナビゲーションに置き換えるものです。
重要な部分でレイテンシが短縮されます。FinanceBench 全体を通じて、ナビゲーション検索ツールの導入によりレイテンシが改善し、p90 が 255 秒から 154 秒へ、平均レイテンシも 108 秒から 71 秒へと低下しました。一般的に、検索専用ループは広範囲な検索を繰り返す傾向がありますが、ナビゲーション機能によりモデルが証拠をより迅速に特定できるようになります。
OfficeQA Pro:696 件の財務省ブリーフィング、133 の質問
OfficeQA Pro は、過去の米国財務省ブリーフィングを対象とした検証可能な数値ベンチマークです。スキャンされた表の多い政府財務関連 PDF を含む、696 ドキュメント、約 89,000 ページにわたる大規模なコーパスを分析対象としています。今回は、133 の質問からなる「Pro」サブセットに対する最初の結果を発表します。
.svg)
調査で明らかになったのは次の通りです。
エージェント型検索と、その中核となる「エージェントループ+ナビゲーション」の組み合わせが、より困難で検証可能なベンチマークにおいて成功を収めました。OfficeQA Pro は数値回答を必要とし、スキャンされた PDF や複雑な表の参照を含むテストです。この厳しい環境でも、フルエージェント型ループを採用することで、ワンショット RAG(Retrieval-Augmented Generation)と比較して精度が劇的に向上しました。GLM-5.2 では 51.9% に達し、従来比で +45.6 ポイント の改善が見られました。同様に MM 3.5 でも +27.1 ポイント の向上を果たしています。
ナビゲーション機能の導入は、精度を高めつつ無駄なリソース消費を抑える効果があります。エージェントループとナビゲーションをフル活用することで、MM 3.5 では最大 35.6%(+7.5 ポイント)、GLM-5.2 では +8.3 ポイント(19.0%)の精度向上を実現しました。同時に、トークン消費量も削減され、拒否されたリクエスト数は MM 3.5 で最大 7.0%、GLM-5.2 で 2.3% 減少しています。
ベンチマークが困難になるほど、検索ループの重要性は増大します。OfficeQA Pro はスキャン文書内の数値回答と表データに焦点を当てたテストです。ワンショット RAG では対応しきれないケースが多い中、エージェント型ループではモデルが反復的に検索を行い、証拠を検証することで、精度を大幅に向上させることが可能になります。
ツールスタックの選定は、ドキュメントインテリジェンスと検索性能に決定的な影響を与えます。Kimi の研究 によると、Claude Code ハーネスを使用した場合、GLM-5.2 は OfficeQA Pro で 41.4% のスコアでした。しかし、Mistral のハーネスに切り替えると同じモデルで 51.9% を記録し、+10.5 ポイント の差がつきました。
はじめに
エージェント型検索の詳細は ドキュメント でご確認ください。クラウド環境およびオンプレミス環境のどちらでも、以下の方法で利用を開始できます:
Mistral Search Toolkit を使えば、独自のエージェントやワークフロー、顧客向けデプロイメントにアジェンティック検索を簡単に統合できます。
また、Libraries を利用すれば、検索システムをゼロから構築する必要なく、Studio や Vibe でアジェンティック検索をすぐに活用可能です。
Search Toolkit のテストには、Search Starter App が最も手軽です。デフォルト設定で独自のデータセットにローカルインデックスを作成するため、検索の専門家である必要なくアジェンティック検索を試せます。使い方をカスタマイズする際は、以下の手順に従ってください。
- インgestion の設定: データやファイルタイプに合わせてパーサー、チャンキング戦略、埋め込みモデル、抽出器を選択します。
- インデックスとランキングの調整: Vespa スキーマ、インデックス動作、関連性プロファイルを管理します。
- 検索機能の拡張: 検索パイプラインにクエリ再書き換え、再ランク付け、ハイブリッド検索を追加できます。
原文を表示
Thinking
Summary
Mistral Agentic Search delivers more accurate search results while reducing turns, token use, and latency against FinanceBench and OfficeQA Pro benchmarks. Agentic Search is the retrieval layer that enables AI systems to navigate, read, and verify information inside even the most complex documents. Available through Mistral Search Toolkit and Libraries.

Mistral Agentic Search helps enterprises get better results from their AI systems by letting models search and navigate their organization’s most complex data and documents. Agentic Search introduces a multi-step retrieval loop for finding, inspecting, and verifying information across data sources, wherever it is stored. Agentic Search is available through Mistral Search Toolkit, built into Libraries in both Studio and Vibe, and gives you:
- Support for sensitive domain-specific data. Mistral’s portable and open tooling helps you unlock value from your data without crossing your isolation boundaries in the cloud or on-premises.
- Improved search results. Your models can search and navigate your data beyond retrieved chunks–inside long, dense documents or across multiple sources.
- Access to existing indexes. Agentic Search builds on your existing search index using five tools: search, open, navigate, read, and grep.
- Higher accuracy. Agentic Search delivers to 3x correctness on financial filings, from 26.7% to 86%, based on FinanceBench. On table-heavy, multi-doc questions of the OfficeQA Pro benchmark, we measure a +45.6 point gain (6.3% to 51.9%).
- Lower latency and token use. Targeted navigation enables Agentic Search to reduce p90 latency up to 39.6%. Fewer repeated searches reduce token consumption by up to one-third.
Data creates competitive advantage
Competitive edge is built upon years of real-world operations–your data, your processes, and your domain expertise. Proprietary knowledge is both critical to your success and highly confidential, meaning it lives behind isolation boundaries, segmented deployments, and self-hosted platforms. It accumulates in financial filings, legal contracts, internal resources, and government records–long, dense documents that traditional search methods can’t navigate effectively.
Agents that learn and improve continuously can help you compound your competitive advantage, but these agents are often separated from confidential data and proprietary knowledge for security reasons. Getting real impact from AI means pairing frontier reasoning with retrieval tools that can safely reach your most sensitive material.
Traditional RAG falls short
Traditional, one-shot RAG retrieves a fixed set of text chunks and asks a model to answer in a single pass. This works when the answer appears in one of the top results, but falters when the model must navigate a long report, follow references, compare multiple documents, or verify the underlying evidence.
The limitation is more pronounced on dense, complex data and documents. The information needed to answer a question may be spread across documents or buried in a particular table, footnote, or clause. One-shot RAG-based search fails to use the full power of frontier AI and to provide reliable answers for three reasons:
- Retrieval without reasoning: The model must answer from the chunks selected during the initial retrieval, even when they are incomplete or not relevant. It cannot decide that it needs a different document, another section, or more context before responding, which limits the impact of the model’s reasoning.
- Chunk-level limit: Critical data is often held in complex multi-modal documents. When asked, “What was the company’s effective tax rate in Q3?” an index may find the correct document but cannot open it, navigate to the table, read the surrounding context, or verify the answer.
- No iteration: Many questions need more than one retrieval pass to get the correct answer. The model may need to refine its search, inspect a promising document, follow a reference, compare multiple sources, keep track of what it has seen, and try a new route when the first results are insufficient. One-shot RAG provides no way to take these next steps.
With Agentic Search
Using specifically only the reported values for all individual calendar months in 1953, what is the total sum of these values of expenditures for U.S. national defense and associated activities (in millions of nominal dollars)?
Trajectory 3 tool_calls (2× search → read)
search("national defense expenditures monthly 1953") → per-month bulletins (partial year)
search("…1953 November December 1954 to date") → surfaces treasury_bulletin_1954_02.pdfp.15 (Table 3, all 12 months of 1953)
read(treasury_bulletin_1954_02.pdf, p.15) → pulls the complete Table 3
Monthly Values for 1953
Table 3, in $millions
| Jan | Feb | Mar | Apr | May | Jun | Jul | Aug | Sep | Oct | Nov | Dec |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 3,632 | 3,501 | 3,789 | 3,891 | 3,746 | 4,056 | 3,890 | 3,519 | 3,787 | 3,647 | 3,540 | 3,465 |
Sum = 44,463.
How Agentic Search works
Mistral Search Toolkit provides open modules for ingesting, embedding, and indexing critical and complex data in the cloud or on-premises. Agentic Search builds on this index by giving the model five tools that resemble familiar file-system operations:
- search finds relevant documents across the corpus using the existing index.
- open opens a specific document.
- navigate moves to a page, section, or region within it.
- read retrieves the content at that location.
- grep finds a pattern within an open document.
Rather than answering only from the initial top-*k* results, the model can inspect what it finds, refine its search, open relevant documents, navigate to specific sections, and read the source material before answering. The index identifies likely sources; Agentic Search determines what to inspect within and across them.
One-shot RAG

Agentic Search

These tools do not require fine-tuning or model-specific training. As models get better at reasoning and tool use, retrievals get better without infrastructure changes. This is a key property: retrieval quality scales with model capability instead of being capped by your chunking strategy.
Use Agentic Search for
- Long documents. Filings, contracts, manuals, technical specifications, and reports where the answer may appear on a particular page or in a specific table, clause, figure, or footnote.
- Questions across multiple sources. Research that requires the model to find, compare, or reconcile evidence from several documents before reaching an answer.
- Answers that must be verified. Financial figures, legal clauses, regulatory references, and operational data, where the response can be referenced in a stable and specific document location.
- Tables and structured documents. Financial statements, government records, and scanned PDFs where meaning depends on rows, columns, page position, or surrounding context–not narrative text alone.
Indexed retrieval is the right starting point for
- Direct lookups. Short, clean documents where the answer is likely to appear in one of the first retrieved chunks.
- High-volume search. Keyword or semantic lookups that need to return relevant passages without reasoning over or navigating through them.
- Simple, predictable questions. Use cases where the likely source and location of the answer are known in advance and additional retrieval steps are unlikely to improve the result.
One-shot RAG is often sufficient for these searches. Add Agentic Search when questions require the model to move beyond the initial results and investigate the source material. A well-configured index remains the right foundation in both cases.
More relevant results, faster
We benchmarked Agentic Search on two industry-standard evaluations, using the out-of-the-box Mistral Search Toolkit stack: default chunking, default ranking, no tuning. These results are floors, not ceilings, meaning you can further improve result quality with use-case-specific tuning.
With these benchmarks, we tested two models using the Mistral Search Toolkit: Mistral Medium 3.5 (MM 3.5) and Z.ai GLM-5.2 (GLM-5.2), showcasing performance of a smaller model (MM 3.5) and a larger model (GLM-5.2).
Benchmark results are consistent: the agentic loop delivers substantive quality improvements and navigation tools increase accuracy while reducing wasted tokens, turns, and latency. We observe the same performance patterns across first- and third-party models, which indicates that Agentic Search is model-agnostic, and that search quality should improve with new models.
FinanceBench: 368 SEC filings, 150 questions
FinanceBench (Islam et al., 2023) tests financial question-answering over 368 SEC filings (10-K / 10-Q / 8-K), averaging ~147 pages each, ~53,900 pages total: long, table-heavy financial documents. Answers scored by an LLM judge calibrated against human labels.
financebench_-one-shot-rag-vs-agentic-search-glm-1.svg)
We found:
- The search-only Agentic loop is the biggest quality lever. Moving from one-shot RAG to a search-only loop lifts accuracy by +47.3pp for MM 3.5 and +52.6pp for GLM-5.2–a ~3x improvement for both models. Because models can search iteratively, they can recover from weak first results, refine queries, and use the index as an active tool.
- Navigation adds accuracy. Adding open, navigate, read, and grep lifts accuracy again (+8.7pp for MM 3.5, +6.7pp for GLM-5.2). This means a targeted drill-in search beats repeated broad search in complex documents.
- Token and performance efficiency improve with better retrieval tools. The full loop with Navigation answers more questions correctly while using fewer tokens than the search-only loop (MM 3.5: -23.9% token usage, GLM-5.2: -33.7%). The retrieval tools are not additional overhead–they replace wasted search retries with precise navigation.
- Latency goes down where it matters. Across FinanceBench, adding navigation retrieval tools improves latency: p90 drops 255s → 154s and mean latency drops 108s → 71s. In general, we see the search-only loops conduct repeated broad searches, while navigation helps the model identify evidence more quickly.
OfficeQA Pro: 696 Treasury Bulletins, 133 questions
OfficeQA Pro is a verifiable numeric benchmark over historical U.S. Treasury Bulletins: scanned, table-heavy government-finance PDFs across a 696-document, ~89,000-page corpus. We report the first pass for the 133-question "pro" subset.
.svg)
We found:
- Agentic Search and the Agentic loop + Navigation are successful against a harder, verifiable benchmark. OfficeQA Pro has numeric answers, scanned PDFs, and deep table lookups. Even here, the full agentic loop lifts accuracy materially from one-shot RAG, reaching 51.9% for GLM-5.2 (+45.6pp) and increasing +27.1pp for MM 3.5.
- Navigation improves quality while cutting waste. Using the full loop (Agentic loop + Navigation) improves accuracy by up to 35.6% (+7.5pp, MM 3.5; +8.3pp, 19.0% GLM-5.2), while reducing token consumption. Turns declined by up to 7.0% (MM 3.5, 2.3% GLM-5.2).
- The harder the benchmark, the more important the retrieval loop becomes. OfficeQA Pro is built around numeric answers in scanned, table-heavy documents. One-shot RAG barely gets started, while the agentic loop allows the model to search iteratively, inspect evidence, and deliver substantial accuracy improvements.
- The tooling stack drives substantial impact on document intelligence and search performance. Per Kimi research, GLM-5.2 scores 41.4% on OfficeQA Pro with the Claude Code harness, compared with 51.9% on the Mistral harness–+10.5pp on the same underlying model.
Getting started
Learn more about Agentic Search in the documentation. You can get started across cloud and on-premises deployments using either:
- Mistral Search Toolkit. Integrate Agentic Search into your own agents, workflows, and customer deployments.
- Libraries. Use Agentic Search out-of-the-box in Studio and Vibe, without building the retrieval system yourself.
The fastest way to test Search Toolkit is with the Search Starter App. It creates a local index for your own corpus using a default configuration, so you can try Agentic Search without needing to be a search expert. When you’re ready to configure your use case, you can:
- Set up ingestion. Select parsers, chunking strategies, embedding models, and extractors for your data and file types.
- Tune indexing and ranking. Manage Vespa schemas, indexing behavior, and relevance profiles.
- Extend retrieval. Add query rewriting, reranking, or hybrid retrieval to the search pipeline.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み