企業 AI データガバナンス、文脈層整備が回答精度に直結
本文の状態
日本語全文を表示中
詳細モードで約26分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
VentureBeat AI
ベンチャーベイトの調査によると、企業はAIエージェントに信頼できる文脈を提供する「統制されたセマンティック層」を構築しているにもかかわらず、むしろ失敗を検知する頻度が二倍になるという逆説的な結果が示された。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 17:11
AI深層分析
キーポイント
文脈欠如による誤答の常態化
調査対象企業の68%が過去半年に、モデル自体のエラーではなくビジネス文脈の不足や不整合によって「自信満々だが間違った回答」を生成したと報告しており、その37%は再発している。
セマンティック層の逆説的効果
統制されたセマンティック層を構築・運用する企業は、それを持たない企業の2倍以上の頻度で文脈エラーを検知しており、これは層が失敗を引き起こしているのではなく、失敗を可視化していることを示す。
アーキテクチャ設計の迷走
ハイブリッド検索とユースケース別マルチアーキテクチャが拮抗しており、業界全体として文脈解決のための統一されたアーキテクチャ基準が存在しない状態である。
プロバイダー依存への抵抗
ベクトルデータベースやネイティブ検索機能は広く利用されているものの、企業は文脈層の管理をクラウドプロバイダーに委ねることに消極的であり、独自構築志向が強い。
アーキテクチャの分散とプロバイダー依存の回避
ハイブリッド型やユースケース別選択など特定のアーキテクチャが多数派となっておらず、企業は単一プロバイダーへの依存を避けベスト・オブ・ブリードまたは明示的なミックスを選好している。
重要な引用
The central finding is that the context failure is no longer an incident; it is a condition.
Enterprises that have built or are building a layer report recurring context failures at 50%, against 21% for those without one.
The layer isn't causing the failures — it's catching them, which makes it the most useful finding in the wave.
Access control and permissions is now tied with ease of data ingestion as the top selection factor at 24% each, and response correctness is the primary success metric for 38% of enterprises.
編集コメントを表示
編集コメント
この調査結果は、企業がAI導入において「文脈の質」を過小評価していた現実を浮き彫りにしており、単なるモデル選定からデータ基盤のガバナンスへ視点を移す必要性を示唆している。セマンティック層がエラー数を減らすのではなく検知する機能を持つという発見は、実務における期待値管理と投資判断に重要な示唆を与える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
101 社の企業を対象とした調査では、AI エージェントに提供されるコンテキストが頻繁に、そして繰り返し失敗していることが明らかになりました。過去 6 ヶ月間で、自信満々ながら誤った回答を返したエージェントの原因が「不足している、あるいは矛盾するビジネスコンテキスト」にあると特定したのは、企業の 68% に上ります。しかも、その原因は一度きりではなく、「複数回発生している」というケースが最も多いのです。
さらに興味深いのは、この問題を報告している企業の特徴です。自社固有の定義や関係性を管理する「ガバナンスされたセマンティックレイヤー(semantic layer)」を構築・運用している企業では、そのレイヤーを持たない企業に比べて、同様の失敗が 2 倍以上も頻発していると回答しています。
現在、このコンテキストの問題を解決するために整備が進められているインフラは、むしろ「どれほど多くの企業が不適切なコンテキストに直面しているか」を浮き彫りにする結果となっています。また、根本的な解決策となるアーキテクチャについては合意が得られておらず、「ハイブリッド型検索」と「完全な多様性(pluralism)」のどちらを採用すべきかで意見が割れ、回答者の差はわずか 1 人という接戦状態です。
今回の VentureBeat Pulse Research は、企業の RAG(Retrieval-Augmented Generation)とコンテキストレイヤーに焦点を当てます。AI エージェントにどのようなビジネスコンテキストを提供しているのか、企業が運用する検索システムの種類、その導入・測定方法、アーキテクチャの将来像、そして何よりも示唆に富む点として「コンテキストがすでに失敗している頻度」について掘り下げます。
中核的な発見は、コンテキストの失敗がもはや単なるインシデントではなく、組織に定着した「状態」になっている点です。過去 6 ヶ月間に AI エージェントが自信満々ながら誤った回答を生成し、その原因がモデル自体の欠陥ではなく、不足しているか矛盾するビジネスコンテキストにあると特定した企業は 68% に達します。
注目すべきは失敗の総数よりも、その「形状」です。37% の企業が同様の失敗を繰り返していると報告しており、一度きりだったのは 32% です。回答可能な企業の範囲に限れば、社内のデータ上でエージェントを実行する際にもっとも頻繁に遭遇するのは、モデルとは無関係の理由で繰り返し誤った結果を出すという経験です。
業界が解決策として採用したのが、ガバナンスされたセマンティック層(またはコンテキスト層)です。これは AI エージェントと BI ツールがデータの共通理解を持つための基盤であり、すでに大規模に構築され始めています。32% の企業が本番環境で稼働させており、31% がパイロット運用や構築中、さらに 20% が評価段階にあります。
しかし、詳細な分析結果は不愉快な事実を浮き彫りにしました。コンテキスト層を構築済み、あるいは構築中の企業では、同様の失敗が再発する割合が 50% に達しています。一方、層を持たない企業ではその割合は 21% です。この層自体が失敗を引き起こしているわけではありません。むしろ、この層があるからこそ失敗を「検知・特定」できているのです。これが今回の調査で得られた最も有用な知見です。セマンティック層こそが、コンテキストの欠陥を追跡可能にする鍵となります。層を持たない組織は、失敗そのものが減っているわけではなく、単に失敗の原因を特定できていないだけなのです。
基盤となる技術スタックは、想定とは異なる形で不安定化しています。検索機能(Retrieval)が主要なコンテキストソースとして引き続き 31% を占めていますが、プロバイダーネイティブの検索機能—OpenAI のファイル検索(46%)や Google Vertex AI Search(41%)—は、専用ベクトルデータベースを大きく引き離して首位を維持しています。しかし、期待されていたアーキテクチャに明確な多数派はありません。ハイブリッド型検索が 30%、「ユースケースに応じて複数のアーキテクチャを選択する」アプローチが 29% と僅差で並び、明確な勝者が出ていません。また、企業はコンテキスト層をプロバイダーに委ねることに依然として消極的です。単一のモデルプロバイダーのネイティブなコンテキストスタックへ統合すると回答したのはわずか 12% に過ぎず、37% が「各分野で最も優れた製品を組み合わせたベスト・オブ・ブリード」を選択し、残りの 37% は明示的なミックス構成を計画しています。
購買基準において、商業的な失敗の兆候が明確になり始めています。アクセス制御と権限管理は、データ取り込みの容易さと並び、24% ずつで最優先選択要因となりました。また、回答の正確性が企業の主要な成功指標であるという回答も 38% に達しています。企業はコンテキストを移動させる機能ではなく、コンテキストを統制する機能を備えた検索ソリューションを購入し始めています。
調査手法
VentureBeat は、継続的な「Pulse Research」シリーズの一環として本調査を実施しました。今回の調査は、エンタープライズ向けの RAG(Retrieval-Augmented Generation)インフラと、AI エージェントに情報を供給するコンテキスト層——検索システム、セマンティックレイヤー、およびコンテキストソースに焦点を当てています。
回答者は従業員数 100 人以上の組織から集められ、有効回答数は 101 社です。データはすべて 2026 年 7 月に実施された単一の調査ウェーブに基づくため、このレポートは横断的な分析であり、月ごとの推移を推測するものではありません。また、複数の選択肢に回答できる設問については、結果を「全回答数に対する割合」ではなく「回答者数に対する割合」として報告しています。そのため、項目の合計が 100% を超える場合があります。
組織規模別に見ると、サンプルは中堅企業に集中しています。具体的には、従業員数 101〜250 人が 34%、1,001〜5,000 人が 25%、251〜1,000 人が 25% を占めています。これらに続くのは、10,001 人以上が 12%、5,001〜10,000 人が 5% です。
役職別では、マネージャー層が 39%、個人貢献者(Individual Contributors)が 29%、VP やディレクターが 22%、C スート(経営陣)が 9% を占めています。購買権限については信頼性が高く、最終決定権を持つ層が 38%、推奨や影響力を行使する層が 43% に達しています。
業界別では、テクノロジー・ソフトウェア分野が 31% で最も大きく、次いでヘルスケア・ライフサイエンス(14%)、小売・EC(10%)、製造業(9%)となっています。
コンテキスト失敗のベースに関する注記:回答者 101 名のうち、10 名(5%)はエンタープライズデータ上でエージェントを実行していないか、そのレベルで根本原因を追跡していません。失敗に関する主要な統計は全 101 名を対象に報告されていますが、「Finding 2」のサブグループ比較では、91 名の回答者(Yes/No で明確に回答できた層)のみを使用しています。なぜなら、失敗を観察できない者を含まれると、計測手段が整っていない側に偏った結果となり、比較が歪んでしまうからです。各サブグループのサンプル数は約 10 名から 62 名の範囲で、相応に粗い数値となっています。また、特定のセルの回答数が 10 名未満の場合は、パーセンテージとして報告していません。
少数の回答者が「その他」を選択し、リストされたカテゴリに該当しない業界(6%)や役職(2%)を自由記述しました。これらの割合は付録において「未記載」として扱われ、再配分されることはありません。
回答者数 101 名というサンプル規模は小規模であり、これは精密な測定値というよりは方向性を示すシグナルとして捉えるべきです。また、この調査は自己申告によるものであり、確率サンプリングではありません。したがって、これは大規模事業者の視点ではなく、RAG(検索拡張生成)やコンテキストインフラを積極的に構築・運用している組織からの見解と解釈するのが最も適切です。
Finding 1: 自信満々だが誤りであり、繰り返される
「一度きり」が最も多い答えではない
過去 6 ヶ月間に、エンタープライズにおいて、モデルの誤りではなく、不足または不整合なビジネスコンテキストに起因する、自信満々だが誤ったエージェント回答を追跡した経験があるか尋ねました。多くの組織でその事象は確認されており、さらにその多くが「再発」を経験しています。
これは本レポートの最も重要な数値です。企業の68%が、AIエージェントが自信満々に誤った回答を出力した経験があり、その原因は文脈の不備(指標定義の誤り、古くなったデータ、欠落したドキュメントなど)に遡れます。こうした事象は、単発的なケース(32%)よりも反復するケース(37%)の方が頻度が高く、問題視されています。このような失敗を経験していないと回答したのはわずか22%でした。
なお、本レポートでは原因の特定と追跡が可能な91社に限定して分析しています。この対象群において、76%が何らかの形で失敗を経験しており、そのうち41%は反復的な失敗に見舞われています。
この失敗モードは、一見すると失敗には見えない点こそが、極めて危険な理由です。モデルが視覚的に幻覚を起こしているわけではありません。むしろ、供給される文脈が薄く、古く、あるいは矛盾していたために自信満々に誤った回答を出力し、それが正しい回答と同じ権威を持って提示されてしまうのです。
重要なのは、この「典型的な体験」が単発の事象ではなく反復するものであるという点です。これは単なる見出しの数値以上の意味を持ちます。一度きりの失敗は修正すべきインシデントですが、反復する失敗は、ビジネス上の文脈がエージェントにどのように届くかという仕組み自体に構造的な欠陥があることを示しています。
本レポートの他のすべての要素——企業が何を取得し、どうガバナンスし、今後何を構築しようとしているか——はすべて、この根本的な問題に起因する下流の結果です。
発見2:セマンティック層が失敗を修正する前に明らかにする
ガバナンスされたレイヤーを構築している企業ほど、反復する失敗を経験している。むしろ、少ないわけではない
エンタープライズが、エージェントや BI ツールに対してデータの共通理解を提供するために、統制されたセマンティック層またはコンテキストレイヤーを活用しているかどうかを調査しました。多くの企業がその道を進んでおり、この回答と「発見 1」での失敗をクロス集計すると、最も直感に反する結果が浮かび上がります。
統制されたコンテキストレイヤーへの関与は広範囲に及んでいます。63% の企業が、すでに本番環境で稼働させている(32%)か、パイロット運用や構築中(31%)であり、さらに 20% が積極的に評価を進めています。つまり、5 社中 4 社以上が何らかの形でこのアイデアに関与していることになります。関与を表明しないのはわずか 14% です。
ここで興味深いのがクロス集計の結果です。失敗に関する質問に回答できた 91 の企業のうち、セマンティック層を構築済みまたは構築中の企業では、文脈の失敗が再発する割合が 50% に達しています。一方、層を持たない企業(評価中または計画なし)ではその割合は 21% です。この差は統計的な有意水準(p=0.01)を明確に超えており、技術が謳う効果とは逆の方向を示しています。
さらに、本番環境で稼働している層を持つ企業に焦点を絞ると、同様の傾向が見られますが、今回のサンプルでは統計的有意性までは持ちません。再発率は 53% で、他の企業群(34%)との差は有意ではありませんが、方向性を示すものとして解釈すべきです。
これを因果関係として捉えるのは不自然です。管理された定義レイヤーが誤った回答を生み出すわけではありません。しかし、検出の観点からなら、これは今回の調査で最も有用な結果と言えます。
自信を持って間違った回答を返す原因を、特定の文脈欠陥(具体的には「二つの方法で定義される指標」や「陳腐化したテーブル」、あるいは「エージェントが参照できなかったドキュメント」など)に特定するには、セマンティック・レイヤー(意味層)が提供する共有された管理された定義が必要です。これがなければ、同じ失敗が繰り返され、それがモデルの問題なのかユーザーのミスなのか、あるいは何らかの原因もないまま記録されるだけになってしまいます。
なお、因果関係は逆方向にも一部働いている可能性が高いです。つまり、過去に何度も痛い目を見た企業こそが、自らの手でこのレイヤーを構築したのです。
組織規模によるデータの違いも同じ傾向を示しており、本番環境での違いよりも重要な意味を持ちます。従業員数が 1,000 人以上の企業では、文脈に関する失敗が再発する割合が 55% に達しています。一方、101〜1,000 人の企業では 30% です(p=0.02)。これは、大規模組織の方が本番環境でセマンティック・レイヤーを導入している可能性が低いこと(24% 対 37%)を考慮してもなお顕著な差です。大企業ほど監視体制や監査体制が整っており、「なぜ数値が間違っているのか」を問う担当者が多く存在します。
読者にとっての現実的な示唆は、少し不愉快ですが明確です。「文脈失敗率が低い」という報告は、文脈レイヤーが健全である証拠にはなりません。むしろ、誰も監視していないことの証拠である可能性の方が高いのです。
発見 3:RAG が文脈ソースとして首位に立つ一方、そこでも多くの失敗が発生している
検索拡張生成(RAG)は他のどの手法よりも多くのエージェントで採用されていますが、その分だけ多くの失敗も引き起こしています。
企業が AI エージェントにデータを理解させる際、主に何を利用しているのかを尋ねました。その結果、検索(Retrieval)がトップになりましたが、以前ほど圧倒的な差ではありません。
検索は依然として企業コンテキストの基盤であり、31% を占めています。これに対し、管理されたセマンティック層(Semantic Layer)は 19%、混合アプローチは 17% です。しかし、注目すべきは「その他」の割合が増えている点です。長文コンテキスト(Long-context)を主要な情報源とする企業は 13% に達し、5% の企業ではエージェントが企業データ層を使わず、モデルの汎用的知識だけで動作しています。この二つを合わせると、約 5 社に 1 社の割合で、エージェントへのビジネス文脈提供が「無理やりコンテキストウィンドウを広げるか、あるいは全く行わない」かのどちらかに偏っていることがわかります。
第 1 の調査結果とクロス集計したデータを見ると、情報源によって失敗の発生状況は均一ではありません。主要な情報源として検索を利用している企業では、文脈に起因する失敗が 87% で報告され、そのうち 48% が再発しています(このグループは 31 社という最大規模のサンプルです)。管理されたセマンティック層に依存する企業ではそれぞれ 79% と 53%、混合アプローチでは 79% と 36%、直接ライブシステムへのクエリでは 40% と 30% です。一方、長文コンテキストを利用するグループは逆の傾向を示す外れ値となっており、あらゆる失敗が 64% で報告されるものの、再発率はわずか 9% でした。
これらのサブグループの数字は、発見 2 で示された検出に関する注意点を必ず考慮して読む必要があります。各グループは、誤った回答を文脈の欠陥に帰属させる能力と、実際にその欠陥に直面する頻度の両面で異なります。また、今回の調査における各セルの有効回答数は 10 から 31 件です。検索グループで示された 87% という数値は、RAG を多用する企業において文脈の失敗が実際に発生しており、かつその失敗を認識していることを示す証拠と捉えるべきであり、相対的な信頼性を単純に測定した結果として解釈すべきではありません。
注意点を除いた上で残る構造的なポイントは以下の通りです。企業の文脈情報の多くが検索プロセスを経由し、さらにサンプルの中で最も広範な基盤を担っているのは検索機能であるため、検索の質そのものが回答の質を決定します。RAG がデフォルトの情報源となっている場合、不完全な検索こそが主要な失敗要因となります。
発見 4:モデル支援型およびハイパースケラーによる検索は、ベクトルデータベースを上回る
OpenAI のファイル検索と Google の Vertex AI Search は、すべての専用システムを凌駕しています
現在、企業が本番環境で運用している検索システムについて尋ねたところ、その回答は引き続き、専門ベンダーよりもモデルプロバイダーやハイパースケラー企業に有利な結果を示しました。
専用ベクトルデータベースは、RAG(検索拡張生成)スタックの中心ではありません。OpenAI のファイル検索機能(46%)と Google Vertex AI Search(41%)が、他の目的別ソリューションを 3 倍以上引き離して首位を争っています。
専門特化型ツールに絞っても、企業がすでに他用途で運用しているものが最も使われています。Elasticsearch/OpenSearch が 20%、pgvector が 15% です。一方、ベクトルデータベースというカテゴリそのものを定義する純粋なベンダー製品(Pinecone、Weaviate、Milvus、Qdrant など)は、それぞれ 7% から 12% のシェアにとどまっています。
さらに興味深いのは、カスタム社内検索スタックが 18% を占め、すべての純粋なベクトルデータベースベンダーを上回っている点です。
どのシステムを「主軸」とするかで、単なる利用状況とインフラとしての位置づけは明確に区別されます。
使用頻度だけでは、この格差の実態を十分に表せません。実際には、多くの企業が複数のシステムを並行して運用しているからです。そこで今回は、「どのシステムが主軸か」を尋ねてみました。各システムのユーザー自身が「これが主たる検索プラットフォームだ」と答えた割合で市場を整理すると、明確な違いが見えてきます。
Elasticsearch と pgvector は広く導入されていますが、主軸として使われることは稀です。これらのユーザーの 5 人中 4 人は、実際には別のシステムを通じて検索を行っています。これらは企業がすでに運用していたインフラであり、RAG スタックの中心が他にある中で、端役として活用されているのです。
一方、モデル依存型やハイパースケーラー提供の検索機能は、単にリストの中で最も一般的なシステムであるだけでなく、導入企業の多くにとって「公式の記録システム」として機能しています。カスタム社内スタックも同様の傾向を示します。企業が独自に構築する場合、それは副次的なプロジェクトではなく、主軸となるシステムとして設計されることがほとんどです。
主要なプラットフォームに関する質問は単一選択形式でしたが、101 人中 18 人が複数の選択肢を選んだため、上記の割合は各システムのユーザー数に対する比率として計算されています。正確に一つの回答をした 83 人の場合、ランキングに変化はありませんでした。OpenAI のファイル検索が 28%、Vertex AI Search が 23%、独自開発のスタックが 12%、純粋なベクトルデータベースを使用している企業は 8% を超えていません。
ここで注目すべき比較対象は、RAG(Retrieval-Augmented Generation)専門分野に残された状況です。専門インフラを基盤としたカテゴリにおいて、単一の専用ベクトルデータベースを採用するよりも、多くの企業が独自の検索スタックを開発しています。また、モデルプロバイダーや既に契約しているクラウドサービスにバンドルされた検索機能を利用している企業は、その約 4 倍に達します。生産環境で RAG を全く導入していないのはわずか 7% であり、これは早期採用の物語ではありません。重要なのは、企業がどこから検索機能を調達しているかという点です。
発見事項 5:特定のアーキテクチャが合意形成されているわけではない
ハイブリッド型検索と「ユースケースによる」が同率首位
2026 年末までに、どの検索アーキテクチャが生産環境の RAG システムを支配すると企業は考えているかを尋ねました。単一の回答が過半数に達するものはありませんでした。また、トップを争う二つの選択肢の間には、たった一人の違いしかありませんでした。
ハイブリッド型検索が 30% で首位に立ち、単一のアーキテクチャが即座に支配するとは考えられておらず、29% がこれに続きます。この差は回答者 1 名分であり、このサンプルサイズでは誤差の範囲内です。実態としては、両者が同率で並んでいると捉えるのが妥当でしょう。合わせて 58% の企業が該当し、両者を結びつける共通点は、単一の検索手法ではなく層を成すパイプラインを採用している点にあります。また、カテゴリーの起源となった純粋なベクトル検索アプローチだけで本番環境を支えると期待する企業もありません。
より明確な信号を送っているのは、2 つの小さな回答です。15% の企業が、専用のベクトル層を設けずにツールファーストまたはロングコンテキスト型検索が支配的になると予想しており、これはカテゴリーの前提そのものに挑戦する立場です。一方、12% は依然としてベクトル検索のみが主流になると考えています。ベクトルデータベース構築に 3 年間を費やしてきた業界において、反ベクトル派が純粋なベクトル派をわずかに上回っている(回答者 3 名差で明確な優劣はつけられない)という事実は、興味深い逆転現象と言えます。さらに、15% が不確実であるか大規模な RAG を期待しないことを加えると、市場全体として「ベクトル検索単体では不十分」という点では合意しているものの、それを代替する手法についてはまだ合意に至っていない状況が浮かび上がります。
発見 6:企業は特定のベンダーにこの層を委ねるのを拒否する
ベンダーのネイティブなコンテキストスタックへの集約は、ほとんど注目を集めていない
モデルプロバイダーが検索機能、メモリ管理、オーケストレーションをプラットフォームに統合する中、企業はどのように対応していくのか。その意図と実際の利用状況には明確な対立が見られます。
このスタックの核心にあるのは緊張関係です。プロバイダーネイティブな検索機能が実際の利用で圧倒的な優勢(ファインディング 4)を示している一方で、企業の 12% しかがプロバイダーのネイティブなコンテキストスタックへの集約を意図していません。トップには「ベスト・オブ・ブリード」のスタンドアロンツールと、明示的なミックス構成が同率 37% で並び、6% が自前で構築・所有する方針です。つまり、企業の 79% はコンテキスト層の一部を特定のプロバイダーに依存させないことを望んでいます。
企業が実際に運用しているものと、口にする理想の間のギャップこそがこのカテゴリにおける戦略的課題です。企業は既に購入済みのツールが付属しているため束ねられた検索機能を採用しつつも、独立性の維持を主張しています。ファインディング 2 を踏まえると、この声明された選好にはベンダー政治以外の合理的な理由があります。企業が解決しようとしているのは、統制され、一貫性があり、アクセス権限を意識したビジネスコンテキストにおける失敗です。まさにその層こそが、彼らが最も外部委託したくない部分なのです。束ねられた利便性と理想の選好がぶつかった際にどちらが生き残るかは、今後の数回の波で決まるでしょう。
ファインディング 7:アクセス制御が購買決定に直結する
ガバナンス機能がシステム選択の理由として、データ取り込み(イングレス)と結びつくようになった
企業が検索システムを選定する際に最も重視する要素と、運用開始後に成功の主要指標として何を捉えているかを調査しました。
選定の基準はガバナンスへとシフトしています。アクセス制御や権限管理(24%)がデータ取り込みの容易さ(24%)と同率で首位に立ち、検索精度やレイテンシ、パフォーマンス(いずれも15%)、運用の簡便さ(14%)を上回りました。これで初めて、ガバナンス関連の特性が購入決定の最優先事項となりました。これは「自信満々だが誤った回答」という失敗、つまりエージェントが閲覧すべきでない情報を提示したり、閲覧すべき情報を見逃したりする問題と最も密接に関連する特性です。
システムが稼働した後は、正しさへの重視は明白です。回答の正確さが主要な成功指標となっているのは全企業の38%で、これは2番目に多いセキュリティやアクセス制御(19%)の倍に相当します。回答の関連性(17%)、レイテンシ(13%)、運用安定性(11%)はこれらに次いでいます。合計すると、56%の企業が検索システムの評価を「速いか」「安定しているか」ではなく、「回答が正しいか」「権限設定が適切か」という点に基づいて行っています。
現在のシステムに対する満足度は中程度に良好です。5 段階評価において、総合的な満足度が平均 4.13、コストパフォーマンスが 4.01、導入の容易さが 3.98 となっています。これは、この層における直近の証拠では約 4 割の企業で誤答が繰り返されているにもかかわらず、得られたスコアとしては妥当なものです。これはつまり、企業がこれらのツールを「正解を得る結果」に対して評価しているのではなく、「検索インフラが果たすべき役割」という期待値に基づいて評価していることを示唆しています。
発見事項 8:市場の半分が動き出している
Vertex AI Search が検討セットをリードする一方、不確実性もまた顕著です
企業側が検索プロバイダーの変更や追加を検討しているか、あるいは現在どの選択肢を検討しているかを尋ねました。その結果、検討対象となる候補は現在の技術スタックよりも広範囲に及んでいることがわかりました。
検索スタックはまだ確定した状態ではありませんが、激しい入れ替わりがあるわけでもありません。約半数の企業には変更計画がなく、残りの半分(101 社中 52 社)は 12 ヶ月以内にプロバイダーの切り替えや追加を予定しています。そのうち 5 分の 1 は、より短期間で実施する意向です。
原文を表示
Across 101 enterprises, the context feeding AI agents is failing often and repeatedly. Sixty-eight percent have traced a confident but wrong agent answer to missing or inconsistent business context in the past six months, and the single most common answer is not "once" but "more than once." The counterintuitive part is which companies report it. Enterprises building or running a governed semantic layer (a layer of company-specific definitions and relationships) report recurring failures at more than twice the rate of those without one.
The infrastructure built to fix bad context is, so far, mostly revealing how much bad context there is. Meanwhile the architecture meant to solve the problem commands no consensus at all: hybrid retrieval and outright pluralism finish one respondent apart, in a dead heat.
This wave of VentureBeat Pulse Research examines the enterprise RAG and context layer: what feeds AI agents their business context, which retrieval systems enterprises run, how they buy and measure them, where the architecture is heading, and — most revealingly — how often that context is already failing them.
The central finding is that the context failure is no longer an incident; it is a condition. Sixty-eight percent of enterprises say that in the past six months their AI agents produced confident but wrong answers they traced to missing or inconsistent business context rather than to model error. More striking than the total is its shape: 37% report the failure recurring, against 32% who saw it once. Among enterprises in a position to answer at all, the most prevalent experience of running agents on company data is being wrong repeatedly for reasons that have nothing to do with the model.
The remedy the industry has settled on — a governed semantic or context layer giving agents and BI a shared understanding of the data — is being built at scale: 32% run one in production, another 31% are piloting or building one, and 20% more are evaluating. But the cross-tabs deliver an uncomfortable result: Enterprises that have built or are building a layer report recurring context failures at 50%, against 21% for those without one. The layer isn't causing the failures — it's catching them, which makes it the most useful finding in the wave. The semantic layer is what makes a context defect traceable. Organizations without one are not having fewer failures so much as attributing fewer failures.
Underneath, the stack is unsettled in a way it was not expected to be. Retrieval remains the leading primary context source at 31%, and provider-native retrieval — OpenAI's file search (46%) and Google Vertex AI Search (41%) — still runs well ahead of every dedicated vector database. But the expected architecture has no majority behind it: hybrid retrieval (30%) and "multiple architectures, chosen by use case" (29%) are separated by a single respondent. And enterprises remain firmly unwilling to hand the context layer to a provider — just 12% intend to consolidate onto a single model provider’s native context stack, against 37% holding to best-of-breed and 37% planning an explicit mix.
The buying criteria are where the failure is starting to register commercially. Access control and permissions is now tied with ease of data ingestion as the top selection factor at 24% each, and response correctness is the primary success metric for 38% of enterprises. Enterprises are beginning to buy retrieval for the properties that govern context rather than the properties that move it.
Methodology
VentureBeat fielded this survey as part of its ongoing Pulse Research series. This survey focused on enterprise RAG infrastructure and the context layer — the retrieval systems, semantic layers, and context sources that feed AI agents. Responses are filtered to organizations with more than 100 employees (n=101). All responses are from a single July 2026 wave, so the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select; those shares are reported as a percentage of respondents, not of total selections, so they can sum to more than 100%.
By organization size the sample concentrates in the mid-market: 101–250 employees (34%), 1,001–5,000 (25%), and 251–1,000 (25%) lead, with 10,001+ (12%) and 5,001–10,000 (5%) above them. By role it spans managers (39%), individual contributors (29%), VPs and directors (22%), and the C-suite (9%); on purchasing authority it is buyer-credible, with 38% final decision-makers and another 43% recommenders or influencers. Technology/Software is the largest industry at 31%, followed by Healthcare/Life Sciences (14%), Retail/E-commerce (10%), and Manufacturing (9%).
A note on the context-failure base: Of the 101 respondents, 10 either do not run agents on enterprise data (5%) or do not trace root cause at that level (5%). Headline shares for the failure question are reported on the full 101; the subgroup comparisons in Finding 2 use the 91 respondents who were able to give a yes-or-no answer, since including those who cannot observe the failure would bias the comparison toward whichever group is less instrumented. Subgroup cells run from roughly 10 to 62 respondents and are correspondingly coarse; where a cell falls below 10 it is not reported as a percentage. A small number of respondents selected "Other" and gave a write-in industry (6%) or role (2%) that didn't map to a listed category; those shares appear as not stated in the appendix rather than being redistributed.
At 101 respondents this is a modest sample and should be read as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It is best read as the view from organizations actively standing up RAG and context infrastructure rather than from the largest operators.
Finding 1: Confident, wrong, and repeating
The most common answer isn't "once" but "more than once"
We asked whether, in the past six months, enterprises had traced a confident but wrong agent answer to missing or inconsistent business context rather than to model error. Most had — and most of those had seen it happen again.
This is the report’s defining number. Sixty-eight percent of enterprises have had an AI agent produce a confident, wrong answer they traced to bad context — wrong metric definitions, stale data, missing documents — and the recurring case (37%) outweighs the one-off (32%). Only 22% report no such failure. Restricted to the 91 enterprises able to observe and attribute the failure at all, 76% have experienced it and 41% repeatedly.
The failure mode is specific and dangerous precisely because it does not look like a failure. The model is not visibly hallucinating; it is confidently wrong because the context feeding it was thin, stale, or inconsistent — and it delivers that wrong answer with the same authority as a right one. That the modal experience is recurrence rather than a single incident matters more than the headline share: a one-time failure is an incident to be fixed, while a repeating one indicates a structural defect in how business context reaches the agent. Everything else in this report — what enterprises retrieve, how they govern it, and what they plan to build — is downstream of this problem.
Finding 2: The semantic layer reveals the failure before it fixes it
Enterprises building a governed layer report more recurring failures, not fewer
We asked whether enterprises use a governed semantic or context layer to give agents and BI a shared understanding of their data. Most are on the path — and cross-tabbing that answer against the failure in Finding 1 produces the wave’s most counterintuitive result.
Engagement with the governed context layer is broad. Sixty-three percent of enterprises either run one in production (32%) or are piloting and building one (31%), and a further 20% are actively evaluating, meaning more than four in five are engaged with the idea in some form. Only 14% have no plans.
The cross-tab is where it gets interesting. Among the 91 enterprises able to answer the failure question, those who have built or are building a semantic layer report recurring context failures at 50%, while those without one — evaluating or with no plans — report them at 21%, a gap that clears conventional significance thresholds (p=0.01) and runs in the direction opposite to what the technology is sold to do. Narrowing to enterprises with a layer specifically in production points the same way but does not carry statistical weight on this sample: 53% recurrence against 34% for everyone else, a difference that does not reach significance and should be read as directional only.
Read as causation, this is implausible — a governed definition layer does not manufacture wrong answers. Read as detection, it is the most useful result in this wave. Tracing a confident wrong answer to a specific context defect — a metric defined two ways, a stale table, a document the agent could not see — requires exactly the shared, governed definitions a semantic layer provides. Without one, the same failure occurs and gets logged as a model problem, a user error, or nothing at all. The causation almost certainly also runs backwards in part: enterprises that have been burned repeatedly are the ones who went and built the layer.
The size split points the same way, and carries significance where the production split does not. Enterprises above 1,000 employees report recurring context failures at 55%, against 30% of those between 101 and 1,000 (p=0.02) — despite the larger organizations being less likely, not more, to have a semantic layer in production (24% against 37%). Larger enterprises have more instrumentation, more auditing, and more people whose job is to ask why a number was wrong. The practical implication for readers is uncomfortable but clear: a low reported context-failure rate is not evidence of a healthy context layer. It is at least as likely to be evidence that nobody is looking.
Finding 3: RAG leads as the context source — and carries the failures
Retrieval feeds more agents than anything else, and fails a large share of them
We asked what an enterprise’s AI agents primarily use to understand its data. Retrieval leads, but no longer by the margin the category assumes.
Retrieval remains the backbone of enterprise context at 31%, ahead of a governed semantic layer (19%) and mixed approaches (17%). But the tail has thickened in a way worth noting: long-context loading is now the primary source for 13% of enterprises, and 5% let agents run on the model’s general knowledge with no enterprise context layer at all. Between them, nearly one in five enterprises is feeding agents business context either by brute-force context window or not at all.
Cross-tabbed against Finding 1, the sources do not fail equally. Among enterprises whose primary context source is retrieval, 87% report a context-traced failure and 48% report it recurring — on the largest base of any group, 31 respondents. Those relying on a governed semantic layer report 79% and 53%; mixed approaches 79% and 36%; direct live-system queries 40% and 30%. The long-context group is the outlier in the other direction, reporting 64% any failure but only 9% recurrence.
These subgroup figures should be read with the detection caveat from Finding 2 firmly attached. Groups differ in how well they can attribute a wrong answer to a context defect as much as in how often they suffer one, and the cells here run from 10 to 31 respondents. The retrieval group’s 87% is best read as evidence that RAG-heavy enterprises both experience and notice context failures, not as a clean measurement of relative reliability.
What survives the caveat is the structural point. Because so much enterprise context flows through retrieval, and because retrieval carries that load on the widest base in the sample, the quality of retrieval is the quality of the answer. When RAG is the default source, incomplete retrieval is the main point of failure.
Finding 4: Model-backed and hyperscaler retrieval still leads the vector databases
OpenAI's file search and Google's Vertex AI Search top every purpose-built system
We asked which retrieval systems enterprises run in production today. The answer continues to favor the model providers and hyperscalers over the specialists.
The dedicated vector database is not the center of the RAG stack. OpenAI’s file search (46%) and Google’s Vertex AI Search (41%) lead by better than three to one over any purpose-built alternative. Among the specialists, the most-used remain the ones enterprises already run for other reasons — Elasticsearch/OpenSearch at 20% and pgvector at 15% — while the pure-play vector databases that define the category (Pinecone, Weaviate, Milvus, Qdrant) each sit between 7% and 12%. Custom in-house retrieval stacks, at 18%, outrank every pure-play vendor.
Which system is actually primary separates retrieval from infrastructure
Usage counts alone understate the gap, because enterprises run several of these systems at once. We also asked which one is primary. The share of each system’s own users who name it their primary retrieval platform divides the field cleanly.
Elasticsearch and pgvector are widely present and rarely primary: four in five of their users retrieve mainly through something else. They are infrastructure the enterprise already ran, pressed into service at the edges of a retrieval stack whose center is elsewhere.
Model-backed and hyperscaler retrieval is not merely the most common system on the list; for most of the enterprises that adopt it, it is the system of record. Custom in-house stacks behave the same way — when an enterprise builds one, it is usually the primary, not a side project.
The primary-platform question was fielded as a single-select and 18 of 101 respondents selected more than one option, so the shares above are computed as a proportion of each system’s users rather than of the full sample. On the 83 respondents who gave exactly one answer, the ranking is unchanged: OpenAI's file search 28%, Vertex AI Search 23%, custom in-house stack 12%, and no pure-play vector database above 8%.
The comparison worth sitting with is what this leaves for the RAG specialists. In a category built around specialist infrastructure, more enterprises have written their own retrieval stack than run any single dedicated vector database — and roughly four times as many use retrieval that arrived bundled with a model provider or cloud they already buy from. Only 7% run no production RAG at all, so this is not a story about early adoption; it is a story about where retrieval gets acquired.
Finding 5: No architecture commands a consensus
Hybrid retrieval and "It depends on the use case" finish in a dead heat
We asked which retrieval architecture enterprises expect to dominate their production RAG systems by the end of 2026. No single answer comes close to a majority — and the two front-runners are separated by one respondent.
Hybrid retrieval leads at 30%, with the expectation that no single architecture will dominate at all immediately behind at 29%. The gap is one respondent, far inside the margin on a sample this size, and the honest reading is that these two finish level rather than that either is in front. Together they account for 58% of enterprises, and what unites them is more instructive than what separates them — both describe layered pipelines rather than a single retrieval technique, and neither expects the pure vector-search approach that launched the category to carry production on its own.
Two smaller answers carry the sharper signal. Fifteen percent expect tool-first or long-context retrieval to dominate without a dedicated vector layer at all — a direct challenge to the premise of the category — while 12% still expect vector-only retrieval to prevail. That the anti-vector position now edges the pure-vector one, on a three-respondent margin that is itself too narrow to call, is a notable inversion for an industry that spent three years building vector databases. Add the 15% who are unsure or expect no large-scale RAG, and the picture is of a market that agrees vector search alone is insufficient and has not agreed on what replaces it.
Finding 6: Enterprises decline to hand the layer to a provider
Consolidation onto a provider's native context stack barely registers
We asked how enterprises will respond as model providers bundle retrieval, memory, and orchestration into their platforms. Their stated intent cuts sharply against their current usage.
Here is the tension at the heart of the stack. Provider-native retrieval leads actual usage by a wide margin (Finding 4), yet just 12% of enterprises intend to consolidate onto a provider’s native context stack. Best-of-breed standalone tools and an explicit mix are tied at the top at 37% each, and 6% intend to build and own the layer themselves — meaning 79% of enterprises expect to keep at least part of the context layer outside any single provider.
The gap between what enterprises run and what they say they want is the strategic question of the category. They are adopting bundled retrieval because it arrives with tools they already buy, while asserting they will preserve independence. Read against Finding 2, the stated preference has a rationale beyond vendor politics: the failures enterprises are trying to fix are failures of governed, consistent, access-aware business context, and that is precisely the layer they are least willing to outsource. Whether the preference survives contact with the convenience of the bundle is what the next several waves will decide.
Finding 7: Access control climbs into the buying decision
Governance now ties ingestion as the reason a system gets chosen
We asked what matters most when enterprises choose a retrieval system, and what they treat as the primary measure of success once it is running.
The selection criteria have moved toward governance. Access control and permissions (24%) is now exactly tied with ease of data ingestion (24%) at the top, ahead of retrieval accuracy and latency and performance (15% each) and operational simplicity (14%). That puts a governance property at the top of the purchase decision for the first time in this series — and it is the property most directly implicated in the confident-but-wrong failures of Finding 1, where an agent surfaces something it should not have seen or misses something it should have.
Once systems are running, the emphasis on correctness is unambiguous: response correctness is the primary success metric for 38% of enterprises, twice the next answer, security and access control (19%). Answer relevance (17%), latency (13%), and operational stability (11%) trail. Taken together, 56% of enterprises measure their retrieval system primarily on whether its answers are right or properly permissioned, rather than on whether it is fast or stable.
Satisfaction with current systems is moderately positive: on a five-point scale, overall satisfaction averages 4.13, value for money 4.01, and ease of implementation 3.98. That is a respectable set of scores for a layer that, on this wave’s evidence, is producing recurring wrong answers in nearly four in ten enterprises — which suggests enterprises are rating the tools against expectations of what retrieval infrastructure does, not against the outcome of getting the answer right.
Finding 8: Half the market is in motion
Vertex AI Search leads the consideration set — and so does uncertainty
We asked whether enterprises plan to change or add a retrieval provider, and which they are considering. The consideration set is broader than today’s stack.
The retrieval stack is not settled, but it is not churning, either: about half of enterprises have no plans to change, while the other half — 52 of 101 — intend to switch or add a provider within twelve months, a fifth of them within
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み