評価の失敗が企業を焦らせた:人間を排除する傾向と自動評価への信頼上昇
本文の状態
日本語全文を表示中
詳細モードで約25分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
VentureBeat AI
VentureBeat の調査によると、自動評価への信頼は高まったが実世界の失敗率は変わらず、過去の失敗経験を持つ企業ほど自律化を加速させていることが判明した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 17:11
AI深層分析
キーポイント
評価と実績の乖離
7 月の調査では自動評価への信頼度が上昇したが、内部テストに合格しても顧客で失敗する事例は前月と同様の約半数に留まり、評価精度には改善が見られない。
経験による信頼の二極化
自動評価を完全に信頼するのは失敗経験がない企業の 24% であり、過去に「誤った自信」による失敗を経験した企業ではその割合は 4% に過ぎない。
失敗が自律化を加速させる
過去の失敗で苦しんだ企業の 85% がゼロ人間でのデプロイを許可しているかその方向で取り組んでおり、失敗が自律化への撤退ではなく加速要因となっている。
評価ツールの市場定着
専用評価ツールを導入しない企業の割合が減り、Braintrust や DeepEval といった専門プラットフォームの利用が増加し、選定基準もコストから統合の容易さへ移行している。
ベンダー市場の成熟と選定基準の変化
専用評価ツールの未導入企業が減少し、専門プラットフォームの利用が増加している。価格よりも統合の容易さが最優先事項となり、企業は適合性を重視して購入を決定するようになった。
重要な引用
Trust is concentrated almost entirely among enterprises that have not experienced a false-confidence failure
Getting burned does not slow the march to autonomy — it speeds it up.
Confidence improved; correctness did not.
Enterprises are done shopping on price and have started buying on fit.
編集コメントを表示
編集コメント
企業の実践データは、評価ツールの信頼性と自律化の心理的メカニズムについて重要な示唆を与えている。この調査結果は、単なるツールの導入ではなく、組織文化や過去の失敗経験が技術判断にどう影響するかを浮き彫りにしている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
108 社の企業を対象とした調査で、7 月に自動化されたエージェント評価への信頼が急上昇しました。しかし、その評価が予測すべきはずの失敗率は全く変化していません。
組織が自動化評価を「完全に信頼する」と答えた割合は、6 月の 5% から 13% に約 3 倍に増えました。また、「評価結果と実世界の成果が一致していない」という不満も 10 ポイント減少しました。しかし、先月と同じく半数未満の企業が、評価には合格したものの顧客の前で失敗するエージェントをリリースしています。
この矛盾の原因は、クロス集計データを見れば明らかです。新たに信頼を寄せたのは、まだ「火傷」を負ったことのない企業に集中しています。「火傷」を経験した企業のうち、自動化評価を信頼するのはわずか 4% ですが、未経験の企業ではその割合が 24% に達します。
そして、「火傷」を負っても自律化への歩みが止まるどころか、むしろ加速しているのです。
これは VentureBeat Pulse Research のエージェント信頼性トラッカーの第 2 波調査です。前月と同じ測定器を用いて実施されたため、7 月のデータは「位置」ではなく「方向性」を捉えた最初の読みとなります。何が動き、何が維持され、その変化が何を意味するのか。
動いたのは「信頼感」です。6 月には自動化評価を完全に信頼する企業はわずか 5% で、最も頻繁に指摘された課題は「評価結果と実世界の成果の整合性が低いこと」(29%)でした。しかし 7 月には、完全信頼する企業が 13% に増え、「整合性」への不満も 19% に低下して最上位の懸念事項ではなくなりました。これらの変化は統計的なノイズではなく、明確な実態として捉えるに十分な規模です。
失敗が明らかになったのは、信頼の根拠である。過去1年間に内部評価を通過したエージェントやLLM機能を導入した組織のうち、約半数(49%)が顧客Facing の障害を引き起こしました。これは6月の50%と統計的に差がない結果です。また、24%の組織ではこの失敗が二度以上発生しています。
信頼感は向上しましたが、正答率は改善していません。7月のギャップとは、6月のように自律性と信頼性の間の問題ではなく、「信頼」と「その根拠となる証拠」の間の乖離を指します。
クロス集計データから、新たな信頼感がどこから来ているかがわかりますが、それは評価精度の向上によるものではありません。信頼は、誤った自信による失敗を経験したことがない企業に集中しています。そのような企業の24%が自動評価を完全に信頼している一方、失敗経験がある企業ではその割合はわずか4%です。つまり、信頼度の曲線は「経験不足」によって押し上げられているのです。
一方で、失敗で痛い目を見た企業たちは自律性を放棄していません。彼らの85%がすでにゼロ人間でのデプロイを許可しているか、あるいはそれを指向したエンジニアリングを進めています。これに対し、失敗経験のない企業では61%です。全体としての自律化の軌道は横ばいの67%ですが、その中身にある組織構成は、「評価が物事を逃す」という最も直接的な証拠を持つ企業へとシフトしています。
一方、ベンダー市場はようやく落ち着きを見せています。評価ツールを全く導入していない企業の割合は 17% から 12% に低下し、専門プラットフォームのシェアは拡大しました。Braintrust は約 2 倍の 15% に達し、DeepEval も 17% を記録しています。また、製品変更を検討する企業は減少し、現状維持を望む企業の割合は 36% から 44% に上昇しました。
選定基準も変化しました。コストよりも「導入の容易さ」が最優先事項となり、その割合は 27% から 39% へと急増しています。企業は価格競争から脱却し、自社の環境に適合する製品を選ぶ段階に入ったのです。
調査手法
VentureBeat は、継続的な「パルス・リサーチ」シリーズの一環として本調査を実施しました。今回の調査(エージェントの信頼性と評価ツールに関する追跡調査)では、技術リーダーがどのようにエージェントのパフォーマンスと信頼性を評価しているかに焦点を当てています。
回答は従業員 100 人以上の組織に限定され(対象数 n=108)、2026 年 7 月に実施されたデータに基づいています。7 月の調査項目は 6 月と同じであるため、このレポートでは必要な場合に前月との比較を行っています。なお、複数回答が可能な設問については、各選択肢の割合を合計すると 100% を超える場合があります。
6 月(サンプル数 n=157)との比較では、統計的に有意な差が認められるのは限定的です。具体的には、自動化評価への完全な信頼が 5% から 13% に上昇したこと、実世界での整合性に関する不満が 29% から 19% に減少したこと、選定要因としての統合の容易さが 27% から 39% に跳ね上がったこと、そして Braintrust が主要プラットフォームとして選ばれた割合が 8% から 15% に増加したことの 4 つです。
一方、本報告で「横ばい」と記述されている項目、つまり失敗率、自律性の推移、生産モニタリングの構成比率、投資優先度の順位などは、調査間に変化が見られず統計的に有意な差はありません。この安定性こそが重要な発見です。その他の数ポイントの差異は、トレンドではなくサンプルの変動として捉えるべきでしょう。
役割別の内訳を見ると、AI 購入に関する最終決定権を持つ層が 44%、推奨者や影響力を持つ層が 25% を占め、全体的に 6 月よりややシニアな構成となっています。具体的な役職では、プロダクトマネージャーとプログラムマネージャー(18%)、コンサルタントおよびアドバイザー(12%)、CIO/CTO/CISO(11%)、エンジニアリング・IT ディレクター(11%)が主要層を占め、その他機能の担当者が 30% を占めています。
組織規模別では、中堅企業に偏った構成です。従業員数 100〜499 人が 33%、500〜2,499 人が 30% で最も多く、次いで 2,500〜9,999 人(23%)、10,000〜49,999 人(9%)、50,000 人以上(5%)となっています。
信頼性に関する調査結果の解釈には、サンプル構成の変化に注意が必要です。業界別の内訳が前回と変化しており、テクノロジー・ソフトウェア分野は 6 月の 23% から 7 月には 14% に減少しました。一方、小売・消費財セクターは 15% から 19% に増加し、現在ではトップとなっています。
テクノロジー依存度の低いサンプルでは、エージェント評価の実践的な経験が相対的に少ない可能性があり、今月の信頼感の上昇の一部は「何が変化したか」よりも「誰が回答したか」に起因している可能性があります。ただし、調査結果 2 で報告された「被害を受けた企業とそうでない企業の違い」という傾向は、7 月のサンプル内でも依然として成立しています。
今回のサンプル数は 108 社で、方向性を示す結論を導くには十分ですが、精密な数値測定として扱うべきではありません。これは自己選択型のサンプリングであり、確率抽出によるものではないためです。ここで報告されているクロス集計は、40〜68 社の小規模グループに基づいているため、粗い概算値にとどまります。
調査結果 1:失敗率は変化していない
評価を通過しても顧客に不具合をもたらすエージェントを出荷している企業が依然として半数近く存在します
過去 12 ヶ月間に、組織が社内評価を通過したエージェントや LLM 機能を導入したが、その後顧客 facing で失敗を引き起こした事例があるかどうかを尋ねました。その答えは先月の結果と変わりません。
内部評価をクリアしたにもかかわらず、顧客の前で失敗する AI 機能をリリースした組織は 49% に上ります。具体的には、誤った出力や破綻したワークフロー、品質インシデントなどが発生しました。これは 6 月の 50% とほぼ同率です。また、24% の組織が「二度以上」同じ失敗を経験しており、この数値も前月と変わっていません。2 回の調査(2 ウェーブ)で 265 の企業を対象に集計した結果、評価を通過したエージェントが実際に失敗する確率は、誤差の範囲内で安定していることがわかりました。
この「安定性」こそが、その後のすべての議論の基盤となります。今月起きた他の動き——信頼感の上昇やツールの統合、購入基準の変化など——はすべて、「評価を通過しても失敗する確率が一向に変わらない」という事実を背景に読み解く必要があります。6 月から 7 月の間に企業側が何らかの対策を講じたとしても、それが「合格した評価が実は間違っていた」という事態の頻度に影響を与えたわけではありません。
発見点 2:信頼は向上したが、「被害を受けたことのない」層に限られる
自動化されたエージェント評価への「完全な信頼」がほぼ 3 倍に増え、実装との不整合に関する不満は 10 ポイント減少しました。
私たちは現在、自動化されたエージェント評価に対する信頼を最も低下させている要因が何かを尋ねました。その回答の分布は 6 月と比べて大きく変化しています。
2 つの大きな動きが同時に起こりました。まず、自動化された評価への「完全な信頼」が 5% から 13% とほぼ 3 倍に増加しました。一方、最も直接的に「過信による失敗(偽りの自信)」を指す不満である「現実世界の結果との整合性が低い」という回答は、29% から 19% に減少し、トップの座を明け渡しました。
現在、評価バイアスと一貫性の欠如(22%)が首位に立ち、データ漏洩への懸念(22%)と同率となっています。表面的には、評価レイヤーがようやくその役割を果たし始めたように見えます。
クロス集計結果は、その逆を示しています。組織が実際に「過信による失敗」を経験したかどうかでサンプルを分割すると、信頼の度合いはほぼ完全に二極化します。
評価テストに合格して顧客先に導入したが、その後で失敗した 53 の企業のうち、自動化された評価を「完全に信頼している」と答えたのはわずか 4% です。一方、そのような失敗を経験したことのない 41 の企業では、24% がそう答えています。これは 6 倍の開きであり、このデータセットにおける最も明確な分岐点です。
失敗モードに直接触れる経験こそが、信頼を失わせる要因なのです。
これが今月の核心的な警鐘です。感情の改善は、評価手法そのものが向上した証拠ではありません。ファインディング 1 で示された失敗率によって、それは否定されます。これはむしろ、「あまり火傷を負っていない」組織群がサンプルに追加され、自らの事前信念を報告した際に現れる信頼曲線の姿に過ぎません。
自社の自信の高まりを評価スタックの妥当性の証明だと読み取る企業は、実は「経験不足」を測定している数値を見ているに過ぎないのです。
ファインディング 3:火傷を負うことが自律化を抑制するのではなく、むしろ加速させる
火傷を負った組織の 85% が「人間をゼロにする」道を選んでいる一方、残りの組織では 61% です。
私たちは、組織が自動化された評価結果のみを根拠に、人間の検証(ヒューマン・イン・ザ・ループ)を経ずに自律型エージェントにコードやシステム変更の生産環境へのデプロイを任せるかどうかを尋ねました。集計値全体では一定の結果でしたが、その構成要素は大きく異なっていました。
結論から言うと、状況に変化はありません。低リスクのエージェントに対して「人間を介在させない」展開をすでに許可している組織が 37%、今後 1 年以内にその体制を整えようとしているのが 30% で、合計 67% です。これは 6 月の調査結果と同じです。一方、「当面は絶対に認めない」とする割合は、22% から 18% に低下しました。
しかし、詳細を見ると直感とは逆の傾向が浮かび上がります。「評価をパスしたエージェントをリリースしたが、顧客で失敗した経験がある」企業では 85% が自律化の道を進んでいます。一方、「そのような失敗を経験していない」企業でも 61% が同様の軌道上にあります。
つまり、評価プロセスに穴があり、それが実際にコストのかかる失敗として現れたという直接的な証拠を持つ組織ほど、人間によるチェックを外す傾向が強いのです。「燃えた経験がある」企業のうち、完全自動化を「絶対に認めない」と答えたのはわずか 11% でした。これに対し、「燃えていない」企業ではその割合は 24% です。このパターンは、一度失敗しただけの企業でも、何度も失敗した企業でも同じです。
最も妥当な説明は、無謀さではなく成熟度によるものです。顧客への影響が出るほどの大量のエージェントをリリースできる規模に達している組織こそ、自動化に必要なパイプラインも十分に整備されています。彼らは失敗を「止める理由」ではなく、「事業運営に伴うコスト」として捉えているのです。
これは合理的な解釈です。しかし同時に、これが「安定した失敗率」を「絶対的な事故数の増加」へと変えるメカニズムでもあります。なぜなら、最も失敗しやすいのは、レビューなしで展開能力を拡大している企業だからです。
ある調査結果は、6 月のデータでは再現されませんでした。先月までは大企業の方が中小企業よりも自律化の道を進んでいるように見えました(70% 対 64%)。しかし 7 月のデータでは両者の差がほぼ消え、従業員数 2,500 人以上の組織で 65%、それ以下の組織で 68% と逆転し、失敗率も 48% と 50% でほとんど同じでした。これは 6 月に見られた差が企業規模によるものではなく、単なるサンプルの変動だったことを示唆しています。企業が積極的に導入を進めるかどうかを分けるのは会社規模ではなく、過去の失敗経験なのです。
調査結果 4:技術スタックの統合が始まる
専門ツールが台頭し、「何もない」という選択肢は縮小
企業は今、どのエージェント信頼性や評価プラットフォームを主に利用しているのかを尋ねました。市場はまだ混雑していますが、トップで「何も使っていない」状態が続いているわけではありません。
最も注目すべき変化は、減少した数値です。6 月では、専用のエージェント評価ツールを持っていないという回答が 17% で最多(同率)でしたが、7 月には 12% に減り、5 位となりました。企業は評価ツールの導入を進めており、その動きの大半を専門ベンダーが取り込んでいます。Braintrust は主要利用シェアをほぼ倍増させ 15% に達し、DeepEval も 17% を記録しました。一方、プロバイダーネイティブなツール(OpenAI が 18%、Anthropic が 12%)は横ばいでした。つまり、成長の源泉がモデルプロバイダーから奪われたのではなく、「何もしない」状態からの移行によるものなのです。
プラットフォーム固有の評価ツールではなく、あらゆる利用形態をカウントすると、その浸透範囲はより広がり、順位も概ね同様です。OpenAI 独自の評価ツールが企業の 31% で採用され、次いで DeepEval が 27%、Braintrust が 22%、Anthropic 独自ツールが 20%、社内開発のカスタムツールが 14%、Weave と Langfuse がそれぞれ 11% です。また、19% の企業はスタック全体で専用ツールの利用を報告していません。6 月には二桁のシェアを持つ独立した候補者がいなかったカテゴリにおいて、現在は 3 つの有力な競合が台頭しており、評価レイヤーがようやく形になりつつあることを示す最初の証拠となっています。
発見 5:本番環境での監視は依然として「間違ったもの」を見ている
半分がエージェントの稼働状況を確認し、3 分の 1 以下が正答性を確認している
AI エージェントの本番環境監視には、大きく分けて 2 つの異なる対象があります。一つはシステムの機能状態です。エージェントが起動しており応答できているか、各リクエストが完了したか、その速度やコストはどうだったか、エラーが発生していないかなどを監視します。もう一つは、エージェントの出力内容が正しいかどうかです。これは、回答が生成されるたびにその内容を自動的に検証するチェック機能です。
自信を持って間違った答えを出すケースは、前者の監視では検知できません。リクエストは正常に完了し、応答速度も速く、エラーも発生せず、すべての機能指標が健全を示すからです。現在、本番環境での監視がどちらを目的として構築されているのかを問いました。
監視対象を何に絞っているかで分類すると、6 月の調査結果とほぼ同じ傾向が見られます。組織の 50% はエージェントが動作しているかどうかだけを監視しており、26% が回答の正しさを自動チェックしています。臨時のレビュー担当者や「不明」と答えたケースを含めると、生産環境における出力の正しさをリアルタイムで評価する自動化された仕組みを導入していない組織は約 75% に達します。
インライン品質アサーションとトランザクショントレースログが、106 社をベースにそれぞれ 26% で並んで最も一般的なアプローチとなっています。監視体制において単一の支配的なパターンはありません。
これは今月の信頼感の高まりと真っ向から矛盾する発見です。自動評価への信頼は 8 ポイント上昇しましたが、評価自体が間違っていることを検出する実行時の能力は全く向上していません。すでに人間を介さないデプロイを許可している企業に限っても、生産トラフィックに対してインライン品質チェックを実施しているのはわずか 28% です。つまり、人間の判断を排除した組織の多くが、その後の出力品質を見守る仕組みも導入していないことになります。ゲートは自動化されているのに、アラームは設置されていない状態です。
発見 6:価格ではなく適合性を重視して購入
統合の容易さがコストを上回る最重要選択基準に
企業が評価ベンダーを選ぶ際に最も影響を受けた要因と、成功の主要指標として何を捉えているかを尋ねました。一方の答えは大きく変化しましたが、他方は全く変わりませんでした。
統合のしやすさが12ポイント上昇して39%となり、コストが主要な選定基準からその座を奪いました。これはデータ上最も明確な購買シフトの一つです。
評価精度は小幅に28%まで向上しましたが、コストは23%に低下しました。また、観測範囲の広さ(6%)やベンダーのロードマップ(2%)といった要素は依然として限定的な要因にとどまっています。この傾向は「発見4」と併せて読むと明確になります。初めて専用の評価ツールを導入する企業は、抽象的な安さや機能の多さを求めるのではなく、今四半期に既存のパイプラインにスムーズに組み込めるものを最適化しているのです。
これは市場が評価から導入へと移行した際に見られる姿です。
一方、ツールをインストールした後に企業が何を求めているかについては変化がありません。評価の一貫性が38%で主要な成功指標であり、6月の36%とほぼ同等の水準を維持しています。これは、失敗削減(20%)、実験速度(18%)、本番環境での可視性(16%)、コンプライアンス(7%)といった他の項目を大きく引き離す結果です。
優先順位は依然として「再現性」、つまり同じ行動に対して毎回同じ判断が下されることにあります。これは、バイアスや不整合性が「発見2」で信頼性の最大の制限要因として最も多く指摘されていることを考えると特筆すべき点です。
企業は統合のしやすさを求めて導入し、安定性を測る指標として評価の一貫性を見ていますが、まだ後者については十分な成果を得られていないようです。現在のツールに対する満足度は全体的に中程度で、総合的な満足度、実装の容易さ、コストパフォーマンスを合わせた平均スコアは5段階評価で3.9です。これは6月の3.8とほとんど変わりません。
発見7:人的レビューが最優先事項となる
そして、失敗した経験を持つ企業ほど、この分野への投資を強化しています。
次年度における信頼性向上と評価(エバリュエーション)への投資拡大について尋ねたところ、「人的レビュー」が首位となりました。
人的レビューワークフロー(31%)とプロダクション観測ツール(30%)の順位が入れ替わりましたが、この変化自体はノイズの範囲内です。しかし、その背後にある傾向は 6 月の調査で指摘されたものと同じであり、さらに強まっています。企業は、人的レビューを代替する自動化評価パイプライン(19%)への投資拡大よりも、人的レビューへの予算増額を優先しています。同時に、企業の 3 分の 2 がデプロイメントの判断プロセスから人間を排除しようとしています。予算が据え置かれると回答したのはわずか 6% で、前年の 8% から減少しました。
クロス集計の結果、この傾向は明確になりました。顧客に対して評価に合格したエージェントを提供したが失敗した経験がある企業では、38% が人的レビューを最優先の投資先としています。一方、そのような失敗経験がない企業ではその割合が 24% に留まり、彼らは観測ツールへの投資を好んでいます。つまり、「失敗した経験を持つ」企業群は、両方の施策を同時に推進しています。すなわち、自律化(ゼロ人間化)に対して最も積極的であり(調査結果 3 より 85% がこの道を選んでいる)、一方で人的レビューの予算確保にも最も力を入れているのです。これは矛盾ではなく、明確な戦略です。「デプロイメントの判断は自動化し、自動化が逃したミスを人間が拾う」という仕組みです。これがスケーラブルかどうかは未だ不明ですが、エージェントの稼働量が増加してもコストが下がらない「人的レビュー」こそが、このスタックの中で唯一のボトルネックだからです。
調査結果 8:スイッチング(移行)の波は冷めていく
変更を計画しない企業の割合が3割台からほぼ半数に上昇
企業側が新たな評価プラットフォームの導入、追加、または代替を検討しているか、また具体的にどの選択肢を検討しているかを尋ねました。先月と比較すると、購入検討中の企業は減少しています。
今後12ヶ月以内に新しい、あるいは追加・代替の評価プラットフォームを導入する予定だと答えた企業の割合は過半数(56%)に達していますが、これは先月の64%から低下した数字です。また、より短期間で導入を検討している層も31%から24%へと縮小しました。一方で、現状のシステムを維持する方針を示す企業は36%から44%へと増加しています。これらの数値変動がそれぞれ単独で統計的な有意水準に達しているわけではありませんが、両方の動きは同じ方向性を示しており、これは「発見4」の結果とも整合しています。つまり、企業が実際にツールを導入し始めると、まだ導入を検討中の企業群の規模は縮小していく傾向にあるのです。
検討対象となるプラットフォームの順位も入れ替わりました。変更を計画している60社の中で、現在最も評価されているのはOpenAIのネイティブな評価機能で20%です。次いでBraintrustが18%、Weights & Biases Weaveが12%、DeepEvalが10%となっています。さらに10%は積極的に検討中ですが、候補リストを固めていません。先月の6月にはDeepEvalが20%でトップの検討対象でしたが、その後多くの関心が実際の利用へと転換しました。これは「検討から採用へ」という典型的な移行プロセスです。現在、BraintrustはかつてDeepEvalが占めていたポジション、つまり「導入済みベースよりも先行して高い関心を集めている状態」を占めており、次の波で注目すべきベンダーとなっています。
結論:信頼性は高まったが、正解率は変わっていない
6 月、あるギャップが浮き彫りになった。企業がエージェントに与えている自律性の度合いと、それを統制するために設けた評価への信頼の間に存在する乖離だ。しかし 7 月には、その隙間が誤った方向から埋まりつつあることが判明した。
信頼は高まった。自動化された評価に対する完全な信頼はほぼ 3 倍に増え、「評価が現実を捉えきれていない」という不満も 10 ポイント減少した。だが、本来ならその信頼が追うべき指標となるものは、全く変化していない。依然として企業の約半数は、評価には合格するものの顧客の前では失敗するエージェントをリリースし続けている。先月と状況は変わらない。
クロス集計結果から、新たな自信の所在が明確になった。それは評価自体にあるわけではないのだ。一度も「過信による失敗」を経験したことがない企業では 24% が自動化された評価を完全に信頼している一方、経験済みの企業ではその割合はわずか 4% に過ぎない。この市場における信頼は、証拠ではなく暴露(エクスポージャー)によって決まるのだ。
そして、暴露が警戒心を生むわけではない。むしろ失敗を経験した層こそが、サンプルの中で最も自律性の高いグループであり、すでに 85% が人間のレビューなしに展開しているか、その方向へ進んでいる。彼らが代わりに取るのは「対抗策」だ。同じ組織ほど、人間によるレビューワークフローへの資金投入を最優先し(38%)、一方で展開のゲートから人間を排除する動きも同時に加速させている。
ベンダー市場は、今月の本当に明るい話題となりました。専用ツールの提供を行わない企業の割合が 17% から 12% に低下し、専門家が初めて実質的なシェアを獲得しました。また、購入者の関心も価格から統合の適合性へと移り、導入が完了したことで乗り換え意向は冷却されました。ようやく評価層が形成されつつあります。しかし、ランタイムの状況はこれに追いついていません:企業の半数はまだ...
原文を表示
Across 108 enterprises, trust in automated agent evaluation rose sharply in July — and the failure rate it is supposed to predict did not move at all. The share of organizations that fully trust automated evaluation nearly tripled, from 5% in June to 13%, and the complaint that evaluations don’t match real-world outcomes fell 10 points. Yet the same share as last month — just under half — shipped an agent that passed its evals and then failed a customer. The reason is visible in the cross-tabs: the new trust belongs almost entirely to enterprises that have not yet been burned. Among those that have, 4% trust automated evaluation; among those that haven’t, 24% do. And getting burned does not slow the march to autonomy — it speeds it up.
This is the second wave of the VentureBeat Pulse Research agent reliability tracker, and the first fielded on an instrument identical to the month before it. That makes July the first read on direction rather than position: what moved, what held, and what the movement means.
What moved is confidence. In June, only 5% of enterprises said they fully trusted automated evaluation, and the most-cited limitation was that evaluations align poorly with real-world outcomes (29%). In July, 13% fully trust automated evaluation and the alignment complaint has fallen to 19%, no longer the leading objection. Both shifts are large enough to read as real rather than noise.
What held is the failure. Just under half of organizations (49%) deployed an agent or LLM feature in the past year that passed internal evaluations and then caused a customer-facing failure — statistically indistinguishable from June’s 50% — and a quarter (24%) have seen it happen more than once. Confidence improved; correctness did not. That is the July gap: not between autonomy and trust, as in June, but between trust and the evidence for it.
The cross-tabs explain where the new confidence comes from, and it is not from better evaluations. Trust is concentrated almost entirely among enterprises that have not experienced a false-confidence failure: 24% of them fully trust automated evaluation, against 4% of those that have. The trust curve is being lifted by inexperience. Meanwhile the enterprises that have been burned are not retreating from autonomy — 85% of them already allow zero-human deployment or are engineering toward it, against 61% of those that have not been burned. Overall the autonomy trajectory is flat at 67%, but the population inside it has shifted toward the organizations with the most direct evidence that evaluations miss things.
The vendor market, by contrast, is finally showing signs of settling. The share of enterprises running no dedicated evaluation tooling fell from 17% to 12%; specialist platforms gained, with Braintrust nearly doubling to 15% and DeepEval reaching 17%; and switching intent cooled, with those planning no change rising from 36% to 44%. Selection criteria moved with it: ease of integration overtook cost as the top factor, jumping from 27% to 39%. Enterprises are done shopping on price and have started buying on fit.
Methodology
VentureBeat fielded this survey as part of its ongoing Pulse Research series. This wave — the agentic reliability and evals tracker — examines how technical leaders evaluate agent performance and reliability. Responses are filtered to organizations with 100 or more employees (n=108), drawn from a July 2026 fielding. Because the July instrument is identical to June’s, this report makes month-over-month comparisons where they are warranted; where questions were multiple-select, shares can sum to more than 100%.
Comparisons against June (n=157) are tested for significance, and only a handful of the month’s movements clear a conventional threshold: the rise in full trust in automated evaluation (5% to 13%), the fall in the real-world-alignment complaint (29% to 19%), the jump in ease of integration as a selection factor (27% to 39%), and the gain in Braintrust as a primary platform (8% to 15%). Movements described in this report as flat — the failure rate, the autonomy trajectory, the production monitoring mix, the investment ranking — are statistically indistinguishable between waves, and that stability is itself the finding. Differences of a few points elsewhere should be read as sample variation, not trend.
By role the sample is senior and buyer-credible: 44% are final decision-makers for AI purchases and another 25% recommenders or influencers, a slightly more senior mix than June. Product and program managers (18%), consultants and advisors (12%), CIOs/CTOs/CISOs (11%), and directors of engineering/IT (11%) lead the named titles, alongside a large “Other” function (30%). By organization size the sample is again mid-market-weighted: 100–499 (33%) and 500–2,499 (30%) employees lead, with 2,500–9,999 (23%), 10,000–49,999 (9%), and 50,000+ (5%) above them.
One composition change is worth flagging because it bears on the trust finding. The industry mix shifted between waves: Technology/Software fell from 23% of the June sample to 14% in July, while Retail/Consumer rose from 15% to 19% and now leads. A less technology-weighted sample plausibly carries less hands-on exposure to agent evaluation, and some of the month’s rise in trust may reflect who answered rather than what changed. The burned-versus-unburned split reported in Finding 2 holds within the July sample regardless, but readers should treat the headline trust movement as directional.
At 108 respondents the sample is large enough to support directional conclusions but should not be treated as a precise measurement; it is self-selected and is not a probability sample. Cross-tabs reported here rest on subgroups of 40 to 68 respondents and are correspondingly coarse.
Finding 1: The failure rate did not move
Just under half still ship agents that pass evals and fail customers
We asked whether, in the past 12 months, organizations had deployed an agent or LLM feature that passed their internal evaluations but then caused a customer-facing failure. The answer is the same as last month.
Forty-nine percent of organizations shipped an AI feature that cleared internal evaluations and then failed in front of a customer — an incorrect output, a broken workflow, or a quality incident — against 50% in June. A quarter (24%) have seen it happen more than once, unchanged. Across two waves and 265 enterprises, the rate at which evaluations certify agents that then fail is stable to within a percentage point.
That stability is the anchor for everything that follows. Every other movement this month — rising trust, consolidating tooling, shifting purchase criteria — has to be read against a failure rate that has not responded. Whatever enterprises did between June and July, it did not change how often a passing evaluation turns out to be wrong.
Finding 2: Trust rose — among those who haven’t been burned
Full trust nearly tripled, and the alignment complaint fell ten points
We asked which limitation most reduces trust in automated agent evaluations today. The distribution shifted materially from June.
Two things moved together: Full trust in automated evaluation nearly tripled, from 5% to 13%, and the objection that most directly describes a false-confidence failure — poor alignment with real-world outcomes — fell from 29% to 19%, surrendering the top spot to evaluation bias and inconsistency (22%), now tied with data-leakage concerns (22%). On the surface this reads as an evaluation layer beginning to earn its keep.
The cross-tab says otherwise. Splitting the sample by whether an organization has actually experienced a false-confidence failure, trust divides almost completely. Among the 53 enterprises that shipped an agent which passed evals and then failed a customer, 4% fully trust automated evaluation. Among the 41 that have had no such failure, 24% do — a six-fold difference, and the sharpest split in the dataset. Direct contact with the failure mode is what removes the trust.
This is the month’s central caution. The improvement in sentiment is not evidence that evaluations got better; the failure rate in Finding 1 rules that out. It is what a trust curve looks like when a cohort of less-burned organizations enters the sample and reports its priors. Enterprises reading their own rising confidence as validation of their evaluation stack are reading a number that measures inexperience.
Finding 3: Being burned accelerates autonomy rather than restraining it
85% of the burned are on the zero-human path, against 61% of the REST
We asked whether organizations would let an autonomous agent deploy a code or system change to production on automated evaluation results alone, with no human-in-the-loop validation. The aggregate held; the composition did not.
At the top line, nothing changed: 67% of organizations either already allow zero-human-in-the-loop deployment for low-risk agents (37%) or are actively engineering their pipelines to permit it within a year (30%), against 67% in June. The share ruling it out for the foreseeable future slipped from 22% to 18%. The autonomy ceiling stopped rising, but it did not come down.
Underneath, the picture inverts the intuitive one. Among enterprises that have shipped an evaluation-passing agent that then failed a customer, 85% are on the autonomy trajectory. Among those that have not, 61% are. Organizations with direct, expensive evidence that their evaluations miss things are substantially more likely to be removing the human check, not less — and only 11% of them rule out full automation, against 24% of those that haven't been burned. The pattern is identical for those burned once and those burned repeatedly.
The most plausible mechanism is not recklessness but maturity: the organizations that ship agents at enough volume to hit a customer-facing failure are the same ones with pipelines sophisticated enough to automate, and they are treating the failure as a cost of operating rather than a reason to stop. That is a defensible read. It is also precisely the dynamic that turns Finding 1’s stable failure rate into a growing absolute number of incidents, since the enterprises most likely to fail are the ones scaling their capacity to deploy without review.
One June finding did not replicate. Last month, larger enterprises appeared slightly further down the autonomy path than smaller ones (70% versus 64%). In July the two converge — 65% for organizations with 2,500+ employees against 68% below that, with near-identical failure rates (48% and 50%) — which suggests the June gap was sample variation rather than a size effect. Company size is not what separates the aggressive adopters; experience of failure is.
Finding 4: The stack begins to consolidate
Specialists gain, and the “Nothing at all” share shrinks
We asked which agent reliability or evaluation platform enterprises primarily use today. The field is still crowded, but it is no longer tied at the top with nothing.
The most consequential number is the one that fell. In June, having no dedicated agent-evaluation tooling was tied for the most common answer at 17%; in July it is 12% and fifth. Enterprises are acquiring evaluation tooling, and the specialists are capturing most of that movement: Braintrust nearly doubled its share of primary usage to 15%, and DeepEval reached 17%. Provider-native tooling held roughly flat — OpenAI at 18%, Anthropic at 12% — meaning the growth came at the expense of running nothing rather than at the expense of the model providers.
Counting any use rather than primary platform, the footprints are wider and the ordering is similar: OpenAI native evals reach 31% of enterprises, DeepEval 27%, Braintrust 22%, Anthropic native evals 20%, custom in-house tooling 14%, and Weave and Langfuse 11% each. Nineteen percent still report using no dedicated tooling anywhere in their stack. The category now has three plausible independent contenders where in June it had none with double-digit primary share — the first evidence in this series of an evaluation layer starting to take shape.
Finding 5: Production monitoring still watches the wrong thing
Half monitor whether the agent runs; under a third monitor whether it’s right
Production monitoring for an AI agent can watch two very different things. It can watch whether the system is functioning — is the agent up and responding, did each request complete, how fast, at what cost, with any errors. Or it can watch whether the agent’s output is correct — automated checks that evaluate the content of each answer as it goes out. A confidently wrong answer is invisible to the first kind: the request completes, the response is fast, no error is thrown, and every functioning-metric reads healthy. We asked which kind live production monitoring is built for today.
Grouped by what is actually being watched, the split is essentially June’s: 50% of organizations monitor only whether the agent is functioning, while 26% run automated checks on whether its answers are right. Counting ad-hoc reviewers and don’t-knows, nearly three-quarters of organizations have no automated, real-time evaluation of output correctness in production. Inline quality assertions and transaction trace logging are tied as the most common approach at 26% each on a base of 106 — no single monitoring posture leads.
This is the finding that most directly contradicts the month’s rising confidence. Trust in automated evaluation went up eight points while the runtime capacity to detect an evaluation being wrong went nowhere. Among enterprises that already permit zero-human deployment, only 28% run inline quality checks on production traffic — which means the majority of organizations that have removed the human from the deployment decision have also not replaced that human with anything watching output quality afterward. The gate is automated and the alarm is not installed.
Finding 6: Bought on fit now, not on price
Ease of integration overtakes cost as the top selection factor
We asked what most influenced enterprises’ choice of an evaluation vendor, and what they treat as their primary measure of success. One answer moved sharply; the other did not move at all.
Ease of integration jumped 12 points to 39% and displaced cost as the leading selection criterion, the clearest purchasing shift in the data. Evaluation accuracy rose modestly to 28%, cost fell to 23%, and breadth of observability (6%) and vendor roadmap (2%) remain marginal. Read alongside Finding 4, the two move together: enterprises adopting their first dedicated evaluation tooling are optimizing for what will slot into an existing pipeline this quarter, not for what is cheapest or most capable in the abstract. That is what a market looks like when it stops evaluating and starts installing.
What did not move is what enterprises want from the tool once installed. Evaluation consistency remains the primary success metric at 38%, essentially identical to June’s 36%, well ahead of reduction in failures (20%), speed of experimentation (18%), production visibility (16%), and compliance (7%). The priority is still repeatability — the same verdict on the same behavior every time — which is notable given that bias and inconsistency is now the top-cited trust limitation in Finding 2. Enterprises are buying for integration and measuring for stability, and are not yet getting the second. Satisfaction with current tooling remains moderate, averaging 3.9 on a five-point scale across overall satisfaction, ease of implementation, and value for money, barely changed from June’s 3.8.
Finding 7: Human review becomes the top line item
And the enterprises that have been burned fund it hardest
We asked which reliability and evaluation investment will grow most over the next year. Human review edged into first place.
Human review workflows (31%) and production observability (30%) swapped positions at the top, a change small enough to be noise on its own — but the underlying pattern is the same one June identified and it has strengthened. Enterprises plan to grow spending on human reviewers faster than on the automated evaluation pipelines (19%) that would replace them, at the same moment two-thirds are engineering the human out of the deployment decision. Only 6% report a flat budget, down from 8%.
The cross-tab makes the hedge explicit. Among enterprises that have shipped an evaluation-passing agent that failed a customer, 38% name human review as their fastest-growing investment; among those that have not, 24% do, and they favor observability tooling instead. So the burned cohort is doing both things at once: it is the most aggressive on autonomy (85% on the zero-human path, per Finding 3) and the most committed to funding human reviewers. That is not a contradiction so much as a strategy — automate the deployment decision, and pay people to catch what the automation misses. Whether that scales is the open question, since human review is the one part of the stack that does not get cheaper as agent volume grows.
Finding 8: The switching wave cools
Those planning no change rise from a third to nearly half
We asked whether enterprises plan to adopt a new, additional, or replacement evaluation platform, and which they are considering. Fewer are shopping than last month.
A majority (56%) still intend to adopt a new, additional, or replacement platform within twelve months, but that is down from 64%, and the near-term cohort thinned from 31% to 24%. The share standing pat rose from 36% to 44%. Neither movement clears a significance threshold on its own, but both point the same direction, and they point it consistently with Finding 4: as enterprises actually acquire tooling, the population still looking for it shrinks.
The consideration set has reordered, too. Among the 60 enterprises planning a change, OpenAI’s native evals lead what they are evaluating (20%), followed by Braintrust (18%), Weights & Biases Weave (12%), and DeepEval (10%), with a further 10% actively evaluating but holding no shortlist. DeepEval led June’s consideration set at 20%; it has since converted much of that interest into primary usage, which is what a consideration-to-adoption handoff looks like. Braintrust now occupies the position DeepEval held — high interest ahead of installed base — and is the vendor to watch in the next wave.
The bottom line: Confidence moved, correctness didn’t
June found a gap between the autonomy enterprises were granting their agents and the trust they placed in the evaluations meant to govern it. July finds that gap closing from the wrong side. Trust rose — full confidence in automated evaluation nearly tripled and the complaint that evaluations miss reality fell ten points — while the thing that trust is supposed to track held exactly still. Just under half of enterprises still ship agents that pass their evals and then fail a customer, the same as last month.
The cross-tabs locate the new confidence precisely, and it is not in the evaluations. Twenty-four percent of enterprises that have never had a false-confidence failure fully trust automated evaluation; 4% of those that have do. Trust in this market is a function of exposure, not of evidence. And exposure does not produce caution: the burned cohort is the most autonomous in the sample, with 85% already deploying without human review or building toward it. What it produces instead is a hedge — the same organizations fund human review workflows hardest, at 38%, while removing humans from the deployment gate.
The vendor market is the month’s genuinely encouraging story. Running no dedicated tooling fell from 17% to 12%, specialists gained real share for the first time in this series, buyers shifted from price to integration fit, and switching intent cooled as adoption completed. An evaluation layer is finally forming. But the runtime picture has not followed: half of enterprises stil
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み