AI 失敗企業、人間排除へ加速:評価信頼上昇もリスク増大
本文の状態
日本語全文を表示中
詳細モードで約16分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
VentureBeat AI
VentureBeat の調査によると、AI エージェントがテスト合格後に実環境で顧客に問題を引き起こした企業の半数近くは、自動化評価への信頼を高めつつも、人間による監視役の削減を加速させている。
AI深層分析を開く2026年8月19日 01:16
AI深層分析
キーポイント
自動化評価への信頼と実態の乖離
調査対象企業の13%が自動化評価を信頼すると回答し前月比で増加したが、テスト合格後に顧客に問題を引き起こした事例は依然として49%存在する。
失敗経験が評価プロセスへの不信を生む
AI エージェントの失敗を経験した企業では自動化チェックへの完全な信頼が4%にとどまる一方、同様の事故がない企業では24%に達し、6倍の開きが生じている。
評価セットの縮小と監視手法の転換
Fortune 100 を含む大手企業は複雑化するシステムに対応するため評価セットを縮小・後回しにし、従来の検証から異常検知や問題検出ソリューションへの依存を強めている。
調査の方向性と対象者
今回のデータは確率サンプリングではなく自己選択型の調査であり、108社の企業における最終購買権限者や影響力を持つ人物を対象とした方向性の示唆である。
評価の信頼性と失敗率の乖離
内部テストに合格したシステムが顧客前で失敗する事例は約半数で、合格スコアが生産環境での信頼性を証明しない。
重要な引用
We are seeing the great-decline of evals as we know them
The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance.
It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.
85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one
編集コメントを表示
編集コメント
AI エージェントの導入現場では、テスト通過と実稼働での失敗というジレンマが依然として深刻である。企業は評価ツールの進化に安住せず、複雑化するシステムに対応できる新たな監視体制の構築を急ぐべき局面にある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI エージェントが評価試験を通過したにもかかわらず、本番環境で失敗して企業に損害を与えた事例を経験した組織は、自動化された評価への信頼が高まる中であっても、むしろ人的な介入を排除する方向へ急速に進んでいます。これは VB Pulse の最新調査で明らかになりました。
7 月の調査では、108 社を対象とした結果、13% が自動化された評価を信頼していると回答しました。先月(6 月)の 5% から大幅に増加しています。一方、「テストと実世界の成果との整合性が不十分であること」を最大の懸念として挙げる企業は、前月から 29% から 19% へと 10 ポイント減少しました。
しかし、49% の回答者が「社内でのテストをクリアした AI エージェントや LLM 搭載機能が、その後顧客に問題を引き起こした」と回答しており、これは 6 月の 50% とほぼ変わっていません。また、約 4 分の 1(24%)の企業が、この問題のある事象が複数回発生したと答えています。
VentureBeat Intelligence の最新の調査結果は、エンタープライズ向けエージェント導入におけるより深刻な段階を浮き彫りにしました。もはや課題は「企業がいかに多くの自律性をエージェントに委ねるか」と「いかに検証できるか」の間のギャップだけではありません。評価層への信頼と、「その層が失敗を防ぐ能力を高めているという証拠」の間で、新たな乖離が生じているのです。
最も示唆に富むのは、7 月のデータの中に現れたこの分裂です。
AI の新機能をテストして顧客を失望させた企業では、自動化されたチェックに完全に依存しているのはわずか 4% でした。一方、同様のインシデントを検知しなかった企業のうち、24% が完全な信頼を寄せており、これは 6 倍の差があります。
当然のことながら、本番環境でテストをパスしたエージェントが失敗する経験をすれば、自動化されたチェックプロセスへの疑念を抱くのは無理もありません。
この問題に対処しようとする企業、例えば自動エージェントエラーの監視と軽減プラットフォーム「Raindrop.ai」などは、数ヶ月前から市場が劇的に変化していることを実感しています。
「我々が知る評価(evals)は急速に衰退しています」と Raindrop の CTO ベン・ヒラック氏は VentureBeat に直接メッセージで語りました。「Fortune 100 企業では評価セットの縮小が進み、メンテナンスの優先度は低下しています。システムが複雑化し(MCP やサブエージェントなど)、すべての失敗ケースを列挙することが不可能になったためです。その代わり、本番前と本番後の両方で異常検知や問題検出ソリューションに頼るようになっています。」
これは市場全体の調査ではなく、傾向を示す発見です。
VentureBeat は 7 月、従業員数 100 人以上の企業を代表する 108 人にアンケートを実施しました。これは 6 月の 157 人から減少した数字です。
回答者のうち 69% が最終的な AI 購入権限を持つ立場か、その購入に影響を与える立場であると自己評価しています。サンプルは中堅企業に偏っており、63% の回答者が従業員数 100 人から 2,499 人の企業で働いています。
本調査結果は方向性として捉える必要があります。このアンケートは確率標本ではなく自己選択型であり、記事全体で引用されている「被害を受けた企業とそうでない企業の割合」は各グループが 41〜53 名の回答者から成るものであり、他のクロス集計でも 40〜68 名となっています。
業種構成も変化しました。テクノロジーおよびソフトウェア分野の参加者は 9 ポイント減少し、14% に低下した一方、小売・消費者向けセクターは 4 ポイント増加して 19% となりました。
それでも報告書は、月次で注目すべき 4 つの変化を特定しています。すなわち、「完全な自信がある」と答える回答者が増え、「実社会での整合性の欠如」を指摘する者が減り、「導入の容易さ」を決定的な購入要因として選ぶ人が増え、Braintrust が主要プラットフォームとしてのシェアを獲得したことです。
自動評価への信頼は向上したが、結果は横ばい
VentureBeat の 6 月の調査では、企業における評価のギャップが指摘されました。つまり、エージェントに権限を与えるスピードが、それをテストする信頼できる方法を開発する速度よりも速かったのです。
7 月も、この最初の調査で示された重要な数値を維持しています。2 ヶ月間で集められた 265 の企業回答全体において、「顧客を失望させたテスト承認済みシステム」が少なくとも一つあると報告した割合は、1 ポイント以内にとどまりました。6 月は 50%、7 月は 49% です。
この数値は、エージェントの全実行のうち 49% が失敗したことを意味するわけでも、特定の評価製品が 49% の確率で失敗することを示すものでもありません。調査では、「過去 1 年間に AI 機能が社内テストを通過した後、顧客に直接影響を与えるインシデントが発生したか」を問うています。より多くのエージェントを導入している企業ほど、こうしたインシデントに遭遇する機会が多くなるのは当然です。
しかし、この限界が結果の重要性を低下させるわけではありません。社内評価はリリースのゲートキーパーとして機能します。調査対象企業の約半数が、そのゲートを通過したシステムが後に顧客の前で失敗していることを経験しているのであれば、「合格点」を得たことが生産環境での信頼性を証明するものとはみなせません。
クロス集計結果もこの点を裏付けています。テストの欠落が特定された 41 社のうち 10 社は、自動化に対して完全な信頼を寄せていました。
一方、過去に失敗を経験した 53 社の中で同じように答えたのはわずか 2 社です。リリースゲートが機能不全に陥る可能性を示す証拠が最も少ない回答者ほど、自信を持っていることになります。
失敗を経験した企業は、より迅速にゼロ人間での運用へと移行しています
評価のミスをきっかけに企業がどのような行動を取るのかという点は、直感に反する発見です。
全体として 67% の企業が、特定の低リスクケースでは人の承認なしにエージェントがコードをプッシュしたりシステムを変更したりすることを許可しているか、あるいは来年に向けてその運用を可能にするためのパイプラインの改修を進めています。これは 6 月の調査結果と変わっていません。7 月時点ですでに 37% が限定的なケースでこれを許可しており、残りの 30% もそれを構築中でした。
テスト承認済みシステムが顧客を失望させた企業では、85% が承認不要モデルの導入を進めており、同様の事故を経験していない企業の 61% と比較すると高い数字です。また、今後におけるエンドツーエンドのデプロイ自動化に反対する「被害を受けた」回答者はわずか 11% でしたが、「被害を受けていない」回答者の 24% よりも低くなっています。
これを単なる無謀さとして片付けるのは簡単ですが、データは別の可能性を示唆しています。それは「デプロイの成熟度」です。より多くのエージェントを、より高いボリュームで、より重要なワークフローに展開している組織ほど、失敗に直面する確率が高い一方で、自動化されたデプロイを実現するためのエンジニアリング基盤も整っている傾向があります。
この調査では、どちらの説明が支配的かを断定することはできません。しかし、顧客に見える形で事故が発生しても、自律化への動きが止まるわけではないことは明確です。テストで欠陥を見逃す可能性を肌で知っている回答者ほど、そのテスト結果に基づいて本番環境への変更を許可する動きを最も積極的に進めています。
デプロイごとの失敗率が一定のままデプロイ量が拡大すれば、影響を受ける企業の割合が増えなくても、事故の総数は増加し得ます。7 月のデータは事故の発生件数を計測していないため、このパターンが示唆するリスクは、測定された結果というよりは潜在的な懸念として残ります。
リリースゲートは自動化されましたが、本番環境の品質モニタリングはまだ遅れをとっています
デプロイ前の評価と本番環境のモニタリングは、異なる問いに答えるものです。評価では「エージェントがリリース準備ができているか」を確認しますが、本番モニタリングでは「リリース後にエージェントは何をしているか」、そして「ライブでの出力が正しいかどうか」を問います。
7 月の調査対象企業の多くは、システムが機能しているかどうかには注力していますが、回答の正しさについては重視していません。
この質問に対して有効だった 106 の回答のうち、26% がインライン品質アサーション(ライブトラフィックで出力品質の問題を検出する自動判定器やガードレール)を採用していました。また 26% は、インフラのスパン、トークン使用量、生入力・生出力といったトランザクショントレースに焦点を当てていました。さらに 24% が、レイテンシ、エラー、コストといったゲートウェイ指標の追跡を主に行っていました。
トレースデータやゲートウェイデータは、サービス停止や遅延、リクエストの破綻を検出できます。しかし、流暢で高速でありながら自信満々に誤った回答については検知できない可能性があります。各アーキテクチャが実際に何を監視しているかによって報告を分類すると、半数の回答者が「エージェントが機能しているか」を監視していた一方、自動で「本番出力が正しいか」を監視していたのは 4 分の 1 をわずかに上回る程度でした。
この格差は、限定的なケースでは承認なしでのデプロイを許可している 40 の回答者の中で最も顕著です。そのグループでも、ライブの回答の意味や正しさを自動でチェックしているのはわずか 28% です。つまり、少なくとも一部のリリース決定から人間を排除した企業の多くは、本番環境の最終防衛として自動化された意味品質モニタリングを導入していないのです。
データから得られる最も明確な運用上の教訓は、本番導入前のテストスイートやインフラの可観測性だけでは不十分だということです。これらは特定の失敗モードをカバーするものですが、エージェントが実際のユーザー、データ、ツールと相互作用を開始した後に発生する不良出力を検知するには別の手段が必要です。特に、導入決定を誰かが事前にレビューしない場合、その検知機能は不可欠です。
独立したエージェント評価市場の萌芽
ベンダーからのデータはより希望的な兆候を示しています。企業は専用の評価ツールの導入を進めており、専門プラットフォームのシェアが拡大しているのです。
OpenAI のネイティブな評価機能とトレースが 18% で主要プラットフォームとしてわずかに首位に立ちました。次いで Confident AI の DeepEval が 17%、Braintrust が 15% です。
Anthropic の Claude Console と Workbench は 12% を占め、専用評価プラットフォームを使用していない組織と同率でした。内部開発ツール、Promptfoo、LangSmith はそれぞれ 6% で並んでいます。
Braintrust の主要シェアは 6 月の 8% から 7 月には 15% に増加し、これは単一ベンダーによる最大の伸びであり、レポートでも統計的に有意と指摘されています。DeepEval も 12% から 17% に上昇しました。一方、目的別プラットフォームを使用しないケースは 5 ポイント減の 12% となりましたが、この変化だけでは明確なトレンドを断定するほど十分ではありません。
多くの企業が複数のツールを併用しているため、導入範囲はより広範になります。OpenAI のネイティブ評価機能はスタックの 31% で採用され、DeepEval は 27%、Braintrust は 22%、Anthropic のネイティブツールは 20% に達しました。カスタム内部ツールの採用率は 14% です。一方、Weights & Biases Weave とオープンソースの Langfuse はそれぞれ 11% となっています。
これらは製品のパフォーマンススコアではなく、あくまで導入率を示す数字です。この調査結果から「あるベンダーが他社よりも信頼性の高いエージェントを生み出している」とは言えません。それでも、評価機能は単なる内部スクリプトの寄せ集めや、モデルプロバイダのプラットフォーム内でのみ使われる機能ではなく、明確なエンタープライズソフトウェア層として確立されつつあることを示唆しています。
市場の変化に伴い、購入時の優先順位も移り変わっています。統合の容易さを決定的要因と選んだ割合は 12 ポイント上昇し 39% に達しました。これにより、コストが首位から転落し(28% から 23% へ)、評価精度が 28% で第 2 位となりました。満足度、導入の簡便さ、経済的価値を合わせた平均点は 5 段階中 3.9 です。
価格重視から統合重視へのシフトは、企業が既存の開発・監視パイプラインにすぐに組み込めるツールを求める傾向が強まっていることを示しています。しかし、最も重要な成功指標依旧是評価の一貫性(38%)であり、次いで失敗や回帰の減少(20%)が挙げられています。購入者は適合性を重視しつつも、結果の再現性についても厳しく評価しているのです。
乗り換え意向も冷却傾向にあります。今後 1 年以内にプラットフォームを追加または置き換える予定があると回答したのは 56% で、6 月の 64% から減少しました。
維持する割合は 8 ポイント上昇し、44% に達しました。専門家の採用と合わせると、一部の購入者が評価フェーズから実装フェーズへと移行していることが示唆されます。
人的レビューが自動化のミスを防ぐための保険として機能しつつある
予算データは、より高い自律性と信頼性の低い評価という矛盾を企業がどう管理しているかを明らかにしています。
人的中心型のレビューワークフローが、生産環境の観測(Observability)をわずかに上回り、投資増額で最も頻繁に挙げられた領域となりました。割合はそれぞれ 31% と 30% です。
自動化された評価パイプラインは 19% で 3 位、その後に安全性やポリシー準拠のテスト(16%)が続きます。信頼性向上と評価への予算増額を行わないと答えた企業はわずか 6% でした。
テストでのミスを体験した企業のうち、38% が人的中心型のレビューに最も速い投資成長が見られると回答しました。一方、そのような失敗を経験していない組織では 24% です。
これは一見すると矛盾しています:失敗を喫したグループこそがリリースチェックポイントから人を排除する可能性が高く、一方でプロセスの他の部分における人への支出は増やす傾向があるからです。その戦略は「人間によるバックストップ付きの自動化」です。エージェントに高速化を許容し、自動化された評価で見落としが生じた場合に人間がそれを捕捉するという仕組みです。
このモデルがスケーラビリティを持つかどうかは、まだ未解決の課題です。エージェントの導入や自動チェックはソフトウェアの規模に応じて拡大できますが、レビューに要する人件費は同じ割合で減少しません。その結果、企業は人間の承認ステップを、より大規模な下流レビュー機能に置き換えているだけで、人間による監視自体を排除しているわけではない可能性があります。
狭いながらも重要な示唆
7 月のデータは、エンタープライズ向けエージェントの評価が至る所で失敗していることを示すものでもなければ、自動判定者が劣化していることを証明するものではありません。より正確には、測定された故障発生率が改善する前に、信頼性が上昇していたことがわかります。
同時に、評価を取り巻くインフラも成熟しつつあります。専門ツールの導入を進める企業が増え、統合機能が購買の主要基準となっています。また、すでに失敗を経験した企業は、人的レビューへの投資を強化しています。市場はこの課題を認識しており、対策に資金を投入し始めています。
しかし、信頼性に関する中核的な結果は依然として頑強です。調査対象企業のほぼ半数が、「AI 機能が社内チェックを通過した後でも、顧客を失望させる事例があった」と回答しています。そのような経験を持つRespondents は自動化への信頼を低下させ、人的承認なしでの導入へとより迅速に舵を切っています。
このレポートはこれを「不完全な検証モデル」として位置づけています。デプロイ前の合格スコアは監視の開始点を示すものであり、その終わりではありません。ほとんどの企業にとって、このギャップを埋めるための本番環境での品質チェックや評価テストはまだ整備されていません。
原文を表示
Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.
In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.
Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers, essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once.
The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much autonomy companies give agents and how well they can verify them. It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.
The most revealing split appears inside the July data.
Of the enterprises that experienced an AI feature clear testing only to go on to disappoint a customer, 4% placed complete faith in automated checks. Of those that had detected no comparable incident, 24% expressed full confidence — a sixfold difference.
It makes sense: those who experienced test-passing agents failing in live production are, unsurprisingly, more likely to doubt the automated checking process.
Perhaps it makes sense then, that companies geared toward tackling this problem — like automated agent error monitoring and mitigation platform Raindrop.ai — are seeing the market transform wildly from just a few months ago.
"We are seeing the great-decline of evals as we know them," Raindrop CTO Ben Hylak told VentureBeat in a direct message. "The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production."
A directional finding, not a market census
VentureBeat fielded the July wave among 108 people representing companies with workforces of at least 100. This is down from 157 respondents in June.
Of the 108, 69% described themselves as final AI-buying authorities or people who recommend and influence those purchases. The sample skewed toward midsize organizations: 63% worked at companies with 100 to 2,499 employees.
The findings should be read directionally. The survey is self-selected rather than a probability sample, and the burned-vs.-unburned splits cited throughout this piece rest on groups of 41 to 53 respondents, and other cross-tabs in the report range from 40 to 68.
The industry mix also changed: technology and software participation declined nine points, ending at 14%, while the retail and consumer share added four points and ended at 19%.
The report nevertheless identifies four month-to-month changes worth noticing: more respondents professing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.
Confidence in automated evals improved, but outcomes stayed flat
VentureBeat's June research identified an enterprise evaluation gap: companies were granting agents more authority faster than they were developing reliable ways to test them.
July preserves the key number from that first wave. Across 265 enterprise responses over the two months, the proportion reporting at least one test-approved system that disappointed customers stayed within a single percentage point: 50% in June and 49% in July.
This figure does not mean that 49% of all agent runs fail, or that any particular evaluation product has a 49% failure rate. The survey asks whether an organization experienced at least one customer-facing incident in the previous year after an AI feature passed its internal tests. Companies that deploy far more agents have more opportunities to encounter such an incident.
But that limitation does not make the result less important. An internal evaluation serves as a release gate. If roughly half of surveyed organizations have seen that gate approve a system that later fails in front of customers, a passing score cannot be treated as proof of production reliability.
The cross-tab reinforces the point. Ten of the 41 enterprises with no identified testing miss placed complete faith in automation.
Only two of the 53 previously burned enterprises said the same. Confidence is strongest among respondents with the least evidence that the release gate can fail.
The enterprises that got burned are moving faster toward zero-human deployment
The counterintuitive finding is what companies do after an evaluation miss.
Overall, 67% either let an agent push code or change a system without a person's approval in certain low-risk cases, or are modifying their pipelines to support that practice during the coming year. That is unchanged from June. In July, 37% already permitted it in limited cases and another 30% were building toward it.
Among enterprises where a test-approved system had disappointed a customer, however, 85% were pursuing that no-approval model, compared with 61% in the group reporting no comparable incident. Only 11% of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.
It would be easy to read that as recklessness, but the data supports another plausible explanation: deployment maturity. Organizations running more agents, at higher volume and across more consequential workflows, are more likely both to encounter failures and to have the engineering infrastructure needed for automated deployment.
The survey cannot establish which explanation dominates. It does establish that a customer-visible incident does not appear to stop the move toward autonomy. Respondents with firsthand proof that testing can miss defects are also moving most aggressively to let those tests authorize production changes.
If their per-deployment failure rate remains constant while deployment volume rises, the total incident count could grow even without the percentage of affected companies increasing. The July data does not measure incident volume, so that remains a risk implied by the pattern rather than a measured outcome.
The release gate is automated, but production quality monitoring still lags
Pre-deployment evaluation and production monitoring answer different questions. An evaluation asks whether an agent appears ready to ship. Production monitoring asks what the agent is doing after release and whether its live outputs remain correct.
Most companies in the July sample still emphasize whether the system functions, not whether the answer is correct.
Among the 106 valid responses to this question, 26% used inline quality assertions — automated judges or guardrails checking live traffic for output-quality problems. Another 26% focused on transaction traces such as infrastructure spans, token usage and raw inputs and outputs, while 24% mainly tracked gateway metrics such as latency, errors and cost.
Trace and gateway data can reveal outages, slowdowns and broken requests. They may not flag a fluent, fast and confidently wrong answer. Grouped by the report according to what each architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.
The gap is sharpest among the 40 respondents already permitting no-approval deployment in limited cases. Only 28% of that group automatically checked the meaning and correctness of live answers. In other words, most enterprises that have eliminated a person from at least some release decisions have not installed automated semantic-quality monitoring as the production backstop.
This is the clearest operational lesson in the data. A pre-deployment test suite and infrastructure observability are necessary, but they do not cover the same failure mode. Enterprises need a way to detect bad outputs after the agent begins interacting with real users, data and tools — especially when nobody reviews the deployment decision first.
An independent agent-evaluation market begins to take shape
The vendor data offers a more encouraging sign: enterprises are adding dedicated evaluation tools, and specialist platforms are gaining ground.
OpenAI's native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI's DeepEval at 17% and Braintrust at 15%.
Anthropic's Claude Console and Workbench held 12%, tied with organizations reporting no dedicated evaluation platform. Three options each held 6%: internally built tools, Promptfoo and LangSmith.
Braintrust's primary share increased from 8% in June to 15% in July, the biggest gain by one vendor and the one the report flags as statistically significant. DeepEval rose from 12% to 17%. Use of no purpose-built platform declined five points to 12%, although that smaller change does not by itself confirm a trend.
Because many companies use more than one tool, the broader footprints are larger. OpenAI native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic's native tooling in 20%. Custom internal tooling reached 14%, while Weights & Biases Weave and open-source Langfuse each reached 11%.
These are adoption figures, not product-performance scores. The survey does not establish that one vendor produces more reliable agents than another. Still, the results point toward evaluation becoming a distinct enterprise software layer rather than a loose collection of internal scripts or a feature used only inside a model provider's platform.
Purchasing priorities are changing with that market. The proportion choosing integration ease as the decisive factor climbed 12 points to 39%, displacing cost, which fell from 28% to 23%. Evaluation accuracy ranked second at 28%. The combined average for satisfaction, implementation simplicity and economic value was 3.9 out of five.
The move from price toward integration suggests enterprises increasingly want a tool they can install into existing development and monitoring pipelines now. Yet their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%. Buyers are selecting for fit while still judging results on repeatability.
Switching intent also cooled: 56% still expected to add or replace a platform during the coming year, down from 64% in June.
The proportion staying put increased eight points, reaching 44%. Together with specialist adoption, they suggest some buyers are moving from evaluation to implementation.
Human review is becoming the hedge against automated misses
The budget data reveals how enterprises are managing the contradiction between greater autonomy and unreliable evaluation.
People-centered review workflows edged narrowly ahead of production observability as the most frequently cited area for increased investment, 31% to 30%.
Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing.
Among enterprises that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, compared with 24% of organizations that had not been burned.
That produces an apparent paradox: the burned group is most likely to remove people from the release checkpoint and most likely to increase spending on people elsewhere in the process. The strategy appears to be automation with a human backstop — allow agents to move faster, then use reviewers to catch what automated evaluation misses.
The open question is whether that model scales. Agent deployments and automated checks can grow with software volume. Reviewer hours do not fall at the same rate. Enterprises may therefore be replacing a human approval step with a larger downstream review function rather than eliminating human oversight.
The narrow but consequential read
July's data does not show that enterprise agent evaluation is failing everywhere, nor does it prove automated judges are getting worse. It shows something more precise: confidence rose before the measured failure incidence improved.
At the same time, the infrastructure around evaluation is maturing. More enterprises are adopting specialist tools, integration has become the leading purchase criterion and companies that have already experienced failures are increasing investment in human review. The market recognizes the problem and is spending against it.
But the central reliability result remains stubborn. Nearly half of surveyed enterprises still report that an AI feature cleared internal checks before disappointing a customer. Respondents with that experience place less faith in automation — and move faster toward deployments with no human approval.
The report frames this as an incomplete verification model: a passing pre-deployment score marks the start of monitoring, not the end of it. For most enterprises, the production quality checks and evaluation testing that would close that gap still aren't in place.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み