Arize AI、AI エージェントの生産環境移行を評価駆動開発で加速する方法を解説
本文の状態
日本語全文を表示中
詳細モードで約25分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Arize AI Blog
CVS Health の事例を通じて、AI エージェントをパイロットから本番環境へ移行させるには、モデル性能だけでなく評価基盤や観測性などの周辺システム整備が不可欠であると示している。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月14日 01:32
AI深層分析
キーポイント
パイロットの停滞要因
企業向け AI プロジェクトは、モデル自体の改善ではなく、リリース基準や監視体制といったソフトウェア配信の運用基盤が追いつかない点で本番化に失敗する。
CVS Health の加速事例
Matt Turner 氏と Lagan Khare 氏は、明確な仕様書や評価ハッチ、コスト管理指標の導入により、機能開発を半日で完了させ、レガシー書き換えを数ヶ月で達成した。
評価駆動型開発の必要性
機械速度でコードやワークフローが生成される環境では、評価と観測性が開発プロセスの核心となり、品質やリスクを管理しながらエージェントに自律性を付与する根拠となる。
パイロット・パurgatory の回避
有望なシステムが組織側の準備不足で待機状態になる「パイロット・パurgatory」を防ぐには、モデルのバージョンアップを待つ前に周辺システムの成熟度を点検する必要がある。
AIネイティブなソフトウェア開発ライフサイクルの定義
エージェントを開発、調整、検証に活用しつつも、意図・証拠・リリース決定における人間の所有権を維持するプロセスである。
重要な引用
Enterprise AI pilots stall when the surrounding operating system of software delivery cannot evaluate, govern, and absorb work at the pace a model produces it.
When a pilot sits for more than a quarter, Turner argued, teams should inspect the surrounding system before assuming that another model release will rescue the project.
"The spec is as much for the engineer as it is for the agent building it"
"Production becomes an engineering path that teams need to design before the first agent begins making consequential decisions."
編集コメントを表示
編集コメント
本番環境への移行を阻む最大の要因がモデル性能ではなく運用基盤であるという指摘は、現場のエンジニアにとって極めて示唆に富んでいる。Arize AI の事例分析は、単なるデモを超えて持続可能な AI 導入を実現するための具体的なロードマップを提供している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
CVS Health のチームは、ある機能のアイデアから本番環境への移行を約 1 日半で完了させました。また、当初 6〜9 か月と見積もられていたレガシーシステムの書き換えも、わずか 1 ヶ月で終わらせています。
Arize Observe 2026 の登壇者である Matt Turner 氏と Lagan Khare 氏は、仕様定義、評価ハーン(テスト基盤)、AI オブザーバビリティ、ガバナンス、そして成果あたりのコスト指標といった要素が、こうしたスピードを持続可能にする鍵であると解説しています。
エンタープライズにおける AI プロジェクトは、説得力のあるデモから信頼性の高い本番システムへと移行する段階で立ち止まることがよくあります。この局面では、モデルの性能だけでなく、評価、AI オブザーバビリティ、ガバナンス、コスト管理が同等の重みを持ちます。チームはすぐに、これらの機能は構築後に安易に追加できるものではないと気づきます。
CVS Health でアーキテクチャ担当ディレクターを務める Matt Turner 氏と、エンジニアリングおよびイネーブルメント担当エグゼクティブディレクターの Lagan Khare 氏は、このギャップを埋めるための実践を長年培ってきました。Turner 氏によると、チームは機能の開発から本番導入まで約 1 日半で完了させ、6〜9 か月かかるはずだったレガシーシステムの書き換えも 1 ヶ月で成し遂げました。
この劇的な加速を実現した背景には、明確な仕様定義、信頼できるテスト、本番環境でのモニタリング、そして測定可能なビジネス成果によって支えられた、規律ある AI ネイティブなソフトウェア開発ライフサイクルがあります。これこそが、そのスピードを再現可能にするための基盤です。
この経験は、エンジニアリングチームにとってより広範な教訓を示しています。AI の提供速度は、モデルを取り巻くシステムに依存します。開発者が機械的なスピードでコード、プロトタイプ、ワークフローを生成するようになると、評価駆動型開発と観測可能性がプロセスの核となります。これらを組み合わせることで、チームは品質、コスト、リスクの管理を失うことなく、エージェントにより大きな自律性を付与するための根拠を得ることができます。
なぜエンタープライズ AI のパイロットは本番環境まで到達できないのか
ソフトウェアデリバリーの運用システムが、モデルが生み出すペースで作業を評価・ガバナンス・吸収できない場合、エンタープライズ AI のパイロットは本番環境への移行前に立ち止まります。デモ室では、リリース基準や本番監視、所有権、アクセス権、投資対効果といった難しい問いに答えが出るまで、洗練されたインタラクションが能力を証明してしまうため、このミスマッチが隠されてしまいがちです。
ターナー氏はセッションの冒頭、「すぐに役立つように見える AI 支援デモを見たことがある人は?」と尋ねた後、「本番環境で運用され、測定可能な価値を生み出したシステムはありますか?」と続けました。これらの問い合わせによって、彼が「パイロット・プurgatory(パイロット煉獄)」と呼ぶ状態が浮き彫りになりました。これは有望なシステムが、それを取り巻く組織の準備が整うのを待つ状態です。
モデルの改良だけでは、欠落した規律を解消することはできません。評価、ガバナンス、デプロイ制御、そして測定は、意図的なエンジニアリング作業を通じて導入されるものであり、より強力な基盤モデルを用意しても、準備が整っていないプロセスに出力が増えるだけになります。ターナー氏は、パイロットプロジェクトが 1 クォーター以上続いている場合、次のモデルリリースでプロジェクトが救われると考える前に、周囲のシステムを検査すべきだと主張しています。
カーア氏は、この問題を経済的持続性(durability)という観点から捉えました。これは、デモで機能が確立され、チームが出荷できることが示された後も、運用モデルが維持され続ける状態を指します。持続的な価値は、利用規模の拡大、ソースデータのシフト、コストの累積、そして新たに顕在化する障害モードに対してそのモデルが耐え抜いたときに生まれます。
この区別は、開発者が問うべき問いを変えます。プロダクション化とは、エージェントが重要な意思決定を下し始める前に、チームが事前に設計すべきエンジニアリングへの道筋なのです。
AI ネイティブなソフトウェア開発ライフサイクルとは何か?
AI ネイティブなソフトウェア開発ライフサイクルでは、個別の開発、チームの調整、そして生産環境での検証においてエージェントを活用しつつも、意図・証拠・リリース決定における人間の所有権は維持されます。ターナー氏は、規律を構成する 3 つの層について説明しました。それぞれの層が新たなレバレッジを生み出す一方で、新しい種類の失敗をもたらすことになります。
第 1 レイヤー:AI コーディングツールと仕様書による個人開発者の加速化
個人生産性の向上は、コード支援ツールが即座に実感のある成果をもたらすため、AI エージェント導入の入り口としてよく選ばれます。エンジニアは、既存のコードベースを検索し、サービスの骨組みを構築し、動作するプロトタイプを作成し、従来の引き継ぎが行われる前にテスト草案まで作成できます。CVS Health では、汎用的なコーディングツールが、共有された Model Context Protocol (MCP) サーバーやソース管理・チケット管理システムとの連携、そして反復的なプラクティスをコード化した再利用可能なスキルと共に活用されています。
内部で「Clarify」と呼ばれるスキルの一つは、実装が始まる前に製品側とエンジニアリング側の意図を一致させる役割を果たしています。このスキルの目的は、エージェントと人間が評価できる仕様へと向かう大きな転換点を反映するものです。つまり、仕様とは構造化された要件、エッジケース、受入基準、システムコンテキストを含む、永続的な真実の源となるべきものなのです。
「仕様は、それを作るエージェントだけでなく、エンジニアにとっても同じくらい重要だ」とターナー氏は述べています。
この義務が重要な理由は、エージェントがコードをコンパイル可能な状態に出力しても、機能の本質を誤解している可能性があるからです。仕様を作成することで、エンジニアは作業を委任する前に意図した動作を明確に言語化することになり、レビュー担当者は実装内容を判断するための安定した基準を得ることができます。良質な仕様がライブラリとして蓄積されていけば、以前のコンテキストや契約、意思決定が利用可能になるため、新機能の構築もより迅速に行えるようになります。
ターナーは、ステークホルダーがその場にいる状態でチームが生きたプロトタイプを構築する様子も目撃しています。作業中のシステムが会話の最中に完成するため、製品・デザイン・エンジニアリングの各部門は、仮説がバックログ項目やアーキテクチャ上のコミットメントとして固定される前に検証できます。このプロトタイプは集団的な推論のための道具となり、従来の会議後の引き継ぎが行われる前にアイデアの有効性を確認する機会を提供します。
しかし、こうした加速には罠も潜んでいます。バイブコーディング(直感的なコーディング)では、開発者が理解していないコードを承認してしまうリスクがあります。また、エージェントが生成したテストは、実装にすでに埋め込まれた仮説を確認するだけになる恐れもあります。大規模なタスクは利用可能なコンテキストを圧迫し、最終的にエージェントの整合性が失われるため、エンジニアは明確な意図を持つ小さな実行単位へと作業を分解する必要があります。また、チーム間で実践が積み重なるためには、共通のツールに十分な一貫性が必要です。孤立したツールの選択では、あるグループの学習が組織的なインフラとして定着しないからです。
レイヤー 2:共有コンテキストと依存関係マップによるエンジニアリングチームの整合
個々の出力が増加すると、調整こそが進行速度を決定するシステムとなります。製品の意図やデザイン判断、サービス契約、依存関係は、人間の頭の中だけでなく、人間とエージェントの両方がアクセスできるアーティファクトとして明確化されなければなりません。文書化されていない知識は、即座にエージェントのパフォーマンスの限界となるからです。
このレイヤーにおいて、仕様書には第二の機能が生まれます。製品・デザイン・エンジニアリングの各チームが共通の受入基準で結束し、生成されたコードが手戻りになる前に課題を可視化できるのです。
ターナー氏は、すべての仕様書を完璧な記念碑のように磨き上げるべきではないと警告しています。有用な成果物に必要なのは、迅速な反復開発を支えるための十分な構造と、チームもエージェントも明確に推論できる範囲の狭さだけです。
依存関係のマッピングも同様に重要です。エージェントによる変更は、組織の記憶が追跡できる速度よりも速くサービス境界を越える可能性があるからです。「エージェントが変更によって何が壊れるかを知らない場合、それは必ず壊します」とターナー氏は言います。
依存関係マップがあれば、エージェントは契約・データフロー・共有コンポーネントを変更する前に、影響範囲(ブラスト・レイジ)を把握できます。この文脈がないと、生成プロセスで節約された時間が、クリーンアップやインシデント対応、他チームとの調整に消えてしまう恐れがあります。
AI はまた、役割間の距離も短縮します。プロダクトマネージャー、デザイナー、エンジニアが同じプロトタイプと仕様書を共有しながら並行して作業できるようになるのです。ターナー氏は、複数のチームがこの新たなあり方を「ビルダー」と呼んでいると聞いています。この用語は、機能の定義・設計・実装という境界線がより流動的になっている様子を捉えています。
組織上の制約は、開発チームから AI への引き継ぎが終わっても消えるわけではありません。ターナー氏は、開発者が数時間で完了させた機能でも、ユーザー検証や製品レビュー、あるいは依存する他チームの対応が遅れるために、数週間も停止してしまう事例を挙げました。AI は既存の調整システムを増幅させるため、健全なチームでは並行処理能力が高まる一方で、分断された組織(サイロ)はより早期に衝突を起こすことになります。
レイヤー 3:評価と観測性がエージェントの自律性を支える
本番環境での自律性は、システムが正しく動作しているという証拠と、挙動に変化が生じた際に停止できる仕組みによって担保されます。ターナー氏はこのレイヤーを、信頼できるテスト、エージェントの観測性、段階的なロールアウト、行動分析ダッシュボード、そしてロールバック機能といった要素がスピードアップのための必須条件となる地点として位置づけました。
決定論的(データーミスティック)なテストは基盤となりますが、テストが意味のある挙動を反映していない場合、カバレッジ率の数字だけでは安心できません。AI はテストケースの提案やコード草案作成を担いますが、そのテストがビジネスやユーザーにとって重要な失敗を検出できるかどうかを確認するのは依然として人間の責任です。
エージェント型ワークフローには非決定性(ノンデーターミスティック)な要素が含まれるため、出力結果やエンドツーエンドの挙動が意図した成果を満たしているかを検証できる評価ツールが必要です。プルリクエストでは行ごとのレビューで妥当に見えても、本番展開後に予期せぬ挙動を引き起こす可能性があります。そのため、レビューの焦点は「コードそのもの」から、「意図」「システムへの影響」、そしてトレースや評価結果によって裏付けられた証拠へとシフトします。
ユーザー検証は開発プロセスに組み込むべきです。システムが意図した問題を解決しているかをユーザーが確認できなければ、迅速なリリースは価値を生みません。また、フィードバックのサイクルを伴わない高速実装は、効率よく作られた廃棄物へと転落する恐れがあります。
この検証層を構築する前に生成量を増やしてしまうと、発生した不具合の特定が困難になり、収束コストも高騰します。規制環境下では、検出されなかったエラーがコンプライアンス違反事件に発展しかねません。つまり、自律性の上限を決めるのはモデルの能力だけでなく、説明可能性と監査証跡の整備状況にも直結しているのです。
AI 評価ハネスには何が含まれるべきでしょうか?
AI 評価ハネスには、許容される動作を定義する「ゴールデンデータセット」、回帰テスト用の「リグレッションスイート」、そしてドリフトを検知する継続的なプロダクションモニタリングが必要です。これらを組み合わせることで、チームは「リリースに耐えうる品質か」「リリース後もその品質が維持されているか」を反復して検証できるようになります。
Khare は「エージェントを開発する前にハネスを構築せよ」と断言しています。
Khare は生成 AI システムの評価を、従来のソフトウェア工学におけるユニットテストに例えています。このアナロジーは特に、開発の初期段階で評価を組み込む場合に有効です。なぜなら、ケースと閾値を設定することで、プロンプトチューニングやモデル選定が議論を支配する前に品質基準を明確に定義できるからです。
ゴールデンデータセットとリグレッションスイートは、リリース品質を定義します
ゴールデンデータセットは、リリースを不可とするような失敗や既知のエッジケース、期待される動作など、重要なケースを網羅的に記録します。回帰テストスイートは、エージェントやその周辺システムに変更があった際にこれらのケースを再実行することで、得られた教訓を維持・保存する役割を果たします。
この実践により、評価活動が設計プロセスの一部となります。どの事例を恒久的に保護すべきかをチームで決定する際、意図に関する見解の相違が表面化しますが、その段階であれば解決コストは低く抑えられます。結果として得られるスイートは、手動で選ばれたデモ数個に依存せず、リリース判断を裏付ける確かな根拠となる証拠の集合体を提供します。
人間の判断は依然として不可欠です。評価者が弱ければ、テストが許容する浅い行動を報酬として与えてしまうリスクがあるためです。チームは、事例、採点基準、失敗閾値について、エージェントの出力に複数の正解が存在する場合など特に注意が必要ですが、本番コードに対するのと同じ厳格さでレビューを行う必要があります。
継続的な評価では、モデル、データ、ユーザー行動のドリフトを検出します。
継続的評価はハッチネスを生産環境へ拡張するもので、ここではモデル、ソースデータ、ユーザー行動がすべてドリフトする可能性があります。ローンチ時のスコアは特定の瞬間を捉えるに過ぎませんが、運用中の評価プログラムによって、そのスコアが今後 30 日間でどのように変化するのか、またどのトラフィックセグメントが変化の要因となっているかが明らかになります。
この履歴情報により、開発者は広範な回帰現象と、ツール呼び出しの失敗や、新たに曖昧になったユーザー要求、あるいは単一のデータソースにおける品質問題といった局所的な障害を区別できるようになります。トレースや本番環境でのサンプルは新たなケースとしてゴールデンセットにフィードバックされ、システムと共にハッチネス(評価基盤)も成熟していきます。
このフィードバックループこそが、AI エージェントの評価に実用価値をもたらします。評価駆動型開発においては、ハッチネスはチームが元のテストセットでは想定していなかった条件下でエージェントがどのように振る舞うかを学ぶための手段となります。また、新たに観測された失敗一つひとつが、次のリリース判断を鋭くする役割を果たします。
チームは AI エージェントのコストと ROI をどのように測定すべきでしょうか?
チームは「成果あたりのコスト」を通じて AI エージェントの経済性を測定し、その数値をサイクルタイム、エラー率、取引あたりのコスト、解決までの時間といったビジネス KPI と結びつけるべきです。トークン使用量はインフラ分析には有用ですが、信頼性の低いワークフローが引き起こす再試行、手直し、エスカレーション、レビューの負担といった要素は捉えきれません。
成果あたりのコストは、再試行や手直し、レビューを包括する
Khare 氏は、出力に追加的なクリーンアップが必要な場合、安価なモデルであっても運用システム全体のコストが高騰しうるため、「成果あたりのコスト」が意味のある単位であると主張しています。推論時に数ペンス節約できたとしても、人間による検証や修正、再実行、エスカレーションが必要になれば、下流工程で何ドルもの追加費用が発生する可能性があります。
エージェントが担当する業務に合わせた指標を設定する必要があります。サポートワークフローなら「解決した問い合わせあたりのコスト」、ドキュメント処理なら「処理した文書あたりのコスト」、内部自動化なら「完了したタスクあたりのコスト」などを追跡します。これらの指標は、実用的な結果に至るまでの全工程を反映するため、業務の成否と切り離された月次のトークン請求額よりもはるかに有用です。
したがって、本番環境への移行計画には、機能ごとの支出上限、フォールバック経路、そして支出状況の可視化を含めるべきです。これらの制御により、エンジニアリングチームと財務部門は、実際の負荷下でのコスト挙動を理解できます。特にリトライや例外処理が頻発する場面では、パイロット段階での経済性が本番環境で誤解を招く恐れがあるため、こうした対策が不可欠です。
ビジネス KPI は AI 導入プロセスに組み込むべきです
本番候補となるシステムは、開発フェーズに入る時点で関連するビジネス指標を既に持っている必要があります。Khare氏は、AI システムと業務上の価値を結びつける具体例として、サイクルタイム、スループット、ドキュメント処理の turnaround(回転時間)、エラー率、取引あたりのコスト、解決までの時間を挙げています。
導入プロセスで指標が明確になれば、エンジニアリングチームは組織が期待する成果に基づいて評価スイートや本番監視を設計できます。同じ指標は優先順位付け、リリース判断、ローンチ後の分析にも活用され、予算執行後に ROI(投資対効果)を後付けで語るという事態を防ぎます。
Khare のルールはシンプルだった。モデルよりもまず指標が重要だと。
エンタープライズ規模になると経済性が変わるため、財務部門はこの議論に早期から参画すべきだ。パイロット段階では対象ユーザーを限定し、複雑なケースをシステム外へ迂回させることで一見安価に見えるかもしれないが、本番環境ではリトライや人的レビュー、フォールバックモデル、データアクセス、継続的な監視などにかかる真のコストが露呈する。
「成果あたりのコスト」は、財務とエンジニアリングの両者がワークフローを拡大する価値があるかどうかを判断するための共通単位となる。
どの部分がモデル進化に耐えるのか?
エージェント・ハネスの中で、モデルの能力向上後も価値を保ち続けるのは、データ、ワークフロー設計、評価スイート、権限管理、そして信頼基盤だ。チームは、カスタムオーケストレーション、検索機能、ルーティング層を厳しく検証すべきだ。なぜなら、将来的なモデルや API が現在では広範なインフラを必要とする機能を吸収してしまう可能性があるからだ。
「モデルがあなたのスタックを食う」と Khare は語った。これは、機能が周囲のインフラからモデル層へと急速に移行する様子を指している。
この洞察は、モデルが変化しても価値を持ち続けるシステムの一部へのエンジニアリング投資の方向性を示す。ドメインの意図を捉え、アクセスを制限し、トレースを記録し、成果を測定できるワークフローなら、基盤となるモデルが変わっても価値を生み出し続けることができる。一方、一時的な制約を取り巻いて作られた脆い層は、短期間でメンテナンスコストに転化しかねない。
データ品質、ワークフロー設計、評価、そして信頼こそが、長く持続する要素だ。
エージェントは規模の拡大に伴い曖昧さを増幅させるため、データ品質がクリティカルパスに位置づけられます。熟練したアナリストなら文脈から解釈できる不揃いな列でも、エージェントにとっては確信を持った誤りへと誘導し、その結果生じた回答がミスの原因が誰にも気づかれる前にワークフロー全体に波及する恐れがあります。
そのため、クリーンなソースデータ、関連するコンテキスト、明確に定義されたセマンティクスは、本番環境のエンジニアリングにおいて引き続き不可欠です。この作業には新しいモデルのデモのような派手さはないかもしれませんが、システムがそのモデルを安全に活用できるかどうかを決定づける重要な要素です。
ワークフロー設計もまた、モデルが組織の運用コンテキストを自動的に継承しないため存続します。チームは依然として、どの課題に自動化を適用すべきか、責任がシステムと人の間でどのように移行するか、承認はどこで行うべきか、行動にはどのような証拠を添えるべきかを決定しなければなりません。評価スイートはこれらの判断を実行可能な期待値として保存し、信頼性のファブリックはそれらを実行するために必要な権限管理と監査可能性を担保します。
なぜ AI ワークフロー自動化が下流のボトルネックを生むのか?
AI ワークフロー自動化は、ある工程が加速しても、その出力を受け取るチームやシステムのキャパシティがそれに伴って増強されない場合、下流にボトルネックを生み出します。AI は大規模なプロセスの一部しか担当しないことが多いため、局所的な処理量は向上しても、エンドツーエンドの完了時間が逆に悪化することがあるのです。
Khare氏は、チームAのエージェントが連続稼働してチームBをタスクで埋め尽くしてしまった事例を紹介しました。最初のグループはより高速なローカルプロセスを指し示しましたが、組織全体としては単にキューの場所を変えただけで、処理能力が低い場所に押しやったに過ぎませんでした。
このパターンは、下流チームが追加作業を受け入れられる低負荷のパイロット段階では見落としがちです。しかし、本番環境に移行するとシステム全体の形状が変わります。24時間365日ケースやドキュメント、コード変更、レビュー要求を生成するワークフローは、まだ少数の専門家に依存している工程を圧倒してしまいます。
自動化する前に、フルワークフローのマッピングを行う
自動化の対象を選ぶ前に、チームはアップストリームからの入力、ダウンストリームのキュー、承認ポイント、例外処理パスを含む完全なワークフローをマッピングする必要があります。このマップにより、負荷増加がどこに集中するか、別の工程が新たなボトルネックになるかが明らかになります。
この演習は優先順位の決定も変えます。最も簡単な手作業であっても、限られたリソースを持つレビュアーに依存する工程であれば、システム全体としての価値は低くなります。一方、一見すると目立たない介入でも、ハンドオフを排除したり、ソースデータを改善したり、エスカレーションが必要なケース数を減らしたりできる可能性があります。開発者は、ユーザー、レビュアー、コンプライアンス機能、依存チーム、その他のソフトウェアを含むエンドツーエンドのワークフロー全体を通じて自動化を評価する必要があります。
このシステム全体の視点こそが、ビジネス指標を守ることにもつながります。チームがワークフロー全体を通じてサイクルタイムや完了タスクあたりのコストを測定する場合、局所的な速度向上だけで全体的なプロセスが遅くなっていることを隠すことはできません。
規制環境下で AI エージェントのガードレールはどのように機能するのでしょうか?
規制環境における AI エージェントのガードレールは、エージェントが物理的に実行できる行為を制限する機械的な制御と、意図した行動をレビュー可能にするテキストベースの制御を組み合わせたものです。Khare 氏は、両方のレイヤーがアーキテクチャに不可欠だと主張しています。その理由は、失敗に伴うコストが非対称だからです。消費者向けアプリケーションで単なる再試行を招くだけの不具合でも、規制されたワークフローでは監査リスクの露呈や通知義務の発生、あるいはデータ侵害につながる可能性があります。
機械的なガードレールは生産アクセスを制限する
機械的なガードレールは、データ、ツール、および行動の周りに境界線を設けるために決定論的な制御を使用します。これには、情報へのハードチェックが含まれます
原文を表示
CVS Health teams moved a four-week feature from idea to production in roughly a day and a half, then completed a legacy rewrite in one month after estimating six to nine months. In their Arize Observe 2026 talk, Matt Turner and Lagan Khare explain how specifications, evaluation harnesses, AI observability, guardrails, and cost-per-outcome metrics can make that speed durable.
Enterprise AI projects often stall at the point where a convincing demo has to become a dependable production system. Evaluation, AI observability, governance, and cost controls suddenly carry as much weight as model performance, and teams quickly discover that these capabilities cannot be added casually after the build.
At CVS Health, lead director of architecture, Matt Turner, and executive director of engineering and enablement, Lagan Khare, have spent years developing the practices required to close that gap. Turner said their teams had taken a feature from idea to production in roughly a day and a half, and completed a legacy rewrite in one month after estimating six to nine months. Although the acceleration was dramatic, making it repeatable required a disciplined AI-native software development lifecycle supported by clear specifications, trusted tests, production monitoring, and measurable business outcomes.
Their experience points to a broader lesson for engineering teams: AI delivery speed depends on the systems surrounding the model. As developers generate more code, prototypes, and workflows at machine speed, evaluation-driven development and observability become core parts of the development process. Together, they provide the evidence teams need to grant agents greater autonomy without losing control of quality, cost, or risk.
Why enterprise AI pilots stall before production
Enterprise AI pilots stall when the surrounding operating system of software delivery cannot evaluate, govern, and absorb work at the pace a model produces it. The demo room tends to conceal that mismatch because a polished interaction can prove capability before anyone has settled the harder questions about release criteria, production monitoring, ownership, access, or return on investment.
Turner began the session by asking how many people had seen an AI-assisted demo that looked immediately useful, then followed with a question about systems that had reached production and produced measurable value. Together, those questions framed what he called “pilot purgatory,” a state in which a promising system waits for the organization around it to become ready.
Model improvements do little to resolve the missing discipline. Evaluation, governance, deployment controls, and measurement arrive through deliberate engineering work, while stronger base models can simply deliver more output into an unprepared process. When a pilot sits for more than a quarter, Turner argued, teams should inspect the surrounding system before assuming that another model release will rescue the project.
Khare approached the same problem through durability, which she described as the operating model that continues to hold after a demo has established capability and a deployment has shown that the team can ship. Durable value emerges when that model survives growing usage, shifting source data, accumulating costs, and newly visible failure modes.
That distinction changes the question developers should ask. Production becomes an engineering path that teams need to design before the first agent begins making consequential decisions.
What is an AI-native software development lifecycle?
An AI-native software development lifecycle uses agents across individual development, team coordination, and production validation while preserving human ownership of intent, evidence, and release decisions. Turner described three layers of discipline, each of which creates a new form of leverage and a new class of failure.
Layer 1: AI coding tools and specifications accelerate individual developers
Personal productivity is the usual on-ramp because coding assistants produce immediate, tactile gains. An engineer can explore a codebase, scaffold a service, generate a working spike, and draft tests before a conventional handoff would have begun. At CVS Health, general coding tools sit alongside shared Model Context Protocol (MCP) servers, source-control and ticketing integrations, and reusable skills that encode repeatable practices.
One internal skill, called Clarify, helps align product and engineering intent before implementation starts. Its purpose reflects a larger shift toward specifications that agents and humans can evaluate: the specification becomes a durable source of truth containing structured requirements, edge cases, acceptance criteria, and system context.
“The spec is as much for the engineer as it is for the agent building it,” Turner said.
That obligation matters because an agent can produce code that compiles while misunderstanding the feature. A specification forces the engineer to articulate the intended behavior before delegating the work, which gives reviewers a stable artifact against which they can judge the implementation. As a library of well-formed specifications grows, new features become faster to scaffold because previous context, contracts, and decisions remain available.
Turner has also seen teams prototype live while stakeholders remain in the room. Because the working system arrives while the conversation is still active, product, design, and engineering can test assumptions before they harden into a backlog item or an architectural commitment. The prototype becomes an instrument for collective reasoning, allowing the team to validate the idea before a conventional post-meeting handoff would return.
The same acceleration creates traps. Vibe coding encourages developers to approve code they do not understand, while agent-generated tests may confirm the assumptions already embedded in the implementation. Large tasks can overwhelm the available context until the agent loses coherence, so engineers need to decompose work into smaller, executable units with clear intent. Shared tooling also needs enough consistency for practices to compound across teams, since isolated tool choices prevent one group’s learning from becoming organizational infrastructure.
Layer 2: Shared context and dependency maps align engineering teams
Once individual output rises, coordination becomes the pacing system. Product intent, design decisions, service contracts, and dependencies must leave people’s heads and enter artifacts that humans and agents can both access, because undocumented knowledge becomes an immediate ceiling on agent performance.
Specifications gain a second function at this layer. They align product, design, and engineering around the same acceptance criteria, while revealing gaps before generated code turns them into rework. Turner cautioned against polishing every spec into a monument because the useful artifact only needs enough structure to support fast iteration, with a scope small enough for both the team and the agent to reason about clearly.
Dependency mapping becomes equally important because agentic changes can cross service boundaries faster than organizational memory can track them. “If an agent doesn’t know what its change is going to break, it will break it,” Turner says.
A dependency map gives the agent a view of the blast radius before it changes a contract, a data flow, or a shared component. Without that context, the time saved during generation can disappear into cleanup, incident response, and cross-team negotiation.
AI can also compress the distance between roles, allowing product managers, designers, and engineers to work in parallel around the same prototype and specification. Turner heard several teams describe this emerging identity as “builders,” a term that captures how the boundaries between defining, designing, and implementing a feature can become more porous.
The organizational constraint does not vanish with the handoff. Turner described features that developers completed in hours and then watched stall for weeks because user validation, product review, or dependent teams could not move at the same speed. AI amplifies the coordination system already in place, so healthy teams gain more parallel capacity while brittle silos collide earlier.
Layer 3: Evaluation and observability earn agent autonomy
Production autonomy depends on evidence that the system behaves correctly and can be stopped when its behavior changes. Turner framed this layer as the point where trusted tests, agent observability, staged rollouts, behavioral dashboards, and rollback mechanisms become prerequisites for speed.
Deterministic tests provide the foundation, although coverage percentages alone offer little comfort when the tests fail to represent meaningful behavior. AI can propose cases and draft test code, while humans remain responsible for verifying that those tests catch failures the business and the user would care about.
Agentic workflows add nondeterminism, which requires evaluators that can inspect whether outputs and end-to-end behavior satisfy the intended outcome. A pull request may look reasonable line by line while producing an unexpected behavior after deployment, so review shifts toward intent, system impact, and the evidence supplied by traces and evaluations.
User validation belongs in the same loop. Faster shipping creates value only when users can confirm that the system solves the intended problem, and a rapid implementation without rapid feedback can turn acceleration into efficiently produced waste.
When teams increase generation volume before building this layer, the resulting failures become harder to locate and more expensive to contain. In a regulated environment, an undetected error can also become a compliance event, which means explainability and audit trails influence the ceiling on autonomy as directly as model capability does.
What should an AI evaluation harness include?
An AI evaluation harness should include golden datasets, regression suites, and continuous production monitoring that together define acceptable behavior, catch regressions, and reveal drift. The evaluation harness gives teams a repeatable way to answer whether an agent is good enough to ship and whether it remains good enough after release.
“Build the harness before the agent,” Khare asserts.
Khare compared evals for generative AI systems to unit tests in conventional software engineering. The analogy becomes especially useful when teams place evaluation at the beginning of development, because the cases and thresholds force them to define quality before prompt tuning and model selection consume the conversation.
Golden datasets and regression suites define release quality
Golden datasets encode the cases that matter, including expected behavior, known edge cases, and failures that would make a release unacceptable. Regression suites then preserve those lessons by rerunning the cases whenever the agent or its surrounding system changes.
This practice turns evaluation into a design activity. When a team has to decide which examples deserve permanent protection, it exposes disagreements about intent while those disagreements remain inexpensive to resolve. The resulting suite also gives reviewers a body of evidence that can support a release decision without relying on a handful of manually selected demos.
Human judgment remains essential because a weak evaluator can reward the same shallow behavior that a weak test would permit. Teams need to review the cases, scoring criteria, and failure thresholds with the same seriousness they apply to production code, especially when an agent’s output has several acceptable forms.
Continuous evaluation detects model, data, and user drift
Continuous evaluation extends the harness into production, where models, source data, and user behavior can all drift. A launch score captures a moment; an operational evaluation program reveals whether that score changes over the next 30 days and which slices of traffic account for the movement.
That history allows developers to separate a broad regression from a localized failure, such as a tool-call issue, a newly ambiguous user request, or a data-quality problem in one source. Traces and production samples can then feed new cases back into the golden dataset, which lets the harness mature alongside the system.
This feedback loop gives AI agent evaluation its production value. In evaluation-driven development, the harness becomes the mechanism through which teams learn what the agent does under conditions the original test set did not anticipate, while each newly observed failure sharpens the next release decision.
How should teams measure AI agent cost and ROI?
Teams should measure AI agent economics through cost per outcome and connect that figure to a business KPI such as cycle time, error rate, cost per transaction, or time to resolution. Token usage remains useful for infrastructure analysis, although it cannot capture the retries, rework, escalations, and review burden created by an unreliable workflow.
Cost per outcome captures retries, rework, and review
Khare argued that the meaningful unit is cost per outcome because a cheaper model can produce a more expensive operating system when its outputs require additional cleanup. A model that saves pennies during inference may add dollars downstream when a human has to verify, correct, rerun, or escalate the work.
The unit should match the job the agent performs. A support workflow might track cost per resolved inquiry, a document workflow might track cost per processed document, and an internal automation might track cost per completed task. Those measures incorporate the full path to a usable result, which makes them more informative than a monthly token bill detached from task success.
Production planning should therefore include per-feature spending ceilings, fallback paths, and observability on spend. These controls allow engineering and finance to understand how costs behave under real volume, where retries and exceptions can make pilot economics misleading.
Business KPIs belong in the AI intake process
A production candidate should enter development with a business metric already attached. Khare named cycle time, throughput, documentation turnaround, error rate, cost per transaction, and time to resolution as examples that can connect an AI system to operational value.
When the metric appears during intake, engineering can design the evaluation suite and production monitoring around the result the organization expects. The same metric then supports prioritization, release decisions, and post-launch analysis, which prevents ROI from becoming a story assembled after the budget has already been spent.
Khare’s rule was simple: the metric comes before the model. Finance belongs in this conversation early because enterprise scale changes the economics. Although a pilot may look inexpensive while serving a limited audience and routing difficult cases around the system, production exposes the true cost of retries, human review, fallback models, data access, and ongoing monitoring. Cost per outcome gives finance and engineering a common unit for deciding whether the workflow deserves to expand.
Which parts of the agent harness survive better models?
The durable parts of an agent harness are the data, workflow design, evaluation suites, permissions, and trust fabric that remain valuable as model capabilities improve. Teams should scrutinize custom orchestration, retrieval, and routing layers because a future model or API may absorb functions that currently require extensive scaffolding.
“The model will eat your stack,” Khare said, describing how quickly capabilities can migrate from surrounding infrastructure into the model layer.
The observation redirects engineering investment toward the parts of the system that retain value as the model changes. A workflow that captures domain intent, constrains access, records traces, and measures outcomes can continue to create value after the underlying model changes. A brittle layer built around a temporary limitation may become maintenance work with a short half-life.
Data quality, workflow design, evals, and trust remain durable
Data quality moves onto the critical path because agents amplify ambiguity at scale. A messy column that an experienced analyst can interpret through context may lead an agent toward a confident error, and the resulting answer can propagate through a workflow before anyone notices the source of the mistake.
Clean source material, relevant context, and well-defined semantics therefore remain part of production engineering. The work may lack the spectacle of a new model demo, although it determines whether the system can use that model safely.
Workflow design also survives because a model does not inherit the organization’s operating context. Teams still have to decide which problem deserves automation, how responsibilities move between systems and people, where approval belongs, and what evidence should accompany an action. Evaluation suites preserve those decisions as executable expectations, while the trust fabric enforces the permissions and auditability required to operate them.
Why can AI workflow automation create downstream bottlenecks?
AI workflow automation creates downstream bottlenecks when one step accelerates without a corresponding increase in the capacity of the teams and systems that receive its output. Because AI often performs only a small part of a larger process, local throughput can rise while end-to-end completion time becomes worse.
Khare described an example in which Team A’s agent ran continuously and buried Team B under the resulting tasks. The first group could point to a faster local process, yet the organization had merely moved the queue to a place with less capacity.
The pattern is easy to miss during a low-volume pilot, which allows downstream teams to absorb the extra work, whereas production changes the shape of the system. A workflow that generates cases, documents, code changes, or review requests around the clock can overwhelm a step that still depends on a small group of specialists.
Map the full workflow before automating a step
Teams should map the complete workflow, including upstream inputs, downstream queues, approval points, and exception paths, before choosing the automation target. That map reveals where increased volume will land and whether another step will become the new constraint.
The exercise also changes prioritization because the easiest manual task may offer little system-level value when it feeds a scarce reviewer, while a less obvious intervention could remove a handoff, improve source data, or reduce the number of cases that require escalation. Developers need to evaluate the automation across an end-to-end workflow containing users, reviewers, compliance functions, dependent teams, and other software.
This systems view also protects the business metric. When teams measure cycle time or cost per completed task across the full workflow, a local speedup cannot disguise a slower overall process.
How do AI agent guardrails work in regulated environments?
AI agent guardrails in regulated environments combine mechanical controls that constrain what an agent can physically do with text-based controls that make its intended behavior reviewable. Khare argued that both layers belong in the architecture because failures carry asymmetric costs, and a defect that would cause a retry in a consumer application can trigger audit exposure, notification obligations, or a breach in a regulated workflow.
Mechanical guardrails constrain production access
Mechanical guardrails use deterministic controls to enforce boundaries around data, tools, and actions. They include hard checks on infor
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み