AI 開発の加速に伴う ChatGPT の高速維持に関する発表
本文の状態
日本語全文を表示中
詳細モードで約47分の本文を読めます。
OpenAI ではアジェンティックなワークフローが導入され、コード変更量が劇的に増加していることが示された。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月8日 19:01
AI深層分析
キーポイント
エージェントワークフローによる負荷増大
OpenAI ではアジェンティックなワークフローが導入され、コード変更量が劇的に増加していることが示された。
GPU 以外の隠れた性能コスト
急速なリリースサイクルにおいて、GPU リソース以外にもシステム全体の性能を低下させる隠れたコストが存在すると指摘されている。
常時 AI エージェントの活用
プロファイリング、回帰検出、継続的最適化を自動化するために、常時稼働する AI エージェントをデプロイしている。
ユーザーベースの急成長と開発ワークフローの変化
企業は以前よりもはるかに速く数百万人のユーザーを獲得しており、エージェント型コーディングなどの新技術により製品リリースの速度も劇的に変化している。
パフォーマンスエンジニアリングの重要性
開発ワークフローの変化とリリース頻度の向上は、製品の高速性と効率性を維持するためのパフォーマンスエンジニアリングに直接的な影響を与える。
重要な引用
Martin Spier explains how agentic workflows dramatically increase code change volume at OpenAI.
He discusses the hidden systemic performance costs of rapid shipping beyond GPUs
deploying always-on AI agents automates profiling, regression detection, and continuous optimization
Today I want to tell you guys a story about two accelerations that are happening somewhat at the same time.
編集コメントを表示
編集コメント
この発表は、単なる技術の紹介ではなく、開発プロセス自体が AI エージェントによって変容する現代におけるインフラ管理の在り方を示唆している。OpenAI の事例は、他社も直面するスケーラビリティ課題に対する具体的な解決策の一つとして注目される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
InfoQ ホームページ
プレゼンテーション一覧
AI 開発の加速に伴い、ChatGPT の高速性を維持する方法
プレゼンテーションを見る
所要時間:50:14
スライド画像/presentations/openai-performance-engineering-agentic-coding/en/slides/slide-1785151851732.jpg)
概要
Martin Spier は、エージェント型ワークフローが OpenAI におけるコード変更の量を劇的に増加させる仕組みを解説します。GPU の性能だけでなく、迅速なリリースに伴う隠れたシステム全体の性能コストについて言及し、常時稼働する AI エージェントを導入することで、プロファイリングや回帰検出、継続的な最適化を自動化できる点を紹介します。これにより、大規模なグローバルスケールでも製品の速度と拡張性を維持することが可能になります。
登壇者紹介
Martin Spier は OpenAI で ChatGPT のパフォーマンス統括を務めています。ChatGPT をより高速で信頼性の高いものにし、エンジニアリングチームがスケーラブルに運用しやすくするための取り組みを主導しています。その活動範囲は、AI やクラウドインフラストラクチャ、パフォーマンスエンジニアリング、観測可能性(オバザビリティ)、信頼性、プラットフォームエンジニアリング、そして開発者生産性に及びます。
会議について
QCon AI は、これらのワークロードを安全にスケールさせるために必要なエンジニアリング分野に特化した、実践者主導のイベントです。参加者は、他社が本番環境で採用しているアーキテクチャのプレイブックや障害指標への直接アクセスを得ることができます。
INFOQ EVENTS
2026 年 8 月 27 日午後 1 時(東部夏時間)
フレームワークの先:なぜエージェントのコンテキストはインフラの問題なのか
登壇者: Boyd Stowe 氏(Tacnode 創設者兼ソリューションアーキテクト)
講演要旨
Martin Spier: 今日は、現在同時に進行している 2 つの加速についてお話しします。1 つ目は成長の加速です。OpenAI や ChatGPT に限らず、企業が最初の数百万人のユーザーを獲得するまでのスピードが以前にも増して速くなっています。
2 つ目は、エージェントによるコーディングや関連技術の影響で開発ワークフローが変容し、その結果として製品をリリースする速度がどう変化しているかという点です。そして、これが製品の高速性や効率性を維持するためのパフォーマンスエンジニアリングにどのような影響を与えているのかについてお話しします。
私の名前はマーティンです。OpenAI で ChatGPT のパフォーマンスチームを率いています。技術キャリアのほとんどはパフォーマンスエンジニアリングに捧げてきました。16 年ほど、あるいはそれ以上ですね。その大半を Netflix で過ごしました。約 9 年間在籍していました。ディオが言及していたクラウド黎明期、私が入社した頃です。当時は「本当にデータをクラウドに預けても大丈夫なのか」という議論がありました。まるで魔法のような時代でした。今まさに同じような瞬間を迎えていると感じています。
その後、Snowflake や Expedia などいくつかの企業でも同様の業務に従事しました。OpenAI に入る前は、Parasail という AI ベンチャーに在籍していました。そこでは推論をサービスとして提供しており、エンジニアリングチーム全体を統率していました。さらにその前には、ブラジルのフィンテック企業 PicPay で、インフラプラットフォームと開発者体験の責任者を務めました。
ご想像の通り、私のキャリアはこうした中核組織に集中しています。私が特に愛する分野です。私はこれを「家の配管」だと表現します。普段は誰もその存在を意識しませんが、詰まったり壊れたりすれば大混乱が起きるものです。会社にとって常にやりがいのある部分でもあります。
成長のペース
3 つの変化が起きました。まず、成長のペースです。私が以前触れた OpenAI だけでなく、企業はかつてないスピードで、最初の 100 万人、1,000 万人、1 億人のユーザーに到達しています。曲線は指数関数的ですが、以前よりもはるかに急峻になっています。また、より多くのコードや機能をリリースするようになりました。エージェント型開発によって、より多くの成果物を世に出せるようになっています。
ここには「AI スロップ(質の低い生成物)」に関する議論もあるでしょうが、私たちはより多くの論理を製品化しています。3 つ目の変化は、以前は「すべての変更内容を理解している人間が存在する」という前提がありました。誰かが実際にコードを書き、アーキテクチャを設計し、リリース前に何が起こっているかを把握しているという考え方です。
しかし、私の見解では、この前提はもはや完全に正しいとは言えません。少しバブル的な側面があるかもしれませんが、抽象化のレイヤーが一段高くなっています。
エージェントに任せる業務が増えるにつれ、外部へ送信されるすべての詳細を把握できなくなっています。これは常に念頭に置いておくべき点です。
もちろん、こうした状況はパフォーマンスへの圧力だけでなく、その周辺全体にも影響を及ぼします。GitHub や GitLab のようなプラットフォームを利用する方々にとって、インフラ全体に大きな負荷がかかっていることは明らかです。リソースの消費量が増え、レイテンシを引き起こすロジックが追加されることで、多くの課題が生じています。
私たちはこの急速な成長と開発スピードに対応し続ける必要があります。そのためには、従来の慣行を進化させることが不可欠です。
ChatGPT の歴史を少し振り返り、私が話している規模感の背景をお伝えしましょう。2022 年末に登場した ChatGPT は、当初から数億人のユーザーを支える本格的な消費者向けアプリケーションとしてゼロから構築されたわけではありませんでした。最初は研究用のプレビュー版としてのリリースでした。
これがその姿です。非常に魅力的なサービスで、ぜひ試してみてください。すぐにユーザーが殺到し、急成長を遂げました。こうしてユーザーが増え続ける中で、研究用プレビューは製品へと進化します。製品となるには、さまざまな要件を満たす必要があります。例えば、応答速度(レイテンシ)の要件や信頼性の要件、サポート体制などです。これらはすべて、大規模なユーザー数に対応し、特に日常業務で不可欠なツールとして利用されるために裏側で支える仕組みです。
その成長ぶりは驚異的なものでした。研究用プレビューからわずか 5 日間で最初の 100 万ユーザーを達成した企業は、ほとんどないでしょう。
これは非常に素晴らしいことです。現在、このサービスには100万人のユーザーがいます。裏方でインフラを支えるチームが、その時期にどれほど大変な思いをしていたかを想像してみてください。驚異的な成長であり、製品にとっても画期的なことでした。しかし、その分、裏方の人々には多大なプレッシャーがかかりました。
この成長は止まりません。大幅に鈍化するどころか、ChatGPTは今もなお、非常に速いペースで拡大を続けています。私が共有できる最新の公式データは今年2月のもので、週次アクティブユーザー数が9億人に達した時点です。つまり、毎週9億人がChatGPTを利用していることになります。
この数字の重みを噛みしめてみてください。これは人類全体の約11%に相当します。膨大な数の人々が利用しています。
ご想像の通り、チャットは計算リソースを必要とする処理であり、決して軽量な作業ではありません。それを支えるためには、裏側で多くの仕組みが不可欠です。また、この成長曲線は滑らかではありませんでした。毎日数百万人のユーザーが増え続けるという単純な予測可能なパターンでもなかったのです。
ご想像の通り、急激な変動が頻繁に起こっていました。私たちは常に新製品をリリースしており、その中には爆発的に人気を集めるものも少なくありませんでした。ユーザー数が急増し、それがさらに普及を加速させるという好循環が生まれました。
例えば昨年の画像生成機能での急増は記憶に新しいかもしれません。あの画像のどれかを作成した方もいるでしょう。開始からわずか 7 日間で、1 億 3000 万人以上のユーザーによって 7 億枚を超える画像が生成されました。画像生成の普及スピードは驚異的でした。
こうした急激な成長は、裏方チームに多大なプレッシャーをかけ、慌ただしい対応を迫ることになります。何がどれほど人気を集めるか、その予測は非常に難しいものです。しかし、これは私たちが望むべきことでもあります。誰もが自社の製品が成長することを願っているのです。それは素晴らしいことです。
チャットは米国だけのものではありません。世界中で利用されています。インフラやアーキテクチャ、そしてそれらを支える仕組みを構築する際、複雑さが増すのは当然です。
例えば、POP(ポイント・オブ・プレゼンス)はどこに配置すべきか?実際のアプリケーションサーバーの場所はどうするか?GPU はどこにあり、互いに近接しているのか?各リージョンで十分な容量があるか?異なるデータセンター間のネットワーク接続はどのようになっているか?データはどのように保存され、複製されているのか?また、世界中に広がるさまざまな制約を考慮すると、データの所在地はどこが最適なのか。
こうした要素すべてを考慮して、世界中のユーザーに質の高い体験を提供できるアーキテクチャを構築するには、非常に多くの複雑さが必要です。単なる成長だけでなく、製品自体も静止しませんでした。単なるローンチやユーザー増加だけではありませんでした。
モデルの変更以外にも、大きな変化がありました。もちろんモデルは極めて重要でしたが、この 4 年未満の期間には数々の新機能やサービスが次々と登場しました。音声対応の GPT やエージェント、画像生成など、多様な機能が追加されました。その中には、他よりも viral(拡散性が高い)なものも含まれていました。
製品は絶えず進化し続けていました。それに対応するため、サポートすべきワークロードの種類が増え、それを支えるアーキテクチャもより複雑化しました。製品は急速に成長する一方で、常に変化を遂げていたのです。
変化の速度
これは「新しい」と呼ぶには少し語弊がありますが、昨年後半以降、明らかに加速しています。特に Codex の登場以来、出荷速度は劇的に変化しました。
価値やそれに関する深い議論についてはここでは触れませんが、システムに流入する変更の量が大幅に増えたという事実があります。
エージェント開発について考える際、真っ先に浮かぶのは「この関数を開発する」「そのアプリケーションを作る」「これをリファクタリングしてあれもリファクタリングする」といったコーディングタスクの自動化です。また、完了までの時間を短縮し、アイデアからプロダクションにマージされる PR までのサイクルを加速させることも目指しています。
私が特に驚いたのは、より高度な新モデルが登場したことでエージェントが複雑なタスクも処理できるようになり、実行時間が長くなるような作業も任せることができるようになった点です。
開発者たちはマルチスレッド化し始めました。つまり、一度に一つのことに集中するのではなく、同時に 7 つ、8 つ、9 つ、あるいは 10 個もの異なるタスクを並行して進めるようになりました。進行状況に応じてブロックや待機が発生することも珍しくありません。
最近、エージェントを使ってどの程度コーディングを行っているでしょうか?「今すぐユーザー入力が欲しい」といった通知が鳴り響いているかもしれません。複数の開発環境を同時に使っているか、あるいは組織の階層構造(org trees)を扱っているケースもあるでしょう。
全体的に、並行して進める作業の量が格段に増えています。単に出荷速度が速くなっただけでなく、より多くのものを並行して出荷する時代になっているのです。
DX のベンチマークをご存じの方もいらっしゃるでしょう。これは数社を対象に、エンジニアあたり週間にどれくらいのプルリクエスト(PR)がマージされているかを調査したものです。会社規模や業種ごとにデータを分類すると、予想通り小規模なテック企業ほど PR 発行量が多く、製品リリースも活発です。p90 の値では、エンジニアあたり週間で約 5 つの PR に達しており、これは非常に速いペースです。
しかし、私たちのペースはさらに速いです。最近ではもっと多くのものをリリースしています。残念ながら、私が共有できる最新のデータは昨年末(10 月〜11 月)のものですが、ご存知の方ならおわかりのように、この時期以降コーディングエージェントの分野には大きな進展がありました。実は昨年 10 月の時点ですでに、エンジニアが毎週マージする PR の量は 70% 増加していました。
現在では、社内のほぼすべてのエンジニアが Codex を週次、あるいは日常的に活用しています。
すべてのプルリクエストは、Codex によって自動的にレビューされます。これはコーディングタスクに限定された話ではありません。先ほど述べた通り、Codex は開発ワークフロー全体で必要な作業を行うための作業台として機能しています。
コードを書くことだけではありません。本番環境でバグを発見したら、そのトラブルシューティングにも Codex を活用します。あるいは、ある指標や回帰現象を理解したい場合、観測ツールにアクセスしてデータを取得し、何が起きているのかを説明するためにも使います。データサイエンスや分析を行う際も同様です。
その他、Slack のスレッド要約など、退屈な作業も Codex に任せることができます。「全部読まなくていいようにまとめて」といった依頼は、日々の生産性を大幅に向上させるタスクです。エンジニアに限った話ではありません。私はチームを管理していますが、最近ではコーディングに費やす時間はあまりありません。それでも、Codex はあらゆる場面で活用しています。
1 対 1 の面談の準備や、ドキュメントの要約・作成にも使います。今回のプレゼンテーション自体も、Codex でアウトラインを確認し、何を伝えたいか、どう構成するかを議論しながら始めました。生成された画像やビジュアルもすべて Codex を通じて作られています。
つまり、Codex は純粋なコーディングタスクだけでなく、日々のあらゆる業務で活用されているのです。
ソフトウェアエンジニアリングのワークフローが進化し、私たちはより速く成果物をリリースできるようになりました。製品が世に出るまでの時間が短縮され、並行して取り組むプロジェクトも増えています。また、タスクの委任も進み、エージェントによる開発によってスループットと並列性が向上しました。
その結果、システムに投入される変更の量が増大しています。ここで言う「変更」とは、価値や機能に限らず、リリースされるあらゆるものを指します。新機能、バグ修正、リファクタリング、あるいは設定ファイルの変更など、何らかの変化が常に発生しているのです。
しかし、すべての新機能や新しいコード行には、見えないコストが伴います。理想的には、このコストが価値を生み出すべきです。つまり、私たちはそれを望んでおり、リリースしたいと考えているものです。新機能やバグ修正は確かに必要ですが、どこかに代償が発生します。
そのコストは当初は極めて小さく、目に見えない場合もあります。最初のうちは全く気づかないこともあるでしょう。
この変化の速度がさらに加速すると、状況は悪化します。これは新しい現象ではありませんが、以前よりも速く進行しています。
外部への出力が増え、追加された条件分岐(if 文)が増え、ネットワークリクエストが一つ増えるたびに、あるいはメモリに保持するデータ構造が増えるたびに、それは共有された予算を消費することになります。遅延は一定の限界まで悪化し、ついにユーザーがサービス利用を諦めるに至ります。
ロジックを追加すれば処理が増え、チェック項目が増え、何らかの変更を加えれば、それだけ必ずわずかな遅延が発生します。これが積み重なり、体験が悪化して「もうこの製品は使いたくない」という状態に陥るのです。
ハードウェアリソースについても同様です。メモリや RAM の量は有限です。「クラウドなら無限に拡張できる」と言われることもありますが、それは誤りです。クラウドであってもリソースには限界があります。
「弾力性がある」と思っている人は多いですが、実際はそうではありません。むしろ、効率をわずかに下げるだけです。これは共有された予算からリソースを消費している状態です。
私が言いたいのは、デプロイ前後で明らかに劇的な変化が起きるような大規模な回帰(regression)の話だけではありません。もっと些細な話です。例えば、「重要な何かをチェックするために追加した小さな if 文」のようなものです。データベースにおいては、こうした小さな変更は「違いがない」と思われがちですが、実際には比較対象となる前後で差が生じます。
これらはすべて蓄積していきます。少しずつ、また少しずつ積み上がっていくと、やがて大きな問題へと発展します。それが懸念材料となり、システムに余裕(headroom)がなくなってしまう恐れもあります。結果として、ユーザーがアプリで快適な体験を得られなくなるほど遅延(latency)が悪化する可能性さえあります。
変化の速度が増すにつれて、こうした悪影響はより急速に蓄積されていくのです。
これらの小さな変更や、時には大きな後退によるダメージは、すぐに現れるわけではありません。しかし、その影響は後に表れます。
ユーザーにとっては、アプリ自体の動作が重くなる形で現れるでしょう。効率性については、同じ負荷を処理するコストが増大します。ログを増やし、処理すべき項目が増えても、実際の負荷量は変わらないのです。結果として、支払う費用だけが増えることになります。
スケーラビリティ(拡張性)の観点では、成長のための余地が次第に狭まっていきます。ご想像の通り、ある時点でスケーラビリティの問題は信頼性の問題へと転じます。急激な負荷増加に対応できなくなったり、さらに規模を拡大できなくなった瞬間です。
こうしたダメージは後から現れますが、可能な限り早く対応する必要があります。
パフォーマンスエンジニアリング
これは新しい話ではありません。私たちはキャリアを通じて機能を開発し、機能を追加し、ロジックを追加してきました。
片方には機能をリリースする人たちがいます。もう一方には、効率や遅延を気にし、可能な限り効率的にすることに関心を持つパフォーマンスエンジニアリングチーム、あるいは数人のパフォーマンスマインドを持った個人がいるかもしれません。これらは対立する二つの力であり、バランスを保つ必要があります。
問題は、変化の速度が増すにつれて、方程式の片側が非常に速くなってしまったことです。私たちはより多くの機能をリリースしていますが、パフォーマンスエンジニアリングやパフォーマンス問題に取り組む人々は適応する必要があります。追いつく必要があります。
その答えは言うまでもなく、パフォーマンスエンジニアをさらに雇いたくないということです。そもそもパフォーマンスエンジニアを見つけるのは難しいのです。これも最善の解決策ではありません。私たちのワークフローも速くする必要があります。
パフォーマンスエンジニアリングの側でも、サイクルを加速させる必要があります。検知速度を劇的に高めるための反応フローやループを高速化し、問題の原因をより早く特定できるようにプロファイリングを行うべきです。
それだけでなく、エージェントにコードベースを理解させ、解決策を提案・実装し、リリースしてベンチマークとテストを実施させることも可能です。実際に通知が来ても「すでに修正済みです」と即座に対応でき、ベンチマーク結果も再び良好になるでしょう。これにより作業を継続できます。
能動的な側面では、複数のエージェントを並列で稼働させ、潜在的な最適化ポイントを探索することも可能です。異なるスキルを持つ多数のエージェントが、それぞれ異なる種類の最適化に取り組むことができます。例えば、レイテンシのボトルネックとなっているパスに注目したり、CPU を最も多く消費しているメソッドを特定して最適化を試みたり、メモリ割り当てやバンドルサイズの改善を継続的に実施したりできます。
これらはすべて専門性を発揮しつつ、非同期で絶え間なく動作し、最適化の可能性を探り続けることができます。直列に処理する必要はありません。
私たちは適応し、もっと速く動く必要があります。もし必要なスピードで適応できなければどうなるでしょうか?もちろん、裏側では多くの混乱が生じるでしょうが、それ以上にユーザーが最初に感じることは何でしょうか。
まず挙げられるのは、一般的に「遅さ」です。ロジックを追加すればするほど処理は重くなり、やがて問題となります。するとユーザーは離れ始め、アプリの利用を停止し、プランの解約に至ります。最悪の場合、エラーが増え、システムダウンが長引き、復旧にも時間がかかるようになります。その間、ユーザーは苦しむことになります。
特に ChatGPT のような大規模な消費者向けアプリケーションでは、またはおそらく皆さんが携わっているプロジェクトでも、こうした事態は避けたいはずです。パフォーマンスはビジネスの収益に直結します。ユーザーの定着率や新規獲得にも影響を及ぼすのです。私は Netflix でこのモデルを実証しました。
私は OpenAI でそのモデルを構築しました。貴社の皆様も同様にこの課題に取り組めば、処理の遅延とエラー増加、そしてユーザーによるプラン解約との相関関係が明確に浮かび上がるはずです。これはビジネスへの直接的な影響です。
私が特に気になっているのは、AI アプリケーションにおいて誰もが推論(inference)に注目している点です。確かに推論は難所であり、常時監視すべきリソースです。チームは「最初のトークンまでの時間」や「秒間生成トークン数」、「スループット」といった推論コンポーネントの指標に集中しています。これらのコンポーネント指標は非常に重要です。
しかし、推論は問題の一部ではあっても、全体像ではありません。ユーザーが実際に何を感じているかを測定し、継続して監視する必要があります。ユーザーが製品を利用する際、彼らは基本的に課題解決やタスク実行を目的としています。そのプロセスの中で、ユーザーはさまざまな感情を抱くことになります。
ユーザーにはそれぞれ異なる期待があります。それを捉え、さらに掘り下げる必要があります。
例えば、チャットアプリを想定したケースを見てみましょう。考え方は同じです。私はNetflixでも同様のアプローチを採用しましたし、Chromeも同様だと言われています。
アクションの過程において、ユーザーは段階ごとに異なる期待を抱きます。メッセージを入力して「送信」ボタンをクリックすると、ユーザーは即座に何らかの反応が返ってくることを強く望みます。「ただ待っているだけではない」という安心感や、「処理中である」というフィードバックです。ローディングスピナーが表示されたり、「考え中」のアニメーションが始まったりします。重要なのは、何かが進行中であり、ユーザーが待つ必要が生じる点です。
そして次に現れるのが、私が「初回可視値(first visible value)」と呼ぶものです。
ユーザーが本来やりたいことを続けられるのはいつでしょうか。例えばチャットリクエストの場合、ユーザーは回答を読み始めるタイミングを求めています。最初の数トークンが表示されるまでの時間と、その後のストリーミング速度が重要です。スループットやストリームの cadence が遅すぎると、ユーザーは読み進める途中で止まってしまい、イライラしてしまいます。そのため、ユーザーがブロックされたり挫折したりすることなく読み続けられるよう、ストリームの cadence を十分に良くしておく必要があります。
また、メッセージ全体が完成するまでの時間やタスク完了までの時間も考慮する必要があります。これにも多くのニュアンスが含まれています。具体的にいつタイマーを停止すべきかという点も重要です。チャットレスポンスの場合、最初のトークンが表示された時点で止めるのか、それとも最初の数単語が出揃った時点で止めるのか。ユーザーが実際に価値を感じられるのはどの時点でしょうか。これは製品によって異なります。
思考レスポンスの処理において、タイマーはいつ停止すべきでしょうか?思考トークンが完了した時点で止めるのか、それともユーザーが実際に最終メッセージを受け取れる時点まで計測するのか。画像生成の場合も同様で、高解像度の完成画像が表示された時点で止めるのか、あるいは低解像度の初期画像が表示された時点で止めるのか。これらの判断には多くのニュアンスがあり、正確な数値を算出するためには製品側の視点から多角的に検討する必要があります。
さらに重要なのは、ユーザーの意図(インテント)によって期待される応答速度が異なるという点です。例えば「フランスの首都はどこか」といった単純な質問であれば、検索エンジンですぐに回答を得られるため、ユーザーも即座のレスポンスを期待します。こうしたケースでは、より高速な対応が求められます。
一方で、エージェント型のループを実行して複数の情報源を確認したり、コードの編集を行ったりする必要がある場合、処理には時間がかかります。同様に画像生成も計算負荷が高いため、ユーザーはそれに見合った応答時間を期待します。このように、タスクの種類やユーザーの目的によって、求められるレスポンス速度の基準は大きく異なるのです。
さまざまなタイマー、壁時間(wall times)、異なる意図、そして多数の次元を把握しています。これらに基づいて行動を起こすためには、まずそれをシステムのエビデンスへと分解し、追跡したい特定のコンポーネントや、全体としての壁時間を構成する要素に落とし込む必要があります。
私はこのチャートを活用するのが好きです。もちろん選択肢は多々ありますが、私のチームも同様に推奨しています。レイヤーケーキ型と呼ばれる遅延内訳チャートがその一例で、横軸に分布を、縦軸に実際の遅延値を取り、状況に応じて異なる層(レイヤー)を表示するものです。
重要なのはチャートの見た目自体ではなく、推論(inference)以外の領域でも多くの処理が行われているという事実です。もちろん、これは数値のスケールを示すものではなく、概念を説明するための図解に過ぎません。
実際には、クライアント側の作業やネットワーク通信にかかる時間、データ取得、シリアライゼーション、トークン化など、推論のみで完結しない多様な処理が背後で行われています。
このようにレイヤーごとに分解することで、取り組みの焦点を明確にできます。また、先ほど述べた「意図」も、重点を置くべき領域を見つける手助けになります。例えば、全体の処理時間の大部分が推論そのものではなく、データ取得にあることがわかれば、そこに集中して改善を図ることができます。全体の流れを可視化できることは極めて重要です。そうでなければ、誰もが最初に思い浮かべる「推論」にばかり目が向きがちです。推論は問題の一部に過ぎません。
例えばチャットアプリケーションを考えてみましょう。チャットはあえてシンプルに設計されています。ユーザーがメッセージを入力し、送信します。そして少し待って、応答を受け取ります。このようなシステムを開発した経験がある方は多いでしょうか?しかし、単に入力されたメッセージを推論エンジンへ転送するだけでは不十分です。推論エンジンは、会話の全履歴や適切な応答に必要な状態情報を保持していないからです。
推論エンジンへメッセージを送信する前には、多くの処理が行われています。もちろんクライアント側での作業も含まれますが、ここでは必要な要素をすべて組み立てたり、問題点を修正したりします。
データが戻ってきたら、まずユーザーの身元を確認する必要があります。さらに、クライアント側の身元の検証や、ユーザーのプラン内容、利用枠の確認、実際のユーザー状態のチェックも行います。これらはすべて欠かせないステップです。
次に、文脈情報の取得というやや重い処理に入ります。会話履歴は非常に長くなることもあり、そのすべての履歴を取得しなければなりません。また、ユーザーが PDF や画像などのファイルをアップロードしている場合も、それらをすべて読み込む必要があります。プロジェクト内で作業を行っている場合は、プロジェクトごとの文脈情報も取得する必要があります。
推論エンジンへ実際に送信されるデータを組み立てるまでには、多くのデータ取得プロセスを経る必要があるのです。
データはすべて揃いました。必要な情報は一通り揃っています。エンコーディングもトークナイゼーションも、すべてのモジュールが個別にトークン化されています。
ここで重要なのは、コンテキストウィンドウのサイズです。ユーザーの利用可能なコンテキストウィンドウを超えていないか確認し、必要に応じて切り詰め(truncation)や圧縮(compaction)を行う必要があります。その後、推論エンジンへ送るためのデータを組み立てます。推論エンジンにデータが届く前には、これほど多くの処理が行われているのです。
また、メッセージをストリーミングして返す際にも、後で会話履歴を保存する必要があり、そこでも多くの処理が発生します。
ユーザーが送信したメッセージは、小さな文字列や短い質問といった、全体のプロセスにおける一つの要素に過ぎません。特に長い会話では、データ量が膨大になります。会話の長さが数メガバイト、数十メガバイト、あるいは数百メガバイトに達することさえあります。
システム内では膨大なリクエストとデータが移動しています。ご想像の通り、これらは先ほど挙げたすべてのリソースを消費します。具体的には、データベースへのアクセス、Blob ストレージの利用、I/O 処理によるデータの転送などが頻繁に発生します。また、データをシリアライズ・デシリアライズする際の CPU 負荷も無視できません。さらに、メモリ上でのデータ保持にも多くの RAM が必要です。
つまり、AI と聞いて真っ先に GPU を思い浮かべる方が多いですが、GPU は確かに不可欠で極めて重要です。しかし、それ以外の CPU ワークロードやその他のリソースも同様に重要なのです。全体としてスムーズな体験を実現するには、これらの多様なリソースをすべて適切に管理する必要があります。そのため、GPU 以外の部分にも十分な注意を払う必要があります。
この話題を取り上げたのは、コードの変更や改善、推論エンジン自体の処理速度向上などがある一方で、製品全体はより複雑で多くのコンポーネントが関わっているからです。CPU やメモリ、I/O といった他のコンポーネントにも多大なコード変更が及ぶことがあり、場合によっては GPU 以上に影響を与えることもあります。
なぜこれが重要なのか。どんなに高速なモデルや推論エンジンを持っていても、その後の処理経路が遅ければ意味がありません。最終的にユーザーが得るのは遅い製品であり、利用を断念するでしょう。私たちは全体の流れ、つまりリクエストの全経路に焦点を当て、すべての部分が速く動くようにする必要があります。そうすることで、ユーザーは製品使用中に「快適だ」と感じるはずです。
Human in the Loop
一言でまとめると、製品の方向性はより広がり、推論だけでなくデータやコンテキスト、トークン化、ストリーミングなどすべてを含んでいます。この道筋は、エージェントによるコーディングの増加や私たちが行っているあらゆる変更によって、ますます重くなっています。
リソース消費が増大しており、その対象は GPU だけではありません。そのため、パフォーマンスエンジニアリングに対する考え方を進化させる必要があります。もちろん、社内のエンジニアに速度を落とさせたり、逆効果となる障壁を作ったりしたくはありません。現在、すべての AI プルリクエストは 5 人の人間によるレビューが必要ですが、これにさらに手間がかかる要素を加えれば、AI エージェントを活用して得られる最大のメリットが損なわれてしまいます。
実際には、パフォーマンスエンジニアリングにおけるさまざまなループを加速させる必要があります。具体的には、何らかの regress が発生した際に修正を行う「リアクティブ・ループ」です。
私たちは、処理速度をさらに向上させる必要があります。また、実際のコードや問題の発生源に近づけること、そしてアクティブなループ自体をより並列化することも求められています。
通常、私たちのリアクティブ・ループにはかなりの時間がかかります。ここで、過去にプロファイリングや最適化に取り組んだパフォーマンスエンジニアの方はいらっしゃいますか?このループは非常に遅いのです。あるリリースから次のリリースへ移行する際に性能が低下していることに気づいた場合、CPU の使用率が数パーセント低下したとします。次に何をすべきでしょうか。
まずはプロファイリングを行う必要があります。CPU プロファイルを取得し、実際に CPU で最も時間を費やしているスタックを特定する必要があります。できれば継続的なプロファイリングを行っていれば、前後の比較が容易になり、どこで性能が低下したかをすぐに発見できます。もし継続的なプロファイリングが行われていない場合、状況はさらに複雑になります。コードの変更を確認し、それが実際に CPU で実行されているスタックとどう関連しているかをマッピングして、どの部分が劣化したのかを特定し、最適化の糸口を見つける必要があります。
最適化を考え、実際に実装する必要があります。その後、リリースしてテストし、問題が解決したかを確認します。これらはすべて一人の人間が行うシリアルな作業であり、どうしても時間がかかりすぎます。
私がチーム内で模索している道、そして物事が向かっている方向性は、常時稼働させることです。エージェントを常に起動し、反応ループと能動ループの両方で非停止で動作させたいと考えています。
自動化の対象は、大きな回帰(regressions)だけではありません。これは、あるリリースから次のリリースへの変化を容易に特定できるような問題です。一般的には、その回帰を検知するとエージェントが自動的にトリガーされる仕組みがあります。エージェントはプロファイリングを実行したり、2 つのプロファイル同士を比較したりできます。この機能はすでに Codex で利用可能です。これらの要素がシステム内で連携していれば、実装は非常にシンプルになります。
エージェントはコードベースにアクセスでき、PR による変更前後を比較することも可能です。修正案を提案し、システムが十分に安全であれば、エージェント自身がデプロイして、その修正が期待通りの効果をもたらしたかを実証することもできます。
ドリフト(性能の緩やかな低下)についても同様です。小さな回帰現象については、コード複雑度やベンチマーク実行中のネットワークリクエスト量など、細粒度な指標を設けることで検出可能です。これらの微小な変化を捉え、即座に対応する仕組みも用意できます。
もちろん、私が以前言及した「アクティブな最適化ループ」の実現が最終的なゴールです。理想としては、本番環境への展開前、できればマージ前のすべての段階で、こうした作業を CI が自動的に処理し、問題を検知・対応できるようにしたいと考えています。
基本的にはすでに長く実施されている取り組みであり、特別に新しいことではありません。しかし、パフォーマンスに関する多くの問題は、特定のユーザー負荷が組み合わさる本番環境に至らないと顕在化しないため、事前のチェックでは捉えきれないケースも多々あります。そのため、本番環境で問題が発生した際に迅速に検知し、修正できる仕組みが必要です。そのためには、継続的なプロファイリング(性能計測)が不可欠です。
いくつか具体例を挙げましょう。すべてのプルリクエストに対してマイクロベンチマークを常時実行できます。エージェントがこれらのマイクロベンチマークの開発自体を支援することも可能です。
さまざまな事象を検出できます。特に、パフォーマンスの低下やメソッドの回帰(regression)に注目できます。レイテンシクリティカルパスに含まれるすべてのメソッドを監視し、特定の箇所で回帰が発生していないかを確認できるのです。また、メモリ割り当てのプロファイルも可能です。特定の関数が従来より多くのメモリを確保していないかを把握し、必要に応じて対応策を講じられます。
バンドルサイズの推移も追跡できます。特にモバイル向けに公開する場合は、バンドルサイズが増加していないか、予算の範囲を超えていないかを監視することが重要です。
ここで伝えたいのは、こうした活動は従来から行われてきたものの、それを自動的に実行・対応できるようになった点です。さらに踏み込んで、メトリクスやコードベースを分析することも可能です。システムに関するアーキテクチャ図や既存情報をすべて確認し、エージェントが修正のためのアイデアや方法を提案させることもできます。
実際に実装して、理想を言えばベンチマークを実行し、同じテストを前後で比較することも可能です。ただし、こうした修正が常に有効なわけではありません。問題が解決しない場合は、その案は捨てて別の方向へ進めればよいのです。
複数のスレッドを並列に動かして、さまざまなアイデアを試すことがポイントです。もちろん、これは最終的な状態であり、そこに到達するまでに多くのステップが必要です。Codex には、このプロセスを加速させるためのスキルセットを用意することもできます。手動でプロファイリングをトリガーしたり、データを収集・分析したりする代わりに、エージェントにその情報を集め、比較まで行わせることが可能です。これは十分に実現可能なことです。
エージェントは大量のデータを処理し、比較作業を行うのが得意です。必要であれば、そこからフラムグラフ( Flame Graph)も簡単に生成できます。
マイクロベンチマークでは再現が容易なため、比較的 straightforward です。本番環境でも同様で、十分な観測性(observability)を確保し、発生するあらゆる回帰現象を検知して対応することが求められます。
ここで重要なのは安全性です。エージェントに本番環境へのデプロイを任せる場合、ユーザーが不具合や痛みを感じさせないよう配慮する必要があります。そのためには、実施前に多くのベストプラクティスを導入しておく必要があります。
実際には、ほぼ完璧な状態まで到達可能です。例えば、パフォーマンス最適化の実験を実装するプルリクエスト(PR)をエージェントに提出させることもできます。これは十分に実現可能なことです。
完全に信頼できる状態で、監督なしで本番環境へのデプロイを行える段階に至っていなくても、その手前まではほぼ達成でき、非常に良好な結果が得られるはずです。
私が最も楽しみにしているのは、常時稼働するループです。私のチームには優秀なパフォーマンスエンジニアが多数在籍しており、彼らのスキルセットをエージェントに継承し、環境内で継続的に最適化を探させる仕組みを作ろうとしています。エージェントはコードやトレース、ログなどを常に見つめ、最適な改善点を見つけ出します。
面白いことに、特定の分野で卓越したスキルを持つ人々は、そのスキルを持つエージェントに自分自身の名前をつける習慣があります。私にも「ベン」や「ブレンダン」といった名前のエージェントがおり、多くの人が自分の名前を冠したエージェントを持っています。そうすると、いつの間にかそれらを人間のように扱うようになります。常時稼働して機会を探し続けるこのループは、不思議な感覚です。
実証済みのベストプラクティス
これまで効果的だったいくつかのベストプラクティスをご紹介します。エージェントが完全なループを完遂できるようになれば、システムはより自律的に動作するようになります。
一度ループが完成すれば、複数のエージェントが同時にさまざまなタスクを実行することが可能になります。つまり、イベントや継続的なサイクルからループを開始し、アイデアを実現するところから始めて、その変更が実際にプラスの効果をもたらしたかを測定しながら最後まで処理を完結させることができます。このプロセスは人間の介入なしに継続可能です。
状況によっては簡単ではないこともありますが、いずれにしても、人間の手を介さずにエージェントが完全なループを安全に実行できる仕組みづくりに注力すれば、処理速度は劇的に向上します。監視に必要な人的リソースも大幅に削減でき、ボトルネックとなる要因も激減します。
結局のところ、ユーザーに「同意」ボタンを連打させるのではなく、YOLO モード(試行錯誤モード)としてエージェントに任せるという選択肢もあります。また当然のことながら、長年続けてきた基本的な取り組みほど重要性を増しています。
エンジニアリングの成熟度が問われるようになります。テストカバレッジが十分に高いことは必須です。そうでなければ、エージェントは自分の行動が正しかったか、システムが機能しているかをどうやって判断できるでしょうか?ベンチマークも同様です。実装した改善策が有効だったかどうかを、エージェントがどのように測定できるのでしょうか。こうしたベンチマークとテストの整備は、以前から重要でしたが、今やその重要性はさらに高まっています。
システムの各部分間で明確な契約(コントラクト)と期待値を定義することも必要です。そして何より、観測性(オバザビリティ)が極めて重要です。エージェントがシステム内の状況を正しく理解できるよう、システム全体を網羅的に監視できる状態である必要があります。もし盲点があれば、人間が持つような直感がないため、エージェントは間違った方向へ進んでしまう可能性が高まります。
そのため、安全なロールアウトや自動化されたカナリア分析、ブルーグリーンデプロイメントといったセーフティ対策も不可欠です。バグが含まれていたとしてもユーザーに影響が出ないよう、コードを安全にリリースする仕組みが求められます。
結局のところ、パフォーマンス最適化のすべては「検索問題」として捉えることができます。エージェントはさまざまな代替案を試したり、ベンチマークを実行したり、結果を比較したりできます。これらが特に魅力的なのは、これらの作業を従来の何倍もの速度で、かつ並列に実行できる点です。
ただし、エージェントには明確なフィードバックシグナルが必要です。これにより、「正しい方向に進んでいるか」「行った行動が適切だったか」「一部の作業を破棄すべきか、それとも活用すべきか」を判断できます。そのためには、さまざまな指標(メトリクス)が必要不可欠です。
もし実際のプロダクション環境でのワークロード再現を試みるのであれば、エンジン自体が信頼性の高い実 workload を正確に再現できる必要があります。そうでなければ、本物の運用環境と合致しない何らかの最適化を行ってしまうリスクがあります。これは以前からパフォーマンスチューニングに取り組む際に直面していた課題と同じです。
また、迅速なフィードバックループも重要です。フィードバックが速ければ速いほど、エージェントは待機時間を減らしながらより多くの作業を遂行できるようになります。
フィードバックループに12時間かかるようであれば、エージェントが本当に価値を生んでいるとは言い難いでしょう。人間の介入による遅延がそれほど大きくないからです。一方、数分以内にフィードバックを得られれば、自分が正しい方向に進めているかすぐに判断でき、意思決定と実行を迅速に行えます。その結果、プロセス全体が劇的に加速します。
重要なのは、適切な指標を設定することです。指標はエージェントの行動を導く羅針盤となります。間違ったものを測定すれば、エージェントも間違ったものを最適化してしまい、製品自体の改善にはつながりません。ユーザーがプロセスを通じて実際にどう感じているかを把握し、それを正しく定量化することが何より重要です。
先ほど述べた通り、エージェントは人間よりもはるかに多くの情報を処理できます。ただし、コンテキストウィンドウを適切に管理し、送信する情報の質と量をコントロールすることは不可欠です。
システムの詳細なアーキテクチャ定義を文書化している場合、その情報は非常に重要です。コードベースやデプロイに関するすべての詳細、そして観測データ(observability)など、あらゆる情報を提供できます。これらはすべて、エージェントが何をすべきかを判断するためのコンテキストとして機能します。
人間とは異なり、ここでは属人的な知識(tribal knowledge)に依存する必要はありません。確かにチーム内の暗黙知や記憶は存在しますが、一般的には、エージェントが正しい行動をとれるよう、こうした文脈を明確に提供することが望ましいです。
我々が目指しているのは、人間が方向性を示すだけで、すべての作業、特に意思決定や善悪の判断に関わる退屈な業務をエージェントが担う世界です。
成長が加速しているのはもちろん事実ですが、開発のフロー自体が大きく変化しています。以前よりもはるかに速いペースでリリースや変更が行われるようになり、これが「良いのか悪いのか」という議論を呼んでいます。
変更頻度の増加は、パフォーマンスだけでなく、開発者体験全体に波及します。CI エンジンのサポートやコードリポジトリの管理に関わるすべての領域が影響を受けます。この変化のスピードに対応するためには、多岐にわたる分野での支援体制を整える必要があります。
AI といえばまず GPU が思い浮かびますが、それだけがすべてではありません。推論処理のためにデータが GPU に送られる前段階でも、多くの最適化が行われています。そして何より、パフォーマンスを維持するためのワークフロー自体も進化させなければなりません。変化のスピードに対応できなければ、システム全体のバランスは崩れてしまいます。
詳細な資料については、字幕付きプレゼンテーションをご覧ください。
録画日:2026 年 8 月 8 日
原文を表示
Keeping ChatGPT Fast as AI Development Accelerates
View Presentation
Speed:
50:14
/presentations/openai-performance-engineering-agentic-coding/en/slides/slide-1785151851732.jpg)
まとめ
Martin Spier explains how agentic workflows dramatically increase code change volume at OpenAI. He discusses the hidden systemic performance costs of rapid shipping beyond GPUs, and shares how deploying always-on AI agents automates profiling, regression detection, and continuous optimization to maintain product speed and scalability at massive global scale.
Bio
Martin Spier leads ChatGPT Performance at OpenAI, where he works on making ChatGPT faster, more reliable, and easier for engineering teams to operate at scale. His work spans AI and cloud infrastructure, performance engineering, observability, reliability, platform engineering, and developer productivity.
About the conference
QCon AI is a practitioner-led event focused entirely on the engineering discipline required to scale these workloads safely. It provides direct
access to the architectural playbooks and failure metrics that peer organizations use in production.
INFOQ EVENTS
- August 27th, 2026, 1 PM EDT
Below the Framework: Why Agent Context Is an Infrastructure Problem
Presented by: Boyd Stowe - Founding Solutions Architect at Tacnode
Transcript
Martin Spier: Today I want to tell you guys a story about two accelerations that are happening somewhat at the same time. First one is around growth, how we're actually growing user base faster than ever. Not just OpenAI, ChatGPT, but in general companies are reaching their first few million users a lot faster than before. The second one is how our development workflows are changing given agentic coding and all those things, and how that is changing the rate of shipping of things we're getting out of the door, and how that is actually affecting how we keep our products fast, efficient, so overall performance engineering.
My name is Martin. I lead the ChatGPT performance team at OpenAI. Almost my whole technical career was in performance engineering. I think I was doing that for 16 years or so, or more. The bulk of that time was at Netflix, where I spent, I think, almost 9 years. Dio mentioned the early cloud days, that's around when I joined, when people thought, will I actually leave my data in the cloud? That's like magical thing there. It feels like the moment is similar today. I spent some time at Snowflake, Expedia, and a few other companies doing that as well. Prior to OpenAI, I was with another AI company called Parasail. We did inference as a service. I was leading the whole engineering team there. Prior to that, I was at a fintech out of Brazil called PicPay, where I was leading whole infrastructure platform and developer experience. As you can imagine, my whole career is really in these central organizations. It's an area I really like. I call it the plumbing of the house. No one sees that it's there, but once it blocks and it breaks, all hell breaks loose. It's always a fun part of the company to be at.
Growth Pace
Three things changed. The first one is growth pace. Not just OpenAI, like I mentioned. Companies are reaching their first million, first 10 million, first 100 million, maybe more users a lot faster than before. The curves are exponential, but a lot faster than before. Also, we're shipping more code and more things out. We're getting more things out of the door with agentic development. There's probably some talk here around AI slop and everything that's going on. We are shipping more logic out the door. The third thing is, previously, we always had this assumption that there was a human that understood all the changes that were going out. It's like someone actually coded that and architected it and know what's going on before they decide to push something out. That is not entirely true anymore, at least from my point of view today. I know there might be a slight bubble, but the abstraction layer is a bit higher.
People are delegating more things to the agents, so they don't know the details of everything that might be going out the door. It's something we need to keep in mind. Of course, with all of that, it puts pressure not just on performance. Performance is one of those problems, but everything else that goes around it. Anyone from GitHub, GitLab here, or anything like that? It's putting a lot of pressure on that infrastructure, but performance too. We're just using more and more resources. We're adding logic that adds latency. It introduces a lot of challenges, and we need to keep up. We need to evolve our practices so we can keep up with that fast pace of growth and the fast pace of development.
Going into ChatGPT, a bit of the history, just to give you a bit of context about the scale I'm talking about. ChatGPT was not launched in late 2022, as a full-fledged consumer-based application that was built from scratch to support hundreds of millions of users. It was launched as a research preview. Here it is. It's pretty cool. You guys should try it out. Quickly, users started coming. Once you have those users coming and coming pretty fast, that research preview becomes a product, and product has requirements. You have latency requirements. You have reliability requirements. You have support requirements and everything that goes behind the scenes to make that work for a large volume of users, especially if they're relying on that for their day-to-day on the things that they need to do. The growth was quite impressive. I think very few companies can tell that they reached their first million users in five days from launch in a research preview.
This is pretty cool. We have a million users here. You can imagine how behind the scenes the team supporting the infrastructure were at those times. Incredible growth, incredible for the product, a lot of pressure on everyone behind the scenes. That growth, it did not stop. It did not slow down much. ChatGPT continued to grow and it's still growing pretty fast. The last official number I can share is from February this year, I believe, when we reached 900 million weekly active users. That's 900 million people using ChatGPT every week. Just to let it sink in, that's almost 11% of the whole human population. It's a lot of people. As you can imagine, chat, it's not the lightest thing to do compute-wise. There's a lot that is necessary on the back to make that happen. That growth was not smooth. It was not just adding a few million users every day and keeps getting pretty predictable.
It was full of spikes, as you can imagine. We were launching products all the time. Some of those products became quite viral. Huge spikes on users, and that, of course, drove adoption. Just this one, you might remember the image generation spike from last year. Maybe you generated one of those images that I cannot say the name right now. Within just the seven first days, over 700 million images were generated by over 130 million users. Image generation, pretty quick. That was putting a lot of pressure on the teams behind the scenes, scrambling. This is the thing that's quite hard to predict when something like that becomes hugely popular. Again, it's something we want. Everyone wants their product to grow. This is really great.
Chat is not used only in the U.S. It's used all across the globe. With that, if you support infrastructure, architecture, all those things, you know that adds complexity to things. Like, where are my POPs? Where are my actual application servers? In our case, where are our GPUs? Are they closely together? Do we have enough capacity in all those regions? Or, how is the network connectivity between all those different data centers? The architecture, where is data? How is data being replicated? Where is data located, given all the different constraints you have across the globe? It adds a lot of complexity to the architecture to support a good experience, or to be able to do and provide a really good experience to all our users across the globe. It was not just growth, but the product did not stand still either. It wasn't just a launch or just users joining.
There was a lot of change beyond all just the model changes. The models were extremely important, but there were a lot of launches along those four years, less than four years now. There were GPTs that were voice. There were agents, all different things, image generation, a lot of product launches, some a lot more viral than others. The product was changing constantly. A lot more things, different workloads that had to be supported, more complex architectures behind it to support all of that. The product was changing while it was growing really fast as well.
Change Velocity
This is, I wouldn't call it new, but it definitely accelerated more since late last year. Especially since the launch of Codex, our shipping rate changed dramatically. I'm not discussing value or anything like that. That's a way deeper discussion, but it's just the volume of change entering our systems increased a lot. When everyone thinks about agentic development, the first thing that comes to mind is automating the proper coding tasks, like develop this function or that application, or refactor this, refactor that, and reduce the overall time it takes to get that done. Shorten the time to get from an idea to a PR that gets merged to production. The part that really surprised me is not just that, especially with newer models where the agents are a lot more capable, they can take more complex tasks. We can delegate things that will run for a lot longer.
Developers started going multi-threaded. They started actually not working on a single thing at a time. Developers are working with maybe 7, 8, 9, 10 different things at the same time. They keep on blocking them as things go. I don't know how much you're coding these days with agents. You're probably hearing the pings saying, I need the user input right now. They have multiple development boxes, or maybe you're working with org trees. In general, a lot more parallel things going on. It's not just faster to ship, but shipping more things in parallel.
Some of you might be familiar with this benchmark from DX, where basically they study a few companies, got some numbers about the volume of PRs that get shipped, that merge per engineer per week. Then you split that into different size of companies and the type of company. As you can expect, smaller tech companies are shipping more PRs, they're getting more things out. At p90, you're almost at five PRs a week per engineer, which is quite fast. We are a lot faster than that. We're shipping a lot more these days. Unfortunately, the latest number I can share is from late last year, so October, November. If you've been following, you know there was quite a huge development since late last year in coding agents. Even back then, back in October '25, the volume of PRs that our engineers are shipping every week increased by 70%. Almost every engineer at the company today is using Codex on a weekly or probably a daily basis.
Every single PR gets automatically reviewed by Codex. It's not just the coding tasks. Like I mentioned, Codex became the workbench to do everything you need to do during that whole development workflow. It's not just coding, but I found a bug in production, go troubleshoot it. Or, I want to understand a metric, a regression, go talk with the observability tools and fetch the metrics and explain to you what's going on. Or to do data science and data analysis, I use Codex to do that as well. All the other boring things like, please summarize that Slack thread for me because I don't want to read it all. A lot of productivity tasks that just make your day-to-day a lot easier. Not just engineers. I manage a team. I probably don't spend too much time coding these days anymore, but I use Codex for everything. To prepare my one-on-ones, to summarize and write documents. Even this presentation, I started it on Codex to go over the outline and to discuss what I actually wanted to present and how to slice that, and all the images that got generated, all the visuals in this presentation. It's being used for everything on your day-to-day, not just your pure coding tasks.
With that evolving of the software engineering workflow, we're shipping things faster. It takes less time for things to get out. We're working on more things in parallel. We're delegating more. That agentic development increased throughput and parallelism. That increased the volume of changes going into our system. I'm not saying value, feature, I think just changes, things that are going out. It might be a feature, might be a bug fix, might be a refactor, might be something else, might be a config thing, but changes going out. Shipping all those things, every new feature, every new line of code, it has a hidden cost. Ideally, that hidden cost adds value. It's something you want. It's something you want out. It's a new feature, a bug fix, but it has a cost somewhere. The cost might be tiny initially. It might be invisible. You might not notice it in the first place.
That increased rate of change just exacerbates that. It's nothing new, but it's happening a bit faster. All that change going out, all those new extra if conditional statements, every extra network request you're making, every data structure that you decide to hold in memory, it's consuming from a shared budget. Your latency, it can get so bad to a certain extent until your users decide to stop. If you're adding logic, you're doing more things, you're adding more checks, whatever you're doing, that always adds a little bit of latency just because you're doing a bit more. It just expands to the point where, this experience is not good, and I don't want to use your product anymore. Same goes for hardware resources. You only have a finite amount of memory. You only have a finite amount of RAM. Everyone will raise, cloud can expand forever. No, it cannot.
It's not as elastic as you think. It just makes things a bit less efficient. It's consuming from a shared budget. I'm not talking purely about those big regressions, like things that you clearly see from one deploy before and after that something drastic changed. I'm talking about even small things. Like I said, that small extra if statement that you added to check something that is important. Small things that on a database, they don't make a difference. There's no difference at all when you're comparing before and after. All of that compounds. You keep building up a little bit, little bit, little bit, until it becomes a bigger problem. It becomes a concern. You might be running out of headroom. It might be making your latencies too bad for users to have a pretty good experience in the app. It just compounds. With that increased rate of change, it's just compounding a lot faster.
We're shortening that time window until we have to actually start worrying about those things. That damage from all those small changes and sometimes big regressions, it will show up later. For the users, it will show up in slowness in the app itself, as you can imagine. For efficiency, no, you will be paying more to run exactly the same workload. You add more logs, you do more things, you still have the same workload. You're just paying more for it. For scalability, you have less and less headroom to grow. As you guys probably can imagine, a scalability problem will become a reliability problem at some point when you cannot absorb that spike anymore, when you cannot grow anymore. The damage will appear later. You need to address it as fast as possible.
Performance Engineering
This is nothing new. We've been developing features and adding features and adding logic our whole careers. In one side, you have people shipping features and shipping things. On the other side, you might have a performance engineering team, or you might have a few perf-minded individuals that care about efficiency and care about latency and making things as efficient as possible, balancing that. You have two opposing forces and keeping things in balance. The problem is that with that increased rate of change, one side of the equation got a lot faster. We're shipping more and perf engineering or folks working in performance problems, they need to adapt. We need to keep up. The answer to that is, of course, we don't want to hire more performance engineers. It's hard to find perf engineers in the first place. Again, not the best answer. We need to make our workflows faster as well.
We need to speed up the cycle on the perf engineering side too. We need to speed up our reactive flows, our reactive loops, where we detect things a lot faster. We can profile that and we can get to the root cause of the problem faster. Not just that. Also, have the agents understand the codebase, propose a solution, implement that thing, and get it out and get benchmarked and get tested. When we actually get a notification, "Great, I already have a fix for that." The benchmarks are looking good again. We can continue. Also, on the active side, we can have a lot of agents in parallel looking for possible optimizations. We can have a lot of different skills working on different types of optimizations. You can be looking at the latency hot path. You can be looking at what methods are actually taking most CPU and try to optimize those. You can be looking at allocations and continuously improve that, or bundle size. They all specialize looking at different things, but working nonstop trying to find those optimizations. You don't have to do that in serial.
We need to adapt. We need to move a lot faster. What happens if we don't adapt as fast as we need? Of course, there'll be a lot of chaos behind the scenes, but beyond that, what's the first few things your users will feel? First things are generally slowness. You add logic, things get slow until it becomes a problem. Your users stop coming. They stop using the app. They cancel their plans, or even worse. Maybe got bad to a point where you're just throwing more errors. You're down for longer periods of time. Things take longer to get fixed, but your users are suffering. Especially for a large consumer-based application like ChatGPT, and it might be the case for where you guys work, you don't want that. That has a deep effect in the business bottom line. Performance affects user retention and acquisition. I modeled that at Netflix.
I modeled that at OpenAI. I'm pretty sure if you guys try to model that as well in your companies, you'll see the correlation between things getting slow and you're throwing more errors and users just canceling plans. Direct impact in the business. One thing I've noticed too, especially for AI applications, is everyone is very focused on inference. Inference is the hard part. Inference is the resources we need to watch all the time. They're really focusing on inference components, like your time to first token or your tokens per second, your throughput, which are important. Your component metrics are super important. Inference is a big chunk of the problem, but it's not the whole journey. It's not the whole problem. We need to measure and keep watching what the users actually feel. When the users come to your product, they're generally trying to solve a problem. They're trying to do something, perform an action, and during that process, while they're trying to perform that action, they have different feelings.
They have different expectations. We need to be able to capture that, and then we need to drill down into. Take, for example, this. This is a bit more focused on a chat app, but the idea is the same. I used the same thing at Netflix. I think Chrome uses the same thing as well. The users have different expectations along the action journey. When he types in a message and clicks enter or clicks submit, he generally expects something to come back really quickly, instant, just feedback to that action, just to know that I'm not hanging here or anything. I got a feedback. I might get a spinner. I might get a thinking thing going on there, but something is going on. It's not on me anymore. The product itself, it's working. I just need to wait a little bit. Then comes something I call first visible value.
When can the user continue doing what he wanted to do in the first place? In our case, if it's a chat request, the user wants to start reading that answer. The time it takes to get the first few tokens, and the user can start reading that. You don't want the users to stop reading because the throughput, the stream cadence is slow. You need to make sure that the stream cadence is pretty good so the user can continue reading without getting blocked and frustrated again. Then, the time it takes to finish that message, the time it takes to complete the task itself. Even that has a lot of nuance. Where do you stop the timer specifically? For a chat response, do you stop at the first token? Do you stop at the first few words? When can a user actually extract some value? It depends on the product.
When you have a thinking response, where do you stop the timer? Do you stop at the thinking tokens, or do you actually stop when the user can do the actual final message? In an image, do you stop the timer when the final image is generated, or the first low-resolution image is generated? There's a lot of nuances, and there's a lot of thinking you need to do around product to actually have those numbers right. Having those intents is super important as a dimension as well because the expectations, they also vary depending on what the user is trying to do. He might be asking a very simple question like, what's the capital of France? I can go to a search engine and get a response pretty quickly, so my expectation is that the response should be pretty quick as well. Great. I need to respond faster on those cases.
I actually fired up an agentic loop that needs to go and check a lot of sources. It needs to do some code edits. It needs to do a lot of things, so I expect that to take a bit longer. The expectation for that is longer as well. Image generation is heavier, the same thing. Users have different expectations.
I have all those different timers, those different wall times, different intents, a lot of different dimensions. Now to be able to act on that, then I go and decompose that into system evidence, into specific components I want to track, things that compose that large wall time. I like to use this chart. There's a lot of different options. My team likes this as well. Your latency breakdown chart or layer cake where you have your distribution on the x-axis and actual latency on the y-axis, and the different layers depending on what's going on. The chart is not the important part here, but there's a lot that happens beyond inference. Of course, this is not to scale numbers or anything. It's just to demonstrate. Yes, there is client work. There is networking time. There is time to fetch data, to do serialization, to do tokenization. There's a lot that goes on beyond just inference.
That breakdown into layers, it helps us actually focus our efforts. Also, the intent I mentioned helps us focus our efforts. If we see that the large chunk of time is not inference itself, it's maybe data fetching, we can focus on that and try to improve that. Having visibility the whole path, it's super important. Otherwise, you just default to inference, which is the first thing that comes to mind to everyone. Inference is just part of the problem, as you can imagine. Take, for example, our chat application. Chat is very simple. The interface is simple on purpose. You go, you type your message. You send that message. You wait for a bit, you get the response back. I don't know if many of you developed this sort of system before, but we can't just forward that message to the inference engine. It does not keep that state of your whole conversation and everything that it needs to properly respond to that.
There's a lot that happens before a message gets sent to the inference engine. There's client work on, of course, assembling everything that needs to be assembled at the client side, fixing things there. Once things come back, we need to verify the user identity. We need to verify the client identity. We need to check the user's plan. We need to check quotas. We need to check the actual user state. There's a lot that needs to happen. Then comes a fairly heavy part of it, which is, we need to fetch context. Conversations, they can get pretty long. We need to fetch all that conversation history. The user might have uploaded files like PDFs, images, whatever, to that conversation. We need to fetch those files as well. If you are in a project, we need to fetch that project context as well. A lot of data that needs to be fetched until we start assembling what actually gets sent to the inference engine.
Great. I have all the data. I have everything I need. There's encoding. There's tokenization, and of course every module is like different tokenization. Everything is tokenized. We need to look at the context window, like, is it exceeding the context window the user has? Maybe we need to do truncation. Maybe we need to do some compaction. Then we need to assemble everything to be sent to the inference engine. There's a lot that goes on even before something gets to the inference engine. Then there's a lot that goes on after, we're streaming the message back, because we need to store the conversation somewhere as well later too. The message, as you can expect, what the user sent, that may be a small string, a small question, is just one small ingredient of that whole equation. Especially for long conversations, they are pretty data heavy. You might have conversations that are megabytes long, tens of megabytes long, maybe hundreds of megabytes long.
There's a lot of requests. There's a lot of data moving around our systems. That data, as you can imagine, it consumes all those resources I mentioned before. We are, of course, consuming a lot of database requests, Blob storage, there's I/O, moving things around. There's a lot of CPU serializing, deserializing things. There is a lot of RAM, because you probably need to hold things in memory for a little bit as well. It consumes all those traditional resources that are not GPU. A message I want to give here is just, when people think about AI, the first thing that comes to mind is GPUs. We need GPUs. GPUs are really important. Don't get me wrong, it's super important. Everything else, all your CPU workloads, they are extremely important as well. There's a lot of resources that are taken to make the whole experience work the way it works, so we need to watch out for those things.
Bringing this up to you, because, as you can imagine, there are code changes, there are improvements, there are shipping rate changes to the actual inference engine, but the rest of the product is a lot larger, as there's a lot more components. There's a lot more code change going to the other components that affect CPU and memory, I/O, all those things, sometimes even more than GPU, so bringing that up. Why is that important? It doesn't matter that I have an extremely fast model, the fastest model, the fastest inference engine, if the rest of the path is slow. The user will get a slow product at the end of the day anyway, and will decide not to use it. We need to focus on the whole experience, the whole path of the request, to make sure things are moving faster. The users get that pretty good feeling when they're using the product.
Human in the Loop
Summing things up a little bit, the product path is getting broader, and it's more than inference. It includes all that data, context, tokenization, streaming. That path is getting heavier and heavier with all the agentic coding, and all the changes we're making, and that's just accelerating. We're consuming more resources, and those resources are not just GPU. Keeping in mind, we need to evolve the way we think about perf engineering. Again, lots of changes need to speed things up on our side as well. We, of course, don't want to ask our engineers across the company to slow down or add roadblocks that make things slower. Every AI PR that is open needs to be reviewed by five humans, so we don't want to add things that will slow down, that will just take away the whole advantage we have of using AI agents. We actually need to speed up the different loops that we use in perf engineering, both the reactive loop when something regresses and we need to fix.
We need that to be faster. We need that to be closer to the actual code, or where the problem originated in the first place, and the active loop itself that needs to be a lot more parallel. Usually, our reactive loop, it takes quite a bit of time. Are there any perf engineers around here that did profiling and optimizing things in the past? The loop is quite slow. You find a regression, maybe from one release to another, you see that your CPU regressed a few percent. What do you do next? You need to do some profiling. You need to go and capture your CPU profiling, which stacks are actually in CPU the most time. Hopefully, you have continuous profiling so you can actually compare a before and after and easily spot where things regressed. If you don't, then things become even a bit more complex. You need to look at the code changes, try to map that to the stacks that are running CPU, and see what regressed so we find an optimization.
Then you need to think of optimizations, and you need to implement those optimizations. Then you need to go and ship it, test it, and see if things fix. All of that is serial, being done by one person, and it just takes a bit too long. The path I'm exploring within a team and the direction I'm seeing things go is having things always on, having our agents always on, working non-stop in both loops, both reactive and also the active loop. We want to automate not only the large regressions, so things you can spot easily from one release to another, where generally you can have an agent triggered automatically by that regression. The agent could go and fire a profile. The agent could actually compare two profiles. We have that today. We have performance skills on Codex to do that. Pretty straightforward if you have all those things working in your systems.
The agent has access to the codebase. He can compare before and after what PRs got into that code change. He can go and propose a fix, and hopefully your systems are safe enough to the point where maybe the agent can deploy that and test if that fix actually had the expected effect. Same goes for drift. Small regressions, we can have fine-grained metrics of maybe code complexity, maybe the volume of network requests that gets executed during a benchmark. We can have small measures to find that drift and act on that drift as well. Of course, the active optimization loop that I mentioned before. The idea here, this is more of an end scenario, that's where we want things to go to. There's a lot of intermediary steps on that. We want everything that happens pre-production, so ideally before a merge, or before something gets deployed, all the boring things be caught by a CI.
Pretty straightforward. We've been doing that for a long time. Nothing new there. As you can imagine, there's a lot of things, especially on the perf side, that don't manifest until they get to production, until they get a specific combination of user workloads. There's no way of catching that pre-production in the first place, so we need to catch that and we need to fix that once we see. Again, active profiling.
A few examples. We can have microbenchmarks running all the time on every single PR. Agents can actually help develop those microbenchmarks. We can find a lot of different things. We can find though that performance, that method regression. We can actually monitor all the methods that are part of our latency critical path, and find specific regressions on those. We can profile memory allocations, and we can see if any specific function is allocating more memory, and we can act upon that. We can watch bundle sizes. Especially if you publish into mobile, you can watch bundle size if these things are increasing and if that exceeds your budget or not. The message here is we've probably been doing that for a long time, but we can automatically act upon it. We can go deeper. We can look at the metrics. We can look at the codebase. We can check all our architecture diagrams and all the information we have about the system running, and the agent can propose a few ideas, a few ways of fixing that.
We can go and implement that. Ideally, you can actually benchmark, run the same benchmarks again, and compare before and after. Yes, many times that fix will not work. It will not fix the problem, and you can just throw that away and continue doing. You can fire multiple threads with multiple ideas. That's the point. Of course, this is end state. Ideally, there are a lot of steps to get to that in the first place. You can actually have a set of skills in Codex that just make that process faster. Instead of manually triggering profiles, collecting things, and analyzing, you can have an agent collect that information and do that comparison for you. It's something that can be easily achieved. The agents are pretty good at chugging a lot of data and comparing things. Pretty easy, you can generate a flame graph out of it if you want.
Fairly straightforward, especially with microbenchmarks where you can reproduce things pretty easily. Same goes for production. You want pretty good observability in production. You want to catch all those different regressions that you see. Again, act upon those. Same thing, you get the same information. The difference here is the safety. Things are a bit trickier if you're asking an agent to deploy something to production and you don't want your users to feel any issues, feel any pain. There's a lot of best practice that you need implemented even before you could do that. You can get almost all the way there. You can actually have an agent submit a PR that implements an experiment of a perf optimization. You can do that. If you're not at a stage where you can fully trust an agent to deploy to production without any supervision, you can get almost all the way there, and it should work really well.
The always-on loop is the one that probably excites me the most. In my team we have a lot of really great perf engineers, and we're in the process of how can I actually translate all that skill set into a set of skills to my agents that can continuously be looking for optimizations in my environment. They could be looking at the code. They could be looking at my traces. They could be looking at logs, whatever, but continuously finding optimizations. It's funny too, the habit people have when they have a very specific niche skill set, they name the agents after themselves. I have a Ben agent. I have a Brendan agent. I have lots of agents named as people. Then you start treating them as people at some point too. It's weird, that always-on loop trying to find opportunities.
Proven Best Practices
A few best practices of things that have been working well for us. When you can get agents to complete full loops, things become a lot more autonomous. You can have a lot of agents doing a lot of things at the same time, if they can complete the loop. Meaning, I start the loop from some event or a continuous loop, I can have an idea. The agent can implement that, and it can do the whole thing all the way through actually measuring if that change you made was positive and it worked. We can continue doing that without any human intervention. Sometimes it's not that easy, and at other times it's a bit easier. If you can focus on trying to build all the safety to allow agents to complete full loops without any intervention, things will become a lot faster. A lot less supervision. A lot less bottleneck on people having to monitor.
Then just have people not clicking accept, accept, accept, or just do YOLO mode and let the agents do their thing. As you can imagine too, the basic things we've been doing for a long time, they become more and more important. Your engineering maturity becomes more important. Having really good test coverage, it's important. Otherwise, how would the agents know what they did was right or it's working? It's functionally working or not? Benchmarks, the same way. How can the agents measure if the improvement they implemented was good or not? Having those benchmarks, having those tests, extremely important, became even more important now. You need your contracts expectations between the different parts of the system. Observability, super important. You need really good coverage about your system so the agent can understand what's going on. If there's blind spots, it's very likely that it might go in the wrong direction because it doesn't have that intuition about the system that the humans might have. That safe rollout and maybe an automated canary analysis or a blue-green push, all those safety measures to get the code out even if maybe there's a bug there, but that shouldn't affect users. Roll out in a safe way.
At the end of the day, all those perf optimizations, that's a search problem. The agents can try the different alternatives. They can try the benchmarks. They can compare things. It's really cool because they can do all of that a lot faster, and they can do all those things in parallel. The agents, they do need a clear signal, again, to know if they're moving in the right direction, if what they did is right or wrong, or if they can discard some work or actually use it. They need all those metrics. If you're actually trying to reproduce a production workload, you need actual reliable workloads to be reproduced by the engine itself. Otherwise, you might be optimizing something that doesn't really match production. Same problem you had before as you are working with perf. Fast feedback is important too. The faster the feedback loop, the more the agent can do without waiting.
If your feedback loop might be 12 hours, then maybe your agent doesn't really add that much value because the human interruption is not that much. If you can have a fast feedback, within a few minutes I can know if I'm in the right direction so I can make a decision and move, in one way or another, things become a lot faster. That measurement you have, it just guides the agent. Having the right measurements is super important. If I'm measuring the wrong thing, the agent will optimize the wrong thing and I'm not actually improving the product in the first place. Defining the metric correctly, making sure you're capturing that. That's why I focus a lot in what the user is perceiving during the whole process. Focus on that. Agents, like I mentioned, they can chug through a lot more information than humans. Of course, you need to manage the context window and what you're actually sending to be useful.
You can send a lot of information, all your observability, all the details about your deploys, all your codebase. If you have actual written down definitions of the architecture of how the system works, it's extremely important. It's all context that the agent uses to define things. Different than humans, you don't have tribal knowledge here. You might have tribal knowledge and memories, but, in general, you want to provide that context so agents can do the right thing. General idea, we are working towards a world where humans are just setting the direction and our agents are just doing all the work, at least the boring work that is not making the decisions and defining what is good or bad.
Key Insights
It's not just growth, growth is increasing, of course, but the development flow changed significantly. We're shipping and making changes a lot faster than we were before. There's arguments around if it's good or bad. We are making a lot more changes than we were before, and that has implications in other areas of development too. Not just performance, but all your developer experience, whoever is supporting your CI engines or your code repositories. There's a lot of impact in different areas to support that rate of change. When you're thinking about AI, the first thing that comes to mind is GPUs. It's not just GPUs. There's a lot of optimizations, a lot of things that happen even before anything gets sent to be processed by a GPU for inference. Last but not least, our perf workflow needs to change. All the different workflows that are affected by that different rate of change, that increased rate of change, it needs to evolve to keep up. Otherwise, our balance will be wrong.
See more presentations with transcripts
Recorded at:
Aug 08, 2026
同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み