マイクロソフト、AI エージェントの大量運用手法を公開
マイクロソフトが数千の生産用 AI エージェントを運用する具体的なアーキテクチャと運用戦略を明らかにし、AI エージェントの実装における実用的なベストプラクティスを提示した。
キーポイント
大規模エージェント運用の基盤
マイクロソフトは単一のモデルに依存せず、数千もの異なるタスクを処理する生産用 AI エージェントを並列で運用しており、その背後には堅牢なオーケストレーション層が存在する。
信頼性と安全性の確保
膨大な数のエージェントを管理するために、厳格なガバナンスフレームワークと監視システムを導入し、誤動作やハルシネーションを防ぐ仕組みが構築されている。
実用的なアーキテクチャパターン
複雑なタスクを処理するために、専門化されたエージェント(Specialized Agents)と調整役のエージェント(Coordinator Agents)を組み合わせた階層型アプローチを採用している。
重要な引用
We ship thousands of production AI agents
Reliability is not an afterthought; it's the foundation
The key to scaling isn't just more compute, it's better orchestration
影響分析・編集コメントを表示
影響分析
この記事は、AI エージェント技術が理論や実験室から実社会の大規模運用へと移行した重要な転換点を示しています。マイクロソフトのような大手テック企業が「数千」という具体的な数値で生産用エージェントを公開することは、業界全体におけるエージェントベースのアーキテクチャの実用性と信頼性を裏付ける強力な証拠となり、他社や開発者に対する実装の指針となるでしょう。
編集コメント
マイクロソフトが「数千」という具体的な規模で生産用 AI エージェントを運用しているという事実は、AI エージェント技術の実用化スピードが我々の予想を上回っていることを示唆しています。特に信頼性確保のためのアーキテクチャ設計は、自社でエージェント導入を検討する企業にとって極めて参考になる実例です。
SSO のデバッグ、ユーザー管理、認証ポリシーの調整、ブランディングの設定など、これまですべての設定作業は人間しか操作できない UI の背後で行われてきました。
WorkOS MCP サーバーを使えば、エージェントにもダッシュボードへのログインと同じアクセス権限を付与できます。実行時に検出可能な数百もの操作に対応可能です。OAuth を介して 1 コマンドで接続し、マスター API キーの代わりにスコープ限定トークンを使用します。マーケティングサイトのスクリーンショットを渡して「ログインページに合わせて調整して」と指示するだけで、以前は人間にしかできなかった作業を今やエージェントが実行できます。
Microsoft は驚異的な規模で事業を展開しています。現在、8 万社以上の企業が Microsoft Foundry を基盤として AI エージェントやアプリケーションの構築・デプロイ・運用を行っています。Microsoft 自身のコパイロットも同プラットフォーム上で稼働しており、その代表例が Microsoft 365 Copilot です。この製品だけで 2,000 万人以上のユーザーに提供されており、自社製エージェントの月間アクティブ利用数は今年に入って前年比で 6 倍に成長しています。
そのような規模でエージェントを実際にリリースするために何が必要なのかを理解するため、Microsoft Core AI のプロダクト担当バイスプレジデントである Marco Casalaina 氏にお話を伺いました。氏は、本番環境での運用から得た知見や、それに伴うエンジニアリング上の課題、そしてエンタープライズ AI が次に目指す方向性について詳しく解説してくださいました。
この記事では以下を学びます:
なぜプロトタイプのエージェントは本番環境で生き残れないのか
本番環境のエージェントには何が含まれており、なぜマイクロソフトは「文脈」が鍵だと考えるのか
Foundry の背後にある 2 つのエンジニアリングアイデア:サブエージェントとしての検索機能、そしてエージェントに固有のアイデンティティと活動領域を与えること
マイクロソフトが評価基準に基づく評価(ルブリック・エバル)と自動改善ループを用いて本番環境のエージェントをどう評価しているか
他チームへの教訓と今後の展望
マイクロソフトがエンタープライズ規模で AI エージェントを運用する仕組み(全体像)
プロトタイプでは見えない理由で、本番環境のエージェントは失敗します。問題となるのはモデルそのものではなく、むしろモデルを取り巻くすべてです。エージェントが参照するデータや呼び出すツール、実際のユーザーへの対応方法、そして周囲の環境変化に伴う品質の低下など、システム全体が影響を受けます。
今年、エンタープライズ企業が AI エージェントを本番導入しようとしている際、彼らが直面しているのは昨年の課題とは異なるエンジニアリングの問題です。
本番環境のエージェントは単なるモデルではありません。システムの大部分を占めるのは、そのモデルを取り巻いて構築されたインフラや仕組みです。
なぜそうなのかを理解するには、エンタープライズ企業が今、何を目指して変化しているのかから始める必要があります。マコ氏はこの転換点を以下のように定義しています。
「私たちは AI の『質問応答』フェーズを脱却しようとしています。2026 年には、音声インターフェースを採用する顧客が劇的に増加しており、これはつまり『チャットボット』の時代も終わりを迎えることを意味します。」
従来の形はチャットボットでした。ユーザーが入力し、エージェントが返信するだけで、質問に答えることしかできません。新しい形は、ユーザーの代わりに実質的な作業を行うエージェントです。会議の予約を行い、分析を実行し、メールを送信し、チケットを提出します。ユーザーは一切入力する必要がない場合もあります。フロントエンドが音声インターフェースになるためです。例えば、Foundry の Voice Live を使えば、既存のテキストベースのエージェントを再構築することなく、音声対応のエージェントに変換できます。
チャットボットからエージェントへ、質問への回答から作業の実行へという転換点
この変化こそが、エンジニアリング上の課題を根本的に変える要因です。チャットボットが誤った回答を返すのは「使いにくい体験」で済みますが、エージェントが誤った行動をとれば「ビジネスインシデント(事故)」になります。リリースに耐えうる品質の基準は、以前とは全く異なるレベルへと引き上げられたのです。
ここから、プロトタイプと本番環境のエージェントとの間に大きな隔たりが生じます。最初のプロトタイプを作るのは簡単です。午後に「雰囲気」でコードを書き上げることも可能です。モデルが賢く、テスト用のプロンプトも機能し、デモは印象的であれば、わずか 1 週間でパイロット版をリリースできます。
プロトタイプで想定されるリクエスト数
本番環境こそが、システムにひび割れが生じる場所です。実際のユーザーは、予期していなかった質問を投げかけます。エージェントが依存するドキュメントが古くなり、評価データセットには存在しなかった新たなエッジケースが発生します。モデルの更新によってエージェントの挙動が微妙に変化しても、顧客から苦情が届くまで誰も気づきません。
ID 制御がないと、エージェントは共有されたシステム権限として実行され、問題発生時に監査証跡が残されません。ガードレール(安全装置)がなければ、本来言うべきではないことを自信満々に発言してしまいます。可観測性(オバザビリティ)がなければ、品質が向上しているのか劣化しているのかを判断できません。
これらの問題はプロトタイプ段階では一切現れません。すべてが本番環境で顕在化するのです。
本番環境でのみ表面化し、プロトタイプでは決して現れない失敗について伺いました。
Foundry チームがこうしたシステムを大規模に運用する中で得た最も大きな教訓は何かと Marco に尋ねると、その答えは「モデルと同じくらいハルネス(運用基盤)が重要だ」というものでした。
ハルネスとは、モデルを取り巻くすべての要素のことです。ランタイム、ツール、文脈の取得、ID レイヤー、ガードレール、評価器、デプロイパイプラインなどが含まれます。モデルは常に変化しており、データベースのバージョン管理のように扱うことはできません。Postgres の場合、バージョンを変更すればほぼ確実に動作します。しかし、モデルはそうではありません。それぞれが異なる特性を持っており、ハルネス側でそれらに適応させる必要があります。
Anthropic が Claude Opus 4.8 をリリースした際、Microsoft の GitHub Copilot CLI チームも、これを本番環境に投入する前にハルネスの再調整と評価の実施を余儀なくされました。
新しいモデルリリースに伴うハッチの再チューニング
AI はインフラを再構築していますが、多くのチームが予想していたような形ではありません。最も成果を出しているのは、最新技術をいち早く取り入れるチームではなく、プラットフォーム、ガバナンス、自動化されたパイプラインを活用し、AI 生成によるインフラ変更を安全に本番環境へ移行できるチームです。
2026 インフラ自動化レポートでは、406 名のインフラおよびプラットフォームエンジニアリングのリーダーを対象に調査を行いました。その結果、AI の恩恵を最も受けている「パイオニア層」の特徴が明らかになりました。このレポートでは以下が学べます。
- AI 導入においてなぜプラットフォームエンジニアリングが不可欠なのか
- 従来の指標に加え、監視すべき新しい AI 固有のシグナルとメトリクス
- AI マチュアリティ指数における「パイオニア層」に到達するための 5 つの実践的なステップ
ハッチ(基盤)がモデルと同じくらい重要であるなら、次なる疑問は「その中に何が含まれているのか」という点です。ハッチを下から上へと紐解いていくと、各レイヤーが存在する理由と、それを欠いた場合に何が崩壊するかを理解できます。
本番環境の AI エージェントを構成する 5 つのレイヤー
最下層は「推論レイヤー」です。これはハッチがモデルにアクセスするために使用する単一のインターフェースです。モデル自体はハッチの外側にあり、いつでも差し替え可能です。エージェントごとに必要なモデルは異なり、最適なモデルも数週間に一度変わります。Foundry は OpenAI、Anthropic、xAI、DeepSeek、そしてマイクロソフト独自の MAI ファミリーなど、1 万 1000 以上のモデルをサポートしています。
数千もの差し替え可能なモデルを扱う推論レイヤー
モデルの上層に位置するのが「エージェントランタイム」です。これは単なるモデルを実用的なエージェントへと変換する役割を担います。
このランタイムは、オーケストレーションループの制御、ツール呼び出しの実行、会話状態の管理、そしてハネス全体が使用するプロトコルの処理を一貫して担当します。
ただし、ループ内のすべてのステップをモデルに委ねる必要はありません。適切に設計されたエージェントは、推論が必要な部分だけを大規模言語モデル(LLM)へ送り、それ以外は通常のコードで処理します。データベースの参照や、目的別に特化して構築された抽出モデルの方が、同じ作業を LLM に依頼するよりも高速であり、コストも安く、信頼性も高いからです。
エージェントフレームワークは増え続けており、LangChain、LangGraph、CrewAI といったオープンソースの選択肢から、ベンダーが提供するランタイムまで多岐にわたります。ここで重要なのは「フレームワーク中立性」です。あるフレームワークで構築されたエージェントが、周囲のハネスを書き換えることなく別のフレームワークへも移植可能であるべきです。例えば Foundry では、エージェントを任意のフレームワーク間で自由に切り替えて実行できます。
エージェントランタイム層
エージェントを本番環境に導入したら、組織は可観測性とガバナンスの仕組みが必要です。すべてのプロジェクトで稼働する各エージェントを一元的に把握し、健全性スコアやトークン使用量、レイテンシ指標、ドリフト検知機能、そしてプラットフォームチームが fleet を統制するために不可欠なプロジェクト横断的な集計データを提供できることが求められます。この層がないと、性能の低下が見えなくなり、コストもコントロール不能になります。
Microsoft の場合、これは「Foundry Control Plane」です。同システムはプロジェクト横断的な fleet 可視性を提供し、エージェントからのテレメトリデータを Azure Monitor や Application Insights にルーティングします。これらはすでにインフラストラクチャのアラート処理を担っている同じパイプラインです。
可観測性とガバナンスの層
組織内で実行動を開始するようになると、エージェントには独自のアイデンティティが必要です。各エージェントに固有のロール割り当てと監査証跡を持たせるべきです。なぜなら、振る舞いの悪いエージェントも、同様に問題のある従業員と同じアクセス制御によって制限される必要があるからです。
業界全体でまだ最適な方法が収束しつつありますが、多くのプラットフォームが採用している答えは、並行したシステムを新たに作るのではなく、既存のエンタープライズアイデンティティシステムを拡張して、エージェントを「新たなクラスのプリンシパル」として扱うことです。例えば Foundry は、Microsoft のエンタープライズアイデンティティプラットフォームである Entra を拡張し、このようにエージェントを扱います。これがアクセス制御の基本要素です。後ほど本記事の「エージェントが行動する場所を与える」セクションで、このアイデンティティを取得したエージェントが実際に何を行うのかについて詳しく解説します。

Microsoft は現在、数千もの AI エージェントを実環境で稼働させています。これらは単なる実験段階のプロジェクトではなく、すでに本番環境で重要な役割を果たしています。
この規模を実現するために、Microsoft は「エージェント・プラットフォーム」と呼ばれる独自の基盤を構築しました。これは、個々のエージェントを開発・デプロイ・管理するための包括的なインフラストラクチャです。
同社のアプローチは、従来の「単一の巨大なモデル」に依存するのではなく、「多数の専門特化した小規模エージェント」を協調させるというものです。各エージェントは特定のタスク(例:データ検索、コード実行、ユーザーとの対話)に最適化されており、必要に応じて連携して複雑な作業を遂行します。
このアーキテクチャにより、Microsoft はスケーラビリティと信頼性を両立させています。数千ものエージェントが同時に動作しても、システム全体の安定性は損なわれません。
また、これらのエージェントは人間とのインタラクションを通じて継続的に学習・改善されます。フィードバックループを自動化することで、パフォーマンスの向上が持続的に行われています。
Microsoft のこの取り組みは、AI エージェントの実用化における新たな基準を示すものと言えるでしょう。今後、より多くの産業分野で同様のアプローチが採用されることが予想されます。
データ属性:"src":"https://substack-post-media.s3.amazonaws.com/public/images/67ec5a9d-8304-4487-b3e0-0d8edbb12c3e_2048x1163.png","srcNoWatermark":null,"fullscreen":null,"imageSize":null,"height":827,"width":1456,"resizeWidth":null,"bytes":null,"alt":null,"title":
原文を表示
Debugging SSO, managing users, adjusting auth policies, configuring branding: every configuration task has lived behind a UI that only a human can drive.
The WorkOS MCP server gives agents the same access as your dashboard login. Hundreds of operations, discoverable at runtime. Connect in one command via OAuth, with scoped tokens instead of a master API key. Pass a screenshot of your marketing site and ask your agent to match the login page. If a human had to do it before, an agent can do it now.
Microsoft operates at an enormous scale. More than 80,000 enterprises now build on Microsoft Foundry, the company’s platform for building, deploying, and running AI agents and applications. Microsoft’s own copilots run on the same platform, including Microsoft 365 Copilot, which alone serves over 20 million users and has a monthly active usage of first-party agents growing 6x year-to-date.
To understand what it actually takes to ship agents at that scale, we spoke with Marco Casalaina, VP of Products for Microsoft Core AI. He walked us through what his team has learned from running these systems in production, the engineering challenges that come with it, and where he thinks enterprise AI is headed next.
In this article, you’ll learn:
- Why a prototype agent doesn’t survive in production
- What’s in a production agent harness, and why Microsoft believes context is the key
- Two engineering ideas behind Foundry: retrieval-as-a-subagent, and giving agents their own identity and a place to act
- How Microsoft evaluates production agents with rubric-based evals and an auto-improvement loop
- Lessons for other teams, and what’s next
Production agents fail for reasons that aren’t visible in a prototype. The model is rarely the problem. What breaks is everything around the model, including the data the agent retrieves, the tools it calls, the way it handles real users, and the way quality drifts as the world around it changes. Enterprises trying to ship agents this year are running into a different engineering problem than the one they were solving last year.
To see why, it helps to start with what’s actually changed about what enterprises are trying to build. Marco framed the shift this way:
We are leaving the question-answering phase of AI. In 2026, we are seeing a huge increase in the number of our customers that are using voice as a front-end, so we’re also leaving the chatbot era of AI.
The old shape was a chatbot. The user types, the agent types back, and it can only answer questions. The new shape is an agent that does meaningful work on the user’s behalf. It books the meeting, runs the analysis, sends the email, files the ticket. The user might not type at all because the front end can be voice. For example, Foundry’s Voice Live lets a team turn an existing text agent into a voice agent without rebuilding it.
This shift is what makes the engineering problem different. A chatbot returning a wrong answer is a bad experience. An agent taking the wrong action is a business incident. The bar for what’s good enough to ship has moved.
That’s where the gap between a prototype and a production agent opens up. The first prototype is easy. You can vibe-code one in an afternoon. The model is smart, your test prompts work, the demo is impressive, and the pilot ships in a week.
Production is where the cracks open. Real users ask things you didn’t anticipate. Documents the agent depends on go stale. New edge cases emerge that never appeared in your eval dataset. A model update changes the agent’s behavior subtly, and nobody notices until a customer complains. Without identity controls, the agent runs as a shared system principal and there’s no audit trail when something goes wrong. Without guardrails, it confidently says something it shouldn’t. Without observability, you can’t tell whether quality is improving or degrading. None of these problems showed up in the prototype. All of them show up in production.
When we asked Marco for the single biggest lesson the Foundry team has learned from running these systems at scale, his answer was “the harness matters as much as the model.”
The harness is everything around the model. The runtime, the tools, the context retrieval, the identity layer, the guardrails, the evaluators, the deployment pipeline. Models change constantly, and you cannot treat them like database versions. With Postgres, you change versions and you pretty much expect it to straight up work. Models are not like that. Each one has different properties that the harness has to adjust to. When Anthropic released Claude Opus 4.8, Microsoft’s GitHub Copilot CLI team had to re-tune their harness and re-run their evaluations before they could ship it.
AI is reshaping infrastructure, but not how most teams expected. The teams getting the most from it aren’t adopting the fastest. They’re taking advantage of the platform, governance, and automated pipelines that allow AI-generated infrastructure to move safely to production.
The 2026 Infrastructure Automation Report surveyed 406 infrastructure and platform engineering leaders to reveal how Pioneers are benefiting the most from AI. You’ll learn:
- Why platform engineering is the critical component of AI adoption
- New AI-specific signals and metrics that need monitoring beyond traditional metrics
- 5 tactical steps to reach the Pioneer segment of the Al Maturity index
If the harness matters as much as the model, the next question is what’s in it. Walking the harness from the bottom up shows why each layer exists and what breaks if you try to do without it.
At the bottom is the inference layer, the single interface the harness uses to reach the models. The models themselves sit outside the harness and stay swappable. Different agents need different models, and the right one changes every few weeks. Foundry supports more than 11,000 of them, from OpenAI, Anthropic, xAI, DeepSeek, and Microsoft’s own MAI family.
Above the model is the agent runtime, which is what turns a model into an agent. The runtime handles the orchestration loop, the tool calls, the conversation state, and the protocol the rest of the harness speaks.
Not every step in that loop should run through the model. A well-built agent sends only the parts that need reasoning to the LLM and leaves the rest to ordinary code, since a database lookup or a purpose-built extraction model is faster, cheaper, and more reliable than asking a model to do the same work.
The number of agent frameworks keeps growing, from open-source options like LangChain, LangGraph, and CrewAI to vendor-built runtimes. The principle that matters is framework neutrality. An agent built on one framework should be portable to another without rewriting the surrounding harness. Foundry, for example, lets agents run on any framework interchangeably.
Once agents are in production, an organization needs observability and governance. It needs a single view of every agent running across every project, with health scoring, token usage, latency metrics, drift detection, and the kind of cross-project rollups that let a platform team govern a fleet. Without this layer, regressions are invisible and cost is uncontrolled. In Microsoft’s case, this is Foundry Control Plane, which provides cross-project fleet visibility and routes agent telemetry into Azure Monitor and Application Insights, the same pipeline that already handles infrastructure alerts.
Once agents start taking real actions inside an organization, they need their own identities. They need their own role assignments and their own audit trails, because a misbehaving agent has to be bounded by the same access controls that bound a misbehaving employee. The industry is still converging on how to do this well, and the answer most platforms are landing on is to extend an existing enterprise identity system to treat agents as a new class of principal rather than inventing a parallel system. For example, Foundry extends Entra, Microsoft’s enterprise identity platform, to treat agents this way. This is the access-control primitive; later in the article, the section on giving agents a place to act shows what an agent does once it has one of these identities.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み