Azure、クラウド信頼性向上のためのAIシステム「Brain」を発表
本文の状態
日本語全文を表示中
詳細モードで約19分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Azure AI Blog
Microsoft Azure は、クラウドの信頼性を向上させるための AI 駆動型インテリジェンスシステム「Brain」を発表し、これはデジタルツインと統合された AIOps として機能して顧客への通知や障害対応を加速している。
AI深層分析を開く2026年8月4日 12:09
AI深層分析
キーポイント
Brain の定義と仕組み
Brain は Azure Resource Graph を基盤とし、プラットフォームのテレメトリデータ、AI/ML モデル、サービス依存関係、顧客への影響を統合してクラウド全体のリアルタイムな健康状態を示すデジタルツインとして機能する。
既存の信頼性ワークフローへの適用
同システムは既に顧客リソースの健康通知、デプロイメント safeguards(安全装置)、および障害宣言の決定プロセスを支えており、問題の検知速度と範囲特定精度を向上させている。
大規模インフラにおける必要性
80 以上のリージョンや 50 万キロを超えるケーブル網を有する巨大なクラウド環境において、顧客からの報告よりもシステム側が先に問題を発見することは極めて困難であり、Brain はこのギャップを埋めるために構築された。
次世代の自律型 AI への基盤
Brain は単なる監視ツールではなく、将来的にクラウド運用を再定義する「エージェント型 AI(agentic AI)」の基盤として位置づけられており、洞察を即座に行動に変換する自動化の土台となる。
信頼性の限界要因はツール不足ではなく理解不足
クラウドの信頼性を制限しているのはツールの不足ではなく、膨大な信号を人間が処理しきれないという理解の問題である。従来のダッシュボードやアラートの増加は解決策にならず、リアルタイムで状況を説明する仕組みが必要となる。
重要な引用
Brain is an AIOps-powered cloud health intelligence system that operates as an intelligent layer on top of Azure Resource Graph (ARG)
If you run on Azure, Brain is already changing three things you can notice: How fast we tell you when something is wrong.
This post starts a multi-part series on what Brain is, how we built it, what we've learned operating it at scale, and where it goes next.
"That gap between what we measure and what we know is the limiting factor on cloud reliability today. It is not a tooling problem. We have plenty of tools. It is a comprehension problem."
編集コメントを表示
編集コメント
クラウド運用における AI の役割が、単なる分析から自律的な対応へと進化していることを示す重要な事例である。Microsoft は自社の巨大なインフラを維持するために独自に構築したシステムの詳細を公開しており、業界全体のパターンを示唆している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
本記事では、Azure の AI 駆動型信頼性インテリジェンスシステム「Brain」の仕組みについて解説します。
なぜ Brain が必要なのか、また Brain とは何か(クラウド信頼性を支える Azure の集中型 AIOps)についても触れます。さらに、クラウドヘルスを実現するためのデジタルツインの基盤や、このクラウドインテリジェンスシステムを運用することの意味、アジェンティック AI とクラウド運用の未来、そして Azure 信頼性と Brain の今後の展望について取り上げます。
Brain は、Azure Resource Graph を上位層として統合し、プラットフォームのテレメトリデータ、AI/ML モデル、サービス間の依存関係、顧客への影響を一つの視点に融合させる、AI 駆動型のクラウド信頼性インテリジェンスシステムです。これは、すべてのサービス、リージョン、ワークロードのパフォーマンスを継続的に更新された形で可視化する役割を果たします。
すでに、顧客向けの Azure リソースヘルス通知やデプロイメントの安全装置、障害宣言の決定などを実行しており、Azure の運用方法を変革しつつあるアジェンティック AI の基盤となっています。本シリーズでは、Brain が何であるか、どのように構築されたか、大規模運用で得た教訓、そして今後の展望について多角的に掘り下げていきます。
Azure の信頼性を支えるのは、自社の健康状態をデジタルツインとして再現するシステムです。その中心にあるのが「Brain」で、AIOps(人工知能を活用した運用管理)を駆使したクラウドヘルスインテリジェンスシステムです。Brain は Azure Resource Graph (ARG) の上に構築された知的レイヤーとして機能し、両者が一体となることでデジタルツインを実現しています。
プラットフォームからのテレメトリデータ、AI/ML モデル、そしてデータエンジニアリング技術を統合することで、Azure 上のサービス、リージョン、顧客のワークロードがどのように稼働しているかというリアルタイムな状況を絶えず維持・強化し続けています。この共有された状況把握は、次第に「インサイトをアクションに変える」より高度な自動化基盤へと進化しています。
現在、Brain は Azure 全体で重要な信頼性ワークフローを支えています。具体的には、顧客リソースのヘルス通知、デプロイ時のセーフガード機能、障害宣言などです。Azure を利用している方なら、すでに Brain の影響を以下の3点で実感できるはずです。
- 問題発生時に、どれだけ迅速に通知が届くか
- 影響範囲を顧客のリソースに正確に特定できるか
- 適切な担当エンジニアがどれだけ早く対応を開始するか
この記事では、Brain がどのように機能し、それによって何が変わるのかについて解説します。
本シリーズは複数回に分けて連載する予定で、今回は Brain の概要、構築プロセス、スケールした Azure 運用から得られた知見、そして今後の展望について取り上げます。まずは基礎となる部分から始めましょう。
Azure の信頼性に関する詳細はこちら
Azure は 80 以上のリージョンにまたがる数百のサービス、500 を超えるデータセンター、そして 80 万キロを超える光ファイバーや海底ケーブルを運用しており、世界最大級のクラウドインフラの一つです。しかし、これらの Azure サービスが世界中で生み出し、管理し、処理する膨大な活動量にもかかわらず、静かに劣化が進む日でも、時には自社のシステムが問題に気づく前に顧客から報告が入ることがあります。顧客にとって、これは最悪のインシデントです。なぜなら、障害の原因が Azure 側にあると判明するまで、顧客自身がアプリケーションのデバッグを余儀なくされるからです。
私たちが計測していることと、実際に把握できていることの間のギャップこそが、現在のクラウド信頼性を制限する要因です。これはツール不足の問題ではありません。ツールは十分に揃っています。問題は「理解」にあります。超規模クラウドが生み出す信号の量が、人間がそれを読み解く能力を超えてしまっているのです。従来の解決策である「ダッシュボードを増やす」「アラートを増やす」「オンコールローテーションを強化する」といった対応は、単なる運動会でしかありません。追加されるダッシュボードはオペレーターにとって新たな窓に過ぎず、本当に必要なのは、何を見ているのかを即座に理解させ、行動に移せるような仕組みです。
このギャップを埋めるためには、これまで作ったことのないものを構築する必要がありました。より優れたダッシュボードや、より賢いアラートではありません。プラットフォームの健全性をリアルタイムで推論し、プラットフォームが要求する規模でその結論に基づいて自動的に行動する、継続的に更新されるモデルです。
Brain とは何か?クラウド信頼性を実現する Azure の集中型 AIOps
Brain は、Azure の中央集権型 AIOps 基盤のクラウドヘルスインテリジェンスシステムです。エージェント AI やデータエンジニアリングを含む AI/ML を活用し、Azure の健全性を常時モデル化するとともに、その分析結果に基づいて自動的に信頼性向上アクションを実行します。現在では Azure 本番環境でも広く利用されており、プラットフォーム全体のリソース健康状態の判定を担っています。
Brain の核心は、3 つの要素によって形作られています。それは「何を入力するか」「何を出力するか」、そして「その出力が何を引き起こすか」です。

Brain の概要。
Brain は、3 つの異なるソースクラスからシグナルを取り込みます。
- 標準化されたサービスレベル指標(SLI):Azure の顧客や運用担当者が、すでに信頼性ダッシュボードで慣れ親しんでいる指標です。
- ドメイン固有のモニター:各サービスチームが独自に構築し、Brain に登録したモニター。さらに、デプロイメント情報、サポート件数、サービス間の依存関係を示すシグナルなどを含む広範なテレメトリストリームも含まれます。
- サードパーティ指標:Azure のあらゆる運用を取り巻く外部からの指標です。
それぞれの経路は異なる目的を果たしており、これらを組み合わせることで、単一の経路では不可能な包括的なカバレッジを実現しています。
入力内容に関わらず、Brain はすべての対象(サービス、リージョン、デプロイメントユニット、顧客リソース)を評価し、4 つの出力を返します。それは「健康状態」「深刻度」「影響範囲」、そして「結論に至った理由」です。標準的な用語を用いた標準的な出力形式により、下流システムすべてが共通の言語で通信できるようになります。これにより、「影響」という言葉の意味がチーム間で曖昧になるような齟齬はなくなります。
Brain が生成する洞察は、以下のような一連の自動化された信頼性対策を駆動しています。
アウトエージ宣言:影響範囲(ブラスト・レイジ)に基づいて行われます。
顧客への通知:影響を受けるサブスクリプションとリージョンにターゲットを絞って配信されます。
インシデントのルーティング:適切なサービスチームへ転送されます。
デプロイメントゲート:有害なロールアウトを一時停止する機能です。
関連インシデントのリンク:関連する事象同士をつなぎます。
診断ツール:エンジニアが問題調査を行うのを支援します。
クラウドヘルスにおける Azure のデジタルツインの基盤
「インテリジェンスシステム」が単なる「ダッシュボード」とどう違うのかを理解するには、その基盤に何が含まれているかを見る必要があります。Brain が構築する Azure の表現には、少なくとも以下の要素が含まれます。
トポロジー:Azure Resource Graph によって有効化されたすべてのサービス、リージョン、アベイラビリティゾーン、デプロイメントユニット、依存関係グラフが、ライブモデルとして表現されます。サービスのスケールアップや依存関係の変化、新コンポーネントのオンライン化に伴い、このモデルはリアルタイムで更新されます。Azure サービスの健全性と下流への影響を可視化することで、顧客はアプリケーションの問題をより迅速に理解・診断できるようになり、Azure 上で構築されたアプリケーションの信頼性が向上します。
サービスカタログ:各サービスの機能、所有者、ティア(等級)、期待される動作、およびサービスレベル目標(SLO)が定義されています。
ランタイム状態:エラーレート、レイテンシ、スループット、リソース利用率、顧客間でのエラー分布など、すべてのコンポーネントの現在の挙動を示すライブ指標が含まれます。
「Intent(意図)」とは、現在起こるべきこと、進行中のデプロイメント、実施中の計画されたオペレーション、そしてスケジュールされている容量変更のことです。
「History(履歴)」は、過去のインシデント、その原因、それを緩和した対策、およびそれらに先行するシグナルを指します。これは、Azure が以前どのように不健康な状態になったか、また何が有効だったかをシステムが記憶している作業用メモリのようなものです。
「顧客の視点」では、各テナントが現在どのような状況を経験しているかが示されます。プラットフォームが発信しているデータだけでなく、実際に顧客のアプリケーションに到達している情報も含まれます。顧客が目にするエラー、体感するレイテンシ、そしてトラフィックが成功または失敗しているリージョンです。
これら個別の要素自体は新しいものではありません。あらゆるクラウドプラットフォームにはそれぞれのバージョンが存在します。Brain の真価は、これらを 12 の異なるツールにある 12 の別々のダッシュボードに散在させるのではなく、オペレーターが時間的プレッシャーの中で頭の中でつなぎ合わせる必要なく、単一の統合された AI ドライブンな表現として集約する点にあります。
Brain が「サービスが劣化している」と判断したとき、それは単に閾値を超えたという信号ではありません。トポロジー、ランタイムの状態、現在の意図、過去の傾向、そして顧客側の証拠を同時に推論して下された結論です。これは数値がアラートを発するのではなく、知能システムが判断を下しているのです。その判断にかかる時間は秒単位であり、人間が別々のツールから情報を集めて同じ状況を把握するのに必要な数分とは比較になりません。この速度の差が、顧客体験に直結します。インシデントの短縮、通知の精度向上、そして迅速なルーティングが可能になるのです。
クラウド知能システムに対して運用することの意味
これは Azure の顧客にとって全てを変える動きであり、「デジタルツイン」を比喩としてではなく、実際のシステムとして捉えなければ見落としがちな点です。
デプロイメントに起因する劣化が、異なる 2 つの環境でどのように解決されるかを考えてみましょう。
共有された知能システムが存在しない世界では、作業は「再構築」になります。ロールアウトが進行中です。あるリージョンのエラー率が徐々に上昇し始めます。
そのサービスを担当するチームは、自らのダッシュボードでこのズレを確認します。
上流の依存関係を担当するチームは、別の指標のズレを自らのダッシュボードで見ます。
デプロイメントシステムを担当するチームは、自らのダッシュボードではロールアウトが正常に進んでいることを確認します。
当初、3 つのチームには全体像がありません。彼らは橋を渡り、断片を組み立てていきます。その間、顧客への影響は拡大し続けます。ロールアウトと依存関係、そして顧客が直面するエラーのつながりを、人間が圧力の中でインシデント発生中に手作業で特定した頃には、ロールアウトはさらに多くのリージョンに到達しており、顧客からの問い合わせチケット数は増加しています。その結果、本来よりもはるかに困難な対応を迫られることになります。
知能システムが存在する世界では、この作業は「消費」に過ぎません。ロールアウトは知能システム内にあり、「Brain(脳)」はその進行中であることを把握しています。何を変更しているのか、どのリージョンに到達しようとしているのか、そして本来何をすべきなのかを知っています。エラー率のドリフトもシステム内にあります。「Brain」はこれをロールアウトと相関関係にあると認識し、依存グラフに対して重み付けを行い、「小さな揺らぎ」と「実際の劣化」がそれぞれどのように見えるかという過去の傾向と比較評価します。
影響を受ける顧客もまたシステム内にいます。彼らのテナントは、上流の依存関係によって影響を受けているプラットフォームリソースにマッピングされており、その依存関係自体もロールアウトの影響下にあります。「Brain」は単一の判断を下します。「このリージョンでは、ロールアウトが顧客に視認可能な影響を与えている。期待される解決には、ロールアウトを一時停止する必要がある」という結論です。
その判断は、即座に、それに基づいて行動すべきすべてのシステムへ伝達されます。判断が有効な間、デプロイメントシステムはロールアウトを一時停止します。これにより、「Brain」が次々と影響を与えるはずだった顧客たちは、全く影響を受けずに済みます。
インシデント管理システムは、上流の依存関係を特定した単一のインシデントを作成します。混乱する 3 つのチームから重複する 3 つのインシデントが作成されるのを防ぎ、適切なエンジニアが最初に正しい問題にたどり着けるようにします。
顧客向けコミュニケーションシステムでは、対象テナントのスコープと、一般の人にもわかりやすい英語による説明を含む通知文を自動生成します。これにより、影響を受けた顧客は Microsoft からより早く、実際に役立つ情報を含んだ更新を受け取ることができます。
Azure の顧客にとって、こうした調整作業が可視化されることはありません。顧客が目にするのは、短縮されたインシデント対応期間と、人間ではなく自動化システムに届く正確なアラートです。また、オンコール担当者が通話を開始した時点で診断名も確定しています。Brain のリソースヘルス評価が生産環境で稼働しているサービスでは、サービスに影響を与える事象の検出精度が大幅に向上し、対象範囲内のインシデントへの対応カバレッジも拡大し続けています。
過去 1 年間で、Brain と連携した障害の大半は影響を受けた顧客へ自動的に通知されました。これらのケースでは、手動で通知を出す場合と比較して、通知までの時間が大幅に短縮されています。
これらの下流システムはそれぞれ独自に調査を行うわけではありません。すべてが、インテリジェンスシステムから同じ結論を、共通の用語体系と裏付け証拠を用いて受け取ります。これが「インテリジェンスシステムに対して動作する」という意味であり、今日 Azure と結びつけられるエージェント型 AI の取り組みを実現可能にする前に、まず構築すべき最初の要素でもあります。
これは Azure の信頼性向上に寄与するだけでなく、Azure 上でアプリケーションを構築している顧客にも恩恵をもたらします。サービスの健全性に関する透明性を提供し、タイムリーなコミュニケーションを実現できるためです。
エージェント型 AI とクラウド運用の未来
今年、クラウド業界全体で、単に観察するだけでなく行動する AI システム、つまりエージェント型 AI について活発な議論が行われています。マイクロソフトもその議論の一員です。しかし、この議論には、注目されるべきほどには取り上げられていない静かな非対称性があります。
エージェントが「エージェンシー(主体性)」を発揮するためには、何らかの基盤が必要です:
依存関係グラフを知らないトリアージエージェントは、何をトリアージできるでしょうか?
過去のインシデント履歴にアクセスできない診断エージェントは、根本原因について推論できるでしょうか?
実際に影響を受ける顧客を把握していないコミュニケーションエージェントは、彼らへメッセージを送れるでしょうか?
これらのシステムが意味ある自律性を持つことはありません。もしそれぞれが、毎回、生データから現実の状況を探る必要があるなら、それらが信頼に値するはずもありません。
これが、ヘルスインテリジェンスシステムを「デジタルツイン」と呼ぶ所以です。これは、この規模でのエージェント型オペレーションにおける結果ではなく、前提条件なのです。断片化されたデータの上にまずエージェントを構築すれば、互いに矛盾する自信満々なシステムの連合が生まれるだけです。一方、モデルを先に構築すれば、エージェントはコンポーザブル(組み換え可能)になります。同じ視覚情報から推論を行うため、その情報は監査可能な一つの共通の図として機能します。
これが、今日から始める本シリーズの一貫したテーマです。Brain は、次世代クラウドエージェントが不可欠とするクラウドヘルスインテリジェンスシステムです。組織でクラウド、アプリケーション、あるいはインフラストラクチャといった運用機能においてエージェント型 AI の導入を検討しているなら、Brain が示すアーキテクチャパターンは慎重に検討すべき対象です。エージェントが注目されるべき headlines ですが、その背後にあるインテリジェンスシステムの構築こそが本質的な作業なのです。
Azure の信頼性と Brain の今後の展望
システムは完成しました。そして、このシステムには明確な判断基準があります。あるリージョン内のサービスが劣化しているという状況です。
しかし、何と比較して劣化しているのでしょうか?誰の定義における「健全」なのでしょうか?2 つのチームが自社のサービスの健全性について意見が割れている場合、どちらが正しいとされるべきでしょうか?プラットフォーム全体は劣化しているものの、まだ個別の顧客には影響が出ていない段階で、私たちは実際にどのような状態にあると言えるのでしょうか?
これらは哲学的な問いではありません。システムが判断を下すためには、まずそれを構築する人々が「判断とは何か」について合意する必要があります。しかし、業界の多くは最近まで、この根本的な点を誤解したまま進めてきました。
シリーズの次回の投稿では、具体的に何が間違っていたのか、そして過去 10 年にわたり業界が運用してきた不具合のあるクラウドヘルスの用語体系をどう置き換えたのかをお伝えします。新着記事は「Advancing reliability」ブログタグでフォローできます。
Azure の信頼性向上
大規模なクラウドリソースの照会、探索、分析が可能です。
詳細を見る

謝辞
本稿は、Brain AIOps チーム、Microsoft Research(MSR)、および Azure サービスチームに所属する多くのエンジニアや研究者の貢献によるものです。
「Meet Brain: The AI system behind Azure reliability」という記事は、Microsoft Azure Blog で最初に公開されました。
原文を表示
In this article
How Azure's AI-powered reliability intelligence system works
Why Brain is needed
What is Brain? Azure’s centralized AIOps for cloud reliability
Foundations of Azure’s digital twin for cloud health
What it means to operate against a cloud intelligence system
The future of agentic AI and cloud operations
What's next for Azure reliability and Brain
Azure reliability
Takeaway: Brain is Azure’s AI-powered cloud reliability intelligence system: an AIOps system that sits as an intelligent layer on top of Azure Resource Graph and fuses platform telemetry, AI/ML models, service dependencies, and customer impact into a single, continuously updated view of how every service, region, and workload is performing. It already powers customer Azure resource health notifications, deployment safeguards, and outage declaration, and it is the foundation for agentic AI now reshaping how Azure operates. This post starts a multi-part series on what Brain is, how we built it, what we’ve learned operating it at scale, and where it goes next.
How Azure’s AI-powered reliability intelligence system works
Azure runs on a digital twin of its own health. Brain is an AIOps-powered cloud health intelligence system that operates as an intelligent layer on top of Azure Resource Graph (ARG); together, they form this digital twin. It integrates platform telemetry, AI/ML models, and data engineering to continuously maintain and enrich a real-time view of how services, regions, and customer workloads are performing across Azure. Over time, that shared picture is becoming the foundation for a more automated reliability surface: one that can turn insight into action.
Today, Brain already powers important reliability workflows across Azure, such as health notifications for customer’s resources, deployment safeguards, and outage declaration. If you run on Azure, Brain is already changing three things you can notice:
How fast we tell you when something is wrong.
How accurately we scope it to your resources.
How quickly the right engineer gets on it.
This post is about how and what it lets you do differently.
We’re starting a multi-post series with this one to take you through what Brain is, how we built it, what it has learned operating Azure at scale, and where it goes next. Today, the foundation.
Learn more about Azure reliability
Why Brain is needed
Azure runs hundreds of services across more than 80 Azure regions, over 500 datacenters, and over 800,000 kilometers of fiber and subsea cable, representing one of the world’s largest global cloud footprints. And yet with the massive amount of activity these Azure services create, manage, and process worldwide, on a quietly degrading day, we will sometimes still learn about an issue from a customer before our own systems do. For customers, that gap is the worst kind of incident; the one where they are debugging their own application before they learn the fault was ours.
That gap between what we measure and what we know is the limiting factor on cloud reliability today. It is not a tooling problem. We have plenty of tools. It is a comprehension problem. The amount of signal a hyperscale cloud produces has outgrown the human ability to read it, and the conventional answer: more dashboards, more alerts, more on-call rotations. It’s a treadmill, not an answer. Every additional dashboard gives an operator another window to look through; what’s missing is something that tells them what they’re looking at, in time to act.
Closing that gap meant building something we hadn’t built yet: not better dashboards, not smarter alerts, but a continuously updated model of the platform’s health that reasons across every signal in real time, and acts on those conclusions automatically at the scale the platform demands.
What is Brain? Azure’s centralized AIOps for cloud reliability
Brain is Azure’s centralized AIOps-powered cloud health intelligence system that uses AI/ML, including agentic AI and data engineering, to continuously model Azure’s health and to automatically take reliability actions based on it. It has been utilized in Azure production generating resource health determinations across the platform.
At its core, Brain is shaped by three things: what goes in, what comes out, and what those outputs drive.

Brain at a glance.
Brain ingests signals from three classes of source:
Standardized service-level indicators: the SLIs Azure customers and operators already know from their reliability dashboards.
Domain-specific monitors that individual service teams have built and registered with Brain, and the broader telemetry stream including deployments, support volume, and cross-service dependency signals.
Third-party indicators that surround every Azure operation.
Each path serves a different purpose; together, they give Brain coverage that no single path could.
Regardless of the input, Brain evaluates every subject (service, region, deployment unit, or customer resource) and returns four outputs: health state, severity, impact, and the reason for its conclusion. Standard outputs in standard vocabulary mean every downstream system speaks the same language; no more disconnect in what “impacted” means across teams.
The insights generated by Brain power a growing set of automated reliability actions, including:
Outage declarations based on blast radius.
Customer notifications targeted to affected subscriptions and regions.
Incident routing to the appropriate service team.
Deployment gates that pause harmful rollouts.
Linking related incidents.
Diagnostic tools that help engineers investigate issues.
Foundations of Azure’s digital twin for cloud health
To understand what makes “the intelligence system” different from “a dashboard,” it helps to look at what’s actually in its foundation. Brain’s representation of Azure carries, at minimum:
Topology: every service, region, availability zone, deployment unit, and dependency graph enabled by Azure Resource Graph is represented as a live model that updates as services scale, dependencies change, and new components come online. This transparency into Azure service health and downstream impact helps Azure customers understand and diagnose application issues more quickly and improves the reliability of applications built on Azure.
Service catalog: what each service does, who owns it, what its tier is, what its expected behavior looks like, and what its service-level objectives are.
Runtime state: live indicators of how every component is currently behaving, including error rates, latency, throughput, resource utilization, and error distributions across customers.
Intent: what’s supposed to be happening right now, which deployments are in flight, which planned operations are underway, and which capacity changes are scheduled.
History: prior incidents, what caused them, what mitigated them, and which signals preceded them. The system’s working memory of how Azure has gotten unhealthy before, and what worked to fix it.
The customer’s view: what each tenant is currently experiencing. Not just what the platform is emitting, but what’s actually arriving at the customer’s application. Errors customers see, latency customers feel, and regions where their traffic is succeeding or failing.
None of these are novel on their own: every cloud platform has versions of each. Brain brings them together into a single, unified, AI-driven representation instead of scattering them across twelve separate dashboards in twelve separate tools that an operator has to mentally connect under time pressure.
When Brain says a service is degrading, that statement is not a threshold being crossed. It is a determination made by reasoning across topology, runtime state, current intent, historical patterns, and customer-side evidence simultaneously. It is the intelligence system speaking, not a metric firing. And it is the speed of that determination measured in seconds, not in the minutes a human would take to assemble the same picture from separate tools that translates directly into customer experience: shorter incidents, sharper notifications, and faster routing.
What it means to operate against a cloud intelligence system
This is the move that changes everything for an Azure customer, and it’s the one most easily missed if you read “digital twin” as a metaphor rather than as a system.
Consider how a deployment-driven degradation typically resolves in two different worlds.
In a world without a shared intelligence system, the work is reconstruction. A rollout is in flight. A region’s error rate begins to drift.
The team that owns the service sees the drift in their dashboard.
The team that owns the upstream dependency sees a different metric drift in their dashboard.
The team that owns the deployment system sees the rollout proceeding normally from their dashboard.
None of those three teams initially have the picture; they get on a bridge and assemble it from fragments. While they assemble, the customer impact spreads. By the time the connection between the rollout, the dependency, and the customer-visible errors is made, by humans, under pressure, mid-incident, the rollout has reached more regions, the customer ticket queue has grown, and the resolution is now harder than it had to be.
In a world with the intelligence system, the work is consumption. The rollout is in the intelligence system, Brain knows it’s in flight: what it’s changing, what regions it’s reaching, what it’s supposed to do. The error-rate drift is in the system: Brain sees it correlated to the rollout, weighted against the dependency graph, evaluated against historical patterns of what “small wobble” looks like versus what “real degradation” looks like.
The affected customers are in the system, their tenants map to platform resources affected by the upstream dependency, which is itself affected by the rollout. Brain produces a single determination: the rollout is causing customer-visible impact in this region; expected resolution requires the rollout to pause.
That determination then flows, at the same moment, to every system that needs to act on it. The deployment system pauses the rollout while the determination is true, so the next set of customers Brain would have impacted aren’t impacted at all.
The incident management system creates a single incident with the upstream dependency identified, not three duplicate incidents from three confused teams so the right engineer reaches the right problem first. The customer communication system drafts a notification with the right tenant scope and the right plain-English description, so the customers who are affected receive updates from Microsoft sooner, with information they can actually use.
For Azure customers, none of that coordination is visible. What’s visible is a shorter incident, an accurate alert that hit their automation instead of a human, and diagnosis that was already named when their on call opened the bridge. On services where Brain’s resource-health evaluation is in production, detection precision for service-impacting issues has improved significantly, and coverage of in-scope incidents continues to expand.
In the past year, a substantial majority of Brain-integrated outages were auto-communicated to affected customers, and on those, time-to-notification improved materially compared to manually issued notifications.
None of those downstream systems are doing their own investigation. They all consume the same determination from the intelligence system, in the same vocabulary, with the same supporting evidence. That is what “operating against an intelligence system” means and it is the first thing we found we had to build before any of the agentic AI work that people associate with Azure today became viable.
This not only helps to improve Azure’s reliability, but also benefits Azure customers who built their applications on top of Azure by providing transparency of service health and timely communications.
The future of agentic AI and cloud operations
There is a larger conversation happening across the cloud industry this year about agentic AI and about AI systems that act, not just observe. Microsoft is part of that conversation. But the conversation has a quiet asymmetry that gets less attention than it deserves.
Agents need something to be agentic about:
A triage agent that doesn’t know the dependency graph cannot triage anything.
A diagnosis agent that cannot reach prior incident history cannot reason about root cause.
A communication agent that doesn’t know which customers are actually affected cannot write to them.
None of these systems are meaningfully autonomous; none of them deserve your trust if every one of them has to do their own investigation of what reality is, every time, from raw signals.
That is what made the health intelligence system “the digital twin”: the prerequisite, not the consequence, of agentic operations at this scale. Build the agents first, on top of fragmented data, and you get a federation of confident systems that disagree with each other in production. Build the model first, and the agents become composable: they reason from the same picture, and the picture is one you can audit.
This is the throughline of the series we’re starting today. Brain is the cloud health intelligence system the next generation of cloud agents will need. If your organization is exploring agentic AI for any operations function: your cloud, your applications, or your infrastructure, the architectural pattern Brain represents is one to look at carefully. The agents are the headline; the intelligence system underneath is the work.
What’s next for Azure reliability and Brain
We have the system. The system has determination. A service in a region is degrading.
However, degrading compared to what? Healthy by whose definition? When two teams disagree about whether their service is healthy, which one is right? When the platform is degrading but no individual customer is impacted yet, what state are we actually in?
Those are not philosophical questions. They are the next engineering questions we have to answer, because a system cannot make determinations until the people building it agree on what determinations actually are. Most of the industry, until recently, has been quietly getting this wrong.
In the next post in this series, we’ll show you exactly how, and what we built to replace the broken vocabulary of cloud health that the industry has been operating on for the last decade. To follow the series as new posts are published, see the Advancing reliability blog tag.
Azure reliability
Query, explore, and analyze your cloud resources at scale.
Learn more
image
Acknowledgments
This work reflects the contributions of many engineers and researchers across the Brain AIOps team, MSR (Microsoft Research), and Azure service teams.
The post Meet Brain: The AI system behind Azure reliability appeared first on Microsoft Azure Blog.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み