マイクロソフト、サイバー防御専用モデル「MAI-Cyber-1-Flash」公開
Microsoft AI はサイバー防御専用モデル「MAI-Cyber-1-Flash」を公開し、MDASH ハーネス内での CyberGym ベンチマークで 95.95% のスコアを達成した。
AI深層分析を開く2026年7月28日 22:04
AI深層分析
キーポイント
サイバー防御専用モデルの発表
Microsoft AI はサイバー防御に特化した最初のモデル「MAI-Cyber-1-Flash」を発表し、これは MAI-Thinking-1 シリーズを基盤とした Transformer 構造を持つ。
MDASH ハーネスでの高性能達成
Microsoft のマルチモデルエージェントスキャンハーネス「MDASH」内で GPT-5.4 と併用することで、CyberGym ベンチマークで 95.95% という業界最高スコアを記録した。
コスト削減のためのルーティング戦略
MAI-Cyber-1-Flash が MDASH タスクの約 90% を処理し、困難な 10% のみを上位モデルへエスカレートさせることで、前構成比で 50% のコスト削減を実現した。
チームと実績の背景
MDASH は DARPA AI サイバーチャレンジ優勝チーム「Team Atlanta」出身者を含む Microsoft の Autonomous Code Security チームが開発し、Windows ネットワークスタックで複数の CVE を発見した実績がある。
防御特化設計によるゼロスコア
ExploitGymでのスコアが0/0/0であるのは欠陥ではなく、マルウェア作成などの攻撃タスクを行わずバグ修正に特化した意図的な設計によるものである。
重要な引用
MAI-Cyber-1-Flash is a transformer with self-attention and sparse Mixture-of-Experts layers.
MDASH running MAI-Cyber-1-Flash alongside GPT-5.4 scores 95.95%.
Replacing 80% of the existing models in MDASH moved the harness from 88.4% to 95.95%.
The straight zeros on ExploitGym are deliberate, not a defect.
編集コメントを表示
編集コメント
Microsoft は単なるモデルの性能向上だけでなく、MDASH という「ハーネス」全体を最適化することで実用面での劇的な成果を挙げており、AI セキュリティの実装アプローチに新たな基準を示している。特にコスト削減と高性能を両立させたルーティング戦略は、大規模なセキュリティ運用における現実的な解決策として注目される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Microsoft AI は、サイバー防御に特化して設計された初のモデル「MAI-Cyber-1-Flash」を公開しました。このモデルはスタンドアロンのエンドポイントとして提供されるものではなく、Microsoft のマルチモデル・エージェント型スキャンハーンである MDASH 内で稼働します。
MAI-Cyber-1-Flash は、自己注意機構とスパースな Mixture-of-Experts レイヤーを備えたトランスフォーマーアーキテクチャです。パラメータ数は総計 137B で、そのうち 5B がアクティブに使用され、コンテキスト長は 256k に達します。入出力はテキストのみに対応しています。
これは、GitHub Copilot や VS Code に既に組み込まれている軽量エージェント型コーディングモデル「MAI-Code-1-Flash」をサイバーセキュリティ分野向けに微調整したものです。また、リリース情報によると、このモデルは MAI-Thinking-1 の系譜から派生したとされています。
ベンチマーク
CyberGym は、188 の OSS-Fuzz プロジェクトから収集された 1,507 の実世界脆弱性再現タスクを網羅した公開スイートです。Microsoft は CyberGym のデフォルト設定(レベル 1)で評価を行いました。この設定では、脆弱なソースコードと高レベルの説明が提供されます。
MDASH で MAI-Cyber-1-Flash と GPT-5.4 を併用した結果、スコアは 95.95% に達しました。Microsoft はこれを Anthropic の Mythos より約 12 ポイント上回る成果として位置付けています。また、発表されたチャートでは、競合する 4 つのシステムが 83.2% から 85.6% の範囲にあることが示されています。
Microsoft が MDASH を初めて詳細に紹介したのは 2026 年 5 月でした。その際、一般利用可能なモデルのみを使用し、CyberGym で 88.45% のスコアを記録していました。これはすでにトップの公開リーダーボードスコアであり、次点(83.1%)よりも約 5 ポイント上回っていました。研究チームは改善点を率直に述べています。MDASH 内の既存モデルの 80% を置き換えたことで、ハッチングのスコアが 88.4% から 95.95% に向上したのです。
なぜルーティングこそが真の製品なのか
MDASH は「準備(Prepare)」「スキャン(Scan)」「検証(Validate)」「重複排除(Dedupe)」「証明(Prove)」の 5 つのステージを通じて、100 以上の専門エージェントを管理します。監査員エージェントが発見事項をフラグ付けし、論客エージェントは不一致をシグナルとして用いて攻撃性の有無を議論します。また、「証明」ステージでは、C/C++ ターゲットに対して ASan を使用してトリガー入力を実行します。
大規模な最先端モデルのコストを抑制するため、MAI-Cyber-1-Flash は MDASH タスクの最大 90% を処理し、最も困難な 10% のみを GPT-5.4 にエスカレーションします。このルーティングにより、以前の構成(GPT-5.4、5.4 mini、5.3 codex)と比較してコストを 50% 削減することに成功しました。
MDASH は、Microsoft の自律型コードセキュリティ(ACS)チームによって開発されました。同チームには DARPA AI サイバーチャレンジで優勝した「Team Atlanta」のメンバーも含まれています。今年 5 月には、MDASH を活用した作業により、Windows のネットワークおよび認証スタックにおいて 16 の CVE(そのうち 4 つは深刻なリモートコード実行脆弱性)が特定されました。
過去に遡って検証すると、clfs.sys における 28 の MSRC ケースの 96%、tcpip.sys における 7 つのケースをすべて、5 年間のデータから回復させることに成功しています。
性能
研究チームは、軽量なターミナルハネスで得られた単独の結果を発表しました。
ベンチマーク | MAI-Cyber-1-Flash
---|---
CVEBench | 0.314
CyberSecEval4 — Threat Intel | 0.553
CyberSecEval4 — Malware Analysis | 0.33
CRSBench (POV=1200) | 0.651
ExploitGym — Kernel / Userspace / Browser | 0 / 0 / 0
ExploitGym のスコアがすべてゼロであることは、欠陥ではなく意図的な設計です。Microsoft チームによると、このモデルはマルウェアの配布といった攻撃的なタスクではなく、バグの修正など防御的なタスクの実行を目的としてトレーニングされています。5B アクティブパラメータながらエクスプロイトを生成できないものの、95.95% の脆弱性発見パイプラインを駆動できるこのモデルは、まさに防御者専用製品に必要な成果物です。
使い方の詳細については、以下の埋め込みコンポーネントをご確認ください。
Microsoft AI が「MAI-Cyber-1-Flash」を発表しました。これは、CyberGym ベンチマークで MDASH のスコアを 95.95% に引き上げた、アクティブパラメータ 50 億のサイバーモデルです。
主なポイント
MAI-Cyber-1-Flash は、総パラメータ数 1370 億のうちアクティブに使用されるのは 50 億というスパースな MoE(Mixture of Experts)アーキテクチャを採用しています。これは、256k のコンテキスト長を扱える「MAI-Code-1-Flash」を微調整したモデルです。
CyberGym で達成された 95.95% というスコアはシステム全体の評価値です。MDASH と新モデル、そして GPT-5.4 を組み合わせた構成によるもので、2026 年 5 月の 88.45% から大幅に向上しています。
このモデルは MDASH タスクの約 90% を処理し、残りの難しい 10% は GPT-5.4 に任せることで、コストを最大 50% 削減できると主張しています。
なお、ExploitGym のスコアがすべて「0/0/0」なのは設計上の意図です。このモデルはバグの修正を行うものであり、エクスプロイト(攻撃コード)を作成するものではありません。
アクセス制限について
原文を表示
Microsoft AI has released MAI-Cyber-1-Flash, its first model built specifically for cyber defense. The model does not ship as a standalone endpoint. It runs inside MDASH, Microsoft’s multi-model agentic scanning harness.
MAI-Cyber-1-Flash
MAI-Cyber-1-Flash is a transformer with self-attention and sparse Mixture-of-Experts layers. It carries 137B total parameters with 5B active, and a 256k context length. Inputs and outputs are text only.
It is a cybersecurity-specialized fine-tune of MAI-Code-1-Flash, the lightweight agentic coding model already embedded in GitHub Copilot and VS Code. The release describes it as derived from the MAI-Thinking-1 lineage.
Benchmarks
CyberGym is a public suite of 1,507 real-world vulnerability reproduction tasks drawn from 188 OSS-Fuzz projects. Microsoft evaluated at CyberGym’s default level 1 configuration, which supplies vulnerable source and a high-level description.
MDASH running MAI-Cyber-1-Flash alongside GPT-5.4 scores 95.95%. Microsoft frames this as roughly 12 points above Anthropic’s Mythos, and the launch chart places the four competing systems between 83.2% and 85.6%.
When Microsoft first detailed MDASH in May 2026, the harness scored 88.45% on CyberGym using only generally available models. That was already the top public leaderboard score, about five points ahead of the next entry at 83.1%. The research team states the improvement plainly: replacing 80% of the existing models in MDASH moved the harness from 88.4% to 95.95%.
Why the routing is the real product
MDASH manages over 100 specialized agents through five stages: Prepare, Scan, Validate, Dedupe, and Prove. Auditor agents flag findings, debater agents argue exploitability (using disagreement as signal), and the Prove stage executes triggering inputs with ASan for C/C++ targets.
To control frontier model costs at scale, MAI-Cyber-1-Flash handles up to 90% of MDASH tasks, escalating the hardest 10% to GPT-5.4. This routing yields a 50% cost saving over the previous configuration of GPT-5.4, 5.4 mini, and 5.3 codex.
MDASH was developed by Microsoft’s Autonomous Code Security (ACS) team, featuring members from the DARPA AI Cyber Challenge-winning Team Atlanta. In May, MDASH-assisted work generated 16 CVEs (including four Critical remote code execution flaws) in the Windows networking and authentication stack. Retrospectively, it recovered 96% of 28 MSRC cases in clfs.sys and 100% of 7 cases in tcpip.sys over a five-year window.
Performance
The research team present standalone results from a lightweight terminal harness:
BenchmarkMAI-Cyber-1-Flash
CVEBench0.314
CyberSecEval4 — Threat Intel0.553
CyberSecEval4 — Malware Analysis0.33
CRSBench0.651 (POV=1200)
ExploitGym — Kernel / Userspace / Browser0 / 0 / 0
The straight zeros on ExploitGym are deliberate, not a defect. Microsoft team states the model was trained to perform defensive tasks such as patching bugs, and not offensive tasks such as deploying malware. A 5B-active model that cannot generate exploits but can drive a 95.95% discovery pipeline is exactly the artifact a defender-only product needs.
How to use it
Key Takeaways
MAI-Cyber-1-Flash is 137B total / 5B active, a sparse MoE fine-tune of MAI-Code-1-Flash with 256k context.
95.95% on CyberGym is a system score — MDASH plus the new model plus GPT-5.4, up from 88.45% in May 2026.
It handles up to 90% of MDASH tasks, escalating the hard 10% to GPT-5.4 for a claimed 50% cost cut.
ExploitGym scores are 0/0/0 by design — the model patches bugs, it does not write exploits.
Access is gated
The post Microsoft AI Releases MAI-Cyber-1-Flash: A 5B-Active-Parameter Cyber Model That Pushes MDASH to 95.95% on CyberGym appeared first on MarkTechPost.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み