Google DeepMind、サイバーセキュリティ向け Gemini 3.5 Flash 発表
Google DeepMind はサイバーセキュリティ分野の脅威検出と分析を高速化し、防御体制を強化するための専用 AI モデル「Gemini 3.5 Flash」を発表した。
キーポイント
専門特化モデルの発表
Google DeepMind がセキュリティ脅威の検出と分析に特化した「Gemini 3.5 Flash」を新リリースした。
高速化による防御強化
このモデルは処理速度を重視しており、迅速な脅威対応を通じて組織の防御体制を強化する目的で設計されている。
実用性の高いセキュリティ対策
従来の汎用 AI モデルに比べ、サイバー攻撃への即応性を高めることで、現場での実運用効果を期待できる。
重要な引用
Gemini 3.5 Flash はセキュリティ脅威の検出や分析を高速化し、防御体制の強化を目指すものである。
影響分析・編集コメントを表示
影響分析
この発表は、AI がセキュリティ分野で汎用性を超越し、特定のタスクに特化した高性能モデルとして進化していることを示しています。企業にとっては、脅威への対応時間を短縮し、防御の質を向上させる新たな手段が手に入ったことになります。
編集コメント
セキュリティ分野における AI の特化は、汎用モデルの限界を補う重要なステップであり、実戦での即応性向上に直結する意義があります。ただし、詳細な技術仕様やベンチマークデータが今後の展開次第で、その真価が問われるでしょう。
2026 年 7 月 21 日
モデル紹介
ラルカ・アダ・ポパ、フォー・フリン
Google は長年にわたりサイバーセキュリティへの投資を続けており、世界規模のコードベースを守るための自動化された脆弱性発見技術で先駆的な役割を果たしてきました。CodeMender に代表されるようなコードセキュリティエージェントは、重要なソフトウェアの脆弱性を自動的に特定し、修正する能力を持っています。しかし、AI エージェントが脆弱性の発見速度を防御側の対応速度を上回るようになると、この世界的な脅威に対処するには、高い性能を持ちながら低コストでスケーラブルなアプローチが必要となります。
そこで本日、私たちは長年の取り組みの一環として、Gemini 3.5 Flash Cyber を発表します。これは「3.5 Flash」を基盤とした軽量なサイバーセキュリティ専用モデルです。脆弱性の発見、検証、パッチ適用を迅速かつ効率的に行うように微調整されており、これらのタスクにおいては Gemini のメインラインである Flash モデルよりも高い効果を発揮します。
Flash モデルが持つ優れたパフォーマンスと効率性は、サイバーセキュリティ分野での基盤として理想的です。この Flash を土台に構築した「3.5 Flash Cyber」は、大規模で高コストな既存のセキュリティモデルに対する、費用対効果が高く高性能な代替案となります。
この技術には二重利用の性質があるため、3.5 Flash Cyber の展開方法については意図的なアプローチを採用しました。限定的なアクセスによるパイロットプログラムの一環として、3.5 Flash Cyber は今後、CodeMender を通じて政府機関や信頼できるパートナーに限定して提供されます。対象範囲は順次拡大していく予定です。これにより、最前線で防御に当たるチームが脆弱性を悪用される前に発見し修正するための有利なスタートを切れるようになります。同時に、広範な誤使用のリスクも軽減できます。
また、Gemini Enterprise Agent Platform を通じて、一般利用可能な Gemini モデルにも CodeMender の基盤機能を直接提供しています。
探索空間の問題:コードセキュリティにおける軽量モデルの優位性
根深い欠陥を見つけるには、膨大な実行探索空間を調査する必要があります。巨大な言語モデルに対して高価な呼び出しを単発で行うだけではボトルネックとなりかねません。3.5 Flash Cyber は、エージェントが広範なコードベースをスキャンし、多数のコードパスを分析する必要がある脆弱性の発見に特に適しています。
CodeMender は 3.5 Flash Cyber を複数回呼び出すことで、エージェントははるかに多くのコードパスを分析し、脆弱性を発見・検証できるようになります。その後、サブエージェントが統合された高品質なレポートを生成します。
その高速性と低コストにより、Gemini 3.5 Flash Cyber は、頻繁なスキャンや時間制約のあるリリースプロセス、大規模なコミットスキャンパイプラインへの統合が容易です。
3.5 Flash Cyber のベンチマーク結果:より大きなセキュリティモデルに対する効率的な代替案
私たちは 3.5 Flash Cyber を多様なベンチマークでテストしました。特に注目すべきは、AI エージェントが数百の実際のソフトウェア脆弱性に対してどう振る舞うかを評価する「CyberGym」での検証です。CodeMender にて 1 つの最終レポート生成のために 3.5 Flash Cyber を最大 5 回呼び出すよう設定し、その低コストを有効活用した結果、このエージェントは CyberGym において、はるかに大きなモデルと比較しても競合する性能を発揮しました。
競合他社の数値は、各プロバイダーが自己申告したスコアに基づくものです。
また、CyberGym の枠を超え、安全装置を無効化した状態でモデルの能力を限界まで試すストレステストも実施しました。Google の Big Sleep チームは独立して評価環境を構築し、Chrome や Safari といった世界で最も複雑なコードベースにおける、重要かつ発見が困難な脆弱性の特定に焦点を当てています。この評価において、3.5 Flash Cyber はメインラインの 3.5 Flash や 3.6 Flash を大きく上回る結果を残しました。
Big Sleep 評価における成功率 (pass@1)
3.5 Flash Cyber は、Google Chrome の本番環境におけるコミットスキャンパイプラインでも評価されました。脆弱性情報は公開されていないため、このベンチマークが Gemini や競合モデルにとって汚染されない状態を維持できています。
その結果、3.5 Flash と比較して 3.5 Flash Cyber は著しい性能向上を示しました。なお、Opus 4.6 より後の最新競合モデルは、組み込まれたセキュリティガードレールによりタスクの遂行を拒否するため、ここでは結果を表示していません。
Chrome 本番コミットスキャンパイプラインにおける成功率 (pass@1)
さらに、3.5 Flash Cyber はメインラインの 3.5 Flash や Claude Opus 4.6 と比較しても、より多くの固有の脆弱性を発見し続けています。固定された呼び出し回数で非常に複雑な V8 JavaScript エンジンにテストした結果、3.5 Flash Cyber は 55 の固有かつ確認済みの問題を特定しました。一方、メインラインの 3.5 Flash が 47 件、Opus 4.6 が 36 件の発見にとどまりました。そのうち 10 件の問題は、他二つのモデルでは検出できていませんでした。
基本的なサイバーセキュリティモデルは、同じ問題を繰り返し見つけてしまうループに陥りやすく、重要な脆弱性を見逃す傾向があります。優れたモデルはより広い範囲を網羅し、より多くの固有の問題を検出します。
呼び出し回数を増やしていくと、3.5 Flash Cyber は新たなコードパスや脆弱性を発見し続けることが確認できました。
Google における実世界での応用と防御のスケール
ベンチマークの結果は物語の一部に過ぎません。CodeMender に搭載された Gemini 3.5 Flash Cyber は、すでに Chrome、Android、Cloud、Ads、YouTube を含む Google の内部コードベースにおいて、脆弱性の発見と修正を現実のものとしています。
軽量モデルが可能にした高速な発見スピードが、明確なインパクトをもたらしています。
例えば、Google Cloud の脆弱性調査チームは、3.5 Flash Cyber を活用して記録的な短時間でシステムの防御を強化しました。わずか 2 時間のうちに、公開 API におけるリモートコード実行の脆弱性と、重要な本番環境サービスでのメモリ破壊脆弱性を特定。さらに、アドレス空間配置ランダム化 (ASLR) や Write XOR Execute (W^X) といった標準的な対策手法を回避する、信頼性 100% のリモートコード実行エクスプロイトも生成しました。
Wiz および Cloud CISO セキュリティエンジニアリングのテストチームからの初期フィードバックでは、メインラインの 3.5 Flash モデルと比較して、3.5 Flash Cyber の能力が大幅に向上していることが確認されています。
スケールした防御者への支援
ソフトウェアセキュリティにおける Google のリーダーシップは、独自の強みとなっています。具体的には、Google が運営する脆弱性情報データベース OSV.dev は 70 万件以上のオープンソース脆弱性を網羅しており、10 年以上にわたる OSS-Fuzz の結果と相まって、最も品質の高い脆弱性の特定を可能にしています。
これにより、合成されたサイバーセキュリティの事例に頼るのではなく、実際のセキュリティ専門家がどのように活動するかをモデルに教えることが可能になります。当社のモデルは、業界標準のツールを操作し、Chromium などの大規模プロジェクトにおける数百万行のコードを読み込み、数時間にわたる継続的で深い分析が必要な複雑なセキュリティタスクを自律的に処理できるようになります。
CodeMender に Gemini 3.5 Flash Cyber を搭載することで、より多くの防衛チームがソフトウェアの保護を強化できるよう支援する、高性能で拡張性が高く、コスト効率に優れたアーキテクチャを提供しています。
原文を表示
July 21, 2026 Models
Raluca Ada Popa and Four Flynn
Google has invested in cybersecurity for years, pioneering automated vulnerability discovery to secure the world’s codebases. Tools like CodeMender, our code security agent, can automatically find and fix critical software vulnerabilities. But as AI agents become more capable at finding vulnerabilities faster than defenders can fix them, addressing this global threat requires a highly capable, affordable, and scalable approach.
Today, we’re expanding our longtime efforts to better prepare defenders by introducing Gemini 3.5 Flash Cyber, our lightweight cybersecurity model built on top of 3.5 Flash and fine-tuned to find, validate, and patch vulnerabilities quickly and efficient, making it more effective at these tasks than Gemini’s mainline Flash models.
Flash’s performance and efficiency makes it an ideal foundation for our cybersecurity model efforts. By building on top of Flash, 3.5 Flash Cyber offers a cost-efficient and highly capable alternative to large, costly cybersecurity models.
Given the dual-use nature of this technology, we have taken an intentional approach to how we deploy 3.5 Flash Cyber. As part of a limited-access pilot program, 3.5 Flash Cyber will be exclusively available to governments and trusted partners via CodeMender soon, expanding over time. This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse.
Separately, we're also bringing CodeMender's foundational capabilities directly to customers with generally available Gemini models through the Gemini Enterprise Agent Platform.
The search space problem: The advantage of lightweight models in code security
Finding deep-seated flaws requires exploring an immense execution search space. Relying on a single, expensive call to a massive language model can create a bottleneck. 3.5 Flash Cyber is particularly suitable for finding vulnerabilities where the agent has to scan a large codebase and analyze a large number of codepaths.
CodeMender invokes 3.5 Flash Cyber multiple times, so agents can analyze vastly more code paths to discover and validate vulnerabilities. The sub-agents then produce a single, high-quality report.
Thanks to its speed and affordability, 3.5 Flash Cyber can be easily integrated into frequent scans, time-sensitive launch processes or commit scanning pipelines at scale.
3.5 Flash Cyber benchmark results: an efficient alternative to larger cybersecurity models
We tested 3.5 Flash Cyber on a variety of benchmarks. In particular, we tested 3.5 Flash Cyber on the CyberGym benchmark, which evaluates AI agents against hundreds of real-world software vulnerabilities. Leveraging the low cost of 3.5 Flash Cyber by configuring CodeMender to call 3.5 Flash Cyber up to five times for a single, final report, the overall agent achieved competitive performance against significantly larger models on CyberGym*.
**Competitor results are sourced from provider self-reported scores*
We also stress-tested the model’s capabilities beyond CyberGym without safety guardrails. Google’s Big Sleep team independently built an evaluation focused on finding critical and hard to find vulnerabilities in some of the world’s most complex codebases like Chrome and Safari. Here, 3.5 Flash Cyber significantly surpassed mainline 3.5 Flash and 3.6 Flash.
Success Rate on Big Sleep Evaluation (pass@1)
3.5 Flash Cyber was also evaluated on Google Chrome’s production commit scanning pipeline. The vulnerabilities were not publicly disclosed, which ensured this benchmark remained free of contamination for Gemini and competitor models.
The results showed a significant uplift from 3.5 Flash Cyber compared to 3.5 Flash. Note: More recent competitor model versions after Opus 4.6 refuse to fulfill the tasks due to built-in safety guardrails, and therefore are not shown.
Success Rate on Chrome Production Commit Scanning Pipeline (pass@1)
Moreover, 3.5 Flash Cyber consistently discovered more unique vulnerabilities compared with mainline 3.5 Flash and Claude Opus 4.6. When tested on the highly complex V8 JavaScript Engine across a fixed number of invocations, 3.5 Flash Cyber found 55 unique confirmed issues, compared to 47 found by mainline 3.5 Flash and 36 found by Opus 4.6, including 10 issues that the other two models tested did not catch.
Basic cybersecurity models can get stuck in a loop, finding the same issue repeatedly while missing critical vulnerabilities. A strong model casts a wider net, finding a higher number of unique issues.
As we scale the number of invocations, we find that 3.5 Flash Cyber continues to discover new code paths and vulnerabilities.
Real-world application and scaling defenses at Google
Benchmarks are only part of the story. 3.5 Flash Cyber in CodeMender is already finding and fixing vulnerabilities in Google’s internal codebases including Chrome, Android, Cloud, Ads, and YouTube.
The speed of discovery made possible by a lightweight model has delivered measurable impact.
For example, Google’s Cloud Vulnerability Research team used 3.5 Flash Cyber to proactively secure our systems in record time. In just 2 hours, the model uncovered remote code execution vulnerabilities in public APIs and found a memory-corruption vulnerability in a sensitive production service. It then generated a 100% reliable remote-code execution exploit that bypassed standard mitigation techniques like Address Space Layout Randomization (ASLR) and Write XOR Execute (W^X) .
Early feedback from Wiz and Cloud CISO Security Engineering testers confirms the significant capability improvement of 3.5 Flash Cyber over the mainline 3.5 Flash model.
Empowering defenders at scale
Google’s leadership in software security gives us a unique advantage. For example, OSV.dev, a vulnerability database run by Google spanning over 700,000 open-source vulnerabilities, and more than 10 years of OSS-Fuzz results, help us identify the most high quality vulnerabilities.
This allows us to move beyond synthetic cybersecurity examples and teach our models how real security professionals work. Our models learn to operate industry-standard tools, read through millions of lines of code in large-scale projects like Chromium, and independently tackle complex security tasks that require hours of continuous, deep analysis.
By powering CodeMender with 3.5 Flash Cyber, we’re providing a highly capable, scalable, and affordable architecture designed to help more defenders secure software.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み