Google DeepMind、新AIモデル「Gemini 3.6 Flash」シリーズを発表
Google DeepMind は、高性能・軽量・セキュリティ特化の 3 つの新モデル「Gemini 3.6 Flash」シリーズを正式発表し、推論速度と専門領域への対応を強化した。
キーポイント
新モデルの多層化戦略
高性能な「Gemini 3.6 Flash」、軽量版の「3.5 Flash-Lite」、セキュリティ特化型の「3.5 Flash Cyber」の 3 つを同時にリリースし、用途に応じた最適化を図った。
セキュリティ特化モデルの登場
「3.5 Flash Cyber」はセキュリティ分野に特化した機能を提供し、AI の安全性と脆弱性対策における専門的な役割を担うことを示している。
Flash シリーズの進化
Gemini 3.x フレームワーク内での「Flash」シリーズが高速推論とコスト効率を追求するラインナップとしてさらに洗練され、実用性が向上した。
重要な引用
Google DeepMind は、高性能な「Gemini 3.6 Flash」、軽量版の「3.5 Flash-Lite」、およびセキュリティ特化型の「3.5 Flash Cyber」という 3 つの新しい AI モデルを正式に発表した。
影響分析・編集コメントを表示
影響分析
今回の発表は、Gemini モデルが単なる汎用性能の向上だけでなく、セキュリティや軽量化といった特定領域での実装ニーズに応える多様性を強めたことを示しています。これにより、開発者はタスクの特性に応じて最適なモデルを迅速に選択・展開できるようになり、AI 導入のハードルがさらに低下すると予想されます。
編集コメント
セキュリティ特化型モデルの独立したリリースは、AI の実社会実装において「安全性」が単なる付加機能ではなく中核的な要件として扱われ始めていることを示唆しています。用途別ラインナップの拡充により、開発現場でのモデル選定がよりシームレスになることが期待されます。
2026 年 7 月 21 日
15 分間の読了時間
最新の Gemini モデルは、大規模な AI エージェントを構築するために必要な効率性、低遅延、そして信頼性を提供します。

このコンテンツは Google AI によって生成されています。生成 AI は実験的な技術です。
開発者や、本番環境で AI エージェントを構築している顧客にとって、トークン効率の向上、遅延の低減、そしてより安定したパフォーマンスが求められています。Flash シリーズのモデルは、効率と品質の最適なバランスを実現し、エージェントワークフローのスケーリングを可能にするために設計されています。
Gemini 3.5 Flash を基盤に、新たに以下の Gemini モデルを発表します:
3.6 Flash:これは当社の主力モデルであり、コーディング能力、知識処理、マルチモーダル性能のすべてにおいて向上を実現しています。Artificial Analysis Index によると、出力トークンあたりの使用量は 3.5 Flash と比較して 17% 削減され、Datacurve の DeepSWE ベンチマークでは最大 65% も削減できることが確認されています。これらはすべて、出力トークンあたりのコストを下げながら達成された成果です。
3.5 Flash-Lite:これは 3.5 クラスの中で最速かつ最もコストパフォーマンスに優れたモデルです。Artificial Analysis Index によると、1 秒間に 350 トークンの生成速度を誇り、エージェントワークフローにおいては以前の Flash-Lite シリーズを大きく上回る性能を発揮します。
3.5 Flash Cyber in CodeMender:効果的なサイバーセキュリティ対策には、モデルとエージェントインフラの慎重な連携が不可欠です。そこで今回紹介するのは、高効率でサイバーセキュリティに特化した新モデルと、CodeMender コードセキュリティエージェントを組み合わせたソリューションです。この組み合わせにより、最先端レベルでの競争力あるパフォーマンスを実現しています。
本日のリリースに加え、Gemini 3.5 Pro は現在パートナー企業とのテストを進めており、準備が整い次第広く提供開始する予定です。並行して、チームは次世代モデルの開発に注力しています。すでに Gemini 4 のための最も野心的な事前学習を開始しており、その進捗に大きな期待を抱いています。
3.6 Flash:3.5 Flash よりも効率的で、品質も向上
Gemini 3.6 Flash は、開発者や顧客からの 3.5 Flash に関するフィードバックを直接反映して構築されました。コーディングや知識作業の性能が一段階アップするだけでなく、トークンの効率性も大幅に改善されています。例えば、Artificial Analysis Index のベンチマークでは、3.6 Flash は 3.5 Flash に比べて出力トークンを 17% 削減しています。また、マルチステップワークフローを完了させるための推論ステップ数やツール呼び出し回数も減少しました。
この効率性の向上は、価格の引き下げとも相まって実現されています。入力トークン 100 万あたり 1.50 ドル、出力トークン 100 万あたり 7.50 ドルという設定により、3.6 Flash はエージェントタスクあたりの総コストを削減し、より低コストでエージェントの構築と運用が可能になります。
OSWorld で検証されたタスク(API)において、3.6 Flash は 3.5 Flash と比較してトークン効率が高く、冗長性が低いことが示されています。
効率的であるにもかかわらず、3.6 Flash はあらゆるユースケースで 3.5 Flash よりも性能が向上しています。
3.6 Flash は、DeepSWE(49% vs. 37%)や MLE Bench(63.9% vs. 49.7%)での評価に見られるように、不要なコード修正を減らし実行ループを短縮することで精度が向上しています。また、ML Research 分野でも大きな進歩が見られます。
OSWorld-Verified(83.0% vs. 78.4%)の結果に示される通り、コンピュータ操作能力も強化されました。この機能は Gemini API や Gemini Enterprise を通じて、クライアントサイドの組み込みツールとして利用可能です。
GDPval-AA v2(1421 vs. 1349)などのベンチマークで示されるように、3.5 Flash よりも知識処理タスクでのパフォーマンスが優れています。Hebbia や Harvey といった顧客からは、ドキュメントの解析やチャート・データ分析、レポート作成といったマルチモーダルなタスクにおいて特に高い能力を発揮していると評価されています。
顧客からは、3.6 Flash がコストと品質の両面で一歩進んだモデルであると報告されています。複雑なワークフローや知識ベースのタスクにおいて、トークン効率性、精度、速度のバランスが優れている点が高く評価されています。
セーフティを重視して構築
3.6 Flash は、化学・生物・放射能・核(CBRN)分野およびサイバー攻撃の悪用に関する Frontier Safety の強化されたセーフガードを搭載してリリースされています。これらの対策により、モデルは jailbreak に対して大幅に耐性を持つようになりました。同時に、有益な用途に対する拒否反応を最小限に抑えるようトレーニングが行われています。
詳細は、3.6 Flash のモデルカードをご覧ください。
3.5 Flash-Lite: エージェントワークフローの拡張に特化
Flash シリーズに加え、Gemini 3.5 Flash-Lite もリリースしました。このモデルは低遅延が求められるタスクや、エージェント検索やドキュメント処理など、開発者ワークフローにおいてスループットが重要な用途のために設計されています。
3.5 Flash-Lite は 3.5 シリーズ中最速のモデルです。Artificial Analysis の測定によると、出力トークンあたりの処理速度は秒間 350 トークンを記録しています。入力トークン 100 万あたり 0.3 ドル、出力トークン 100 万あたり 2.5 ドルという価格設定で、3.1 Flash-Lite と比べて品質も大幅に向上しているため、高スループットの生産環境を運用する開発者や顧客にとって、非常に優れたコストパフォーマンスを実現しています。
3.5 Flash-Lite は、3.5 Flash を下回る遅延時間で大量のタスクを実行できます。
このモデルは、エージェントシステムの効率的な拡張を可能にします。思考レベル(thinking levels)に関わらず、3.1 Flash-Lite よりも著しく高い性能を発揮します。ワークロードに応じて、開発者は最小限または低レベルの思考設定で低遅延・低コストの実行を優先し、大量のタスクを処理することもできます。また、多段階のサブエージェント作業を処理する必要がある場合は、より高度な思考レベルを選択可能です。さらに、これらのエージェントタスクをあらゆる画面で確実にサポートするため、コンピュータ操作機能を組み込みツールとして備えています。
Terminal-Bench 2.1(54% vs 31%)、GDM-MRCR v2(72.2% vs 60.1%)、GDPval-AA v2(1140 vs 642)といった評価結果から、コーディングやエージェントタスクにおける大幅な進化が確認できます。
実際、多くのエージェントおよびコード関連の評価において、3.5 Flash-Lite は 3 Flash を上回る性能を示しています。具体的には SWE-Bench Pro で 54.2% vs 49.6%、OSWorld-Verified で 74.0% vs 65.1% との結果を残し、2.5 および 3 Flash のワークロードにおいて、より高速で高機能な選択肢として注目されています。
3.5 Flash-Lite の初期導入企業からは、エージェントワークフローやデータ処理タスクの拡張に向けた、速度・知能性、コスト効率を兼ね備えた独自の組み合わせが評価されています。
モデルの詳細については、3.5 Flash-Lite のモデルカードをご覧ください。
3.5 Flash Cyber:CodeMender で脆弱性を効率的に発見・修正する
AI モデルは、既存のシステムが対応できる速度よりも速くセキュリティ上の脆弱性を特定できるようになっています。この拡大する脅威に対処するには、高度な能力と効率性を兼ね備えたソフトウェア保護のアプローチが必要です。
Flash のパフォーマンスと効率性は、大規模なコードセキュリティ課題の検出、検証、修正のための理想的な基盤となります。
Gemini 3.5 Flash Cyber は 3.5 Flash をベースに構築され、より大型モデルよりもトークンあたりのコストを抑えつつ、サイバーセキュリティの脆弱性を見つけ、修正することに特化して微調整されています。
複数の 3.5 Flash Cyber エージェントが連携して単一の統合レポートを生成する「CodeMender」においては、3.5 Flash Cyber は人気ベンチマークである CyberGym で最先端レベルと競合する性能を発揮しています。
この技術には二重利用の性質があるため、3.5 Flash Cyber の導入にあたっては意図的なアプローチを採用しました。本モデルは、限定的なアクセスパイロットプログラムの一環として、まもなく政府機関および信頼できるパートナー限定で CodeMender を通じて提供されます。これにより、最前線の防御チームが脆弱性が悪用される前にそれを見つけ、修正する上で有利なスタートを切れるよう支援しつつ、広範な誤使用を防ぐことを目指します。
3.6 Flash と 3.5 Flash-Lite: 今日から利用開始可能
3.6 Flash と 3.5 Flash-Lite は、本日より利用可能です。
Gemini API を通じて Google AI Studio や Android Studio で開発を行う方々には、3.6 Flash が利用可能です。また、Google Antigravity でも 3.6 Flash をお試しいただけます。まずは開発者ガイドをご覧ください。
企業向けには Gemini Enterprise Agent Platform が用意されています。Gemini Enterprise アプリでも 3.6 Flash をご利用いただけます。
一般ユーザーは Gemini アプリを通じてアクセスできます。3.5 Flash-Lite は現在、Google Search でも順次展開中です。
3.6 Flash や 3.5 Flash-Lite の活用を始めるにあたり、今後の Gemini モデル改善のためのフィードバックをお寄せください。また、間もなく 3.5 Pro のリリースも予定しています。
原文を表示
Jul 21, 2026
|
15 min read
Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.

Your browser does not support the audio element.
Listen to article
[[duration]] minutes
This content is generated by Google AI. Generative AI is experimental
Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows. Building on Gemini 3.5 Flash, we’re introducing new Gemini models:
- 3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash, and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token.
- 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index, also significantly outperforming prior Flash-Lite generations in agentic workflows.
- 3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure. We’re introducing a combination of a new, highly efficient, specialized cyber-focused model paired with our CodeMender code security agent that delivers competitive performance at the frontier.
Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready. In parallel, our team is already focusing on building the next generation of models. We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
3.6 Flash: More efficient and better quality than 3.5 Flash
Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency. For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash. It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows.
This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
3.6 Flash shows better token efficiency and reduced verbosity than 3.5 Flash in an OSWorld verified task (API)
Even while being more efficient, 3.6 Flash sees performance gains compared to 3.5 Flash across use cases:
- 3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%), and shows significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%).
- It has improved computer use capabilities as seen in OSWorld-Verified (83.0% vs. 78.4%). Computer use is now a built-in client side tool via the Gemini API and Gemini Enterprise.
- It outperforms 3.5 Flash in knowledge work, as shown by benchmarks like GDPval-AA v2 (1421 vs. 1349). Customers like Hebbia and Harvey have found it particularly capable at multimodal tasks like document parsing, chart and data analysis, and report drafting.
Customers report 3.6 Flash is a step forward in both cost and quality, balancing token efficiency, accuracy, and speed across complex workflows and knowledge-based tasks:
Built with safety
3.6 Flash is shipping with enhanced Frontier Safety safeguards in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense misuses. These safeguards make the model substantially more resistant to jailbreaks. At the same time, the model has been trained to minimize refusals for beneficial uses.
For more information, see the 3.6 Flash model card.
3.5 Flash-Lite: Built to scale agentic workflows
Beyond Flash, we’re also releasing Gemini 3.5 Flash-Lite, designed for both low-latency tasks and tasks where high throughput is critical for developers workflows, like agentic search and document processing.
3.5 Flash-Lite is the fastest model in the 3.5 series. As measured by Artificial Analysis, it runs at 350 output tokens/s. Priced at $0.3/1M input tokens and $2.5/1M output tokens and with significantly better quality than 3.1 Flash-Lite, 3.5 Flash-Lite offers a strong price-to-performance ratio for developers and customers running high throughput production traffic.
3.5 Flash-Lite executes high volume tasks at a lower latency than 3.5 Flash.
3.5 Flash-Lite enables efficient scaling for agentic systems. Across thinking levels, the model significantly outperforms 3.1 Flash-Lite. Depending on the workload, developers can configure the model to prioritize low-latency, low-cost execution for high-volume tasks with the minimal and low thinking levels, or engage higher thinking levels to process multi-step subagent workloads. The model now also has computer use as a built-in tool to reliably support these agentic tasks across surfaces.
It’s a significant step up in coding and agentic tasks as seen in Terminal-Bench 2.1 (54% vs 31%), long context as seen in GDM-MRCR v2 (72.2% vs. 60.1%), and real-world task execution as seen in GDPval-AA v2 (1140 vs. 642).
In fact, on many agentic and coding evals, 3.5 Flash-Lite even outperforms 3 Flash, including on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), making it a faster & more capable option for workloads on both 2.5 and 3 Flash.
Early customers of 3.5 Flash-Lite are highlighting its unique combination of speed, intelligence, and cost efficiency for scaling agentic workflows and data processing tasks:
For more information about the model, see the 3.5 Flash-Lite model card.
3.5 Flash Cyber in CodeMender: finding and fixing vulnerabilities efficiently
AI models have become capable of finding security vulnerabilities faster than current systems can fix them. Tackling this growing threat requires an approach to securing software that is highly capable and efficient.
Flash’s performance and efficiency makes it an ideal foundation to detect, validate, and patch code security issues at scale. Gemini 3.5 Flash Cyber is built on top of 3.5 Flash, and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models.
Within CodeMender, which uses multiple 3.5 Flash Cyber agents working together to produce a single combined report, 3.5 Flash Cyber reaches competitive performance at the frontier on the popular benchmark CyberGym.
Given the dual-use nature of this technology, we have taken an intentional approach to deploying 3.5 Flash Cyber. The model will be exclusively available to governments and trusted partners via CodeMender soon as part of a limited-access pilot program. This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse.
3.6 Flash and 3.5 Flash-Lite: Get started today
3.6 Flash and 3.5 Flash-Lite are available starting today:
- For developers in the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity. Get started with the Developer Guide.
- For enterprises in Gemini Enterprise Agent Platform. 3.6 Flash is also available in the Gemini Enterprise app.
- For everyone via the Gemini app. 3.5 Flash-Lite is also rolling out in Google Search.
As you start building with 3.6 Flash and 3.5 Flash-Lite, we welcome your feedback to improve future Gemini models and look forward to releasing 3.5 Pro soon.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み