Google DeepMind、新モデル「Gemini 3.6 Flash」などを発表
Google DeepMind は、AI エージェントのスケール構築を目的とした効率性、低遅延、高信頼性を備えた新モデル「Gemini 3.6 Flash」および関連シリーズを発表した。
AI深層分析を開く2026年7月27日 22:03
AI深層分析
キーポイント
新モデルの発表と目的
Google DeepMind は、開発者や顧客がプロダクション環境で AI エージェントをスケールさせるために設計された新しい Gemini モデルシリーズを発表した。
Gemini 3.6 Flash の性能向上
「Gemini 3.6 Flash」はワークホースモデルとして位置づけられ、コーディング、知識作業、マルチモーダル処理において従来比の改善が図られている。
Flash シリーズの設計思想
Flash シリーズ全体で効率性と品質の最適なバランス(スイートスポット)を追求し、トークン効率の向上と遅延の低減を実現している。
Gemini 3.6 Flash の効率性とコスト削減
Artificial Analysis Index によると、出力トークン使用量が 3.5 Flash よりも 17%減少し、DeepSWE ベンチマークでは最大 65%の削減を実現する。入力トークンあたり$1.50、出力トークンあたり$7.50という価格設定により、エージェントタスク全体の費用対効果が向上した。
Gemini 3.5 Flash-Lite の高速化
Artificial Analysis Index で 1 秒間に 350 トークンの出力速度を記録し、3.5 クラスモデルの中で最速かつ最もコスト効果が高い。これにより、以前の Flash-Lite 世代よりもエージェントワークフローでのパフォーマンスが大幅に向上した。
重要な引用
Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.
Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance.
"Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows."
"3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency."
編集コメントを表示
編集コメント
Gemini シリーズのフラッシュラインがさらに進化し、エージェント構築の実用性が一段と高まった。特にコーディングやマルチモーダル処理の強化は、開発現場での即戦力としての期待を高める内容だ。
2026 年 7 月 21 日
15 分間の読了時間
最新の Gemini モデルは、大規模な AI エージェントの構築に必要な効率性、低遅延、そして信頼性を提供します。

プロダクション環境で AI エージェントを開発・運用する開発者や顧客は、より高いトークン効率、低遅延、そして安定したパフォーマンスを求めています。Flash シリーズのモデルは、効率と品質の絶妙なバランスを実現し、エージェントワークフローのスケーリングを可能にするために設計されています。
Gemini 3.5 Flash の実績を踏まえ、新たに以下の Gemini モデルを発表します:
3.6 Flash は、コーディング、知識処理、マルチモーダル性能を向上させた主力モデルです。Artificial Analysis Index のデータによると、出力トークンあたりのコストは 3.5 Flash よりも低く抑えながら、トークン使用量を 17% 削減しています。また Datacurve が実施した DeepSWE ベンチマークでは最大で 65% の削減を達成しました。
3.5 Flash-Lite は、3.5 クラスモデルの中で最速かつ最もコストパフォーマンスに優れています。Artificial Analysis Index によると、出力速度は秒間 350 トークンを記録し、エージェントワークフローにおいても以前の Flash-Lite シリーズを大きく上回る性能を発揮します。
CodeMender における 3.5 Flash Cyber の活用事例です。サイバーセキュリティ対策では、モデルとエージェントインフラの連携が不可欠です。そこで私たちは、高効率でサイバーセキュリティに特化した新モデルと、コードセキュリティのエージェント「CodeMender」を組み合わせたアプローチを導入しました。これにより、最先端レベルの競争力あるパフォーマンスを実現しています。
本日のリリースに加え、Gemini 3.5 Pro は現在パートナー企業とのテストを進めており、準備が整い次第広く提供予定です。同時にチームは次世代モデルの開発にも着手しており、Gemini 4 のための最も野心的な事前学習をすでに開始しました。その進捗に大きな期待を抱いています。
3.6 Flash:より効率的で、品質も向上
Gemini 3.6 Flash は、開発者や顧客からの 3.5 Flash へのフィードバックを直接反映して構築されました。コーディングや知識作業の性能が一段階向上するだけでなく、トークンの効率性も大幅に改善されています。例えば、「Artificial Analysis Index」によると、3.6 Flash の出力トークン消費量は 3.5 Flash より 17% 減少しています。また、マルチステップワークフローを完了させるための推論ステップ数やツール呼び出し回数も削減されました。
この効率性の向上は、3.5 Flash よりも低廉な価格設定と相まって、さらに大きなメリットをもたらします。入力トークン 100 万あたり 1.50 ドル、出力トークン 100 万あたり 7.50 ドルという価格により、エージェントタスクあたりの総コストが削減され、エージェントの構築と運用をよりコストパフォーマンス良く実現できます。
3.6 Flash は、OSWorld で検証されたタスク(API)において、3.5 Flash を上回るトークン効率性と、冗長性の少ない出力を示しています。
効率的であるにもかかわらず、3.6 Flash はあらゆるユースケースで 3.5 Flash よりもパフォーマンスが向上しています。
3.6 Flash は、DeepSWE(49% 対 37%)や MLE Bench(63.9% 対 49.7%)などのベンチマークで示される通り、不要なコード編集を減らし実行ループを短縮することで精度が向上しています。また、ML Research 分野での大幅な改善も確認されています。
OSWorld-Verified(83.0% 対 78.4%)の結果に見られるように、コンピュータ操作の能力も強化されました。この機能は Gemini API や Gemini Enterprise を通じて、クライアントサイドの組み込みツールとして利用可能です。
GDPval-AA v2(1421 対 1349)などのベンチマークで示される通り、3.5 Flash よりも知識作業での性能が向上しています。Hebbia や Harvey といった顧客からは、文書解析やチャート・データ分析、レポート作成といったマルチモーダルタスクにおいて特に高い能力を発揮していると評価されています。
顧客からは、3.6 Flash がコストと品質の両面で一歩前進したモデルであり、複雑なワークフローや知識ベースのタスクにおいて、トークン効率、精度、速度をバランスよく実現できているとの報告があります。
セーフティを重視して設計
3.6 Flash は、化学・生物・放射線・核(CBRN)分野およびサイバー攻撃における悪用防止に向けた、強化された Frontier Safety 対策を搭載してリリースされています。これらの対策により、モデルは jailbreak に対する耐性が大幅に向上しています。同時に、有益な用途については拒否反応を最小限に抑えるようトレーニングされています。
詳細は 3.6 Flash のモデルカードをご覧ください。
3.5 Flash-Lite: エージェントワークフローのスケーリングに特化
Flash シリーズに加え、Gemini 3.5 Flash-Lite もリリースしました。低遅延が求められるタスクや、エージェント検索、ドキュメント処理など開発者のワークフローにおいて高スループットが不可欠な用途向けに設計されています。
3.5 Flash-Lite は 3.5 シリーズ中最速のモデルです。Artificial Analysis の測定によると、出力トークン数は秒間 350 に達します。入力トークン 100 万あたり 0.3 ドル、出力トークン 100 万あたり 2.5 ドルという価格設定で、3.1 Flash-Lite よりもはるかに高品質を実現しています。これにより、大規模な生産環境のトラフィックを処理する開発者や顧客にとって、非常に優れたコストパフォーマンスを提供します。
3.5 Flash-Lite は、3.5 Flash を凌駕する低遅延で大量タスクを実行できます。
エージェントシステムにおける効率的なスケーリングを実現。思考レベルに応じて、3.1 Flash-Lite を大幅に上回る性能を発揮します。ワークロードの状況に合わせて、開発者は最小限または低い思考レベルを設定して低遅延・低コストでの実行を優先し、大量タスクを処理することも可能です。また、多段階のサブエージェント作業を処理する必要がある場合は、より高い思考レベルを活用できます。さらに、これらのエージェントタスクをあらゆる画面で確実にサポートするため、「コンピューター操作」機能を標準ツールとして搭載しました。
Terminal-Bench 2.1(54% vs 31%)や GDM-MRCR v2(72.2% vs 60.1%)、GDPval-AA v2(1140 vs 642)での結果が示す通り、コーディングやエージェントタスクにおける性能は大幅に向上しています。
実際、多くのエージェントおよびコード評価において、3.5 Flash-Lite は 3 Flash を上回るパフォーマンスを発揮します。具体的には SWE-Bench Pro(54.2% vs 49.6%)や OSWorld-Verified(74.0% vs 65.1%)でその差が顕著です。これにより、2.5 および 3 Flash のワークロードにおいて、より高速かつ高機能な選択肢として採用できます。
3.5 Flash-Lite の初期導入企業からは、エージェントワークフローやデータ処理タスクのスケールアップにおいて、速度・知能性・コスト効率という独自の組み合わせが評価されています。
モデルの詳細については、3.5 Flash-Lite のモデルカードをご覧ください。
3.5 Flash Cyber:CodeMender で効率的に脆弱性を発見・修正する
AI モデルは、既存のシステムが対応できる速度よりも速くセキュリティ上の脆弱性を特定できるようになっています。この拡大する脅威に対処するには、高度で効率的なソフトウェアセキュリティアプローチが必要です。
Flash の高いパフォーマンスと効率性は、大規模なコードセキュリティ課題の検出、検証、パッチ適用のための理想的な基盤となります。Gemini 3.5 Flash Cyber は 3.5 Flash をベースに構築され、より大きなモデルよりもトークンあたりのコストを抑えながら、サイバーセキュリティの脆弱性を見つけ、修正することに特化して微調整されています。
複数の 3.5 Flash Cyber エージェントが連携して単一の統合レポートを生成する「CodeMender」では、人気ベンチマークである CyberGym において最先端レベルと競合するパフォーマンスを発揮しています。
この技術には二重利用の性質があるため、3.5 Flash Cyber の展開には意図的なアプローチを採用しました。本モデルは、政府機関や信頼できるパートナー限定で、まもなく CodeMender を通じて提供される限定的なパイロットプログラムの一部として利用可能になります。これにより、最前線の防衛チームが脆弱性が悪用される前にそれを発見し修正する上で有利に進められるよう支援しつつ、広範な誤使用を防ぐことが可能となります。
3.6 Flash と 3.5 Flash-Lite:今日から始められます
3.6 Flash と 3.5 Flash-Lite は本日より利用可能です。
Gemini API(Google AI Studio および Android Studio を通じて)を利用する開発者向けには、3.6 Flash が利用可能です。また、Google Antigravity でも 3.6 Flash をお試しいただけます。詳細はデベロッパーガイドをご覧ください。
Gemini Enterprise Agent Platform を活用する企業様向けには、3.6 Flash が提供されています。さらに、Gemini Enterprise アプリでも 3.6 Flash をご利用いただけます。
Gemini アプリを通じて一般ユーザーの皆様にも展開しており、3.5 Flash-Lite は Google Search でも順次配信を開始しています。
3.6 Flash や 3.5 Flash-Lite の活用を始めていただいた皆様からのフィードバックを歓迎いたします。皆様の声を元に、今後の Gemini モデルの改善に努めてまいります。また、間もなく 3.5 Pro のリリースも予定しております。
原文を表示
Jul 21, 2026
|
15 min read
Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.

Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows. Building on Gemini 3.5 Flash, we’re introducing new Gemini models:
- 3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash, and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token.
- 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index, also significantly outperforming prior Flash-Lite generations in agentic workflows.
- 3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure. We’re introducing a combination of a new, highly efficient, specialized cyber-focused model paired with our CodeMender code security agent that delivers competitive performance at the frontier.
Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready. In parallel, our team is already focusing on building the next generation of models. We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
3.6 Flash: More efficient and better quality than 3.5 Flash
Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency. For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash. It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows.
This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
3.6 Flash shows better token efficiency and reduced verbosity than 3.5 Flash in an OSWorld verified task (API)
Even while being more efficient, 3.6 Flash sees performance gains compared to 3.5 Flash across use cases:
- 3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%), and shows significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%).
- It has improved computer use capabilities as seen in OSWorld-Verified (83.0% vs. 78.4%). Computer use is now a built-in client side tool via the Gemini API and Gemini Enterprise.
- It outperforms 3.5 Flash in knowledge work, as shown by benchmarks like GDPval-AA v2 (1421 vs. 1349). Customers like Hebbia and Harvey have found it particularly capable at multimodal tasks like document parsing, chart and data analysis, and report drafting.
Customers report 3.6 Flash is a step forward in both cost and quality, balancing token efficiency, accuracy, and speed across complex workflows and knowledge-based tasks:
Built with safety
3.6 Flash is shipping with enhanced Frontier Safety safeguards in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense misuses. These safeguards make the model substantially more resistant to jailbreaks. At the same time, the model has been trained to minimize refusals for beneficial uses.
For more information, see the 3.6 Flash model card.
3.5 Flash-Lite: Built to scale agentic workflows
Beyond Flash, we’re also releasing Gemini 3.5 Flash-Lite, designed for both low-latency tasks and tasks where high throughput is critical for developers workflows, like agentic search and document processing.
3.5 Flash-Lite is the fastest model in the 3.5 series. As measured by Artificial Analysis, it runs at 350 output tokens/s. Priced at $0.3/1M input tokens and $2.5/1M output tokens and with significantly better quality than 3.1 Flash-Lite, 3.5 Flash-Lite offers a strong price-to-performance ratio for developers and customers running high throughput production traffic.
3.5 Flash-Lite executes high volume tasks at a lower latency than 3.5 Flash.
3.5 Flash-Lite enables efficient scaling for agentic systems. Across thinking levels, the model significantly outperforms 3.1 Flash-Lite. Depending on the workload, developers can configure the model to prioritize low-latency, low-cost execution for high-volume tasks with the minimal and low thinking levels, or engage higher thinking levels to process multi-step subagent workloads. The model now also has computer use as a built-in tool to reliably support these agentic tasks across surfaces.
It’s a significant step up in coding and agentic tasks as seen in Terminal-Bench 2.1 (54% vs 31%), long context as seen in GDM-MRCR v2 (72.2% vs. 60.1%), and real-world task execution as seen in GDPval-AA v2 (1140 vs. 642).
In fact, on many agentic and coding evals, 3.5 Flash-Lite even outperforms 3 Flash, including on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), making it a faster & more capable option for workloads on both 2.5 and 3 Flash.
Early customers of 3.5 Flash-Lite are highlighting its unique combination of speed, intelligence, and cost efficiency for scaling agentic workflows and data processing tasks:
For more information about the model, see the 3.5 Flash-Lite model card.
3.5 Flash Cyber in CodeMender: finding and fixing vulnerabilities efficiently
AI models have become capable of finding security vulnerabilities faster than current systems can fix them. Tackling this growing threat requires an approach to securing software that is highly capable and efficient.
Flash’s performance and efficiency makes it an ideal foundation to detect, validate, and patch code security issues at scale. Gemini 3.5 Flash Cyber is built on top of 3.5 Flash, and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models.
Within CodeMender, which uses multiple 3.5 Flash Cyber agents working together to produce a single combined report, 3.5 Flash Cyber reaches competitive performance at the frontier on the popular benchmark CyberGym.
Given the dual-use nature of this technology, we have taken an intentional approach to deploying 3.5 Flash Cyber. The model will be exclusively available to governments and trusted partners via CodeMender soon as part of a limited-access pilot program. This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse.
3.6 Flash and 3.5 Flash-Lite: Get started today
3.6 Flash and 3.5 Flash-Lite are available starting today:
- For developers in the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity. Get started with the Developer Guide.
- For enterprises in Gemini Enterprise Agent Platform. 3.6 Flash is also available in the Gemini Enterprise app.
- For everyone via the Gemini app. 3.5 Flash-Lite is also rolling out in Google Search.
As you start building with 3.6 Flash and 3.5 Flash-Lite, we welcome your feedback to improve future Gemini models and look forward to releasing 3.5 Pro soon.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み