Google、Gemini 3.6 Flash など新モデル発表
Google は、生産環境でのAIエージェント構築を目的とした効率性、低遅延、高信頼性を強化した「Gemini 3.6 Flash」「Gemini 3.5 Flash-Lite」「Gemini 3.5 Flash Cyber」の3つの新モデルを発表しました。
キーポイント
AIエージェント向け新モデルシリーズの発表
開発者や顧客がスケールするAIワークフローを構築するために、効率性と品質のバランスに優れた「Flash」シリーズの最新3モデル(3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber)が登場しました。
Gemini 3.6 Flash の性能向上
新ワークホースモデルである「3.6 Flash」は、コーディング能力、知識処理、そしてマルチモーダル性能において前世代(3.5 Flash)よりも大幅な改善が図られています。
生産環境における要件への対応
本リリースは、トークン効率の向上、遅延の低減、パフォーマンスの信頼性という、実運用(Production)で求められる3大課題に特化して設計されています。
Gemini 3.6 Flash の効率化とコスト削減
出力トークン使用量が 17% 減少し、DeepSWE ベンチマークでは最大 65% の削減を達成。また、入力 100 万トークンあたり$1.50、出力 100 万トークンあたり$7.50 と、3.5 Flash よりも低価格で提供される。
Gemini 3.5 Flash-Lite の高速化
Artificial Analysis Index によると出力速度が秒間 350 トークンに達し、アジェントワークフローにおいて以前の世代を大幅に上回る性能を発揮する。
セキュリティ特化モデルの登場
「CodeMender」コードセキュリティエージェントと組み合わせた専用サイバーモデル「3.5 Flash Cyber」が導入され、高度なセキュリティ応用における競争力のあるパフォーマンスを実現する。
性能と効率の向上
Gemini 3.6 Flash は、トークン効率と冗長性の削減に成功し、DeepSWE や MLE Bench などのベンチマークで 3.5 Flash を上回る精度を示しています。
重要な引用
Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.
Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance.
Our workhorse model that delivers better coding, knowledge work, and multimodal performance.
At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%), and shows significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%).
Computer use is now a built-in client side tool via the Gemini API and Gemini Enterprise.
影響分析・編集コメントを表示
影響分析
この発表は、生成AIが実験段階から実社会での大規模運用(Production)へと移行する中で、開発者が直面するコストと遅延のボトルネックを解消する重要な一歩です。特に「Flash」シリーズの強化により、複雑なエージェントワークフローの実装ハードルが下がり、より多くの企業がAI自動化を導入しやすくなるでしょう。
編集コメント
Gemini 3.6 Flashの登場は、特にコーディング支援や複雑なマルチモーダルタスクにおいて、実用レベルでの性能向上を示唆しており、開発現場での採用加速が期待されます。ただし、Artificial Analysis Indexなどの第三者評価との整合性を確認しつつ、実際のベンチマーク結果を注視する必要があります。
2026 年 7 月 21 日
開発者や、本番環境向け AI エージェントを構築する顧客にとって、トークン効率の向上、レイテンシの低減、そしてより安定したパフォーマンスが求められています。Flash シリーズは、効率的さと品質のバランスが最適化されたモデルとして設計されており、大規模なエージェントワークフローの実現を支えます。
Gemini 3.5 Flash の成果を踏まえ、新たに Gemini モデルを発表します。

3.6 Flash:これは私たちの主力モデルであり、コーディング能力、知識作業の効率化、そしてマルチモーダル性能を大幅に向上させます。Artificial Analysis Index によると、出力トークンあたりの使用量は 3.5 Flash と比較して 17% 削減され、Datacurve の DeepSWE などの特定のベンチマークでは最大 65% も削減可能です。しかも、出力トークンあたりのコストはさらに低下しています。
3.5 Flash-Lite:これは 3.5 クラスの中で最速かつ最もコストパフォーマンスに優れたモデルです。Artificial Analysis Index によると、1 秒間に 350 トークンを生成可能で、エージェントワークフローにおいては以前の Flash-Lite シリーズを大きく上回る性能を発揮します。
3.5 Flash Cyber in CodeMender:効果的なサイバーセキュリティ対策には、モデルとエージェントインフラの慎重な連携が不可欠です。そこで私たちは、高効率かつサイバーセキュリティに特化した新しい専用モデルと、コードセキュリティエージェント「CodeMender」を組み合わせたソリューションを発表します。これにより、最先端レベルで競合他社を凌駕するパフォーマンスを実現しています。
今回のリリースに加え、Gemini 3.5 Pro は現在パートナー企業とのテストを進めており、準備が整い次第広く提供予定です。同時に、チームは次世代モデルの開発に注力しており、Gemini 4 のための最も大規模な事前学習をすでに開始しました。その進捗には大きな期待を抱いています。
3.6 Flash:3.5 Flash よりも効率的で、品質も向上
Gemini 3.6 Flash は、開発者や顧客からの 3.5 Flash に関するフィードバックを直接反映して構築されました。コーディング能力や知識処理の面で一段階進化するだけでなく、トークンの効率性も大幅に改善しています。例えば、Artificial Analysis Index のベンチマークでは、3.6 Flash は 3.5 Flash に比べて出力トークンを 17% 削減しました。また、マルチステップワークフローを完了させるための推論ステップ数やツール呼び出し回数も減少しています。
この効率性の向上は、価格の引き下げとも相まって実現されています。入力トークン 100 万あたり 1.50 ドル、出力トークン 100 万あたり 7.50 ドルという設定により、3.6 Flash はエージェントタスクあたりの総コストを削減し、より安価にエージェントを構築・運用できるようにしました。
OSWorld で検証されたタスク(API)において、3.6 Flash は 3.5 Flash と比較してトークン効率が高く、冗長性が低いことが示されています。
効率的であるにもかかわらず、3.6 Flash はあらゆるユースケースで 3.5 Flash よりも性能が向上しています。
3.6 Flash は、DeepSWE(49% vs. 37%)や MLE Bench(63.9% vs. 49.7%)での評価に見られるように、不要なコード修正を減らし実行ループを短縮することで精度を高め、ML Research 分野でも大きな改善を示しています。
OSWorld-Verified(83.0% vs. 78.4%)の結果が示す通り、コンピュータ操作能力も向上しました。この機能は Gemini API や Gemini Enterprise を通じて、クライアントサイドの組み込みツールとして利用可能です。
GDPval-AA v2(1421 vs. 1349)などのベンチマークで 3.5 Flash を上回る性能を発揮しており、Hebbia や Harvey といった顧客からは、文書の解析やチャート・データ分析、レポート作成といったマルチモーダルタスクにおいて特に高い能力があると評価されています。
顧客からは、3.6 Flash がコストと品質の両面で一歩前進したモデルであり、複雑なワークフローや知識ベースのタスクにおいて、トークン効率、精度、速度をバランスよく実現できているとの報告が寄せられています。
セーフティを重視して設計
3.6 Flash は、化学・生物・放射能・核(CBRN)分野およびサイバー攻撃の悪用防止に関する Frontier Safety の強化されたセーフガードを搭載して提供されています。これらの対策により、モデルは jailbreak に対して大幅に耐性を持つようになりました。同時に、有益な用途については拒否を最小限に抑えるようトレーニングが行われています。
詳細は 3.6 Flash のモデルカードをご覧ください。
3.5 Flash-Lite: エージェントワークフローの拡張に最適化
Flash の他にも、低遅延タスクや、エージェント検索、ドキュメント処理など開発者ワークフローにおいて高スループットが不可欠な用途向けに設計された「Gemini 3.5 Flash-Lite」をリリースします。
3.5 Flash-Lite は 3.5 シリーズ中最速のモデルです。Artificial Analysis の測定によると、出力トークン速度は秒間 350 トークンを記録しています。入力トークン 100 万あたり 0.3 ドル、出力トークン 100 万あたり 2.5 ドルという価格設定で、3.1 Flash-Lite よりも品質が大幅に向上しているため、高スループットの生産環境を運用する開発者や顧客にとって、非常に優れたコストパフォーマンスを提供します。
3.5 Flash-Lite は、3.5 Flash を下回る遅延時間で大量のタスクを実行できます。
このモデルは、エージェントシステムの効率的な拡張を可能にします。思考レベル(thinking levels)に関わらず、3.1 Flash-Lite を大きく上回る性能を発揮します。ワークロードに応じて、開発者は低コスト・低遅延での実行を優先し、最小限または低い思考レベルで大量タスクを処理させる設定も可能です。また、多段階のサブエージェント作業を処理する必要がある場合は、より高い思考レベルを活用できます。さらに、これらのエージェントタスクをあらゆる画面で確実にサポートするため、「Computer Use」が組み込みツールとして追加されました。
Terminal-Bench 2.1(54% vs 31%)、GDM-MRCR v2(72.2% vs 60.1%)、GDPval-AA v2(1140 vs 642)での結果が示す通り、コーディングやエージェントタスクにおける能力は大幅に向上しています。また、リアルワールドのタスク実行においても同様の成果を上げています。
実際、多くの評価項目において、3.5 Flash-Lite は 3 Flash を上回る性能を発揮しています。SWE-Bench Pro(54.2% vs 49.6%)や OSWorld-Verified(74.0% vs 65.1%)での結果がその証です。これにより、2.5 および 3 Flash を使用するワークロードに対して、より高速で高機能な選択肢として注目されています。
3.5 Flash-Lite の初期利用者は、エージェントワークフローやデータ処理タスクのスケールアップにおいて、このモデルが持つ「速度」「知能」「コスト効率」の優れた組み合わせを高く評価しています。
モデルの詳細については、3.5 Flash-Lite のモデルカードをご覧ください。
3.5 Flash Cyber in CodeMender: 脆弱性の効率的な発見と修正
AI モデルは、現在のシステムが対応できる速度よりも速くセキュリティ上の脆弱性を特定できるようになっています。この増大する脅威に対処するには、高度で効率的なソフトウェアセキュリティアプローチが必要です。
Flash のパフォーマンスと効率性は、大規模なコードセキュリティ課題の検出、検証、修正のための理想的な基盤となります。Gemini 3.5 Flash Cyber は 3.5 Flash をベースに構築され、より大きなモデルよりもトークンあたりのコストを抑えつつ、サイバーセキュリティの脆弱性を見つけて修正することに特化してファインチューニングされています。
複数の 3.5 Flash Cyber エージェントが連携して単一の統合レポートを生成する「CodeMender」においては、3.5 Flash Cyber は人気ベンチマークである CyberGym で最先端レベルと競合するパフォーマンスを発揮しています。
この技術には二重利用の性質があるため、3.5 Flash Cyber の展開については意図的なアプローチを採用しました。本モデルは、限定的なアクセスパイロットプログラムの一環として、まもなく政府機関および信頼できるパートナー限定で CodeMender を通じて提供されます。これにより、最前線の防衛チームが脆弱性が悪用される前にそれを見つけて修正する時間を確保できるとともに、広範な誤用のリスクを軽減できます。
3.6 Flash と 3.5 Flash-Lite:今日から利用可能に
3.6 Flash と 3.5 Flash-Lite は、本日より利用可能です。
Gemini API は Google AI Studio と Android Studio を通じて開発者向けに提供されています。また、3.6 Flash は Google Antigravity でも利用可能です。まずは開発者ガイドをご覧ください。
企業向けには Gemini Enterprise Agent Platform が用意されており、3.6 Flash は Gemini Enterprise アプリでも利用できます。
一般ユーザーは Gemini アプリを通じてアクセスできます。さらに 3.5 Flash-Lite も Google Search で順次展開中です。
3.6 Flash と 3.5 Flash-Lite を活用して開発を進めるにあたり、今後の Gemini モデル改善のためのフィードバックを歓迎します。また、まもなく 3.5 Pro のリリースも予定しています。
原文を表示
Jul 21, 2026
|
13 min read
Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.

Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows. Building on Gemini 3.5 Flash, we’re introducing new Gemini models:
- 3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash, and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token.
- 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index, also significantly outperforming prior Flash-Lite generations in agentic workflows.
- 3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure. We’re introducing a combination of a new, highly efficient, specialized cyber-focused model paired with our CodeMender code security agent that delivers competitive performance at the frontier.
Beyond today’s releases, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready. In parallel, our team is already focusing on building the next generation of models. We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress.
3.6 Flash: More efficient and better quality than 3.5 Flash
Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but it does this while meaningfully improving token efficiency. For example, on the Artificial Analysis Index, we see 3.6 Flash consuming 17% fewer output tokens than 3.5 Flash. It also takes fewer reasoning steps and tool calls to accomplish multi-step workflows.
This enhanced efficiency is also combined with a lower price than 3.5 Flash. At $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash reduces the overall cost per agentic task, making agents more cost-effective to build and run.
3.6 Flash shows better token efficiency and reduced verbosity than 3.5 Flash in an OSWorld verified task (API)
Even while being more efficient, 3.6 Flash sees performance gains compared to 3.5 Flash across use cases:
- 3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%), and shows significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%).
- It has improved computer use capabilities as seen in OSWorld-Verified (83.0% vs. 78.4%). Computer use is now a built-in client side tool via the Gemini API and Gemini Enterprise.
- It outperforms 3.5 Flash in knowledge work, as shown by benchmarks like GDPval-AA v2 (1421 vs. 1349). Customers like Hebbia and Harvey have found it particularly capable at multimodal tasks like document parsing, chart and data analysis, and report drafting.
Customers report 3.6 Flash is a step forward in both cost and quality, balancing token efficiency, accuracy, and speed across complex workflows and knowledge-based tasks:
Built with safety
3.6 Flash is shipping with enhanced Frontier Safety safeguards in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense misuses. These safeguards make the model substantially more resistant to jailbreaks. At the same time, the model has been trained to minimize refusals for beneficial uses.
For more information, see the 3.6 Flash model card.
3.5 Flash-Lite: Built to scale agentic workflows
Beyond Flash, we’re also releasing Gemini 3.5 Flash-Lite, designed for both low-latency tasks and tasks where high throughput is critical for developers workflows, like agentic search and document processing.
3.5 Flash-Lite is the fastest model in the 3.5 series. As measured by Artificial Analysis, it runs at 350 output tokens/s. Priced at $0.3/1M input tokens and $2.5/1M output tokens and with significantly better quality than 3.1 Flash-Lite, 3.5 Flash-Lite offers a strong price-to-performance ratio for developers and customers running high throughput production traffic.
3.5 Flash-Lite executes high volume tasks at a lower latency than 3.5 Flash.
3.5 Flash-Lite enables efficient scaling for agentic systems. Across thinking levels, the model significantly outperforms 3.1 Flash-Lite. Depending on the workload, developers can configure the model to prioritize low-latency, low-cost execution for high-volume tasks with the minimal and low thinking levels, or engage higher thinking levels to process multi-step subagent workloads. The model now also has computer use as a built-in tool to reliably support these agentic tasks across surfaces.
It’s a significant step up in coding and agentic tasks as seen in Terminal-Bench 2.1 (54% vs 31%), long context as seen in GDM-MRCR v2 (72.2% vs. 60.1%), and real-world task execution as seen in GDPval-AA v2 (1140 vs. 642).
In fact, on many agentic and coding evals, 3.5 Flash-Lite even outperforms 3 Flash, including on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), making it a faster & more capable option for workloads on both 2.5 and 3 Flash.
Early customers of 3.5 Flash-Lite are highlighting its unique combination of speed, intelligence, and cost efficiency for scaling agentic workflows and data processing tasks:
For more information about the model, see the 3.5 Flash-Lite model card.
3.5 Flash Cyber in CodeMender: finding and fixing vulnerabilities efficiently
AI models have become capable of finding security vulnerabilities faster than current systems can fix them. Tackling this growing threat requires an approach to securing software that is highly capable and efficient.
Flash’s performance and efficiency makes it an ideal foundation to detect, validate, and patch code security issues at scale. Gemini 3.5 Flash Cyber is built on top of 3.5 Flash, and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models.
Within CodeMender, which uses multiple 3.5 Flash Cyber agents working together to produce a single combined report, 3.5 Flash Cyber reaches competitive performance at the frontier on the popular benchmark CyberGym.
Given the dual-use nature of this technology, we have taken an intentional approach to deploying 3.5 Flash Cyber. The model will be exclusively available to governments and trusted partners via CodeMender soon as part of a limited-access pilot program. This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse.
3.6 Flash and 3.5 Flash-Lite: Get started today
3.6 Flash and 3.5 Flash-Lite are available starting today:
- For developers in the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also available in Google Antigravity. Get started with the Developer Guide.
- For enterprises in Gemini Enterprise Agent Platform. 3.6 Flash is also available in the Gemini Enterprise app.
- For everyone via the Gemini app. 3.5 Flash-Lite is also rolling out in Google Search.
As you start building with 3.6 Flash and 3.5 Flash-Lite, we welcome your feedback to improve future Gemini models and look forward to releasing 3.5 Pro soon.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み