Google、Agentic 向け新 Flash モデル 3 種発表
Google は生産性エージェント向けに、コスト効率と速度を大幅に改善した「Gemini 3.6 Flash」および「3.5 Flash-Lite/Cyber」の3モデルをリリースし、業界標準のベンチマークで性能と価格競争力を強化しました。
キーポイント
Gemini 3.6 Flash の効率化とコスト削減
コーディングや知識作業に最適化され、出力トークン数が前モデル比17%減少し、DeepSWEベンチでは最大65%の削減を実現。出力価格も$9.00から$7.50へ引き下げられ、総コストが大幅に低下しました。
Gemini 3.5 Flash-Lite の高速処理能力
低遅延・高スループットを特徴とし、秒間350トークンの出力速度を達成。Terminal-BenchやGDM-MRCR v2などのベンチで前世代モデルを大幅に上回り、エージェント検索や文書処理に適しています。
セキュリティ強化と新機能の追加
化学・生物・放射線・核(CBRN)およびサイバー攻撃の悪用防止を含む「Frontier Safety」 safeguards を実装。また、Gemini API や Enterprise で組み込み型の「Computer use」ツールが利用可能になりました。
ベンチマークにおける性能向上
DeepSWEやMLE Benchなどの主要評価基準で3.5 Flashを大幅に上回るスコアを記録し、特にコード生成と知識作業の精度が顕著に改善されています。
Flash-Lite の構成可能な思考レベル
開発者は高ボリュームタスクに低コスト・低遅延の「minimal」や「low」を、多段階サブエージェントワークロードには「higher」を選択可能で、組み込みツールとしてコンピュータ操作もサポートします。
Gemini 3.5 Flash Cyber の並列実行アーキテクチャ
コードセキュリティエージェント CodeMender では、単一の巨大モデル呼び出しによるボトルネックを回避するため、複数の 3.5 Flash Cyber エージェントを並列に実行し(最大 5 回)、結果を統合してレポートを作成します。
ベンチマークおよび実世界テストでの卓越した性能
V8 JavaScript エンジン評価では固定回数呼び出しで 55 の固有課題を発見し、他モデルを上回り、実際のクラウド脆弱性調査チームによるテストでも公開 API のリモートコード実行欠陥を 2 時間で特定しました。
重要な引用
Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash.
On the DeepSWE benchmark by Datacurve, Google reports up to a 65% reduction.
Computer use is now a built-in client-side tool through the Gemini API and Gemini Enterprise.
The answer is a cheap model called many times.
It caught 10 issues the other two models missed.
"Gemini 3.6 Flash cuts output tokens by 17% (up to 65% on DeepSWE) and drops the output price from $9.00 to $7.50 per 1M."
影響分析・編集コメントを表示
影響分析
今回のリリースは、AI エージェント開発におけるコスト構造とパフォーマンスのバランスを根本から変える可能性があります。特にトークン効率の劇的な改善により、大規模な自動化タスクの実行コストが低下し、企業の AI 導入スピードが加速すると予想されます。また、セキュリティ機能の強化は、実社会での信頼性を高める重要な一歩となります。
編集コメント
Google は「Flash」シリーズをさらに細分化し、コストと速度の最適化に特化したラインナップを整備しました。これは、単なる性能競争から、実運用における経済性と信頼性を重視する市場の変化を反映した戦略的リリースと言えます。
プロダクション環境のエージェントを構築する開発者にとって、トークンの効率化、レイテンシの低減、そして信頼性の高いパフォーマンスが求められています。今日、Google は 3 つの新モデル「Gemini 3.6 Flash」「Gemini 3.5 Flash-Lite」「Gemini 3.5 Flash Cyber」をリリースしました。
これら 3 つはすべて「Flash タイア」に属します。この階層は、最大限の推論深度よりも、速度、コスト効率、そして高ボリュームのエージェントワークロードに適応するように Google がチューニングしたものです。
Gemini 3.6 Flash:品質向上、トークン削減、価格低下
Gemini 3.6 Flash は、新しいデフォルトの主力モデルです。3.5 Flash をベースに、コーディング、知識作業、マルチモーダルタスクを対象としています。最大のポイントは効率性の向上です。
Artificial Analysis Index によると、3.6 Flash の出力トークン数は 3.5 Flash よりも 17% 削減されています。また、Datacurve が提供する DeepSWE ベンチマークでは、最大で 65% の削減を実現したと Google は報告しています。さらに、マルチステップワークフローにおける推論ステップ数やツール呼び出し回数も減少しました。
効率化に伴い価格も引き下げられています。Gemini 3.6 Flash の料金は、入力トークン 100 万あたり 1.50 ドル、出力トークン 100 万あたり 7.50 ドルです。出力単価は、従来の 3.5 Flash で 9.00 ドルだったものが下がりました。冗長性の低減と出力価格の引き下げにより、エージェントタスクあたりの総コストが削減されます。
効率性の向上に伴い、品質も大きく改善されました。DeepSWE では 3.6 Flash が 49% を達成し、3.5 Flash の 37% を上回っています。MLE Bench では 63.9% に達し、同ベンチの 3.5 Flash(49.7%)を大きく引き離しました。また OSWorld-Verified では 83.0%、知識労働向けベンチマークである GDPval-AA v2 では 1421 というスコアを獲得し、3.5 Flash の 1349 を上回っています。
コンピュータ操作機能は、Gemini API や Gemini Enterprise を通じて、クライアントサイドの組み込みツールとして利用可能になりました。Hebbia や Harvey といった初期導入企業からは、文書の解析やチャート・データ分析、レポート作成において性能向上が報告されています。
Google は、3.6 Flash に強化された「フロンティアセーフティ」 safeguards を搭載して提供しています。これは化学・生物・放射線・核(CBRN)関連の悪用やサイバー攻撃目的での誤使用を防ぐためのものです。詳細は 3.6 Flash のモデルカードをご覧ください。
以下のインタラクティブな解説ツールでは、各モデルを前世代と比較したり、ご自身の利用量に応じたトークンコストを見積もったりできます。
Gemini 3.5 Flash-Lite:3.5 シリーズで最速のモデル
Gemini 3.5 Flash-Lite は、低遅延かつ高スループットが求められるタスクに最適化されています。主な用途はエージェントによる検索や文書処理です。Artificial Analysis の測定によると、出力トークン生成速度は秒間 350 トークンを記録しています。
料金体系は、入力トークン 100 万あたり 0.30 ドル、出力トークン 100 万あたり 2.50 ドルです。
このモデルは、以前の 3.1 Flash-Lite を大幅に上回っています。Terminal-Bench 2.1 では 54% のスコアを記録し、対照的に 3.1 Flash-Lite は 31% です。長文コンテキストを評価する GDM-MRCR v2 ベンチマークでは 72.2% に達し、同様に 60.1% を記録した従来モデルを凌駕しました。また GDPval-AA v2 では 1140 というスコアを獲得し、対照値の 642 を大きく引き離しています。
特筆すべきは、Flash-Lite がより古い 3 Flash の評価項目でも勝っている点です。SWE-Bench Pro では 54.2% で 49.6% を上回り、OSWorld-Verified でも 74.0% と 65.1% を引き離してリードしています。
Flash-Lite は、思考レベルを「最小」「低」「高」の 3 つから選べるように設計されています。開発者は、大量処理が必要なタスクには低コスト・低遅延の実行を優先できますし、多段階のサブエージェント処理には高い思考レベルを活用することも可能です。また、コンピュータ操作機能も標準で組み込まれています。
Gemini 3.5 Flash Cyber in CodeMender: 安価なエージェントがバグを発見して修正する
Gemini 3.5 Flash Cyber は、最も専門特化したリリースです。これは 3.5 Flash を基盤に、ソフトウェアの脆弱性を発見・検証・修正するために微調整されたモデルです。
この設計の前提は「探索空間の問題」にあります。深い欠陥を見つけるには、膨大な実行探索空間を調査する必要がありますが、巨大な単一モデルへの呼び出しがボトルネックになりかねません。
その解決策は、「安価なモデルを複数回呼び出す」というアプローチです。Google のコードセキュリティエージェントである CodeMender 内では、複数の Gemini 3.5 Flash Cyber エージェントが並列で動作します。CodeMender は最大 5 回モデルを呼び出し、各サブエージェントの発見結果を統合して一つのレポートにまとめます。
この構成は、CyberGym ベンチマークにおいて、はるかに大規模なモデルに対しても競合するパフォーマンスを発揮しています。
内部評価の結果は目を見張るものがあります。Google の「Big Sleep」ベンチマークでは、Flash Cyber がメインラインの 3.5 Flash や 3.6 Flash を大きく上回りました。V8 JavaScript エンジンのテストでは、一定数の実行回数で 55 件の確認済み不具合を発見しました。これに対し、メインラインの 3.5 Flash は 47 件、Claude Opus 4.6 は 36 件でした。Flash Cyber は他モデルが逃した 10 件の問題も捉えています。
実際の運用テストでは、Google のクラウド脆弱性調査チームがこのモデルを活用し、公開 API におけるリモートコード実行の欠陥をわずか 2 時間で特定しました。
imagehttps://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
コミュニティの反応は、予想通り二極化しました。開発者たちは価格と効率性を歓迎しましたが、旗艦モデルの遅れについては最も厳しい批判が浴せられました。
Hacker News では、「Google は信頼できる供給能力を超えて過剰に宣伝している」という意見も出され、実際にコードを書くセッションで苛立たしい思いをしたという声もありました。また、Flash Cyber のアクセス制限は、自動的な脆弱性発見ツールの管理を誰が行うべきかという「デュアルユース(二重利用)」の議論を巻き起こしました。
以下のダッシュボードは、プラットフォームごとの議論を集約したものです。これはスクレイピングされたデータセットではなく、質的編集による合成であり、その手法に関する注釈も併記されています。
(function(){
window.addEventListener('message', function(e){
var d = e.data || {};
if(d && d.mtpEmbed === 'sentiment' && typeof d.height === 'number'){
var f = document.getElementById('mtp-gemini-sentiment');
if(f){ f.style.height = Math.max(320, d.height) + 'px'; }
}
});
})();
利用開始について
Gemini 3.6 Flash と Gemini 3.5 Flash-Lite は、本日より利用可能です。開発者は Google AI Studio や Android Studio を通じて Gemini API でこれらのモデルにアクセスできます。また、Gemini 3.6 Flash は「Google Antigravity」にも導入され、GitHub Copilot でも順次展開されています。企業向けには Gemini Enterprise Agent Platform で両モデルを提供しており、Gemini 3.6 Flash は Gemini Enterprise アプリでも利用可能です。一般ユーザーは Gemini アプリを通じて利用でき、Gemini 3.5 Flash-Lite は Google Search での展開も進んでいます。
まずは開発者ガイドをご覧ください。
主なポイント
Gemini 3.6 Flash は出力トークン数を最大で 17% 削減(DeepSWE では最大 65%)し、出力料金を 100 万トークンあたり 9 ドルから 7.5 ドルに引き下げました。
Gemini 3.5 Flash-Lite は秒間 350 トークンの処理速度で、100 万トークンあたり 0.30 ドル(出力)と 2.50 ドル(入力)という価格設定です。SWE-Bench Pro や OSWorld-Verified のベンチマークでは、従来の「Gemini 3 Flash」を上回る性能を発揮しました。
Gemini 3.5 Flash Cyber は、低コストでマルチエージェントによるスキャンを可能にし、「CodeMender」の基盤となっています。その結果、V8 ベースの脆弱性として 55 の固有な問題を検出しました。これは Gemini 3.5 Flash が 47 件、Opus 4.6 が 36 件だったのと比較して優れています。
Flash Cyber は二重利用リスク(軍事転用などの懸念)があるため、政府機関や信頼できるパートナー限定の制限付きパイロットプログラムとして提供されています。
Google が Gemini 3.6 Flash、3.5 Flash-Lite、そして 3.5 Flash Cyber を発表しました。これらはエージェントワークロード向けに設計された、より安価でトークン効率の高い「Flash」シリーズの最新モデルです。
この発表は、MarkTechPost で最初に報じられました。
原文を表示
Developers building production agents need higher token efficiency, lower latency, and more reliable performance. Today, Google has released three new Gemini models. The lineup is Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. All three sit in the Flash tier, which Google tunes for speed, cost, and high-volume agentic work rather than maximum reasoning depth.
Gemini 3.6 Flash: better quality, fewer tokens, lower price
Gemini 3.6 Flash is the new default workhorse. It builds on 3.5 Flash and targets coding, knowledge work, and multimodal tasks. The main point is efficiency. On the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash. On the DeepSWE benchmark by Datacurve, Google reports up to a 65% reduction. The model also takes fewer reasoning steps and tool calls per multi-step workflow.
Pricing moves down alongside efficiency. Gemini 3.6 Flash is priced at $1.50 per 1M input tokens and $7.50 per 1M output tokens. The output rate drops from the previous $9.00 on 3.5 Flash. Lower verbosity and a lower output price reduces the total cost per agentic task.
Quality gains accompany the efficiency gains. On DeepSWE, 3.6 Flash scores 49% versus 37% for 3.5 Flash. On MLE Bench, it reaches 63.9% versus 49.7%. On OSWorld-Verified, it hits 83.0% versus 78.4%. On GDPval-AA v2, a knowledge-work benchmark, it scores 1421 versus 1349. Computer use is now a built-in client-side tool through the Gemini API and Gemini Enterprise. Early customers including Hebbia and Harvey cite gains in document parsing, chart and data analysis, and report drafting.
Google is shipping 3.6 Flash with enhanced Frontier Safety safeguards. These cover Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber-offense misuse. Full details are in the 3.6 Flash model card.
The interactive explainer below lets you compare each model against its predecessor and estimate token cost at your own volume.
(function(){
window.addEventListener('message', function(e){
var d = e.data || {};
if(d && d.mtpEmbed === 'explorer' && typeof d.height === 'number'){
var f = document.getElementById('mtp-gemini-explorer');
if(f){ f.style.height = Math.max(320, d.height) + 'px'; }
}
});
})();
Gemini 3.5 Flash-Lite: the fastest model in the 3.5 line
Gemini 3.5 Flash-Lite is highlighted for low-latency and high-throughput jobs. Target use cases include agentic search and document processing. As measured by Artificial Analysis, it runs at 350 output tokens per second. Pricing is $0.30 per 1M input tokens and $2.50 per 1M output tokens.
The model clears the prior 3.1 Flash-Lite by wide margins. On Terminal-Bench 2.1, it scores 54% versus 31%. On GDM-MRCR v2, a long-context benchmark, it reaches 72.2% versus 60.1%. On GDPval-AA v2, it scores 1140 versus 642. Notably, Flash-Lite also beats the older 3 Flash on some evals. It leads on SWE-Bench Pro at 54.2% versus 49.6% and on OSWorld-Verified at 74.0% versus 65.1%.
Flash-Lite exposes configurable thinking levels: minimal, low, and higher. Developers can prioritize low-cost, low-latency execution for high-volume tasks. They can also engage higher thinking levels for multi-step subagent workloads. Computer use is a built-in tool here too.
Gemini 3.5 Flash Cyber in CodeMender: cheap agents that find and patch bugs
Gemini 3.5 Flash Cyber is the most specialized release. It is built on 3.5 Flash and fine-tuned to find, validate, and patch software vulnerabilities. The design premise is the search-space problem. Finding deep flaws means exploring an immense execution search space. A single call to one massive model becomes a bottleneck.
The answer is a cheap model called many times. Inside CodeMender, Google’s code-security agent, multiple 3.5 Flash Cyber agents run in parallel. CodeMender invokes the model up to five times, then merges the sub-agent findings into one report. On the CyberGym benchmark, this setup reaches competitive performance against much larger models.
The internal evaluations are striking. On Google’s Big Sleep evaluation, Flash Cyber significantly surpassed mainline 3.5 Flash and 3.6 Flash. On the V8 JavaScript engine, it found 55 unique confirmed issues at a fixed number of invocations. That compares to 47 for mainline 3.5 Flash and 36 for Claude Opus 4.6. It caught 10 issues the other two models missed. In one real-world test, Google’s Cloud Vulnerability Research team used it to find remote-code-execution flaws in public APIs within two hours.
imagehttps://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
Community Reaction
Reaction split along predictable lines. Builders welcomed the price and efficiency. The delayed flagship drew the loudest criticism. On Hacker News, some argued Google is over-selling capacity it cannot reliably provision, citing frustrating hands-on coding sessions. The gated Flash Cyber release opened a dual-use debate about who should hold automated exploit-finding tools.
The dashboard below aggregates that discussion by platform. It is a qualitative editorial synthesis, not a scraped dataset, and the method note is embedded.
(function(){
window.addEventListener('message', function(e){
var d = e.data || {};
if(d && d.mtpEmbed === 'sentiment' && typeof d.height === 'number'){
var f = document.getElementById('mtp-gemini-sentiment');
if(f){ f.style.height = Math.max(320, d.height) + 'px'; }
}
});
})();
Availability
Gemini 3.6 Flash and 3.5 Flash-Lite are available starting today. Developers can access them through the Gemini API via Google AI Studio and Android Studio. Gemini 3.6 Flash is also in Google Antigravity and rolling out in GitHub Copilot. Enterprises get both models in the Gemini Enterprise Agent Platform, with 3.6 Flash in the Gemini Enterprise app. Everyone can use them via the Gemini app, and 3.5 Flash-Lite is rolling out in Google Search. Start with the Developer Guide.
Key Takeaways
Gemini 3.6 Flash cuts output tokens by 17% (up to 65% on DeepSWE) and drops the output price from $9.00 to $7.50 per 1M.
Gemini 3.5 Flash-Lite runs at 350 tokens/sec for $0.30/$2.50 per 1M and beats the older 3 Flash on SWE-Bench Pro and OSWorld-Verified.
Gemini 3.5 Flash Cyber powers CodeMender with cheap multi-agent scans; it found 55 unique V8 issues versus 47 and 36 for 3.5 Flash and Opus 4.6.
Flash Cyber is gated to governments and trusted partners under a limited-access pilot due to dual-use risk.
The post Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More Token-Efficient Flash Tier Built for Agentic Workloads appeared first on MarkTechPost.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み