Google DeepMind、ロボット制御モデル「Gemini Robotics ER 2」を発表
Google DeepMind は Gemini Robotics ER 2 を発表し、動画理解、タスクオーケストレーション、複数ロボット協調機能を活用して物理世界でのロボットの有用性を高める技術的飛躍を示した。
AI深層分析を開く2026年7月31日 00:23
AI深層分析
キーポイント
3 つの主要機能の統合
Gemini Robotics ER 2 は動画理解、タスクオーケストレーション、複数ロボット協調という 3 つの機能を統合し、ロボットの能力を向上させる。
物理世界での有用性向上
これらの機能により、ロボットが物理的な環境でより役立ち、実用的なタスクを遂行できる可能性が開かれる。
技術的ステップチェンジ
同社は今回の発表をロボティクス分野における「ステップチェンジ(飛躍)」と位置づけ、既存の枠組みを超える進展であると強調する。
リアルタイム思考と実行の並行処理
Gemini Robotics ER 2は、ロボットが行動を実行する一方で次の行動について「考える」ことを可能にする設計となっている。これにより、物理世界のスピードに合わせた意思決定と即時の対応が可能になる。
動画による進捗追跡と自己修正
連続的な動画フィードを監視することで、ロボットは自身の進捗を追跡し、エラーが発生した場合は適応して次のステップへ移行するタイミングを正確に把握できる。
重要な引用
Gemini Robotics ER 2 represents a step change in powering robots with video understanding, task orchestration, and multi-robot collaboration
making it possible for robots to be more helpful in the physical world
The design of Gemini Robotics ER 2 allows the robot to "think" about what comes next while simultaneously performing its actions.
By watching continuous video feeds, robots can now track their own progress, adapt if something goes wrong, and know exactly when to move on to the next step.
編集コメントを表示
編集コメント
動画理解と複数ロボット協調を一つのモデルで実現する試みは、実社会でのロボット活用において極めて重要な転換点となる。Google DeepMind は従来のシナリオベースの制御から、より高度な推論と協調へとパラダイムシフトを図っている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
2026 年 7 月 30 日
7 分間の読み物
Gemini Robotics ER 2 は、ロボットに動画理解、タスク調整、そしてマルチロボット協働の能力をもたらすことで、物理世界におけるロボットの有用性を飛躍的に高める新たなステップです。
スティーブン・ハンセン(シニアスタッフソフトウェアエンジニア)
ペング・シュウ(スタッフソフトウェアエンジニア)

このコンテンツは Google AI によって生成されています。生成 AI は実験的な技術です。
ロボットが日常環境で人間を支援するためには、正確な空間推論だけでは不十分です。ロボットはまた、物理世界のリアルタイムの速度に合わせて意思決定を行い、迅速に思考する必要があります。
そこで本日、ロボット向けの最も高度な「身体性推論」モデルであるGemini Robotics ER 2を発表します。これはロボットの頭脳に相当する高レベルのモデルで、人間との対話や物理世界の理解、多段階タスクの計画を可能にします。その後、モーターの実行は任意の低階層ビジョン・言語・アクション(VLA)モデルに引き継がれます。また、Google Search のようなツールのネイティブ呼び出しもサポートしており、必要な情報を検索したり、ユーザー定義の関数を実行したりできます。この設計により、ロボットは行動を遂行しながら同時に「次に何をするか」を考えることが可能になります。
Gemini Robotics ER 2 は、Gemini Robotics ER 1.6 と比較して大幅な進化を遂げています。連続する動画フィードを監視することで、ロボットは自身の進捗を追跡し、何か問題が起きた場合に適応し、次のステップへ移るタイミングを正確に把握できるようになりました。さらに、マルチロボット協働機能も導入しました。これにより、複数のロボットが共有空間で連携し、単独では達成できない複雑なワークフローを完了することが可能になります。
Gemini Robotics ER 2 が、開発者向けに公開されました。利用は Gemini API や Google AI Studio を通じて可能です。また、Gemini Enterprise Agent Platform ではプライベートプレビューとして提供されています。
使い始めを支援するため、モデルの構成方法や、より有用な物理 AI タスクを実現するためのプロンプト例を GitHub で公開しています。
物理的なエージェント機能の進化
現実世界のタスクは複雑で、完了までに複数のステップが必要です。Gemini Robotics ER 2 は物理エージェントとして、ロボットの動作をオーケストレーションし、自己修正や未知の状況への一般化を可能にします。
開発者は、Vision-Language-Action (VLA) モデルやナビゲーション API のような低レベル制御インターフェースをツールとして宣言し、マルチモーダルな動画、音声、テキストを直接モデルへストリーミングすることで、エージェント構成を実現できます。
Gemini Robotics ER 2 はこのツールオーケストレーションワークフローを改善します。シミュレーション上でのロボット評価や、実機による制御テスト、さらに人間が遠隔操作するロボットとの連携などを通じて、その性能を検証可能です。
Gemini Robotics ER 2 は、ツール操作において ER 1.6 を一貫して上回ります。これはリアルな VLA(Vision-Language-Action)、シミュレーション環境の VLA、そして人間の遠隔操作という 3 つの制御モードすべてで確認されています。
ロボット工学において、高度な推論は実行速度に依存します。Gemini Robotics ER 2 は Gemini Live API に統合されており、遅延が敏感なタスク向けに最適化された双方向ストリーミングエンドポイントを使用しています。その結果、ジェミニはアクションモデルやロボット API を指揮して、多段階のタスクを「停止して考える」という不自然な中断なしで完了させる、滑らかなオーケストレーションを実現します。
この仕組みを実証するために、パートナーである Boston Dynamics の Spot を活用したデモを作成しました。Gemini Robotics ER 2 がナビゲーションやマニピュレータの動作といった Spot APIs を指揮し、ユーザーのために物を運んでくるインタラクティブなロボットを構築しています。
Gemini Robotics ER 2 で駆動される Boston Dynamics の Spot が、自然言語による指示でポップコーンのスナックを回収する様子です。
このコードは、他のサンプルとともに Github で公開されています。
ロボットタスク完了の確実性を支える時間知能の活用
ロボット工学において最も困難な課題の一つが、「タスクが完了したと判断するタイミング」です。Gemini Robotics ER 2 は、動画理解と進捗追跡能力を大幅に強化し、電球のねじ込みやゴミ袋の結びといった複雑なタスクが仕様通りに完了しているかを検証してから次のステップへ移れるようになりました。
今回のアップデートでは、タスクの進捗を理解するための基盤となる 2 つの機能、「進捗分類」と「重要瞬間の特定」において大きな進展を遂げました。
継続的な進捗分類
進捗分類とは、ロボットがタスク完了に向けた進行状況を追跡する能力のことです。評価では、動画フレームごとに進捗度を 5 つのレベル(0-20%、20-40%、40-60%、60-80%、80-100%)に分類します。タスクの進行状況を数値化することで、Gemini Robotics ER 2 はロボットにリアルタイムでの状況認識を提供し、その場で行動を調整したり、失敗した手順をやり直したりできるようにしました。これにより、ワークフロー全体を最初からやり直す必要がなくなります。
進捗分類タスクにおける Gemini Robotics ER 2 の精度は 57.4% に達し、前世代モデルや競合する最先端モデルを上回る結果となりました。
精密な重要瞬間の特定
重要瞬間の特定(モーメント・ファインディング)とは、重要なイベントが発生する動画フレームを正確に識別する能力を指します。例えば、「コーヒーカップへの注ぎ止め」のタイミングです。Gemini Robotics ER 2 はこの分野でも性能を大幅に向上させ、ロボットがタスク間を精密に切り替えたり、成功を検証したり、必要な修正を提案したりすることを可能にしました。
時間特定タスクにおいて、Gemini Robotics ER 2 は 91.3% の精度と平均絶対誤差 0.96 秒を達成しました。このモデルははるかに大規模なカテゴリの競合他社と互角の性能を発揮しますが、計算コストはそのごく一部で済ませながら、実行速度は 4 倍に向上させています。これは現実世界で物理的なロボットを安全に運用するために実際に必要とされる、1 秒未満のレイテンシを実現したことを意味します。
マルチロボット連携
一つのロボットがすべてのタスクに適しているわけではありません。車輪付きローバーは屋内での作業に優れていますが、二足歩行型ヒューマノイドは不整地での活動に強みを持ちます。Gemini Robotics 2 はマルチロボット連携を可能にし、多様な機械が共通のセマンティック理解を通じて通信することで、複雑なタスクの引き継ぎと完遂を実現します。Gemini Robotics ER 2 が Apptronik の Apollo 2 と Franka F3 Duo を連携させる仕組みについては、こちら でご覧ください。
汎用的な空間知能の向上
Gemini Robotics ER 2 は、3 つのベンチマークで測定される当社の中核となる空間推論能力をさらに発展させました。
成功・失敗の検出機能は、静止画像ではなく生動画フィードを直接処理するよう進化しました。これにより、こぼれや滑り、位置ズレなど、実行中のミスをリアルタイムで捉えることが可能になりました。
計器類の読み取り能力も大幅に拡張され、従来の円形ダイヤルや視認窓だけでなく、デジタル表示、直線スケール、定規、液体温度計など多様な計器に対応しています。実際、10 種類の異なる計器種別でテストを実施し、その精度を確認しました。
空間的なビジュアル質問応答(VQA)も、Gemini のマルチモーダル理解技術の向上により強化されています。
Gemini Robotics ER 2 は、これらの主要機能すべてにおいて一貫して最高レベルの精度を達成しています。特に成功検出(画像・動画)、質問応答(ERQA)、そして汎用的な計器読み取り能力がその成果を象徴しています。
実体化知能の安全性を高める
Gemini Robotics ER 2 は、これまでで最も安全なモデルです。推論タスク中の物理的制約への準拠や、人間の検出における空間認識能力を評価する「Safety Instruction Following」と「Human Proximity」のベンチマークにおいて大幅な性能向上を実現しました。具体的には、Gemini Robotics ER 2 は人間が近くにいると人型ロボットを即座に停止し、周囲が安全になるまで自律的に作業を再開しないという挙動を示します。
物理的エージェントの安全性をさらに高めるため、私たちは新たなベンチマークを導入しました。これは、基盤モデルが安全な VLA(Vision-Language-Action)オーケストレーターとして機能できるかを評価するもので、安全制約の適用、環境監視、物理的な実現可能性の評価、そして必要に応じて人間の確認を求める能力をテストします。
詳細は、安全性に関する技術レポートをご覧ください。
Gemini Robotics ER 2 は、「Safety Instruction Following」と「Human Proximity」のベンチマークにおいて、ER 1.6 や他の最先端モデルを上回る性能を発揮しました。
今後の計画としては、これらのモデルをより複雑なタスクへと進化させ、有用なロボットの開発を加速するとともに、ロボット工学コミュニティへの支援を強化していく予定です。
原文を表示
Jul 30, 2026
|
7 min read
Gemini Robotics ER 2 represents a step change in powering robots with video understanding, task orchestration, and multi-robot collaboration — making it possible for robots to be more helpful in the physical world.
Steven Hansen
Senior Staff Software Engineer
Peng Xu
Staff Software Engineer

Your browser does not support the audio element.
Listen to article
[[duration]] minutes
This content is generated by Google AI. Generative AI is experimental
For robots to assist humans in everyday environments, accurate spatial reasoning is not enough. Robots must also think fast, timing their decisions and reasoning with the real-time speed of the physical world.
That’s why today we’re launching Gemini Robotics ER 2, our most capable “embodied reasoning” model for robotics. Think of Gemini Robotics ER 2 as a high-level brain for robots. It allows robots to chat with humans, understand the physical world, and plan multi-step tasks. It then hands off motor execution to any given lower level vision-language-action (VLA) model. Gemini Robotics ER 2 can also natively call tools like Google Search to find information, or any other user-defined function. The design of Gemini Robotics ER 2 allows the robot to “think” about what comes next while simultaneously performing its actions.
Gemini Robotics ER 2 represents a significant upgrade over Gemini Robotics ER 1.6. By watching continuous video feeds, robots can now track their own progress, adapt if something goes wrong, and know exactly when to move on to the next step. We are also introducing multi-robot collaboration, enabling robots to work together in shared spaces and complete complex workflows a single robot could not do alone.
Gemini Robotics ER 2 is now publicly available to developers via the Gemini API, Google AI Studio, and in private preview on Gemini Enterprise Agent Platform. To help you get started, we’re sharing examples of how to configure the model and prompt it to power more useful physical AI tasks.
Advancing physical agentic capabilities
Most tasks in the physical world are complex and require multiple steps to complete. Gemini Robotics ER 2 is a physical agent, orchestrating steps for the robot and enabling it to self-correct, and generalize to more novel situations. To build an agentic setup, developers can declare low-level control interfaces — like Vision-Language-Action (VLA) models or navigation APIs — as tools, and stream multimodal video, audio, or text directly into the model.
Gemini Robotics ER 2 improves this tool orchestration workflow. We can evaluate its performance with robots in simulation, using real-world robot control, and even pair it with a human controlling the robot remotely.
Gemini Robotics ER 2 consistently outperforms ER 1.6 for tool orchestration across three control modes: real VLA, sim VLA, and human tele-op.
In robotics, high-level reasoning depends on execution speed. Gemini Robotics ER 2 integrates into the Gemini Live API, using a bidirectional streaming endpoint optimized for latency-sensitive tasks. The result is fluid orchestration: Gemini Robotics ER 2 commands action models and robotics APIs to complete multi-step tasks without the jarring “stop-and-think” pauses.
To illustrate this, we’ve built a demo with Spot from our partners at Boston Dynamics. We use Gemini Robotics ER 2 to orchestrate Spot APIs, such as navigation and manipulator movement, creating an interactive robot that fetches objects for you.
Gemini Robotics ER 2 powered Boston Dynamic Spot fetches a popcorn snack up on a natural language command.
The code is available on Github with other examples.
Unlocking temporal intelligence for robust task completion
One of robotics’ hardest challenges is knowing when a task is done. Gemini Robotics ER 2 brings a step-change in video understanding and progress tracking to verify that complex tasks — such as tightening a light bulb or tying a trash bag — are complete to specification before switching to the next task.
In this update, we’ve made progress on two foundational capabilities for task progress understanding: progress classification and moment finding.
Continuous progress classification
Progress classification refers to a robot’s ability to track progress towards task completion. In our evaluations, we assign each frame in a video feed into five levels of progress (0-20%, 20-40%, 40-60%, 60-80%, 80-100%). By quantifying task progress, Gemini Robotics ER 2 provides robots with real-time situational awareness, and allows them to adjust actions on the fly or retry failed steps without restarting an entire workflow.
Gemini Robotics ER 2 achieves 57.4% accuracy on progress classification tasks, outperforming previous generation models and competing frontier models.
Precision moment-finding
Moment-finding measures a model's ability to identify the exact video frame where a critical event takes place (i.e. when to stop pouring coffee into a cup). Gemini Robotics ER 2 achieves significant gains in performance on moment finding, enabling robots to precisely switch between tasks, verify success and suggest corrections.
For moment-finding tasks, Gemini Robotics ER 2 achieves 91.3% accuracy and a 0.96s mean absolute distance. It competes closely with much larger model categories, but delivers this precision at a fraction of the compute cost and 4x the execution speed—the sub-second latency actually required to safely operate physical robotics in the real world.
Multi-robot collaboration
No single robot fits every task — a wheeled rover excels indoors, while a humanoid robot may excel at uneven terrain. Gemini Robotics 2 enables multi-robot collaboration, allowing diverse machines to communicate via a shared semantic understanding to handoff and complete complex tasks. See how Gemini Robotics ER 2 enables Apptronik’s Apollo 2 and Franka F3 Duo to collaborate here.
Improving general spatial intelligence
Gemini Robotics ER 2 advances our core spatial reasoning capability, as measured by three benchmarks:
- Success/failure detection: Now operates on raw video feeds rather than static snapshots to catch mid-execution failures like spills, slips, or misalignments.
- General instrument reading: Extends beyond circular dials and sight glasses to include digital displays, linear scales, rulers, and liquid thermometers. We tested it across 10 different types of instruments.
- Enhanced spatial VQA: Improves Visual Question Answering throughGemini’s advancements in multi-modal understanding.
Gemini Robotics ER 2 consistently achieves the highest accuracy across all core capabilities, with highlights including success detection (image/video), Question Answering (ERQA), and generalized instrument reading.
Advancing safety for embodied intelligence
Gemini Robotics ER 2 is our safest model, achieving significant gains on Safety Instruction Following and Human Proximity benchmarks, which evaluate how a model adheres to physical constraints during reasoning tasks and spatial awareness for detecting humans. We found that Gemini Robotics ER 2 successfully halts a humanoid robot when a person is nearby and autonomously resumes work only once the area is clear. To advance safety for physical agents, we’re introducing a benchmark that evaluates a foundation model's ability to act as a safe VLA orchestrator by testing its capacity to enforce safety constraints, monitor the environment, assess physical feasibility, and seek human clarification. For details, see our safety technical report.
Gemini Robotics ER 2 outperforms ER 1.6 and other frontier models on Safety Instruction Following and Human Proximity benchmarks.
Looking ahead, our plans are to push these models towards even more complex tasks to accelerate the development of helpful robots and support the robotics community.
AI算出
主要ニュースainew評価高い
AI モデルの核心機能(動画理解、タスクオーケストレーション、マルチロボット協働)が明確に定義された画期的な発表であり、既存記事との比較でも独立した新事実として評価できる。検索意図は「Gemini Robotics ER 2」という具体的な製品名で極めて高い。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み