Google、Gemini Robotics ER 2 を発表しロボット制御の高速化を強調
本文の状態
日本語全文を表示中
詳細モードで約9分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
Gemini Robotics ER 2 は、ロボットの動画理解、タスク調整、複数ロボット間の協力を強化し、物理世界での実用性を高める新技術である。
AI深層分析を開く2026年7月31日 23:14
AI深層分析
キーポイント
リアルタイム思考の必要性
日常環境で人間を支援するロボットには正確な空間推論だけでなく、物理世界の速度に合わせた高速な意思決定と推論が不可欠であると指摘している。
動画理解機能の強化
Gemini Robotics ER 2 はロボットの動画理解能力を向上させ、視覚情報をより深く解析して行動につなげることを可能にする。
タスク調整と協調動作
単独での作業だけでなく、複数のロボットが連携して複雑なタスクを調整・実行する能力を備えていることが特徴である。
高速な空間推論とリアルタイム思考
Gemini Robotics ER 2 は物理世界の速度に合わせて意思決定を行う「エンボディード・レーニング」モデルであり、行動しながら次のステップについて考える設計となっている。
自己修正とマルチロボット連携の強化
連続する動画フィードを監視することで進捗を追跡し、失敗時に適応して次のステップへ移行できるほか、複数のロボットが協調して単独では不可能な複雑なワークフローを完了できるようになった。
重要な引用
For robots to assist humans in everyday environments, accurate spatial reasoning is not enough.
Robots must also think fast, timing their decisions and reasoning with the real-time speed of the physical world.
It then hands off motor execution to any given lower level vision-language-action (VLA) model.
By watching continuous video feeds, robots can now track their own progress, adapt if something goes wrong, and know exactly when to move on to the next step.
編集コメントを表示
編集コメント
この発表は、ロボットが物理世界でより効果的に動作するために必要な「速度」の概念を明確に提示している。技術的な詳細は限定的であるものの、次世代ロボットの設計思想を示唆する重要な一歩と言えるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
2026年7月30日
7分間の読み物
Gemini Robotics ER 2 は、ロボットに動画理解能力、タスクの調整機能、そして複数台のロボットによる協調動作をもたらすことで、物理世界におけるロボットの有用性を飛躍的に高める新たなステップです。
Steven Hansen(シニアスタッフソフトウェアエンジニア)
Peng Xu(スタッフソフトウェアエンジニア)

ロボットが日常環境で人間を支援するためには、正確な空間推論だけでは不十分です。物理世界のリアルタイムの速度に合わせて判断と推論を瞬時に行う必要があります。
そこで本日、私たちはロボット向けの最も能力の高い「具現化された推論(embodied reasoning)」モデルである Gemini Robotics ER 2 を発表します。これはロボットのハイレベルな脳のようなものです。人間との対話や物理世界の理解、そして多段階のタスク計画を可能にします。その後、モーターの実行は任意の低レベルのビジョン・言語・アクション(VLA)モデルに引き継がれます。また、Google Search などのツールをネイティブで呼び出して情報を取得したり、ユーザーが定義した他の関数を実行することもできます。Gemini Robotics ER 2 の設計により、ロボットは行動を実行しながら同時に「次はどうすべきか」を考えることが可能になります。
Gemini Robotics ER 2 は、Gemini Robotics ER 1.6 を大幅に強化した新モデルです。連続する動画フィードを監視することで、ロボットは自身の進捗状況を把握し、問題が発生すれば即座に適応し、次のステップへ移るべきタイミングも正確に判断できるようになりました。
また、マルチロボット協働機能も新たに追加しました。これにより、複数のロボットが共有空間で連携して作業を行い、単独では不可能な複雑なワークフローを完遂することが可能になります。
Gemini Robotics ER 2 は現在、Gemini API や Google AI Studio を通じて開発者に公開されています。また、Gemini Enterprise Agent Platform ではプライベートプレビュー版として利用可能です。
活用を始める際の参考として、モデルの構成方法や、より有用な物理 AI タスクを実現するためのプロンプト作成例を examples で公開しています。
物理的なエージェント能力の進化
現実世界のタスクは複雑で、完了には複数のステップが必要です。Gemini Robotics ER 2 は物理的なエージェントとして、ロボットの各ステップを調整し、自己修正機能やより新しい状況への汎化能力を実現します。このエージェント環境を構築する際、開発者は Vision-Language-Action (VLA) モデルやナビゲーション API のような低レベルの制御インターフェースをツールとして宣言できます。また、マルチモーダルな動画、音声、テキストをモデルに直接ストリーミングすることも可能です。
Gemini Robotics ER 2 は、このツールの調整ワークフローをさらに改善します。シミュレーション上でのロボット評価や、実世界のロボット制御によるテスト、さらには人間が遠隔操作するロボットとの連携などを通じて、その性能を検証できます。
Gemini Robotics ER 2 は、リアルな VLA、シミュレーション上の VLA、人間のテレオペレーションという 3 つの制御モードすべてにおいて、ツール調整の面で ER 1.6 を一貫して上回ります。
ロボティクス分野では、高度な推論は実行速度に依存します。Gemini Robotics ER 2 は Gemini Live API に統合されており、遅延が敏感なタスク向けに最適化された双方向ストリーミングエンドポイントを利用しています。その結果、「停止して考える」といった不自然な中断を伴わず、アクションモデルやロボティクス API を指揮して多段階のタスクをスムーズに完了させることが可能になります。
これを具体化するため、パートナーである Boston Dynamics のロボット「Spot」を活用したデモを作成しました。Gemini Robotics ER 2 を用いて Spot の API(ナビゲーションやマニピュレータの動作など)を制御し、ユーザーの指示に応じて物を運んでくるインタラクティブなロボットを実現しています。
自然言語による命令で、Gemini Robotics ER 2 が搭載された Boston Dynamics の Spot がポップコーンのおやつを回収する様子です。
このコードは、他のサンプルとともに Github で公開されています。
堅牢なタスク完了のための時間的知性の活用
ロボティクスにおける最も難しい課題の一つが、「タスクがいつ完了したか」を判断することです。Gemini Robotics ER 2 は、動画の理解と進捗追跡において飛躍的な進化をもたらしました。これにより、電球のネジ締めやゴミ袋の結びなど複雑なタスクが仕様通りに完了しているかを検証し、次のタスクへ移行する前に確実な判断を下すことが可能になります。
今回のアップデートでは、タスクの進捗を理解するための 2 つの基盤的能力において進展がありました。それは「進捗分類」と「特定瞬間の検出(moment finding)」です。
プログレス分類の継続的向上
プログレス分類とは、ロボットがタスク完了に向けた進捗を把握する能力のことです。当社の評価では、動画フレームを 5 つの進捗レベル(0-20%、20-40%、40-60%、60-80%、80-100%)に分類します。タスクの進行度を数値化することで、Gemini Robotics ER 2 はロボットにリアルタイムでの状況認識を提供し、その場で行動を調整したり、失敗した手順をやり直したりできるようにします。全体ワークフローを最初からやり直す必要はありません。
Gemini Robotics ER 2 は、プログレス分類タスクにおいて 57.4% の精度を達成し、前世代モデルや競合する最先端モデルを上回りました。
精密な瞬間特定
「モーメント・ファインディング(瞬間特定)」は、重要なイベントが発生する動画フレームの正確な位置を特定するモデルの能力を測る指標です。具体的には、「コップにコーヒーを注ぐのをいつ止めるか」といった判断が該当します。
Gemini Robotics ER 2 はこのモーメント・ファインディングにおいて大幅な性能向上を実現し、ロボットがタスク間を正確に切り替えたり、成功を検証したり、修正を提案したりすることを可能にしました。
モーメント・ファインディングタスクでは、91.3% の精度と 0.96 秒の平均絶対誤差を達成しています。これははるかに大規模なモデル群と互角の性能ですが、計算コストはごく一部で済み、実行速度は 4 倍です。物理的なロボットを実世界で安全に運用するために必要な、1 秒未満のレイテンシを実現しています。
マルチロボットのコラボレーション
1 台のロボットですべてのタスクをこなせるわけではありません。車輪付きローバーは屋内での作業に優れていますが、二足歩行型ロボットの方が不整地での活動には適しています。Gemini Robotics ER 2 は、異なる種類の機械が共通の意味理解を通じて通信し、複雑なタスクを引き継ぎながら完了させるマルチロボットのコラボレーションを実現します。
Gemini Robotics ER 2 がどのように Apptronik の Apollo 2 と Franka F3 Duo を連携させるかについては、こちら でご覧ください。
汎用的な空間知能の向上
Gemini Robotics ER 2 は、以下の 3 つのベンチマークで測定される核心的な空間推論能力をさらに進化させました。
- 成功・失敗の検出:静止画像ではなく生動画フィードを直接処理することで、こぼれや滑り、位置ずれなど実行中の失敗を検知できるようになりました。
- 汎用的な計器の読み取り:円形のダイヤルや視認窓だけでなく、デジタルディスプレイ、リニアスケール、定規、液体温度計などに対応。10 種類の異なる計器でテストを行いました。
- 高度化された空間 VQA(Visual Question Answering):Gemini の多モーダル理解能力の向上により、ビジュアルクエスチョンアンサーイングの精度が改善されました。
Gemini Robotics ER 2 は、すべてのコア機能において一貫して最高レベルの精度を達成しています。特に成功検出(画像・動画)、質問応答(ERQA)、そして汎用的な計器読み取りの性能が際立っています。
具身知能の安全性を向上させる
Gemini Robotics ER 2 は、私たちが開発した中で最も安全なモデルです。推論タスクにおける物理的制約への準拠度や、人間検出のための空間認識能力を評価する「Safety Instruction Following(安全指示の遵守)」および「Human Proximity(人間との近接性)」というベンチマークで、大幅な性能向上を実現しました。
具体的には、Gemini Robotics ER 2 は、人間が近くにいる場合にヒューマノイドロボットを即座に停止させ、周囲が安全になるまで自律的に作業を再開しないことを確認しています。物理的エージェントの安全性をさらに高めるため、私たちは新たなベンチマークを導入します。これは、基盤モデルが安全な VLA(Vision-Language-Action)オーケストレーターとして機能できるかを評価するもので、安全制約の強制、環境監視、物理的な実行可能性の評価、そして必要に応じて人間の確認を求める能力を検証します。
詳細については、安全性に関する技術レポートをご覧ください。
Gemini Robotics ER 2 は、「Safety Instruction Following」および「Human Proximity」のベンチマークにおいて、ER 1.6 や他の最先端モデルを上回る性能を発揮しました。
今後の計画としては、これらのモデルをより複雑なタスクへと進化させ、有用なロボットの開発を加速するとともに、ロボット工学コミュニティをサポートしていくことです。
原文を表示
Jul 30, 2026
|
7 min read
Gemini Robotics ER 2 represents a step change in powering robots with video understanding, task orchestration, and multi-robot collaboration — making it possible for robots to be more helpful in the physical world.
Steven Hansen
Senior Staff Software Engineer
Peng Xu
Staff Software Engineer

For robots to assist humans in everyday environments, accurate spatial reasoning is not enough. Robots must also think fast, timing their decisions and reasoning with the real-time speed of the physical world.
That’s why today we’re launching Gemini Robotics ER 2, our most capable “embodied reasoning” model for robotics. Think of Gemini Robotics ER 2 as a high-level brain for robots. It allows robots to chat with humans, understand the physical world, and plan multi-step tasks. It then hands off motor execution to any given lower level vision-language-action (VLA) model. Gemini Robotics ER 2 can also natively call tools like Google Search to find information, or any other user-defined function. The design of Gemini Robotics ER 2 allows the robot to “think” about what comes next while simultaneously performing its actions.
Gemini Robotics ER 2 represents a significant upgrade over Gemini Robotics ER 1.6. By watching continuous video feeds, robots can now track their own progress, adapt if something goes wrong, and know exactly when to move on to the next step. We are also introducing multi-robot collaboration, enabling robots to work together in shared spaces and complete complex workflows a single robot could not do alone.
Gemini Robotics ER 2 is now publicly available to developers via the Gemini API, Google AI Studio, and in private preview on Gemini Enterprise Agent Platform. To help you get started, we’re sharing examples of how to configure the model and prompt it to power more useful physical AI tasks.
Advancing physical agentic capabilities
Most tasks in the physical world are complex and require multiple steps to complete. Gemini Robotics ER 2 is a physical agent, orchestrating steps for the robot and enabling it to self-correct, and generalize to more novel situations. To build an agentic setup, developers can declare low-level control interfaces — like Vision-Language-Action (VLA) models or navigation APIs — as tools, and stream multimodal video, audio, or text directly into the model.
Gemini Robotics ER 2 improves this tool orchestration workflow. We can evaluate its performance with robots in simulation, using real-world robot control, and even pair it with a human controlling the robot remotely.
Gemini Robotics ER 2 consistently outperforms ER 1.6 for tool orchestration across three control modes: real VLA, sim VLA, and human tele-op.
In robotics, high-level reasoning depends on execution speed. Gemini Robotics ER 2 integrates into the Gemini Live API, using a bidirectional streaming endpoint optimized for latency-sensitive tasks. The result is fluid orchestration: Gemini Robotics ER 2 commands action models and robotics APIs to complete multi-step tasks without the jarring “stop-and-think” pauses.
To illustrate this, we’ve built a demo with Spot from our partners at Boston Dynamics. We use Gemini Robotics ER 2 to orchestrate Spot APIs, such as navigation and manipulator movement, creating an interactive robot that fetches objects for you.
Gemini Robotics ER 2 powered Boston Dynamic Spot fetches a popcorn snack up on a natural language command.
The code is available on Github with other examples.
Unlocking temporal intelligence for robust task completion
One of robotics’ hardest challenges is knowing when a task is done. Gemini Robotics ER 2 brings a step-change in video understanding and progress tracking to verify that complex tasks — such as tightening a light bulb or tying a trash bag — are complete to specification before switching to the next task.
In this update, we’ve made progress on two foundational capabilities for task progress understanding: progress classification and moment finding.
Continuous progress classification
Progress classification refers to a robot’s ability to track progress towards task completion. In our evaluations, we assign each frame in a video feed into five levels of progress (0-20%, 20-40%, 40-60%, 60-80%, 80-100%). By quantifying task progress, Gemini Robotics ER 2 provides robots with real-time situational awareness, and allows them to adjust actions on the fly or retry failed steps without restarting an entire workflow.
Gemini Robotics ER 2 achieves 57.4% accuracy on progress classification tasks, outperforming previous generation models and competing frontier models.
Precision moment-finding
Moment-finding measures a model's ability to identify the exact video frame where a critical event takes place (i.e. when to stop pouring coffee into a cup). Gemini Robotics ER 2 achieves significant gains in performance on moment finding, enabling robots to precisely switch between tasks, verify success and suggest corrections.
For moment-finding tasks, Gemini Robotics ER 2 achieves 91.3% accuracy and a 0.96s mean absolute distance. It competes closely with much larger model categories, but delivers this precision at a fraction of the compute cost and 4x the execution speed—the sub-second latency actually required to safely operate physical robotics in the real world.
Multi-robot collaboration
No single robot fits every task — a wheeled rover excels indoors, while a humanoid robot may excel at uneven terrain. Gemini Robotics 2 enables multi-robot collaboration, allowing diverse machines to communicate via a shared semantic understanding to handoff and complete complex tasks. See how Gemini Robotics ER 2 enables Apptronik’s Apollo 2 and Franka F3 Duo to collaborate here.
Improving general spatial intelligence
Gemini Robotics ER 2 advances our core spatial reasoning capability, as measured by three benchmarks:
- Success/failure detection: Now operates on raw video feeds rather than static snapshots to catch mid-execution failures like spills, slips, or misalignments.
- General instrument reading: Extends beyond circular dials and sight glasses to include digital displays, linear scales, rulers, and liquid thermometers. We tested it across 10 different types of instruments.
- Enhanced spatial VQA: Improves Visual Question Answering throughGemini’s advancements in multi-modal understanding.
Gemini Robotics ER 2 consistently achieves the highest accuracy across all core capabilities, with highlights including success detection (image/video), Question Answering (ERQA), and generalized instrument reading.
Advancing safety for embodied intelligence
Gemini Robotics ER 2 is our safest model, achieving significant gains on Safety Instruction Following and Human Proximity benchmarks, which evaluate how a model adheres to physical constraints during reasoning tasks and spatial awareness for detecting humans. We found that Gemini Robotics ER 2 successfully halts a humanoid robot when a person is nearby and autonomously resumes work only once the area is clear. To advance safety for physical agents, we’re introducing a benchmark that evaluates a foundation model's ability to act as a safe VLA orchestrator by testing its capacity to enforce safety constraints, monitor the environment, assess physical feasibility, and seek human clarification. For details, see our safety technical report.
Gemini Robotics ER 2 outperforms ER 1.6 and other frontier models on Safety Instruction Following and Human Proximity benchmarks.
Looking ahead, our plans are to push these models towards even more complex tasks to accelerate the development of helpful robots and support the robotics community.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み