メタ、ピッツバーグ大学と協働し支援ロボティクスを革新
メタは同社の AI モデルを活用してピッツバーグ大学の支援ロボティクス研究を再構築する取り組みを発表した。
AI深層分析を開く2026年7月28日 00:10
AI深層分析
キーポイント
エッジコンピューティングによる自律的対応
カメラ画像やセンサーデータをデバイス上で直接処理することで、予測不能な環境下でも即座に対応可能なロボティクスプラットフォームを実現する。
自然言語と画像データに基づく直感的操作
ユーザーが周囲の文脈を自然言語で指示し、ロボットが画像データを解析することで、複雑なインターフェース設計から解放された直接的なコマンド入力が可能になる。
DINOv3 を活用した軽量な視覚脳の実装
DINOv3 をコンパクトで効率的な「視覚脳」として基盤に据え、タスク固有の軽量モジュールを積み重ねることで、省電力かつ高速な動作を実現する。
実環境向けモデル最適化とプロトタイプ開発
RAMMP チームは既存のプロトタイプに機能を統合し、メモリフットプリントの削減や低精度処理などを実装して、バッテリー駆動機器での信頼性を高めている。
重要な引用
Processing camera images and sensor data directly on the device, known as edge computing, empowers robotic mobility platforms to function as responsive tools.
The ability to use natural language combined with image data to query the user's and robot's environment and provide more direct commands directly reduces the cognitive load
DINOv3 serves as a compact, efficient 'visual brain' for devices — a general-purpose foundation on which task-specific, lightweight modules can be layered
編集コメントを表示
編集コメント
ピッツバーグ大学との連携により、基礎研究段階の技術が実社会での課題解決に直結する形で見えてきた。特にバッテリー駆動という厳しい制約下で DINOv3 を「視覚脳」として機能させるアプローチは、今後の自律型ロボットの普及に向けた重要な指針となるだろう。
支援機器に頼る人々にとって、1 秒の遅れが命取りになることもあります。子供が突然歩道を横断したり、段差が現れたり、鍵を落としたりする予測不能な日常環境では、即座の反応が求められます。カメラ画像やセンサーデータをデバイス上で直接処理する「エッジコンピューティング」を活用すれば、ロボット移動プラットフォームは迅速に対応できるツールとして機能します。こうしたツールは、人々が暮らす多様な動的環境において、常に堅牢で安定して動作する必要があります。
一方で、DINOv3 や SAM といった高性能な AI モデルを、限られたバッテリー容量のハードウェアに展開するには、大きな技術的課題が伴います。実際の導入においては、バッテリー持続時間や放熱対策、不安定なネットワーク接続、そして厳格なサイズと重量の要件など、現実的な要素をすべて考慮する必要があります。
しかし、これらの制約を克服しても、新しいロボットシステムを使って日常生活の動作を実行できないなら、エンドユーザーにとって意味はありません。最新の手法では、ユーザーは周囲の環境を文脈として活用し、より自然な方法でロボットと対話できるようになります。これにより、エンジニアもユーザーも、複雑で時間のかかるインターフェースの設計や操作から解放されます。
自然言語と画像データを組み合わせてユーザーやロボットの環境を検索し、より直接的なコマンドを提供できる能力は、テーブルの上のカップを掴むような単純な作業を行う際に必要な認知負荷や文脈の切り替え回数を大幅に減らします。
この機能はすでに RAMMP チームによって最初のプロトタイプへの統合が進められています。DINO をベースに構築されたツールを活用し、ロボットの画像センサーから自動ドアボタン、カップ、歩道縁や地面を検出してナビゲーション支援を行う仕組みです。この機能が実世界でのテスト準備を整えたことで、エンジニアらは音声とタッチ入力に注力しています。ユーザーは周囲の特定のオブジェクトを選択して対話できるようになります。
既存の課題であるモデル出力の精度と時間的整合性の確保に加え、これによりユーザーからのプロンプトや入力に対する堅牢性と予測可能性を担保する新たな課題も生じています。
DINOv3 は、デバイス向けのコンパクトで効率的な「視覚脳」として機能します。これは汎用的な基盤モデルであり、その上にオブジェクト検出や移動追跡といった特定のタスクに特化した軽量モジュールを積み重ねることで、視覚データの再利用を可能にし、電力消費を抑えます。
ロボットシステムの開発において両方のモデルを活用する際、エンジニアはエッジデバイス向けにモデルを最適化します。具体的にはメモリ使用量を削減し、必要に応じて精度を下げて処理し、実世界の条件に合わせてフォーマットを調整して展開します。これにより、ユーザーにとって信頼性の高いリアルタイム動作が実現されます。実際の解像度で動作し、効率的なバッチ処理を行うことで、両モデルはロボットモビリティプラットフォームやアームに搭載されるコンパクトなバッテリー駆動ハードウェア上でも高速かつ安定した動作を維持します。時には、移動中のユーザーが必要とする速度と安定性を確保するために、境界の精度や特徴の詳細さを少し犠牲にすることもあります。この「精度」と「実用性」のバランスこそが、本プロジェクトの哲学の中核です。
「アシストロボットにおいて性能は、ベンチマークでの正確さだけで測られるものではありません。重要なのは、予測不可能な日常の現場でシステムが確実に動作できるかどうかです」と、RAMMP プロジェクトのスタッフチーフであるシヴァシャンカル・シワクンタン氏は語ります。「DINOv3 や SAM といったモデルをオンデバイスで実行することが、接続に依存せず安全性も損なわずに、ユーザーが信頼できるリアルタイム知覚を実現する鍵となります。
RAMMP の知覚システムは、DINOv2 の埋め込みベクトルでファインチューニングされた軽量検出モデル「RF-DETR」を基盤としています。学習データには SAM を活用した自動ラベル付けを採用しており、支援機器が現実世界で遭遇するあらゆる角度や高さ、背景、照明条件に対応した高品質な注釈を迅速に生成できます。さらにデータ拡張とマルチビュー戦略を組み合わせることで、視点によるばらつきを抑え、一貫性を確保しています。
その結果、スマートかつ適応性に優れたシステムが実現し、より安全で確かな移動が可能になりました。
SAM の強力なラベル付け機能と DINOv2 が持つ豊かな視覚表現を、ファインチューニングされた RF-DETR モデルに統合したことで、RAMMP はリアルタイムの 360 度環境認識と適応型オブジェクト検出を実現しています。
原文を表示
For people relying on assistive devices, every second counts. The unpredictable nature of everyday environments, such as a child darting across a sidewalk, the sudden appearance of a curb, or a dropped set of keys, requires an immediate reaction. Processing camera images and sensor data directly on the device, known as edge computing, empowers robotic mobility platforms to function as responsive tools. These tools must be robust and consistent across the wide range of dynamic environments in which people live.
At the same time, deploying powerful AI models like DINOv3 and SAM on limited, battery-powered hardware presents significant engineering challenges. Real-world deployments must consider practical factors, including battery life, heat dissipation, unreliable network connectivity, and strict size and weight requirements.
However, overcoming these constraints mean nothing to the end user if they can't perform their activities of daily living using these new robotic systems. These newer methods allow users to interact with the robot more naturally, using their immediate surroundings as context, freeing both engineers and users from having to design and navigate complex, time-consuming interfaces. The ability to use natural language combined with image data to query the user’s and robot’s environment and provide more direct commands directly reduces the cognitive load and amount of context switching required for a user to do something as simple as picking up a cup off a table.
This functionality is already being integrated by the RAMMP team into their first prototype. Leveraging tools built off of DINO to enable querying the robot’s image sensors to detect automatic door buttons, cups, and curbs/ground for navigation assistance. With this functionality now ready for real-world testing, engineers are focusing on voice and touch input, letting users select and interact with specific objects in their surroundings. In addition to the existing challenges of ensuring accuracy and temporal coherence of the model outputs, this provides the additional challenge of ensuring robustness and predictability across user prompts and inputs.
DINOv3 serves as a compact, efficient 'visual brain' for devices — a general-purpose foundation on which task-specific, lightweight modules can be layered for actions such as object detection or movement tracking, enabling reuse of visual data and conserving power.
Applying both models as part of the development of robotics systems, engineers optimize models for edge devices, reducing memory footprint, using lower precision when appropriate, and deploying in formats tailored for real-world conditions. This ensures reliable, real-time operation for users. By running at practical resolutions and with efficient batching, both models stay fast and dependable, even on the compact, battery-powered hardware used in robotic mobility platforms and robotic arms, — sometimes trading a little bit of boundary precision and/or feature detail for the speed and stability needed by users on the go. This balance between precision and practicality is central to the project's philosophy.
“For assistive robotics, performance is not measured by benchmark accuracy alone, but by whether a system can operate reliably in the unpredictability of everyday life," said Sivashankar Sivakanthan, Chief of Staff to the RAMMP project. “Running models like DINOv3 and SAM on-device is what enables real-time perception that users can trust - without relying on connectivity or compromising safety.”
RAMMP's perception system is built on RF-DETR, a lightweight detection model fine-tuned with DINOv2 embeddings. Training data is auto-labeled using SAM, enabling the team to rapidly generate high-quality annotations across the full range of angles, heights, backgrounds, and lighting situations that assistive devices encounter in the real world. Data augmentations and multi-view strategies further enforce consistency across perspectives. The result is a system that is smart, adaptive, and offers safer and more confident mobility.
By combining SAM's labeling power with DINOv2's rich visual representations in a fine-tuned RF-DETR model, RAMMP achieves real-time 360-degree environmental awareness and adaptive object detection.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み