NVIDIA、手術用ロボティクス向けリアルタイム生成シミュレーション「Cosmos-H-Dreams」を発表
NVIDIA は手術用ロボティクス分野において、リアルタイムで動作する生成型シミュレーション技術「Cosmos-H-Dreams」の導入を推進すると発表した。
AI深層分析を開く2026年7月27日 20:02
AI深層分析
キーポイント
NVIDIA の新プラットフォーム発表
NVIDIA は Hugging Face Blog を通じて、手術用ロボティクス分野に特化したリアルタイム生成シミュレーション「Cosmos-H-Dreams」の導入を発表した。
リアルタイム生成技術の実装
同プラットフォームは、医療現場で必要とされる複雑な手術手順をリアルタイムで生成・シミュレートする機能を備えているとされている。
医療ロボティクスへの応用
この技術は、外科医やロボットシステムが安全かつ効率的に手術を練習・実行するための訓練環境を提供することを目的としている。
リアルタイム手術シミュレーションの実現
教師モデルから学生モデルへの蒸留と自己強制学習により、160fps のリアルタイム動作を可能にした。これにより、物理ロボットを実行せずに政策の評価や合成データ生成が行えるようになった。
FlashDreams による推論最適化
ストリーミング KVキャッシュやCUDA Graphキャプチャなどの技術を用いたFlashDreamsエンジンが、RTX PRO 6000上で高速な推論を実現している。
重要な引用
Bringing Real-Time Generative Simulation to Surgical Robotics
"Cosmos-H-Dreams brings action-conditioned surgical world modeling into the real-time loop, creating a new environment for people and policies to practice, explore, generate data, and evaluate what happens next."
"The immediate next step is to evaluate more than visual quality. A useful surgical simulator must respond correctly to actions, preserve instrument and scene structure over long rollouts, and support conclusions that transfer to the physical robot."
編集コメントを表示
編集コメント
手術用ロボティクスにおけるリアルタイム生成シミュレーションの導入は、医療現場の訓練コスト削減と安全性向上に直結する重要な一歩である。NVIDIA が Hugging Face を通じてこの技術を公開したことは、オープンな技術共有が医療 AI の発展を加速させる好例と言える。
記事一覧に戻る
- 手術用世界モデルからインタラクティブ・シミュレーターへ
- リアルタイム化のための Cosmos-H-Surgical-Simulator の知識蒸留
2.1 手術用の教師モデル
2.2 因果関係によるウォームアップ
2.3 自己強制型知識蒸留
- FlashDreams: リアルタイム推論エンジン
- あなただけのデータへの適応
- 次のステップ:クローズドループ手術用物理 AI へ
- 今日から始めよう
手術用ロボティクスは、遠隔操作から、より高度なビジョン・言語・行動ポリシーへと急速に進化しています。しかし、これらのシステムの評価や訓練には依然として大きな課題があります。物理的なロボットプラットフォームの運用コストは高く、実験の再現に時間がかかり、失敗すれば器具や生体組織を損傷するリスクさえあります。従来のシミュレーターはより安全な代替手段を提供しますが、手術現場のモデル化は極めて困難です。変形する組織、細かな器具同士の相互作用、鏡面反射、縫合糸や針、スモーク(煙)、そして視界の遮断など、あらゆる要素が重要になります。
世界規模の基盤モデルは、異なるアプローチを提供します。すべての物体や物理的な相互作用を手動で記述するのではなく、同期された動画とロボットの運動学データから視覚的ダイナミクスを直接学習します。
NVIDIA の Cosmos-H-Surgical-Simulator はこのアプローチを実証しました。初期の手術シーンと一連のロボット動作を入力として、将来の手術動画を生成するモデルです。これにより、物理的な実験よりも高速な評価が可能になり、Open-H-Embodiment エコシステム全体で合成データの生成を実現しました。
さて、その次のステップとして、手術用ロボティクス向けのリアルタイム・アクション条件付き生成シミュレーター「Cosmos-H-Dreams」をご紹介します。Cosmos-H-Dreams は、Cosmos-H-Surgical-Simulator の機能を因果関係を持つ少数ステップの学生モデルに凝縮し、NVIDIA の高速ストリーミング推論ライブラリである FlashDreams を通じて提供します。単一の NVIDIA RTX PRO 6000 GPU で動作するこのシステムは、人間や学習されたポリシーがクローズドループで制御できるインタラクティブな環境を実現しています。
1. 手術用世界モデルからインタラクティブシミュレーターへ
「Cosmos-H-Surgical-Simulator」は、NVIDIA の Cosmos-Predict2.5-2B を基盤とし、Open-H-Embodiment データセットで追加学習された、アクション条件付きの世界基盤モデルです。手術中の映像フレームと将来のロボット軌道を入力として受け取ると、その行動がもたらす可能性のある視覚的な結果を示した動画を生成します。
これにより、オフラインでのポリシー評価や合成データの生成が可能になります。記録されたデータやポリシーから生成された軌道をモデルに送るだけで、物理的なロボットで繰り返し動作を実行することなく、対応するロールアウト(実行結果)を生成し、その結果を検査したりスコアリングしたりできます。
「Cosmos-H-Dreams」は、このモデルをリアルタイム処理に対応させたものです。Cosmos-H-Surgical-Simulator が学習した多様なロボット形態に関する手術の事前知識を出発点に、dVRK(da Vinci Research Kit)を用いた卓上縫合作業に特化させ、因果関係を持つ学生モデルとして蒸留しました。この学生モデルは、シーン全体を自己回帰的に生成します。
公開されたモデルは、初期の RGB フレームとロボットの運動学データをリアルタイムストリームとして受け取り、次のフレーム群を生成した後に、次の行動ブロックへと処理を進めます。
Cosmos-H-Dreams の汎用性を示すため、CMR Surgical と Cambridge Consultants との共同開発により、Versius 手術支援ロボットのコントローラーとの統合を実現しました。これにより、Versius プラットフォーム上でリアルタイムでの動作が可能となっています。
2. リアルタイム対応に向けた Cosmos-H-Surgical-Simulator の最適化
最大の課題は、有用な外科的なダイナミクスを維持しつつ、生成コストをいかに下げるかです。Cosmos-H-Dreams では、長時間の自己回帰的ロールアウトに適した「教師から生徒へ」の学習パイプラインを採用しています。
2.1. 外科手術に特化した教師モデル
双方向型の教師モデルは、統一された 44 次元の動作表現を用いる Cosmos-H-Surgical-Simulator の Open-H チェックポイントを出発点としています。公開されている dVRK テーブルトップモデルでは、相対的なエンドエフェクターの変位・回転およびグリッパーの状態からなる両腕の動作データが、この共通表現へとマッピングされます。
その後、教師モデルは JHU の dVRK テーブルトップ用データセットで微調整が行われます。これには成功したデモンストレーションだけでなく、針の落下や投球ミス、結び目の失敗といった「失敗事例」や「分布外のエピソード」も含まれます。これらの失敗事例は極めて重要です。ポリシーを評価するためのシミュレーターであれば、理想通りのデモンストレーションだけでなく、不適切な動作が引き起こす結果も再現できなければならないからです。
長期ロールアウト時の安定性を高めるため、トレーニングプロセスでは教師モデルの時間的視野を段階的に拡大します。まずは 12 フレームの視野から開始し、これを 72 フレームまで徐々に引き上げていきます。各ステップで視野を広げる際は、事前学習済みの重みを持つウォームアップ済みモデルを初期値として利用します。
2.2. 因果的ウォームアップ
教師モデルによるノイズ除去の軌道は事前に計算され、キャッシュされます。このキャッシュされた軌道を模倣するように、因果的なアテンション(注意機構)を持つ学生モデルが教師から初期化され、トレーニングされます。このウォームアップ段階を通じて、学生モデルは自身の生成履歴から学習を始める前に、まず因果的なアテンションとストリーミング型のキー・バリューキャッシュを用いて動作する方法を習得します。
2.3. セルフフォーシング蒸留
自己回帰モデルにはよくある問題があります。トレーニング時にはクリーンで正解のコンテキストを見せる一方、展開(デプロイ)時には自身の不完全な出力に条件付けなければならない点です。そのため、小さな誤差が時間とともに蓄積・増幅されてしまう恐れがあります。
Cosmos-H-Dreams は セルフフォーシング蒸留 によってこのミスマッチを解消します。トレーニング中は学生モデル自身が生成したコンテキストを用いてロールアウトを進め、凍結された教師モデルからの分布整合性に基づく監督信号が、その自己生成のロールアウトを実際の手術動画に近づけるよう導きます。これにより、インタラクティブな推論時に遭遇する条件に学生モデルを事前に準備させることができます。
生成されたモデルは、従来の教師モデルが用いる多段階のプロセスではなく、少数ステップの拡散(few-step diffusion)をサポートします。具体的には、潜在フレームごとに最大2回のノイズ除去ステップで動作し、ストリーミングに必要な因果構造を維持しながら、教師モデルが持つ外科手術に関する事前知識も統合しています。
3. FlashDreams: リアルタイム推論エンジン
モデルの蒸留はリアルタイム化の一部に過ぎません。Cosmos-H-Dreams は、自己回帰型の世界モデルや動画モデル向けの高速推論ライブラリ「FlashDreams」を通じて提供されています。
FlashDreams は、ストリーミング KV キャッシュの利用、CUDA Graph のキャプチャ、モデルのコンパイルなど、複数の最適化技術を組み合わせることで、蒸留された学生モデルを低遅延のストリーミングシステムへと変貌させます。
これらの技術により、標準的な Cosmos-H-Surgical-Simulator 推論で約 10 フレーム/秒だった処理速度が、単一の NVIDIA RTX PRO 6000 グラフィックボード上で対話的な操作が可能になる約 160 フレーム/秒へと劇的に向上しました。
Cosmos-H-Dreams はまた、生成を相互作用に変えるための人間と機械のインターフェースも提供します。ブラウザクライアントではキーボード入力を送信し、WebRTC を介して生成されたフレームを受信できます。Meta Quest クライアントでは、追跡されたコントローラーの動きをロボットの動作にマッピングし、合成シーンを WebXR で表示可能です。さらに、学習済みの外科手術ポリシーと接続することで、生成された観測値と予測される行動を閉ループ内で交換することもできます。
4. あなただけのデータへの適応
Cosmos-H-Dreams には卓上縫合用の事前学習済みチェックポイントが含まれていますが、このシステムは特定のロボット環境への拡張を前提に設計されています。ご自身のデータセット向けにリアルタイムの学生モデルを訓練したい場合は、教師モデルのファインチューニングと自己強制蒸留のための完全な手順を、ステップバイステップガイドで提供しています。
5. 次の一手:クローズドループ型外科用物理 AI へ
Cosmos-H-Dreams は、実機のロボットデータから学習した環境を構築し、その環境が十分に反応性を持って「住める」ものとなることで、外科シミュレーションの新たな地平を開きました。
次の段階では、視覚的な品質だけでなく、より多角的な評価が必要です。有用な外科用シミュレータは、動作に対して正しく反応すること、長時間のロールアウトを通じて器具やシーンの構造を維持すること、そして物理ロボットへ転送可能な結論を導き出すことが求められます。これらを踏まえ、ツールチップの到達距離と姿勢精度、グリッパーサイクルの忠実度、アイドル時の安定性、反事実的動作の多様性、長期ホライズンにおけるドリフト、シミュレーションと実機でのポリシー結果の整合性といった、新たなクローズドループ型ベンチマークのファミリーを提案します。
リアルタイムなワールドモデルは、外科用ポリシー開発において能動的なパートナーとしても機能し得ます。必要に応じて稀な失敗事例を生成したり、強化学習や模倣学習のためのスケーラブルな環境を提供したり、貴重なロボットハードウェアに依存せず、新しいポリシーの迅速な評価を可能にします。
さらに先には、リアルタイムシミュレーションが、遅延を考慮した遠隔手術(latency-aware telesurgery)において世界モデルがより安定した表示を維持するのを助けるなど、一連の downstream アプリケーションを実現します。また、対話型の手術リハーサル、手技計画、術中意思決定支援といった用途にも応用可能です。
Cosmos-H-Dreams は診断システムや術中画像の代替、あるいは物理的な手術ロボットのコントローラーではなく、研究開発プラットフォームです。しかしながら、これらの可能性を安全に探求するための基盤を提供するものです。
モデルの忠実度、時間的安定性、ハードウェア効率性がさらに向上すれば、リアルタイム生成シミュレーションは、外科医教育、合成データ生成、ポリシー訓練、そして評価を一つの共有された物理 AI 環境でつなぐ役割を果たすようになるでしょう。
6. 今日から始めよう
Cosmos-H-Dreams の背後にあるモデル、データ、ランタイムを探索しましょう:
- Cosmos-H-Dreams のコードとサンプル:GitHub リポジトリ
- Cosmos-H-Dreams モデル:Hugging Face チェックポイント
- Cosmos-H-Surgical-Simulator:Hugging Face モデルおよび GitHub リポジトリ
- 独自のデータへの適応:教師による微調整や自己強制蒸留のためのステップバイステップレシピ
- Open-H-Embodiment:Hugging Face データセット
- FlashDreams:GitHub リポジトリ
- NVIDIA Cosmos-Predict2.5:GitHub リポジトリ
- Cosmos-Surg-dVRK:世界基盤モデルに基づく手術ロボットポリシー学習の自動化オンライン評価(arXiv 論文)
Cosmos-H-Dreams は、行動条件付きの手術用世界モデルをリアルタイムループに組み込むことで、人々やポリシーが練習・探索・データ生成・次の展開の評価を行うための新たな環境を実現します。
原文を表示
- 1. From Surgical World Model to Interactive Simulator
- 2. Distilling Cosmos-H-Surgical-Simulator for Real Time 2.1. A Surgical Teacher
- 2.2. Causal Warmup
- 2.3. Self-Forcing Distillation
- 3. FlashDreams: The Real-Time Inference Engine
- 4. Adapting to Your Own Data
- 5. What Is Next: Toward Closed-Loop Surgical Physical AI
- 6. Get Started Today
Surgical robotics is moving quickly from teleoperation toward increasingly capable vision-language-action policies. But evaluating and training these systems remains difficult. Physical robotic platforms are expensive to operate, experiments are slow to reproduce, and failures can damage instruments or biological material. Conventional simulators provide a safer alternative, but surgical scenes are exceptionally difficult to model: deformable tissue, fine instrument interactions, specular surfaces, sutures, needles, smoke, and occlusions all matter.
World foundation models offer a different path. Instead of manually authoring every object and physical interaction, they learn visual dynamics directly from synchronized video and robot kinematics. NVIDIA's Cosmos-H-Surgical-Simulator demonstrated this approach by generating future surgical video from an initial scene and a sequence of robot actions. It enabled faster-than-physical evaluation and synthetic data generation across the Open-H-Embodiment ecosystem.
Today, we are introducing the next step: Cosmos-H-Dreams, a real-time, action-conditioned generative simulator for surgical robotics. Cosmos-H-Dreams distills the capabilities of Cosmos-H-Surgical-Simulator into a causal, few-step student model and serves it through FlashDreams, NVIDIA's accelerated streaming-inference library. Running on a single NVIDIA RTX PRO 6000 GPU, the result is an interactive environment that a person or a learned policy can control in a closed loop.
1. From Surgical World Model to Interactive Simulator
Cosmos-H-Surgical-Simulator is an action-conditioned world foundation model built on NVIDIA Cosmos-Predict2.5-2B and post-trained on the Open-H-Embodiment dataset. Given a surgical context frame and a future robot trajectory, it generates video showing the likely visual consequences of those actions. This makes it useful for offline policy evaluation and synthetic data generation. A recorded or policy-generated trajectory can be sent to the model, the corresponding rollout can be generated, and the result can be inspected or scored without repeatedly executing the motion on a physical robot.
Cosmos-H-Dreams moves the model into the real-time regime. Starting from the multi-embodiment surgical priors learned by Cosmos-H-Surgical-Simulator, we specialize the model for da Vinci Research Kit (dVRK) tabletop suturing and distill it into a causal student that generates the scene autoregressively. The released model receives an initial RGB frame and a live stream of robot kinematics, then produces the next chunk of frames before continuing with the following action block.
We have also demonstrated the versatility of Cosmos-H-Dreams by collaborating with CMR Surgical and Cambridge Consultants to integrate it with the Versius surgeon controller, enabling real-time operation on the Versius platform.
2. Distilling Cosmos-H-Surgical-Simulator for Real Time
The key challenge is preserving useful surgical dynamics while reducing the cost of generation. Cosmos-H-Dreams uses a teacher-to-student training pipeline designed for long, autoregressive rollouts.
2.1. A Surgical Teacher
The bidirectional teacher begins from the Cosmos-H-Surgical-Simulator Open-H checkpoint, which uses a unified 44-dimensional action representation. For the released dVRK tabletop model, the dual-arm dVRK action content, consisting of relative end-effector translation, rotation, and gripper state, is mapped into this common representation.
The teacher is then fine-tuned on the JHU dVRK tabletop mixture, including successful demonstrations as well as failure and out-of-distribution episodes such as needle drops, missed throws, and unsuccessful knot ties. These failures are important: a simulator intended to evaluate policies must reproduce the consequences of poor actions, not only ideal demonstrations.
To improve stability during long rollouts, the training process progressively increases the teacher's temporal horizon. We start training on a 12-frame horizon and progressively increase this number until 72 frames. At each horizon bump, we initialize the warmed-up model with pretrained weights.
2.2. Causal Warmup
The teacher's denoising trajectories are first precomputed and cached. A causal student is initialized from the teacher and trained to imitate these cached trajectories. This warmup stage teaches the student to operate with causal attention and a streaming key/value cache before it begins learning from its own generated history.
2.3. Self-Forcing Distillation
Autoregressive models face a familiar problem: during training they may see clean, ground-truth context, while during deployment they must condition on their own imperfect outputs. Small errors can therefore compound over time.
Cosmos-H-Dreams addresses this mismatch with self-forcing distillation. During training, the student rolls forward using its own generated context. Distribution-matching supervision from the frozen teacher then guides those self-generated rollouts toward realistic surgical video. This prepares the student for the same conditions it will encounter during interactive inference.
The resulting model supports few-step diffusion, with as few as two denoising steps per latent frame, rather than the many-step process used by the full teacher. It combines the teacher's surgical priors with the causal structure needed for streaming.
3. FlashDreams: The Real-Time Inference Engine
Model distillation is only part of the real-time story. Cosmos-H-Dreams is served through FlashDreams, an accelerated inference library for autoregressive world and video models.
FlashDreams turns the distilled student into a low-latency streaming system through several complementary optimizations, such as streaming KV cache, CUDA Graph capturing, or model compilation.
Together, these techniques bring the distilled surgical world model fine-tune from the roughly ten-frames-per-second regime of standard Cosmos-H-Surgical-Simulator inference to interactive operation (~160 frames per second) on a single NVIDIA RTX PRO 6000.
Cosmos-H-Dreams also provides the human-machine interfaces that turn generation into interaction. A browser client can send keyboard commands and receive generated frames over WebRTC. A Meta Quest client can map tracked controller motion into robot actions and display the synthesized scene through WebXR. The same model can also be connected to a learned surgical policy, with generated observations and predicted actions exchanged inside a closed loop.
4. Adapting to Your Own Data
While Cosmos-H-Dreams includes a pre-trained checkpoint for tabletop suturing, the system is designed to be extensible to your specific embodiment. To train a real-time student model for your own dataset, we provide a complete recipe for teacher fine-tuning and self-forcing distillation in our step-by-step guide.
5. What Is Next: Toward Closed-Loop Surgical Physical AI
Cosmos-H-Dreams opens a new frontier for surgical simulation: environments learned from real robot data that are responsive enough to be inhabited.
The immediate next step is to evaluate more than visual quality. A useful surgical simulator must respond correctly to actions, preserve instrument and scene structure over long rollouts, and support conclusions that transfer to the physical robot. This motivates a new family of closed-loop benchmarks: tool-tip reach and pose accuracy, gripper-cycle fidelity, idle stability, counterfactual action diversity, long-horizon drift, and agreement between simulated and real policy outcomes.
Real-time world models can also become active partners in surgical policy development. They can generate rare failures on demand, provide scalable environments for reinforcement or imitation learning, and enable rapid evaluation of new policies without tying every experiment to scarce robotic hardware.
Further ahead, real-time simulation enables a series of downstream applications such as latency-aware telesurgery, where the world model helps maintain a more stable display; or interactive surgical rehearsal, procedure planning, and intraoperative decision support. Cosmos-H-Dreams is a research and development platform, not a diagnostic system, a replacement for intraoperative imaging, or a controller for a physical surgical robot. Yet it provides a foundation for exploring these possibilities safely.
As model fidelity, temporal stability, and hardware efficiency continue to improve, real-time generative simulation can help connect surgeon education, synthetic data generation, policy training, and policy evaluation within one shared Physical AI environment.
6. Get Started Today
Explore the models, data, and runtime behind Cosmos-H-Dreams:
- Cosmos-H-Dreams code and examples: GitHub repository
- Cosmos-H-Dreams model: Hugging Face checkpoint
- Cosmos-H-Surgical-Simulator: Hugging Face model and GitHub repository
- Adapt to your own data: Step-by-step recipe for teacher fine-tuning and self-forcing distillation
- Open-H-Embodiment: Hugging Face dataset
- FlashDreams: GitHub repository
- NVIDIA Cosmos-Predict2.5: GitHub repository
- Cosmos-Surg-dVRK: World Foundation Model-based Automated Online Evaluation of Surgical Robot Policy Learning (arXiv paper)
Cosmos-H-Dreams brings action-conditioned surgical world modeling into the real-time loop, creating a new environment for people and policies to practice, explore, generate data, and evaluate what happens next.
AI算出
主要ニュースainew評価高い
AI が中心テーマである新技術の発表であり、比較対象となる直近の記事が存在しないため新規性が高い。ただし日本企業や日本固有の情報が含まれていないため、国内関連性は低めとする。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み