MolmoSpaces:具身AIのためのオープンエコシステム
MolmoSpacesは、具身AI(ロボットなど物理的知能)の開発と共有を支援するオープンエコシステム。開発者がモデルや環境を容易に利用・連携できるプラットフォームを提供し、具身AI分野の標準化と普及を促進する。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
2026年2月11日
Ai2
AI の次の波は物理世界において行動するものとなりますが、この可能性を実現するには、新しい空間にわたって一般化できるロボットが必要です。分野がより一般的なロボットシステムを追求するにつれ、ロボットの環境も歩調を合わせる必要があります。つまり、厳密に制御された設定ではなく、多くのユニークなシナリオにわたってトレーニングおよびテストデータを生成することです。過去の Ai2-THOR ファミリーにおける取り組みを通じて、オープンなシミュレーションがロボットの*ナビゲーション*(navigation)の研究をいかに加速できるかを見てきました。また、次世代のロボット*マニピュレーション*(manipulation)をサポートするためのアセットとインフラストラクチャの構築を進めてきました。
本日、私たちは身体性学習(embodied learning)のための大規模かつ完全なオープンプラットフォームであるMolmoSpacesをローンチします。MolmoSpaces は、物理法則に基づくナビゲーションとマニピュレーションをサポートし、Objaverse および THOR からキュレーションされた 230,000 以上の屋内シーンと 130,000 以上のオブジェクトモデルを、4,200 万を超える注釈付きロボット把持(grasps)データと共に単一のエコシステムに統合しています。これには、大規模なアセット生成と体系的な評価をサポートするために設計された、シーン変換、把持の統合、ベンチマークのためのツールキットが含まれています。さらに、USD 変換スクリプトを通じて、MuJoCo、ManiSkill、NVIDIA Isaac Lab/Sim といった一般的なシミュレータとの互換性も提供します。
高忠実度物理を基盤として
以前のシミュレーション環境は、Unity の簡略化された物理学と「魔法のような把持」に依存していました。これは、把持器の周囲にある球体内に物体が入れば把持されたものとするものであり、現実的な接触力をモデル化していませんでした。一方、MolmoSpaces は、MuJoCo などの物理エンジンを使用し、慎重に検証された物理パラメータを採用しています。
剛体に対しては、シミュレーション値を LLM(大規模言語モデル)が注釈付けした推定値と比較することで質量と密度を検証し、必要に応じて密度を調整します。関節付き物体については、シミュレーションされた Franka FR3 を用いて、ジョイント特性や可動部品の密度を調整するテレオペレーションスイートを使用します。FR3 自体は、既知の重量を持つ立方体の押す・掴む動作の実データからシステム同定を行うことでチューニングされています。
また、安定した接触豊富なシミュレーションのために、コライダーとメッシュの準備を手動で注釈付けしています。コライダーメッシュは CoACD を用いて生成され、すべてのオブジェクトアセットに対してプリミティブコライダーを注釈付けします。テーブルやタンスなど receptacle(収納部)の多い剛体には、メッシュ同士の接触問題を回避するために主にプリミティブを使用し、操作可能な物体には高忠実度を実現するために凸分解(convex decomposition)を採用し、小型または薄い物体には再びプリミティブを使用します。
MolmoSpaces-Bench for evaluating generalization
MolmoSpaces には、体系的かつ制御された変異下での一般化に焦点を当てた、一般化ポリシーの評価のためのベンチマークであるMolmoSpaces-Benchが含まれています。このフレームワークにより、研究者は単一の集約成功率を報告するのではなく、物体の特性(形状、サイズ、重量、可動部)、レイアウト(多部屋、多階層、散乱)、タスクの複雑さ(単一ステップから階層的まで)、感覚的条件(照明、視点)、ダイナミクス(摩擦、質量)、タスクの意味(指示の表現)など、複数の軸にわたるパフォーマンスを測定することが可能になります。この体系的な変異は分布分析を可能にし、分布外における失敗モードの特定にも役立ちます。
タスク定義には、原子操作スキル(把持、配置、開閉)とその組み合わせが含まれ、明確にナビゲーション目標も含まれています。アセットと環境は複数のシミュレータバックエンドでインスタンス化可能であり、共通基盤上での比較を可能にします。
これは、ロボット工学研究分野が長年欠いていたものを提供します:数千の現実的なシーンにおいて、他の要因を固定したまま一度に一つの因子を体系的に変化させる能力です。物体質量に対する把持の頑健性を調べたいですか?照明や雑多な環境に対するポリシーの扱い方を研究したいですか?プロンプトの脆弱性や物体頻度のバイアスを明らかにしたいですか?MolmoSpaces-Bench は、これらの制御実験を可能にするために設計されており、体系的な実世界検証を通じてトレーニングの多様性がシミュレーションから現実への転送にどのように影響するかを測定します。
大規模な資産とシーン
MolmoSpaces は、カスタムオブジェクト資産と厳選された Objaverse 由来のバンクを統合し、MJCF で提供され、MuJoCo、ManiSkill、NVIDIA Isaac Lab/Sim を含むシミュレータ間での移植性を確保するために USD に変換されています。
THOR アセットから、134 のカテゴリにわたる 1,600 以上の剛体把持可能オブジェクトインスタンスを抽出・変換しました。また、関節付きの家庭用オブジェクト(冷蔵庫、電子レンジ、オーブン、食器洗い機、ドア、タンスなど)もライブラリに追加し、ジョイントタイプ(ヒンジまたはスライド)、軸、位置、範囲を注釈付けました。これにより、シミュレータ固有の回避策ではなく、明示的に関節動作が定義されます。
Objaverse のアセットについては、当社のパイプラインは 625,000 個のアセットから開始し、メタデータの完全性、単一オブジェクトの検証、スケール正規化、テクスチャ品質(スコア≧4)、クロスレンダラー忠実度(CLIP 類似度≧0.6)、ジオメトリ効率(<1.5 MB)、および受容体検証のためのフィルタリングを適用します。その結果、約 3,000 の WordNet 語義群にわたる 129,000 個の厳選されたアセットが得られ、これらはトレーニング/バリデーション/テストのサブセットに分割されています。手動生成による家屋の作成においては、追加の配置フィルタにより、自動シーン生成に適した約 92,000 個のアセットが抽出されます。
これらのアセットは、iTHOR-120、ProcTHOR-10K、ProcTHOR-Objaverse、Holodeck など複数のデータセットから抽出されたシーンに配置され、家屋、オフィス、教室、病院、学校、博物館などを含む数十万の屋内環境を網羅しています。シーンの作成モードには、手作業による環境構築、手動で再現されたデジタルツイン、ヒューリスティックな手続き型生成、LLM 支援の手続き型生成が含まれており、厳選された設定と非常に多様な設定の両方での評価をサポートします。
シーンコレクションは広範な検証プロセスを経ます。剛体操作対象物に対しては、小さな外力を加え、2 cm を超えて移動しないオブジェクトは固定されているものとして扱います。関節付きオブジェクトに対しては、関節力を加え、関節範囲の少なくとも 60% を通過できないアセットは却下します。また、環境をシミュレーションしてドリフトや交差を検出し、95% 以上のシーンがこれらのテストに合格しています。占有マップにより、ロボットのための衝突のない初期姿勢が特定されます。
スケーラブルな開発のための把持
MolmoSpaces には、48,111 種類の物体(物体あたり最大約 1,000 個)にわたる 4200 万を超える 6 自由度の把持姿勢が含まれています。把持姿勢は、MJCF 幾何形状から Robotiq-2F85 グリッパーモデルを用いて直接サンプリングされます。関節付き物体の場合、サンプリングは葉部コンポーネント(しばしば取っ手)に制限され、非葉部幾何形状と衝突する把持姿勢は破棄されます。
私たちは、多様性と頑健性の両方を備えた把持姿勢を選択します。これらは完全な 6 自由度のポーズ空間でクラスタリングされ、クラスタ間から均一に選択されます。また、接触点の好ましさ(薄物における指パッド中央 vs. 指先)も考慮されます。剛体に対する把持は、線形および回転摂動を用いてテストされます。関節付き物体については、動作可能性を通じて頑健性を評価します。具体的には、接触を維持しながら有効なジョイント範囲の少なくとも 70% を両方向で安定して動作させることを要求します。
把持可能性の検証は、ランダムにサンプリングされたポーズに配置された浮遊する Robotiq グリッパーを用いて行われ、物体を持ち上げて開閉するテストを実施します。
これらの把持姿勢は、把持ローダーを通じて環境に直接埋め込むことができ、付随する軌道生成パイプラインにより、把持データに基づく再現可能なデモンストレーションが可能になります。これにより、大規模なデータセット作成や模倣学習のサポートが実現されます。
What's included
MolmoSpaces のすべてのコンポーネントはオープンかつモジュール化されています。研究者は基盤となる MJCF を検証・修正したり、把持動作を再生成したり、新しいロボットやコントローラーを組み込んだり、単腕および二腕システムを含む複数のシミュレーターと実装体間で実験の複製や拡張を行ったりできます。
今回のリリースには以下の要素が含まれます:
- アセット: 130K ドル相当以上の MJCF オブジェクトアセットに加え、メッシュ・マテリアル、物理パラメータ、および説明、スケール、質量、シンセット/カテゴリからなる豊富なメタデータ
- シーンデータセット: 手作業で作成された単一室のシーンから、手続き的に生成された 10 部屋以上の商業・住宅空間に至るまで、完全な物理演算対応の環境定義とメタデータが 230K+ 件
- 把持動作: 剛体および可動関節を持つ 48,000 個以上のオブジェクトにわたる 6DoF(6 自由度)把持注釈が 42M 件以上
- ツール: アセットを各シミュレーターで利用するためのローダー/ユーティリティ(Isaac Lab/Sim via USD、ManiSkill ローダーを含む)
MolmoSpaces はまた、Teledex などのモバイルプラットフォームを用いたテレオペレーションベースのデータ収集もサポートしており、研究者はスマートフォンから直接デモンストレーションを収集できます。このインターフェースは DROID や CAP を含む既存の実装体セットアップすべてと互換性があり、特別な設定は不要です。
MolmoSpaces の探索と、その結果得られるデータを用いた一般化ポリシーの訓練をご案内します。コミュニティがどのように活用するかを楽しみにしています。
詳細はこちら、始め方はこちら:
Tech Report | Data | Code | Demo
最新の Ai2 ニュースに関する月次更新を受け取るには、購読してください。
原文を表示
February 11, 2026
Ai2
The next wave of AI will act in the physical world, but realizing this potential requires robots that generalize across new spaces. As the field pursues more general robotic systems, robot environments need to keep pace—generating training and testing data across many unique scenarios rather than tightly controlled settings. Through our past work on the Ai2-THOR family, we've seen how open simulation can accelerate research in robotic *navigation*, and we've been building the assets and infrastructure to support the next generation of robotic *manipulation*.
Today, we're launching MolmoSpaces, a large-scale, fully open platform for studying embodied learning. MolmoSpaces supports physics-grounded navigation and manipulation, unifying over 230,000 indoor scenes and more than 130,000 object models – curated from Objaverse and THOR – with over 42 million annotated robotic grasps in a single ecosystem. It includes tooling for scene conversion, grasp integration, and benchmarking, all designed to support large-scale asset generation and systematic evaluation. As a bonus, our assets are compatible with common simulators including MuJoCo, ManiSkill, and NVIDIA Isaac Lab/Sim through a USD conversion script.
High-fidelity physics as a foundation
Our earlier simulation environments relied on Unity's simplified physics and "magic grasps," where an object is considered grasped once it enters a sphere around the gripper—without modeling realistic contact forces. MolmoSpaces instead uses physics engines (such as MuJoCo) with carefully validated physical parameters.
For rigid objects, we verify mass and density by comparing simulated values to LLM-annotated estimates and adjusting density as needed. For articulated objects, we use a teleoperation suite to tune joint properties and movable-part densities using a simulated Franka FR3; the FR3 itself is tuned via system identification from real cube-pushing and picking trajectories of known weights.
We also manually annotate collider and mesh preparation for stable contact-rich simulation. Collider meshes are generated with CoACD, and we annotate primitive colliders for all object assets. Receptacle-heavy rigid objects (tables, dressers) primarily use primitives to avoid mesh-mesh contact issues; manipulable objects use convex decomposition for higher fidelity, with primitives again for small or thin objects.
MolmoSpaces-Bench for evaluating generalization
MolmoSpaces includes MolmoSpaces-Bench, a benchmark for evaluating generalist policies with a focus on generalization under systematic, controlled variation. The framework enables researchers to measure performance across multiple axes – object properties (shape, size, weight, articulation), layouts (multi-room, multi-floor, clutter), task complexity (single-step to hierarchical), sensory conditions (lighting, viewpoints), dynamics (friction, mass), and task semantics (instruction phrasing) – rather than reporting a single aggregate success rate. This systematic variation allows for distributional analysis and helps identify out-of-distribution failure modes.
Task definitions include atomic manipulation skills (pick, place, open, close) and compositions, and explicitly include navigation objectives. Assets and environments can be instantiated across multiple simulator backends, enabling comparisons on shared foundations.
This gives robotics researchers something the field has long lacked: the ability to systematically vary one factor at a time while holding others fixed across thousands of realistic scenes. Want to probe grasps robustness to object mass? Study how policies handle lighting or clutter? Expose prompt fragility or object-frequency biases? MolmoSpaces-Bench is designed to enable these controlled experiments while also measuring how training diversity affects sim-to-real transfer through systematic real-world validation.
Assets and scenes at scale
MolmoSpaces unifies custom object assets with a curated Objaverse-derived bank, all provided in MJCF and converted to USD for portability across simulators including MuJoCo, ManiSkill, and NVIDIA Isaac Lab/Sim.
From THOR assets, we extracted and converted 1,600+ rigid, graspable object instances across 134 categories. We also expanded the library with articulated household objects – fridges, microwaves, ovens, dishwashers, doors, dressers, and more – annotating joint type (hinge or slide), axis, position, and range so articulated behaviors are defined explicitly rather than through simulator-specific workarounds.
For Objaverse assets, our pipeline starts from 625,000 assets and applies filtering for metadata completeness, single-object validation, scale normalization, texture quality (score ≥ 4), cross-renderer fidelity (CLIP similarity ≥ 0.6), geometry efficiency (< 1.5 MB), and receptacle validation. The result is 129,000 curated assets spanning approximately 3,000 WordNet synsets, split into train/val/test subsets. For procedural house generation, additional placement filters yield ~92,000 assets suitable for automatic scene population.
These assets populate scenes drawn from multiple datasets – iTHOR-120, ProcTHOR-10K, ProcTHOR-Objaverse, and Holodeck – spanning hundreds of thousands of indoor environments across homes, offices, classrooms, hospitals, schools, museums, and more. Scene creation modes include hand-crafted environments, manually reproduced digital twins, heuristic procedural generation, and LLM-assisted procedural generation, supporting evaluation across both curated and highly diverse settings.
The scene collection undergoes extensive validation. For rigid manipulables, we apply small external forces; objects that don't move beyond 2 cm are treated as stuck. For articulated objects, we apply joint forces and reject assets that fail to move through at least 60% of their joint range. We also simulate environments to detect drifting and intersections—more than 95% of scenes pass these tests. Occupancy maps identify collision-free starting poses for robots.
Grasps for scalable development
MolmoSpaces includes over 42 million 6-DoF grasp poses across 48,111 objects (up to ~1,000 per object). Grasps are sampled directly from MJCF geometry using a Robotiq-2F85 gripper model; for articulated objects, sampling is restricted to leaf components (often handles), and grasps that collide with non-leaf geometry are discarded.
We select grasps to be both diverse and robust: they're clustered in full 6-DoF pose space and selected uniformly across clusters, with contact-point preferences (mid-fingerpad vs. fingertip for thin objects). Rigid grasps are tested with linear and rotational perturbations. For articulated objects, we evaluate robustness via actuation feasibility—requiring stable actuation through at least 70% of the valid joint range in both directions while maintaining contact.
We verify graspability using a floating Robotiq gripper placed at randomly sampled poses to lift and open/close objects.
These grasps can be embedded directly into environments via a grasp loader, and an accompanying trajectory-generation pipeline enables reproducible demonstrations conditioned on grasp data—supporting dataset creation and imitation learning at scale.
What's included
Every component of MolmoSpaces is open and modular. Researchers can inspect and modify the underlying MJCF, regenerate grasps, plug in new robots and controllers, and replicate or extend experiments across multiple simulators and embodiments—including single-arm and dual-arm systems.
This release includes:
- Assets: 130K+ USD/MJCF object assets plus meshes/materials, physics parameters, and rich metadata consisting of descriptions, scale, mass, and synsets/categories
- Scene datasets: 230K+ environment definitions and metadata for fully physics-enabled scenes ranging from handcrafted single-room scenes to procedurally-generated 10+ room commercial and residential spaces
- Grasps: over 42M 6-DoF grasp annotations across 48,000+ objects (rigid + articulated)
- Tools: loaders/utilities for using assets across simulators (including Isaac Lab/Sim via USD; ManiSkill loader)
MolmoSpaces also supports teleoperation-based data collection using mobile platforms like Teledex, letting researchers gather demonstrations directly from their phone. The interface is compatible with all of our existing embodiment setups, including DROID and CAP, with no special configuration required.
We invite you to explore MolmoSpaces and train generalist policies on the resulting data. We're eager to see how the community uses it.
Learn more and get started:
Tech Report | Data | Code | Demo
Subscribe to receive monthly updates about the latest Ai2 news.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み