NVIDIA、自動運転向け VLA モデル「Alpamayo 2 Super」を公開
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
NVIDIA は自動運転およびロボタクシー向けに、長尾事象への対応を強化した34Bパラメータのオープンビジョン言語行動モデル「Alpamayo 2 Super」を公開し、商用利用可能なライセンスの下で実用化に向けた基盤を提供する。
AI深層分析を開く2026年8月5日 17:45
AI深層分析
キーポイント
商用展開可能なオープンライセンスの採用
同社は「OpenMDW-1.1」というLinux Foundation の許容ライセンスおよびApache 2.0ソースコードの下でモデルを公開し、ファインチューニングや派生モデルの作成、商業的再配布を許可する。
長尾事象への対応とアーキテクチャ
従来の検出・予測スタックが苦手とする稀な多エージェント状況を対象とし、32BパラメータのVLMバックボーンに2.3Bのパラメータを持つ拡散ベースのアクションデコーダーを組み合わせる。
因果説明と安全性の統合
単なる軌道予測に加え、決定の因果関係を示す「Chain-of-Causation」トレースやメタアクションを出力し、NVIDIA Halos の安全検証ワークフローやISO/PAS 8800への整合性をサポートする。
ベンチマークでの圧倒的パフォーマンス
同社のテストではQwen2.5-VL 72BやGPT-4oなどの主要モデルを大幅に上回り、閉ループ評価でも高いスコアを記録している。
モデル構成とアーキテクチャ
34B VLA モデルは、32B の Cosmos 3 Super Reasoner バックボーンに 2.3B の拡散アクション専門家を組み合わせた構造を持つ。
重要な引用
The stated design target is the long-tail events: rare, multi-agent situations that conventional detection-and-prediction stacks handle poorly.
Yes, and for commercial use from day one.
NVIDIA says the model compresses annotation cycles from months to days.
One pass yields trajectory, Chain-of-Causation trace, meta-action, auto-labels, and grounded VQA.
編集コメントを表示
編集コメント
NVIDIA は単に高性能なモデルを公開するだけでなく、OpenMDWライセンスを通じてエコシステム全体の商用活用を加速させる戦略を示した。特に安全性検証のための因果説明機能(CoC)の標準化は、実社会への導入における信頼性確保の鍵となる要素である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
NVIDIA は、自動運転向けに設計された 34B パラメータのビジョン・言語・アクション(VLA)モデル「Alpamayo 2 Super」を、オープン商用ライセンスの下で公開しました。このモデルが特に狙っているのは、従来の検出・予測スタックでは対応が難しい長尾事象、つまり稀な多車体状況です。
同モデルは、NVIDIA Cosmos 3 Super Reasoner をベースに強化学習でポストトレーニングされた 32B の VLM バックボーンと、拡散モデルに基づく 2.3B のアクションデコーダーを組み合わせます。全周囲カメラからの動画データを一度処理するだけで、計画された軌道、その軌道に対する因果的な説明、そしてメタアクションを出力します。
実用化は可能でしょうか?
はい、可能です。発売初日から商用利用が可能です。重み付けデータは Linux Foundation が提供するオープンモデル配布用の寛容なライセンス「OpenMDW-1.1」の下で公開されており、ソースコードは Apache 2.0 です。このライセンスには、ファインチューニングや派生モデルの作成、そして商業的な再配布も含まれています。
NVIDIA は OpenMDW ライセンスを Alpamayo シリーズ全体に適用しており、以前は研究開発目的で公開されていたリリース版も、追加の許可なしに現在では商用展開が可能になりました。
入力・出力とトレーニングデータについて
入力は、マルチカメラからの RGB 動画、テキスト、タイムスタンプ付きの自己運動(egomotion)履歴です。検証済みのパブリックノートブックのプロファイルでは、6 つのカメラからそれぞれ過去 4 フレーム分のデータを取得します。自己運動データは、3D 変位と 3×3 の回転行列を複数時間ステップで含むものです。
軌道 API は、0.1 秒間隔で 0.1 秒から 6.4 秒までの範囲をカバーする 64 個のウェイポイント(航点)を返します。各ウェイポイントには、自己車両フレームにおける XYZ 座標と 3×3 の回転行列が含まれています。
学習データには、視点運動(egomotion)と経路の注釈が付いたマルチカメラによる運転動画が約 115,000 時間含まれています。これには、運転判断の構造化された因果関係に基づく説明である「Chain-of-Causation (CoC) トレーサ」が約 370 万件含まれており、画像学習データは 10 億枚を超えます。
ベンチマーク結果
LingoQA では、Alpamayo 2 Super は Lingo-Judge スコアで 79.2 を記録し、評価されたほぼ 40 のモデルの中で首位となりました。NVIDIA のテストでは、Qwen2.5-VL 72B よりも 17.0 ポイント、Gemini 2.5 Pro よりも 15.1 ポイント、GPT-4o よりも 23.2 ポイント上回りました。
計画作業において特に重要となる数値がもう二つあります。PhysicalAI-AV-NuRec データセットの 910 シナリオで AlpaSim を用いたクローズドループ評価では、AlpaSim スコアは 1.50 ± 0.13 でした。また、PhysicalAI-AV データセットの 937 の困難なサンプルに対するオープンループ評価では、6.4 秒時点での minADE₆ は 0.911m となりました。
一つのモデルから得られる五つの出力
各運転状況に対して、このモデルは経路、その判断を説明する CoC トレーサ、譲歩や車線変更などのメタアクション、自動ラベル付けされた推論結果、そして 2D グランディングを伴う視覚的質問応答の五つを生成します。
この組み合わせが、運用上の価値を生み出しています。開発者は、モデルが観測した内容と選択した行動を結びつけることができます。CoC トレーサは NVIDIA Halos の安全性検証ワークフローに統合され、ISO/PAS 8800 に準拠した AI セーフティをサポートします。
独自 fleet データに対する自動ラベル付けツールとして使用した場合、NVIDIA によると注釈作成のサイクルを数ヶ月から数日に短縮できるとのことです。
インタラクティブな解説機能
Key Takeaways
34B の VLA モデルは、32B の Cosmos 3 Super Reasoner をバックボーンとし、2.3B の拡散型アクション専門モジュールを備えています。
OpenMDW-1.1 の重みと Apache 2.0 ライセンスのコードが公開されており、商用利用や再配布も追加許可なしで可能です。
LingoQA Lingo-Judge では 79.2 を記録し、約 40 モデル中トップとなりました。AlpaSim では 1.50 ± 0.13 のスコアを達成し、6.4 秒の推論時間で minADE₆ は 0.911m を記録しています。
ワンパスで軌道予測、因果連鎖(Chain-of-Causation)の追跡、メタアクション、自動ラベル付け、そして grounded VQA(視覚言語質問応答)を同時に生成します。
このクラウドスケールのモデルは、H100 80GB を 1 枚使用してテストされ、ピーク時のメモリ使用量は 72,115 MiB です。車載推論用には、これを蒸留(distill)して軽量化することが推奨されます。
原文を表示
NVIDIA has released Alpamayo 2 Super, a 34B-parameter vision-language-action (VLA) model for autonomous driving, under an open commercial license. The stated design target is the long-tail events: rare, multi-agent situations that conventional detection-and-prediction stacks handle poorly. The model pairs a 32B VLM backbone, built on NVIDIA Cosmos 3 Super Reasoner and post-trained with reinforcement learning, with a 2.3B diffusion-based action decoder. From one pass over full-surround camera video it emits a planned trajectory, a causal explanation of that trajectory, and a meta-action.
Is it deployable
Yes, and for commercial use from day one. The weights are released under OpenMDW-1.1, the Linux Foundation’s permissive license for open model distributions; source code is Apache 2.0. The license covers fine-tuning, derivative models and commercial redistribution. NVIDIA is applying OpenMDW across the entire Alpamayo family, so earlier releases introduced for R&D are now deployable commercially without additional permission.
Inputs, outputs and training data
Inputs are multi-camera RGB video, text, and egomotion history with timestamps. The validated public notebook profiles use six cameras and four historical frames per camera. Egomotion is 3D translation plus a 3×3 rotation matrix, multi-timestep.
The trajectory API returns 64 waypoints spanning 0.1 to 6.4 seconds at 0.1-second intervals. Each waypoint carries ego-frame XYZ and a 3×3 rotation matrix.
Training data is roughly 115,000 hours of multi-camera driving video with egomotion and trajectory annotations. It includes about 3,700,000 Chain-of-Causation (CoC) traces — structured, causally linked explanations of driving decisions. Image training data exceeds one billion images.
Benchmarks
On LingoQA, Alpamayo 2 Super records a Lingo-Judge score of 79.2 and ranks first among nearly 40 models evaluated. In NVIDIA’s testing it beat Qwen2.5-VL 72B by 17.0 points, Gemini 2.5 Pro by 15.1, and GPT-4o by 23.2.
Two more numbers matter for planning work. Closed-loop evaluation with AlpaSim on 910 scenarios from the PhysicalAI-AV-NuRec dataset gives an AlpaSim score of 1.50 ± 0.13. Open-loop evaluation on 937 challenging samples from the PhysicalAI-AV dataset gives minADE₆ at 6.4s of 0.911m.
Five outputs from one model
For each driving situation, the model produces a trajectory, a CoC trace explaining the decision, a meta-action such as yield or lane change, reasoning auto-labels, and visual question answering with 2D grounding.
That combination is what makes the release interesting operationally. Developers can tie what the model observed to the action it chose. CoC traces integrate with NVIDIA Halos safety-validation workflows and support AI safety aligned with ISO/PAS 8800.
Used as an autolabeler on proprietary fleet data, NVIDIA says the model compresses annotation cycles from months to days.
Interactive explainer
Key Takeaways
34B VLA model — 32B Cosmos 3 Super Reasoner backbone plus a 2.3B diffusion action expert.
OpenMDW-1.1 weights and Apache 2.0 code; commercial use and redistribution allowed, no extra permission needed.
LingoQA Lingo-Judge 79.2, first among nearly 40 models; AlpaSim 1.50 ± 0.13; minADE₆ 0.911m at 6.4s.
One pass yields trajectory, Chain-of-Causation trace, meta-action, auto-labels, and grounded VQA.
Cloud-scale model tested on 1× H100 80GB at 72,115 MiB peak; distill it for in-car inference.
Check out the NVIDIA blog and Hugging Face model card. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post NVIDIA Releases Alpamayo 2 Super: A 34B Open Vision-Language-Action Model for Robotaxis and Autonomous Driving Under OpenMDW-1.1 appeared first on MarkTechPost.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み