NVIDIA、AI エージェント向け軽量モデルとルーティングライブラリを公開
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
同モデルは 30B の混合専門家(MoE)構造を持ちつつアクティブパラメータが 3B に抑えられ、Mamba-2 とアテンションを融合したアーキテクチャを採用して 1M トークンのコンテキストウィンドウを実現している。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 16:20
AI深層分析
キーポイント
Nemotron 3.5 Lightning の技術的特徴
同モデルは 30B の混合専門家(MoE)構造を持ちつつアクティブパラメータが 3B に抑えられ、Mamba-2 とアテンションを融合したアーキテクチャを採用して 1M トークンのコンテキストウィンドウを実現している。
NeMo Switchyard の機能と目的
このオープンソースのルーティングライブラリは、エージェントワークフローの各ステップを最も適したモデルに振り分けることで、コストとレイテンシの削減を図る。
性能比較と実用性
NVIDIA によると、同規模モデルより最大 4 倍高速で出力し、Qwen3.6 35B と同等の精度で PinchBench のタスク完了が 30% 高速化される。
デプロイ環境とライセンス
モデルは OpenMDW-1.1 ライセンスの下で商用利用が可能であり、単一の DGX Spark や H100 GPU で動作するほか、クラウドやオンプレミスでの展開も想定されている。
高速推論と低コスト化の実現
多トークン予測と専用ドラフトモデル(DSpark/DFlash)、NVFP4量子化により、同規模モデル比で最大4倍の出力速度を達成した。
重要な引用
Nemotron 3.5 Lightning is a lightweight, customizable open model built for high-volume agentic tasks
NeMo Switchyard is an open source routing library that directs each step of an agent workflow to the most capable and efficient model available
NVIDIA reports up to 4x faster output speed than similar-sized models
NeMo Switchyard is an open source library that routes each step of an agent workflow to the most capable and efficient model available.
編集コメントを表示
編集コメント
エージェントアプリケーションの実用化において、コストと速度の課題を解決する「実行層」特化モデルの登場は大きな進展である。特に単一 GPU での動作が可能となる点は、小規模チームやスタートアップが本格的な AI エージェント開発に参入する機会を広げるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
NVIDIA は、専門モデルのシステムから常時稼働する AI エージェントを構築するためのオープン技術を公開しました。発表されたのは 2 つの成果物です。
1 つ目は「Nemotron 3.5 Lightning」で、高ボリュームのエージェントタスク向けに設計された軽量かつカスタマイズ可能なオープンモデルです。もう 1 つは「NeMo Switchyard」というオープンソースのルーティングライブラリで、エージェントワークフローの各ステップを、利用可能な最も能力が高く効率的なモデルへと誘導します。
両者が解決しようとしているのは構造的な課題です。長時間稼働するエージェントは、その時間の大半をツール呼び出し、結果の検証、サブエージェントへの委任に費やしています。これらのすべてのステップを最先端の推論モデルへ送信すると、コストとレイテンシが膨らんでしまいます。
Nemotron 3.5 Lightning は、30B の混合専門家(MoE)モデルであり、アクティブパラメータは 3B です。ハイブリッドな Mamba-2 + MoE + Attention アーキテクチャを採用し、1M トークンのコンテキストウィンドウを備えています。NVIDIA によると、同規模の他モデルと比較して出力速度が最大 4 倍高速化されており、Qwen3.6 35B と同等の精度で 10,000 の PinchBench タスクを完了するまでの時間が 30% 短縮されています。
すでに多くの業界プレイヤーがこのモデルのカスタマイズを進めています。CrowdStrike、Harvey、CodeRabbit、Fastino Labs、Lila Sciences などが、それぞれサイバーセキュリティ、法務、コーディング、金融、ヘルスケアのワークロード向けに活用しています。
導入は可能でしょうか?
はい、可能です。Nemotron 3.5 Lightning は、オープンな重み、トレーニングデータ、レシピを公開しており、寛容な OpenMDW-1.1 ライセンスの下で一般提供されています。NVIDIA は、このモデルが商用利用の準備ができていると明言しています。
利用可能な企業規模は、単一の現代 GPU を持つすべての組織です。NVIDIA によると、1 枚の DGX Spark (GB10) または 1 枚の H100 でシングル GPU デプロイメントが可能です。これにより、個人開発者やシード期のスタートアップが、大企業と同等の土俵で戦えるようになります。中堅市場のチームは Baseten、Together AI、Nebius を経由して提供可能です。規制の厳しい大企業であれば、完全にオンプレミス環境で運用することもできます。
対象業界は、サイバーセキュリティ、法務サービス、ソフトウェアエンジニアリング、金融サービス、ヘルスケア、ライフサイエンスなど、NVIDIA が名指しする顧客リストに含まれる分野です。
活用アプリケーションには、ツール呼び出し、結果の検証、サブエージェントへの委任、コードレビューのルーティング、ログの選別、契約書の解析、そして 100 万トークンという広大なコンテキストウィンドウを跨ぐ長文脈検索などが挙げられます。
実行層に焦点を当てる
長時間稼働するエージェントは、その時間の大半を高ボリュームな「実行」に費やしています。ツール呼び出し、結果の検証、サブエージェントへの委任がトークン予算の大部分を占めます。これらの処理ステップすべてを最先端の推論モデルへルーティングすると、コストとレイテンシが大幅に増大してしまいます。
Nemotron 3.5 Lightning は、まさにこの「実行層」を対象としています。これは 30B の混合専門家 (MoE) モデルで、アクティブパラメータ数は 3B です。ハイブリッドな Mamba-2 + MoE + Attention アーキテクチャを基盤に構築されています。コンテキスト長は 1M トークンに達し、NVFP4 レシピを用いて 20 兆トークン以上で事前学習が行われました。
このモデルは Nemotron 3 ファミリーの中で最も小型のメンバーです。Nemotron 3 Ultra などの最先端モデルがオーケストレーションと計画を担当する一方、Lightning はその下位にある日常的な呼び出し処理を担います。
高速化の仕組み
2 つの主要メカニズムがあります:
まず、推測的デコーディングについてです。マルチトークン予測は専用事前学習段階で組み込まれ、その後 MTP 強化フェーズによってさらに改善されています。NVIDIA は 2 つの外部ドラフトモデルも提供しています。1 つ目は DGX Spark や低同時実行数のデータセンターワークロードに推奨される半自己回帰型の DSpark です。もう 1 つは軽量なブロック拡散モデルを採用した DFlash です。
次に量子化についてです。NVFP4 チェックポイントは BF16 とともに提供され、Blackwell および Hopper ではネイティブに対応し、Ampere へも W4A16 カーネルを通じて拡張されます。
NVIDIA によると、同規模のモデルと比較して出力速度が最大 4 倍向上します。PinchBench のベンチマークでは、Qwen3.6 35B と同等の精度で 10,000 タスクを完了する際、処理速度が 30% 高速化され、精度は 86% を記録しました。
公開されたモデルカードの結果(BF16 / NVFP4)は以下の通りです。MMLU Pro は 81.94 / 81.62、GPQA Diamond は 75.44 / 75.57、SWE-bench Verified は 51.56 / 52.80、Terminal-Bench 2.1 は 24.58 / 23.46、AA-LCR は 52.00 / 49.19 です。推奨されるサンプリングパラメータは温度 1.0、top_p 0.95 です。
NeMo Switchyard
NeMo Switchyard は、エージェントワークフローの各ステップを、利用可能な最も能力が高く効率的なモデルへルーティングするオープンソースライブラリです。
セッションアフィニティを持つ LLM クラスファイアー、直近のツール使用状況を読み取るステージ型ルーター、難易度が持続すれば低コストから高機能へ昇格するエスカレーション型ルーターなど、チューニング不要のルーターを提供します。また、モデルの残差ストリームから学習し、どの候補が成功するかを予測する調整可能なプレフィルルーターも用意されています。参照サーバーは OpenAI、Anthropic、および Responses API からのリクエストに対応しています。
2 つの公開された結果があります。LangChain は 145 のマルチターンエージェントタスクでベンチマークを実施しました。エスカレーション型ルーターを用いて Lightning と Claude Opus 4.8 をルーティングしたところ、最上位モデルのみを使用するベースラインと比較してコストを 74% 削減し、呼び出しの 7% を最上位モデルに振り分けました。ただし、精度は約 6 ポイント低下しています。
Cognition は Devin Desktop で段階的ルーティングを実装しました。FrontierCode Main では Opus 5 と Kimi K2.7 の間でルーティングを行い、平均コスト$3.11 で 50.6% の達成率を記録しました。これは Opus 5 の精度と 2.8 ポイント差でありながら、平均コストは約 28% 低く抑えられています。
インタラクティブな解説
主要ポイント
- 30B のオープン MoE でアクティブパラメータは 3B。コンテキスト長は 1M。ライセンスは OpenMDW-1.1 で商用利用も可能。
- 出力速度が最大 4 倍向上。PinchBench の 10,000 タスクを Qwen3.6 35B より 30% 高速で完了。
- この高速化は、マルチトークン予測に加え、DSpark と DFlash ドラフター、および NVFP4 チェックポイントによるもの。
- DGX Spark 1 基または H100 1 基で動作。Ollama、LM Studio、llama.cpp、Unsloth を介したローカル環境での実行も可能。
NeMo Switchyard は LangChain の 145 タスクベンチマークでコストを 74% 削減しました(精度は約 6 ポイント低下)。
build.nvidia.com または OpenRouter で試すことができます。また、Hugging Face や ModelScope から重み(ウェイト)をダウンロードすることも可能です。
※本記事は MarkTechPost に掲載された「NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router」を基にしています。
原文を表示
NVIDIA introduced open technologies for building always-on AI agents from systems of specialized models. Two artifacts shipped together. Nemotron 3.5 Lightning is a lightweight, customizable open model built for high-volume agentic tasks, and NeMo Switchyard is an open source routing library that directs each step of an agent workflow to the most capable and efficient model available. The problem both address is structural: long-running agents spend most of their time on tool calls, result validation, and subagent delegation, and sending every one of those steps to a frontier reasoning model adds cost and latency. Lightning is a 30B mixture-of-experts model with 3B active parameters, built on a hybrid Mamba-2 + MoE + Attention architecture with a 1M-token context window. NVIDIA reports up to 4x faster output speed than similar-sized models, and 30% faster completion of 10,000 PinchBench tasks than Qwen3.6 35B at comparable accuracy. Many industry players like CrowdStrike, Harvey, CodeRabbit, Fastino Labs, and Lila Sciences are already customizing it for cybersecurity, legal, coding, finance, and healthcare workloads.
Is it deployable?
Yes. Nemotron 3.5 Lightning is generally available under the permissive OpenMDW-1.1 license, with open weights, training data, and recipes. NVIDIA states the model is ready for commercial use.
Which companies: Anyone with a single modern GPU. NVIDIA lists single-GPU deployment on 1x DGX Spark (GB10) or 1x H100. That puts solo developers and seed-stage startups on the same footing as enterprises. Mid-market teams can serve it from Baseten, Together AI, or Nebius; regulated enterprises can keep it fully on-premises.
Industries: Cybersecurity, legal services, software engineering, financial services, healthcare, and life sciences all appear in NVIDIA’s named customer set.
Applications: Tool calling, result validation, subagent delegation, code review routing, log triage, contract parsing, and long-context retrieval across a 1M-token window.
The execution layer, not the planning layer
Long-running agents spend most of their time on high-volume execution. Tool calls, result validation, and subagent delegation dominate the token budget. Routing every one of those steps to a frontier reasoning model adds cost and latency.
Nemotron 3.5 Lightning targets that execution layer. It is a 30B mixture-of-experts model with 3B active parameters, built on a hybrid Mamba-2 + MoE + Attention architecture. Context length reaches 1M tokens. Pre-training covered more than 20 trillion tokens using an NVFP4 recipe.
The model is the smallest member of the Nemotron 3 family. Frontier models such as Nemotron 3 Ultra handle orchestration and planning, while Lightning handles the routine calls beneath them.
Where the speed comes from
Two mechanisms:
First, Speculative Decoding: Multi-token prediction was baked in during a dedicated pre-training stage, then improved with an MTP-boosting phase. NVIDIA also ships two external draft models: DSpark, a semi-autoregressive drafter recommended for DGX Spark and low-concurrency data center workloads, and DFlash, which uses a lightweight block-diffusion model.
Second, Quantization: An NVFP4 checkpoint ships alongside BF16. The same checkpoint serves Blackwell and Hopper natively, and extends to Ampere through W4A16 kernels.
NVIDIA reports up to 4x output speed versus similar-sized models. On PinchBench, it reports 86% accuracy while completing 10,000 tasks 30% faster than Qwen3.6 35B at comparable accuracy.
Published model card results (BF16 / NVFP4): MMLU Pro 81.94 / 81.62, GPQA Diamond 75.44 / 75.57, SWE-bench Verified 51.56 / 52.80, Terminal-Bench 2.1 24.58 / 23.46, AA-LCR 52.00 / 49.19. Recommended sampling is temperature 1.0 and top_p 0.95.
NeMo Switchyard
NeMo Switchyard is an open source library that routes each step of an agent workflow to the most capable and efficient model available.
It offers tuning-free routers, including an LLM classifier with session affinity, a stage router that reads recent tool activity, and an escalation router that starts cheap and promotes on sustained difficulty. A tunable prefill router learns from the model’s residual stream to predict which candidate will succeed. The reference server accepts OpenAI, Anthropic, and Responses API requests.
Two published results: LangChain benchmarked 145 multi-turn agentic tasks. Routing between Lightning and Claude Opus 4.8 with the escalation router cut cost 74% versus a frontier-only baseline, sending 7% of calls to the frontier model, at a roughly 6-point accuracy tradeoff. Cognition implemented staged routing in Devin Desktop. On FrontierCode Main, routing between Opus 5 and Kimi K2.7 reached 50.6% at a $3.11 mean cost, within 2.8 points of Opus 5 accuracy at approximately 28% lower mean cost.
Interactive explainer
Key Takeaways
30B open MoE with 3B active parameters, 1M context, OpenMDW-1.1 license, commercial use permitted.
Up to 4x output speed; PinchBench 10,000 tasks completed 30% faster than Qwen3.6 35B.
Speed comes from multi-token prediction plus DSpark and DFlash drafters, and an NVFP4 checkpoint.
Runs on 1x DGX Spark or 1x H100, and locally via Ollama, LM Studio, llama.cpp, and Unsloth.
NeMo Switchyard cut cost 74% in LangChain’s 145-task benchmark at a ~6-point accuracy tradeoff.
Try it on build.nvidia.com or OpenRouter, and download weights from Hugging Face or ModelScope. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router appeared first on MarkTechPost.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み