Nvidia、AI モデルをタスク中に切り替えるルーター「Switchyard」を発表
本文の状態
日本語全文を表示中
詳細モードで約11分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
VentureBeat AI
Nvidia はコスト削減と速度向上を両立させるため、30B パラメータの専用モデル「Nemotron 3.5 Lightning」とワークフロー段階ごとに最適なモデルを選別するライブラリ「NeMo Switchyard」を同時に公開した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 22:07
AI深層分析
キーポイント
コストと速度の両立アプローチ
Nvidia は単一のモデルやルーターだけでなく、専用モデルとルーティングライブラリの組み合わせにより、エッジケースを含む複雑なタスクでもコストを約3分の1に抑えつつ、Qwen3.6-35B よりも約30%高速な処理を実現すると発表した。
Nemotron 3.5 Lightning の性能
同社によると、この30B パラメータのオープンモデルは比較対象クラスで最大4倍の出力速度を達成し、同等の精度を維持しながら高速なエージェントタスクを処理できるという。
競合環境と市場の文脈
Alibaba や Meta などが相次いで高性能なオープンモデルをリリースする中、Nvidia は単独のモデル競争ではなく、モデル層とルーティング層の両方を自社で制御するシステムとしての優位性を強調している。
Switchyard の独自性と競合
Switchyard は既存のプロバイダーに接続可能であり、Not Diamond や RouteLLM といった他社のルーターとは異なり、モデルとルーティングの両方を一つのオープンライセンスで提供する点に強みを持つ。
動的なルーティング戦略によるコスト削減
エージェントの状態変化に応じて最適なモデルを選択する動的ルーティングにより、タスクあたりのコストを大幅に削減できる。
重要な引用
"That is the power of a system of models, matching the right model to each step of the workflow," Kari Briski, vice president of generative AI at Nvidia, said in a briefing.
Nvidia says the combination holds frontier-level task completion while cutting benchmark costs to roughly a third of running Opus 4.8 alone.
"It has many types of routing strategies," Briski said. "You can have a random router, which is not that great, or you can have an agent state route or a classifier route. Depending on your routing strategy, it wants to choose the best model."
"We are an ecosystem lover, and we want to make sure that we are integrated," Briski said. "We've partnered with OpenRouter, LiteLLM and Kong, and they've already integrated our routing algorithm, so you can pick it up right where you're already using the best tools."
編集コメントを表示
編集コメント
Nvidia は単なるモデルの性能競争から、システム全体の最適化へと戦略をシフトさせた。この発表は、オープンソースモデルが台頭する中で、独自のエコシステムを構築しようとする同社の意図を如実に示している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
常時稼働する AI エージェントを運用する企業は、いつも同じジレンマに直面しています。すべてのタスクを最先端モデルに送ればコストが急増し、簡単なタスクを安価なモデルへルーティングするための独自ロジックを構築すれば、それは新たなエンジニアリングプロジェクトとなり、ワークフローが変更されるたびに維持管理が必要になります。
Nvidia はこの問題の両端を同時に解決する提案を行いました。同社は火曜日に、高ボリュームで専門的なエージェントタスク向けに設計された 300 億パラメータのオープンな混合専門家モデル「Nemotron 3.5 Lightning」と、エージェントワークフローの各ステップを最も適したモデルへルーティングするオープンソースライブラリ「NeMo Switchyard」を発表しました。
注目すべき数値は以下の通りです。Nvidia によると、Lightning は同クラスの他モデルと比較して最大 4 倍の高速な出力を実現し、同等の精度で Qwen3.6-35B よりも約 30% 速くエージェントタスクを完了します。Switchyard と組み合わせることで、最先端レベルのタスク完遂能力を維持しつつ、Opus 4.8 を単独で実行する場合と比較してベンチマークコストを約 3 分の 1 に削減できると同社は主張しています。
このタイミングは、業界が数ヶ月で経験した中で最も活発なオープンウェイトモデルの発表ラッシュの真っ只中にあります。今年春以降、アリババ、Moonshot、Zhipu、DeepSeek といった中国の企業が競合するオープンモデルを相次いで公開しており、その中には米国の大手研究機関が持つ最先端性能に匹敵するものも含まれています。これらはサイズや価格面で米国勢を下回るコストで提供されています。また、Meta が 300 億パラメータ規模のオープンエージェントモデル「Muse Glimmer」をリリースしたことで、この競争圧力はさらに高まりました。ここ数ヶ月でオープンウェイトは差別化要因から必須条件へと急速に変化しており、Nvidia の今回の発表はその変化の中心に位置しています。
重要なのは、この組み合わせにあります。モデル単体ではコスト削減の問題は解決できず、ルーターだけでは効率的にルーティングする対象がありません。Nvidia は、オープンソースをモデル層とルーティング層の両方に適用することが、エージェント AI のコスト構造を実際に動かす鍵であると確信しています。安価なモデル一つや、他社のスタックに取り付けられたスマートなルーター単体では不十分なのです。
Switchyard の真の競合は他のオープンモデルではなく、「Not Diamond」(すでに OpenRouter の Auto モードで採用済み)や、カリフォルニア大学バークレー校と LMSYS が開発したオープンソースフレームワーク「RouteLLM」です。これらはいずれも独自のモデルを保有していません。Nvidia の戦略は、意思決定の両側面(モデルとルーター)を一つのオープンライセンスの下で自社が支配することで、ルーター専用やモデル専用の競合には真似できない優位性を築くことにあります。
「これは、ワークフローの各ステップに最適なモデルをマッチングさせるシステムとしての強みです」と、Nvidia のジェネレーティブ AI 担当バイスプレジデントである Kari Briski はブリーフィングで述べています。
ルーターがワークフローをどう変えるか
モデルのルーティング自体は新しいカテゴリではありません。OpenRouter、LiteLLM、そしていくつかのスタンドアロン型ルーティングスタートアップがすでに存在し、開発者が複数のプロバイダー間でトラフィックを振り分けることを可能にしています。Switchyard はこれら既存のサービスを置き換えるのではなく、それらに接続して機能します。
Switchyard が解決する核心的な問題は、エージェントがタスクを進める過程で最適なモデルが変化するという点です。ツールの結果が返ってきたり、エラーが発生したり、あるいはステップが複雑ではなく単純な作業であることが判明したりすると、エージェントの状態は変化します。一度決めた固定のモデルでは、こうした状況の変化に対応できません。
Briski 氏は、静的なタスクカテゴリではなく、変化する状態に応じて反応するルーティング戦略を説明しました。
「Switchyard には多種多様なルーティング戦略があります」と Briski 氏は述べています。「ランダムに選んでしまう単純なルーターも存在しますが、あまり効果的ではありません。代わりに、エージェントの状態に基づいて判断する『エージェント状態ルート』や、分類器を利用した『クラシファイアルート』を採用できます。使用するルーティング戦略によって、最適なモデルの選択基準が変わります。例えば、非常に効率的なタスクには Lightning といったモデルを選ぶべきケースがあり、その場合、ルーターは設定されたモデルプールの中に Lightning があれば、自動的にそれを選択します。」
コストもまた、後付けの考慮事項ではなく、ルーティングの判断に直接組み込まれています。VentureBeat の質問に対し、Briski 氏は Switchyard は「モデルの冗長性(verbosity)」を評価できると説明しました。これは、特定のタスクに対して各モデルが平均して生成するトークン数を指します。この予測値を活用することで、呼び出しを実行する前に、より安価なオプションへと処理を誘導することが可能になります。
この仕組みが独自の統合プロジェクト化しないようになっているのは、Switchyard の位置づけにあります。Nvidia はパートナーを二つのグループに分類しました。一つは Cognition、LangChain、Nous Research といった Switchyard に直接呼び出すエージェントフレームワーク群です。もう一つは Kong、LiteLLM、OpenRouter といった、自社の製品に Switchyard サポートを組み込んだ LLM ゲートウェイ群です。Kong は「Kong AI Gateway」のネイティブ機能として Switchyard を搭載しています。Briski 氏は、既存のルーティングエコシステムにおけるライブラリの位置づけを説明する際にも、このゲートウェイパートナーの一覧を挙げています。
「私たちはエコシステムを愛しており、統合がしっかり行われていることを確認したいのです」と Briski 氏は語ります。「OpenRouter、LiteLLM、Kong と提携し、すでにルーティングアルゴリズムの統合が進んでいます。そのため、既存で最も優れたツールを利用しているその場で、すぐに Switchyard を利用できます」
Nvidia は Switchyard のテスト結果を 9 社から共有しました。いくつかの結果には具体的な数値も含まれています。LangChain によると、145 件のマルチターン Deep Agents タスクにおいて、呼び出しのわずか 7% を最前線モデルにルーティングするだけでコストが 74% 削減され、精度は 6% 低下するという結果でした。Ramp は「Ramp SWE-Bench」で最前線モデルと同等のパフォーマンスを達成しながら、コストを 58%、実行時間を 33% 削減しました。Cognition は Switchyard の段階型ルーターを内部利用向けに Devin Desktop に統合し、すべての呼び出しを単一の最前線モデルにルーティングした場合と比較して平均コストを 28% 削減しながら、FrontierCode Main でほぼ最前線レベルのパフォーマンスを達成したと報告しています。
Lightning のアーキテクチャとパフォーマンス向上
Nemotron 3.5 Lightning は、汎用利用ではなく高負荷の専門的なエージェントタスク向けに設計された、独立したオープンモデルです。
このモデルは、2025 年 12 月に発表された Nemotron 3 シリーズで導入されたハイブリッド Mamba-Transformer と潜在混合専門家(MoE)アーキテクチャを拡張したものです。Nemotron 3 Super の基盤となる同シリーズであり、Lightning は後処理後の比較において自社のベンチマークとして使用されています。Switchyard などのルーティング構成に組み込まれることを想定しており、意思決定の「最前線」ではなく「高速・低コスト」な側を担当するよう設計されていますが、ルーターに依存せず単体でも動作し、出荷可能です。
Artificial Analysis Intelligence Index(9 つの評価項目を網羅した汎用能力ベンチマーク)によると、Lightning のスコアは 24 で、gpt-oss-120b と同点です。一方、Nemotron 3 Super、Gemma 4 31B、Claude 4.5 Haiku、Mistral Medium 3.5 はすべて 30 を記録しており、Lightning よりも上位に位置しています。Lightning はそのサイズクラスにおいて汎用知能のリーダーではありませんし、Nvidia もそれを主張していません。
実際の主張はより限定的です。Nvidia が提供した PinchBench のデータによると、Lightning は Qwen3.6-35B と同等の精度を約 30% 高速で達成します。また、Gemma 4 26B と比較すると、同程度の完了時間においてより高い精度を示しました。PinchBench はコーディング、リサーチ、ファイル管理といった実世界のエージェントタスクを対象としたベンチマークです。これは能力そのものの向上ではなく、速度と精度のトレードオフにおける優位性を示すものです。
Post-training は、Nvidia が最も大きな効果が見られると主張する領域です。同社は、4 つの早期アクセスパートナーからの事前・事後の数値を公開しました。CrowdStrike の悪意あるコンテンツ検出における Nemotron 3 Super ベースラインとの比較、CodeRabbit のコーディングルーターにおける GPT 5.4 Nano ベースラインとの比較、Harvey と Trajectory の法務タスク完了における Opus 4.6 ベースラインとの比較、そして Lila Sciences のエネルギーシミュレーションにおける Opus 4.8 ベースラインとの比較です。特に CodeRabbit の事例が具体的で、Nvidia によると、標準的な NeMo Auto モデルレシピを 1 エポック訓練して組み込むことで、約 2 時間で稼働するルーターエージェントを 85 ドルで構築できるとしています。
これが企業に与える影響
オープンモデル市場は成長しており、競合製品が不足しているわけではありません。今回の新リリース「Nemotron Lightning」も、組織が検討すべき選択肢の一つとなるでしょう。
モデル比較の観点では、Lightning のベンチマークチャートでは Qwen3.6-35B を直接の対照対象として選定しています。VentureBeat から中国製モデル全般との比較について直接問われた際、Briski 氏は明確なベンチマーク結果を示す代わりに、「オープン性とカスタマイズ性」こそが差別化要因であると指摘しました。
「私たちの価値提案は、単にオープンであることだけではありません。非常にカスタマイズしやすい点も大きな強みです」と Briski 氏は述べています。
エージェントインフラを構築する企業にとって、以下の 3 つのトレンドが際立っています:
ルーティングの判断基準が、固定されたものから動的なものへと変化しています。従来は単一のデフォルトモデルを基盤としたエージェントパイプラインを構築していた企業に対し、設計時に固定された割り当てではなく、エージェントの状態やトークンコストといったリアルタイム信号に基づいたステップごとのルーティングへの移行が迫られています。
オープンソースは、もはや単一層のコスト削減要因ではなく、2 つのレイヤーでコスト競争力を高めるレバーとして機能しています。ベンダーがエンドツーエンドを制御するオープンルーターとオープンモデルを組み合わせるというアプローチは、単に重み(ウェイト)が安価であるという従来の議論よりも新しい視点であり、他の研究機関がこのパターンを追従するかどうかも注目すべき点です。
競争の焦点は「最良のモデル」から「最良のシステム」へとシフトしています。ルーティングライブラリが成熟するにつれ、差別化要因は企業がデフォルトとして採用するモデルそのものではなく、本番環境においてタスクとモデルをいかに適切にマッチングさせられるかというルーティング層の性能へと移っています。これはベンチマークも市場への訴求も、従来よりも困難な課題となっています。
原文を表示
Enterprises running always-on AI agents keep hitting the same tradeoff. Send every task to a frontier model and the bill climbs fast. Build custom routing logic to send easy tasks to cheaper models and that becomes its own engineering project, one that has to be maintained every time a workflow changes.
Nvidia is proposing a fix that touches both ends of that problem at once. The company is out on Tuesday with Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model built for high-volume, specialized agent tasks, alongside NeMo Switchyard, an open-source library that routes each step of an agent workflow to whichever model fits it best.
The headline numbers: According to Nvidia, Lightning delivers up to 4x faster output than comparable models in its class, completing agentic tasks roughly 30% faster than Qwen3.6-35B at matching accuracy. Paired through Switchyard, Nvidia says the combination holds frontier-level task completion while cutting benchmark costs to roughly a third of running Opus 4.8 alone.
The timing puts Nvidia in the middle of the busiest open-weight stretch the industry has seen in months. Alibaba, Moonshot, Zhipu and DeepSeek have all shipped competitive open models out of China since the spring, several landing at or near frontier performance while undercutting US labs on size or price. Meta added to that pressure by releasing its own 30-billion-parameter open agentic model, Muse Glimmer. Open weights have gone from a differentiator to table stakes in a matter of months, and Nvidia's release lands squarely inside that shift rather than ahead of it.
The pairing is the point. A model alone doesn't solve the cost problem, and a router alone has nothing efficient to route to. Nvidia is betting that open source, applied at both the model layer and the routing layer, is what actually moves the cost needle on agentic AI, not a single cheaper model and not a smarter router bolted onto someone else's stack.
Switchyard's real rivals aren't other open models — they're Not Diamond, which already powers OpenRouter's Auto mode, and RouteLLM, the open-source framework from UC Berkeley and LMSYS. Neither ships its own model. Nvidia's bet is that owning both sides of the decision, under one open license, is what a router-only or model-only competitor can't match.
"That is the power of a system of models, matching the right model to each step of the workflow," Kari Briski, vice president of generative AI at Nvidia, said in a briefing.
How the router actually changes the workflow
Model routing isn't a new category. OpenRouter, LiteLLM and a handful of standalone routing startups already let developers point traffic across multiple providers. Switchyard plugs into several of them rather than replacing them outright.
The core problem Switchyard solves is that the right model changes as an agent moves through a task. An agent's state shifts as tools return results, errors show up, or a step turns out to be routine rather than complex, and a fixed model choice can't adapt to any of that.
Briski described routing strategies that respond to that shifting state rather than a static task category.
"It has many types of routing strategies," Briski said. "You can have a random router, which is not that great, or you can have an agent state route or a classifier route. Depending on your routing strategy, it wants to choose the best model. In some cases you want to go with a model like Lightning for really efficient tasks, and the router will actually choose Lightning if it's set up in your pool of models."
Cost enters the routing decision directly, not as an afterthought. In response to a question from VentureBeat, Briski said Switchyard can evaluate model verbosity, meaning how many tokens a given model tends to produce for a task, and use that prediction to steer work toward the cheaper option before the call is made.
The part that keeps this from becoming its own integration project is where Switchyard sits. Nvidia split its partners into two groups: agent frameworks that call Switchyard directly, including Cognition, LangChain and Nous Research, and LLM gateways that have built Switchyard support into their own products, including Kong, LiteLLM and OpenRouter. Kong ships Switchyard natively inside Kong AI Gateway. Briski pointed to that same list of gateway partners when describing how the library fits into the existing routing ecosystem.
"We are an ecosystem lover, and we want to make sure that we are integrated," Briski said. "We've partnered with OpenRouter, LiteLLM and Kong, and they've already integrated our routing algorithm, so you can pick it up right where you're already using the best tools."
Nvidia shared results from nine companies testing Switchyard, several with specific figures attached. LangChain reported a 74% cost reduction across 145 multi-turn Deep Agents tasks by routing just 7% of calls to a frontier model, at a 6% accuracy tradeoff. Ramp said it matched a frontier model's performance on Ramp SWE-Bench while cutting costs 58% and runtime 33%. Cognition integrated Switchyard's staged router into Devin Desktop for internal use and reported near-frontier performance on FrontierCode Main while cutting mean cost 28% relative to routing everything to a single frontier model.
Lightning's architecture and performance gains
Nemotron 3.5 Lightning is a standalone open model in its own right, built for high-volume, specialized agent tasks rather than general-purpose use.
It extends the hybrid Mamba-Transformer, latent mixture-of-experts architecture Nvidia introduced with the Nemotron 3 family in December 2025, the same line behind Nemotron 3 Super, which Nvidia uses as Lightning's own baseline in its post-training comparisons. Positioned within a routing setup like Switchyard, it's built to sit at the fast, cheap end of the decision rather than the frontier end, but it runs and ships independent of any router.
According to the Artificial Analysis Intelligence Index, a general capability benchmark spanning nine evaluations, Lightning scores 24, tied with gpt-oss-120b and behind Nemotron 3 Super, Gemma 4 31B, Claude 4.5 Haiku and Mistral Medium 3.5, all at 30. Lightning isn't a general-intelligence leader in its size class, and Nvidia isn't claiming it is.
The actual claim is narrower: according to PinchBench data supplied by Nvidia, Lightning matches Qwen3.6-35B's accuracy roughly 30% faster and beats Gemma 4 26B's accuracy at a similar completion time on PinchBench, a real-world agent task benchmark spanning coding, research and file management. That's a speed-to-accuracy tradeoff, not a capability win.
Post-training is where Nvidia says the bigger gains show up. The company shared before-and-after figures from four early-access partners: CrowdStrike's malicious-content recall against a Nemotron 3 Super baseline, CodeRabbit's coding router against a GPT 5.4 Nano baseline, Harvey and Trajectory's legal task completion against an Opus 4.6 baseline, and Lila Sciences' energy simulation work against an Opus 4.8 baseline. CodeRabbit's case is the most specific: Nvidia says the standard NeMo Auto model recipe, trained for one epoch, built into a working router agent for $85 in about two hours.
What this means for enterprises
There is no shortage of competitive offerings in the growing market for open models. The new Nemotron Lightning release will be yet another option for organizations to consider.
On the model side, Lightning's own benchmark chart picks Qwen3.6-35B as its direct comparison point. Asked by VentureBeat directly how Lightning compares to Chinese models more broadly, Briski didn't offer a head-to-head benchmark, pointing instead to openness and customizability as the differentiator.
"Our value proposition is not just open and it's very customizable," Briski said.
For enterprises building agentic infrastructure, three trends stand out:
The routing decision is becoming dynamic instead of static. Enterprises that built agent pipelines around a single default model are being pushed toward per-step routing based on live signals like agent state and token cost, not a fixed assignment set at design time.
Open source is now a cost lever at two layers, not one. Pairing an open model with an open router a vendor controls end to end is a newer argument than cheaper weights alone, and worth watching for whether other labs follow the same pattern.
The competitive question shifts from best model to best system. As routing libraries mature, the differentiator moves from which model an enterprise defaults to, toward how well its routing layer matches models to tasks in production, a harder thing to benchmark and a harder thing to market.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み