NVIDIA、AIファクトリ向けスケールアップネットワーク「NVLink」を発表
NVIDIA は AI ファクトリの拡張に不可欠なスケールアップネットワーク技術として NVLink を紹介し、大規模 AI 基盤の構築を支援する。
キーポイント
AI ファクトリにおける NVLink の役割
NVIDIA は NVLink を単なる接続規格ではなく、大規模 AI インフラを構築するための不可欠なスケールアップネットワークとして位置づけた。
大規模基盤の構築支援
この技術紹介は、企業がより大規模で効率的な AI 基盤を構築し、拡張性を確保することを支援するものである。
業界標準としての確立
AI ファクトリという概念の普及に伴い、NVLink が事実上の標準的なスケールアップ技術として認知されつつあることを示唆している。
重要な引用
NVIDIA は AI ファクトリの拡張に不可欠なスケールアップネットワーク技術として NVLink を紹介し
大規模 AI 基盤の構築を支援する
影響分析・編集コメントを表示
影響分析
この記事は、AI インフラストラクチャにおける NVLink の戦略的価値を明確に定義し、業界全体で「スケールアップ」の重要性を再認識させる契機となった。特に大規模モデルやファクトリ型 AI の構築において、ネットワーク技術がボトルネックとならないよう設計する際の指針となる。
編集コメント
AI ファクトリという概念が定着する中、それを支える基盤技術としての NVLink の重要性が改めて浮き彫りになりました。大規模化が進む現代の AI 開発において、ネットワーク設計は単なる接続手段を超えた戦略的要素となっています。
AI への需要は加速し続けています。ワークロードは拡大し、モデルは複雑化しており、AI 計算インフラをこれまで以上に迅速に展開する圧力が高まっています。この渇望を満たすために、データとエネルギーを継続的に知能へと変換するデータセンター規模のシステム「AI ファクトリー」が導入され始めています。
この AI ファクトリーというアプローチは、データセンターのシステム設計と運用を根本から変えました。もはやピーク時のアクセラレータ FLOPS(浮動小数点演算性能)だけでは不十分です。現在、パラメータ数が兆を超えるモデルや混合专家(MoE)アーキテクチャ、長文脈推論、分散型サービスなど、AI ワークロードは多数のアクセラレータを単一の計算ユニットとして連携させることを要求します。これを実現するには、GPU 間通信の高帯域化・低遅延化、コレクティブ演算のための高速なネットワーク内計算、そしてソフトウェアを意識したスケジューリングが不可欠です。
構成要素が膨大である以上、レジリエンス(耐障害性)はデータセンター全体に組み込まれている必要があります。AI の進化スピードに追いつき続けるには、業界のイノベーション速度に合わせて動ける技術とサプライチェーンが求められます。
この理由から、スケールアップ・ネットワークは AI ファクトリーにおける最も重要なアーキテクチャ上の意思決定の一つとなっています。スケールアップ・ファブリックこそが、アクセラレータを単一の計算ユニットとして機能させる鍵であり、トークンがエキスパート間でどのように移動するか、集合演算がどれほど迅速に完了するか、そしてファクトリーがどれだけ有用なスループットを生み出せるかを決定します。これは、オペレーターが新しいプラットフォームを導入する際にどの程度のリスクを負うかにも大きく影響する要素です。
NVIDIA NVLink は、AI ファクトリー向けに設計された専用スケールアップ・ネットワーク・ファブリックです。大規模かつ高速な GPU 間通信を必要とする AI 推論、トレーニング、その他の並列計算ワークロードの加速のために作られています。第 6 世代 NVLink インターコネクトと NVLink 6 スイッチは、スケールアップ・ネットワークにおいて最低遅延で最高帯域幅を実現する全対全トポロジーを提供し、集合演算をオフロードするための SHARP(Scalable Hierarchical Aggregation and Reduction Protocol)のサポートも備えています。さらに、生産環境での AI ファクトリーの稼働率向上のために設計されたラックレベルの耐障害性機能も含まれています。
チップからシステム、ファブリック、ソフトウェアに至るまで、スタック全体を協調して設計・最適化する「極限的な共設計」を通じて開発された NVLink は、NVIDIA の年次 AI インフラストラクチャロードマップに組み込まれた技術です。この成熟し、実証済みで広く展開されている技術が、現代の AI インフラストラクチャの基盤となっています。
スケールアップ・ネットワークが決定する AI ファクトリーの経済性
スケールアウトネットワークはデータセンター内のサーバー間を接続します。一方、スケールアップネットワークはドメイン内の GPU を単一の計算エンジンとして振る舞わせることを可能にします。どちらも不可欠ですが、解決する課題は異なります。
NVIDIA Quantum InfiniBand や NVIDIA Spectrum-X Ethernet といったスケールアウトファブリックは、数千から数十万の GPU に及ぶ大規模クラスターを構築できます。これは、大規模なデータ並列処理による AI 学習や大規模な科学シミュレーションにおいて極めて重要です。一方、スケールアップファブリックは単一のドメイン内のアクセラレータ同士を、高帯域幅と予測可能な低レイテンシ、そして共有された高帯域メモリ(HBM)で接続します。現代の AI 学習や推論においては、レイテンシに敏感な通信パターンが主にこのスケールアップ領域で発生します。
具体例として、MoE モデルを用いた推論を見てみましょう。
効率的な推論実装では、expert parallelism を活用してエキスパートを GPU 間に分散させ、大規模バッチサイズを採用してファクトリーの処理スループットを最大化します。これによりワークロードは並列化されますが、GPU 間の集中的なオール・トゥー・オール通信が発生します。トークンは選択されたエキスパートへ配信され、処理され、集約され、再順序付けされて次の段階へと渡される必要があります。
GPU 間での通信はすべて並列で行われる必要があります。もし専門化されたモデル(エキスパート)が、帯域幅の低いレイヤーや遅延の高いネットワーク構造に囲まれていれば、専門化による効率化のメリットは通信オーバーヘッドによって相殺されてしまいます。この現象は、MoE モデルの学習時にも同様に発生します。
AIファクトリーの性能において、すべてのノード間の帯域幅とレイテンシがいかに重要かがわかります。これらの能力を負荷下やあらゆるワークロードで維持するためには、システム全体と協調設計された専用かつスケールアップ可能なネットワークが不可欠です。
スケールアップ型ネットワークは、ワットあたり、ドルあたり、そしてファクトリーの床面積あたりのトークン生成量を最適化することで投資対効果(ROI)を向上させます。その結果、AIファクトリーでは学習期間の短縮、稼働率の向上、そしてトークンあたりのコスト低下が実現されます。
NVIDIA は、大規模な MoE(Mixture of Experts)モデルにおいて、高帯域・低遅延のスケールアップネットワークがもたらす影響を実証しました。DeepSeek-R1 や Qwen 235B、さらにシミュレーション上の 2T パラメータを持つ LLM を対象とした評価では、NVLink は主要な市販 Ethernet(OTS: Off-The-Shelf)と比較して、デコードスループットを最大 2.3 倍向上させることが確認されています。
image*図 1. 第 6 世代 NVLink は、OTS Ethernet(72 アクセラレータのスケールアップドメイン)と比較して、デコードスループットを最大 2.3 倍向上させます。* *結果はシミュレーションに基づくものです*
スケールアップ技術の評価
仕様書では、ファブリックを比較する際に単純な帯域幅の数値——リンクレートやスイッチの総容量、あるいはデバイスあたりの最大帯域幅など——が用いられることがよくあります。これらの数値は有用ですが、それだけでは不十分です。
今日の AI ワークロードにおけるスケールアップ能力を評価するには、工場全体の視点が必要です。重要なのは、単位時間・電力、そして工場の敷地面積あたりで処理・生成できるトークンの数を決める「実効的なフルシステム性能」です。これには、すべての GPU 間の通信(オールツーオール)を支えるファブリック帯域幅、エンドツーエンドのレイテンシ(各 GPU の HBM メモリからファブリックを経由して他の全 GPU の HBM メモリに至るまでの遅延)、およびネットワーク内での集約計算やその他の集合演算の処理能力が含まれます。さらに、ルーティングの最適化、集合演算の可視化、リンクトラフィックの負荷分散、データ転送のパイプライン化を実現し、現在の AI ワークロードで使われているライブラリやフレームワークを完全にサポートする、あらゆるレベルでの統合ソフトウェアも不可欠です。
工場の稼働が前提となって初めて、実効性能は意味を持ちます。故障なく長時間動作し続ける能力、システムヘルスの継続的な監視、そして工場全体の稼働を維持しながら部品レベルでサービスやメンテナンスを行える仕組みこそが、理論上の性能を実際の生産性(グッドプット)へと変換します。これは、工場の全ライフタイムを通じて安定して生産を続けられる容量を指します。
複雑な技術を通じて多数のワークロード全体に最適な実効パフォーマンスとスループットを達成することは、極めて困難な課題です。これは、広く展開されたスケールアップインフラストラクチャでの長年の学習経験と、堅牢なサプライチェーンによって解決が図られてきました。
証明されていない技術スタックや、スケールアップネットワーク用に特別に設計されていないソリューションを使用すると、工場のパフォーマンスが最適化されず、稼働停止による混乱、そして供給の不安定さを招くリスクがあります。
これらの指標はすべて、AI ファクトリーのライフサイクルを通じて持続可能で信頼性の高い ROI(投資対効果)という、唯一重要となる結果へとつながります。
AI ファクトリーにおけるスケールアップネットワークの主要指標
実効パフォーマンス 工場のレジリエンス プラットフォームの成熟度と証明されたサプライチェーン
AI ファクトリーの性能発揮は、帯域幅とレイテンシに始まりますが、それだけでは不十分です。エンドツーエンドのスケーラブルネットワークの性能、集約計算やその他の集合操作のためのインネット・コンピューティング、そしてソリューション全体に統合された成熟したフルスタックの実装も不可欠です。
実際の性能を有効なスループット(グッドプット)に変換するには、システム全体のレジリエンスが求められます。これには、長時間の稼働継続、システムの健康状態とテレメトリデータの常時監視、そしてファクトリーの他の部分が稼働し続ける中でコンポーネントレベルでのサービスや保守を可能にする仕組みが含まれます。
AI ファクトリーのスケーリングは、多くの依存関係を抱える極めて複雑な課題です。運用担当者は、大規模展開の実績と投資対効果(ROI)の確立という、成熟した技術スタックを活用することでリスクを最小限に抑えたいと考えています。
*表 1. スケーラブル・ネットワーキング・ソリューションを評価する上で、この 3 つの指標が極めて重要です。*
NVLink:性能、レジリエンス、そして成熟度
第 6 世代となった NVLink は、ファクトリーのグッドプットを最大化するために必要な性能とレジリエンスを提供します。これは、10 年以上にわたるスケーラブル技術への投資と、世界有数のハイパースケイラー、クラウドサービスプロバイダー(CSP)、スーパーコンピューティングセンターなどでの実証済み展開の成果です。
世界をリードする性能
「Vera Rubin NVL72」では、第6世代のNVLinkにより、1 GPU あたり 3.6 TB/s の双方向帯域幅と、72 GPU で構成されるドメイン全体で 260 TB/s のラックレベル帯域幅を実現しています。既存のエザーネット製品ベースの代替ソリューションと比較して、GPU 間の転送におけるエンドツーエンドのレイテンシは 3 分の 1 に短縮され、パケットレートは 10 倍に向上しました。
- レイテンシ:3 分の 1(低減)
- パケットレート:10 倍(向上)
- インネット・コンピューティング能力:130 TFLOPS
第6世代のNVLinkは、既存のエザーネット製品ベースの代替ソリューションと比較して、GPU間転送におけるレイテンシの短縮とパケットレートの向上を実現するだけでなく、集合演算のためのインネット・コンピューティング機能も備えています。
NVLink スイッチトレイと 5,000 本のケーブルで構成される NVLink スパインは、単一のオールツーオールトポロジーを形成しており、どの GPU でも他の任意の GPU と均一なレイテンシと帯域幅で通信できます。各トレイには NVLink 6 のスイッチチップが 4 基搭載され、合計帯域幅は 28.8 TB/s、FP8 でのインネット・コンピューティング能力は 14.4 TFLOPS です。単一の「Vera Rubin NVL72」ラック内では、NVLink 6 が集合演算(例:all-reduce、reduce、broadcast など)の高速化のために、合計帯域幅 260 TB/s とインネット・コンピューティング能力 130 TFLOPS を提供します。
NVLink の技術ロードマップには、最大 1,152 GPU にスケールアップできるドメインサイズへの対応と、コパッケージド光学素子を通じた接続サポートが含まれています。
image*図 2. Vera Rubin NVL72 の NVLink スパイン、ラック、および NVLink 6 スイッチトレイは、GPU あたり 3.6 TB/s を含む総帯域幅 260 TB/s と、ネットワーク内計算能力 130 TFLOPS を提供します*。
スケールアップ・ネッティングスタックにおいて、ソフトウェアは不可欠な要素です。これには NVIDIA Dynamo、NVIDIA TensorRT-LLM、NVIDIA Collective Communications Library (NCCL) などがあり、これらはすべて約 20 年前に初めてリリースされた世界をリードする並列コンピューティングプラットフォームである NVIDIA CUDA を基盤として構築されています(プラットフォームの成熟度に関するセクションで詳しく解説します)。
ハードウェアとソフトウェアは、極限までの協調設計 (extreme co-design) を通じて開発されています。このアプローチにより、個別のスタック要素を単独で扱うだけでは達成できない性能向上が実現します。今日の AI ワークロードにおいては、これが直接的にインフラの経済性へと結びつきます。
例えば、NVIDIA Hopper から NVIDIA Blackwell への移行では、NVLink の帯域幅を倍増させ、スケールアップドメインを 8 GPU から 72 GPU に拡大し、分散推論のために Dynamo を導入しました。その結果、ワットあたりの MoE(Mixture of Experts)推論性能が 50 倍向上しました。
NVIDIA の Vera Rubin プラットフォーム はさらに、NVLink の帯域幅とネットワーク内計算能力を両方とも倍増させることで、この進化を推し進めています。
image*図 3. NVLink は、高帯域・低遅延のスケーラブル領域で 72 個の GPU を接続し、ネットワーク内での計算(in-network compute)を実現することで、GB300 NVL72 が H200 と比較してトークンあたりの消費電力を 50 倍向上させることを可能にします*。
モデルアーキテクチャが進化するにつれ、ハードウェア基盤とソフトウェア通信スタックも同時に進化する必要があります。この極端な協調設計(co-design)アプローチでは、スタック全体にわたる緊密な連携が不可欠です。特に大規模展開の圧力下では、開発を分断した「サイロ型」のアプローチでこれを再現するのは現実的ではありません。しかし、単にアクセラレータを追加するのと、実用的なパフォーマンスまでスケーリングできるかの差は、まさにこの点にあります。
速度とインテリジェントな耐障害性のために設計された AI ファクトリー
AI ファクトリーは継続的な稼働が求められます。ラックの密度が高まり、その価値が増大する中で、保守性と障害管理もパフォーマンス計算の一部となります。高帯域を提供しながらも、メンテナンスのためにシステムを停止させる必要がある基盤では、実効容量や収益が損なわれてしまいます。
NVLink 6 では、AI ファクトリーの運用に最適化された管理機能とインテリジェントな耐障害性機能を導入しました。これらの機能により、ラックの保守作業が容易になり、コンポーネントのパフォーマンスや負荷分散に関する可視性が向上します。また、特定のノードでサービスやアップデートが必要な場合でも、ファクトリー全体の稼働を維持できます。
第 6 世代の NVLink スイッチは、アップタイムと実効スループット(goodput)を最大化するためのインテリジェントな耐障害性機能をサポートしています:
- コントロールプレーンの耐障害性
- 部分的にラックを構成しての運用対応
- ホットスワップ可能なスイッチトレイ
- マネージメントコントローラーのフォールバック機能を備えたソフトウェア定義ルーティング
- 動的なトラフィック再ルーティング
- サービス中のソフトウェア更新
- モニタリングと障害原因特定のための高精度リンクテレメトリ
これらの機能は二次的なものではありません。生産環境にある AI ファクトリーでは、リンクやスイッチ、トレイ、マネージメントコントローラーのいずれかが故障しても、ラック全体が停止してはいけません。運用担当者は、インフラの経済的価値に見合った障害分離、テレメトリ、そしてサービス手順を必要としています。NVLink 6 は、スケールアップネットワークを AI ファクトリースタックの残りの部分と同じ運用規範に統合します。
高速なサイクルで成熟した技術
最も強力な技術戦略は、成熟度とスピードを両立させるものです。NVLink はその両方を備えています。
NVLink は現在、目的別に設計されたスケールアップネットワークファブリックの第 6 世代を迎えています。ほぼ 10 年にわたり大規模な生産環境で展開されており、数百万台の NVIDIA チップが NVLink 対応システム上で稼働しています。また、サーバー、ラック、ケーブル、スイッチ、ソフトウェア、運用ツールに至るまで、広範なエコシステムを形成しています。
NVIDIA の年次プラットフォームサイクルは、AI の革新スピードに合わせて新ハードウェア機能を導入します。これにより業界は、世界の AI 需要に応えるために新しいモデルやワークフローを迅速に展開することが可能になります。
ハードウェアに加え、スケールアップネットワークスタックの重要な構成要素として、NVIDIA とコミュニティがソフトウェアもリリースしています。主な内容は以下の通りです:
- Dynamo:マルチノード GPU 環境における生成 AI や推論モデルの拡張を可能にするオープンソースフレームワークです。分離型サービス、動的な GPU アロケーション、KV キャッシュを意識したルーティング、そして NIXL を活用したデータ転送機能を備えています。
- TensorRT-LLM:NVIDIA GPU 上で LLM の高速推論を実現するためのオープンソースライブラリです。最適化されたカーネルや計算・通信のオーバーラップ、NVFP4 や FP8 といった低精度フォーマットを活用してスループットを向上させます。MoE モデルへの対応も含まれています。
- NIXL:GPU メモリ、CPU メモリ、NVMe、リモートストレージ間での高速データ転送を実現するオープンソースライブラリです。分散システム内での KV キャッシュや推論状態の効率的な移動を可能にします。
- NCCL:10 年以上にわたるオープンソースの革新を積み重ねた、NVIDIA のオープンソースライブラリです。高速な GPU 間通信を実現し、トポロジー認識型コレクティブ操作、主要 AI フレームワークとの統合、ネットワーク内での集約処理やその他の集合演算をサポートする SHARP 技術への対応も含まれています。
- ファブリックとメモリ管理:アドレス空間管理 API や拡張可能なメモリスемantics、効率的な HBM アクセスのための組み込みメモリ整合性機能などを提供します。
ソフトウェアの革新はハードウェアリリース後も続き、製品ライフサイクル全体を通じて X 倍もの速度向上をもたらします。
NVLink-C2C がファブリックを CPU へ拡張
AI ファクトリーには、GPU から GPU への帯域幅だけでなく、オーケストレーションやデータ転送、メモリ管理、ストレージサービス、そして計算フェーズが混在するエージェントワークロードに対応するための、高帯域かつ整合性を保った CPU と GPU の接続も必要です。
NVIDIA NVLink-C2C がその道筋を示します。Vera Rubin NVL72 プラットフォームに搭載される Vera CPU を用いれば、NVLink-C2C は CPU と GPU の間で 1.8 TB/s の整合性のある帯域幅を提供します。これは PCIe Gen6 の 7 倍の速度です。これにより、CPU メモリと GPU メモリの間に高速なデータ共有と統一された整合メモリアーキテクチャが実現されます。
インフラチームにとっては、制御・メモリ・計算処理間のボトルネックが減ることを意味します。ソフトウェアチームにとっては、プログラミングモデルが簡素化され、KV キャッシュのオフロードやマルチモデル実行、データ処理、エージェント型オーケストレーションといったワークロードに対応できるようになります。
Vera はこの役割のために設計されており、88 基の NVIDIA カスタム製 Olympus コアを搭載し、AI ファクトリー向けワークロードに最適化された高いメモリ帯域幅と省電力動作を実現しています。NVIDIA のプラットフォームでは、CPU、GPU、NVLink、HBM、システムメモリ、ネットワーク、そしてソフトウェアが協調して設計され、パフォーマンスの最大化を図っています。
NVLink Fusion: 一から作り直さないための半カスタム XPU インフラ
ハイパースケイラーや AI ネイティブ企業は、特定のワークロード向けに専用シリコンを構築し、顧客に選択肢を提供しています。しかし、その展開にはいくつかの課題があります。最先端のスケーラブル・ネットワークの統合、ラック規模全体のアーキテクチャ設計と実装、シリコン供給リスクを踏まえたデータセンター設計へのコミットメント、そして異種混合インフラのサポートなど、それぞれが独自の難問となっています。
NVIDIA NVLink Fusion がその道筋を示します。NVLink Fusion は、カスタムシリコンを NVIDIA の世界最高峰の AI インフラプラットフォームに接続するための、高帯域・低遅延な相互接続技術および IP です。
NVLink Fusion を活用すれば、実証済みの NVLink スケールアップスタックとエコシステムを利用できるため、開発や導入の複雑さを減らし、パフォーマンスを向上させ、セミカスタムの AI ファクトリにおける市場投入までの期間を短縮できます。さらに、単一の統一アーキテクチャに標準化することでデータセンター全体の運用が簡素化され、データセンター容量の柔軟な再割り当てが可能になります。これにより、カスタム AI XPUs は GPU とシームレスに統合された分散型コンピューティングを実現します。
これは、カスタム XPUs と世界クラスの AI プラットフォームのどちらかを選ばなければならないという根本的な制約を解消するものです。また、NVLink Fusion は NVIDIA の技術ロードマップに合わせて進化するため、採用企業も同社の年次ペースに追随できます。
生産向け AI ファクトリの道筋
次の AI ファクトリで勝つのは、単体で最も高速なアクセラレーターを持つ企業ではありません。最も少ないコストでトークンあたりの知能を最大化し、最速のペースで提供し、導入リスクを最小限に抑えられるインフラを持つ企業が勝ちます。
NVLink は、3 分の 1 の遅延と 10 倍のパケットレート、さらにネットワーク内での 130 TFLOPs の計算能力を実現する圧倒的なパフォーマンスを提供します。加えて、ファクトリーレベルの耐障害性と、成熟した実績あるプラットフォームを備えています。
AI ファクトリーは、AI 計算時代の基盤インフラとして確立されつつあります。NVLink は、これらを支えるための高性能、運用性、成熟度、そして継続的なイノベーションを提供するネットワークです。
詳しくは、NVIDIA Vera Rubin Platform、NVLink、および NVLink Fusion のページをご覧ください。
原文を表示
The demand for AI continues to accelerate. Workloads are getting larger, models are becoming more complex, and there is mounting pressure to deploy AI compute infrastructure faster than ever. AI factories—data center-scale systems that continuously convert data and energy into intelligence—are being deployed to meet this insatiable demand.
This AI factory approach to the data center has fundamentally changed system design and operation. Peak accelerator FLOPS are no longer enough. Today’s AI workloads, including trillion+ parameter models, mixture-of-experts (MoE) architectures, long-context reasoning, and disaggregated serving, require many accelerators working together as a single unit of compute. Achieving this requires high-bandwidth, low-latency GPU-to-GPU communication, fast in-network compute for collectives, and software-aware scheduling.
Given the sheer number of components, resiliency needs to be built into the entire data center. Keeping up with the pace of AI requires technology and a supply chain that can move at the speed of industry innovation.
This is why scale-up networking has become one of the most important architectural decisions in the AI factory. The scale-up fabric is what enables accelerators to work as a single unit of compute, determining how effectively tokens move across experts, how quickly collective operations complete, and how much useful throughput the factory can deliver. It is a significant factor in how much risk operators take when deploying new platforms.
NVIDIA NVLink is the purpose-built scale-up networking fabric for AI factories. It’s designed to accelerate AI inference, training, and other parallel computing workloads that require large, fast GPU-to-GPU communications. The Sixth Generation NVLink interconnect with NVLink 6 Switch provides the highest GPU-to-GPU bandwidth at the lowest latency all-to-all topology for scale-up networking, as well as support for SHARP in-network compute for offloading collective operations. It also includes rack-level resiliency features designed for production AI factory uptime.
Developed through extreme co-design, where the entire stack from chips to systems to fabric to software are designed and optimized together, NVLink is part of the NVIDIA annual AI infrastructure roadmap cadence. This mature, proven, and widely deployed technology forms the backbone of modern AI infrastructure.
How scale-up networking determines AI factory economics
Scale-out networks connect servers across the data center. Scale-up networks enable the GPUs inside the domain to behave as a single engine of compute. Both are essential, but they solve different problems.
Scale-out fabrics such as NVIDIA Quantum InfiniBandand NVIDIA Spectrum-X Ethernet enable large clusters spanning thousands to hundreds of thousands of GPUs. This is critical for large, data-parallel AI training workloads (as well as large-scale scientific simulations). Scale-up fabrics connect accelerators in a single domain with high bandwidth, predictable low latency and shared high-bandwidth memory (HBM). For modern AI training and inference, the scale-up domain is where most latency-sensitive communication patterns occur.
As an example, let’s take a look at inference with MoE models.
Efficient inference implementations use expert parallelism to distribute experts across GPUs, and they use large batch sizes to maximize factory throughput. This parallelizes the workload, but it creates intensive all-to-all communication between GPUs. Tokens must be dispatched to the selected experts, processed, gathered, reordered, and passed forward.
All of the GPU-to-GPU communication must happen in parallel. If the experts sit behind a low-bandwidth or high-latency fabric, gains from expert parallelism can be erased by communication overhead. The same phenomenon happens during training of MoE models.
The takeaway is that all-to-all bandwidth *and *latency are critical to AI factory performance. AI factories need purpose-built, scale-up networking that is co-designed with the rest of the system to deliver and maintain these capabilities under load, across all workloads. Scale-up networking increases ROI by optimizing the delivered tokens per watt, per dollar, and per factory square foot. It leads to shorter training runs, higher utilization, and lower cost-per-token for AI factories.
NVIDIA has demonstrated the impact of high-bandwidth, low-latency scale-up networking in large-scale MoEs. For models such as DeepSeek-R1 and Qwen 235B, as well as a simulated 2T parameter LLM, NVLink delivers up to 2.3X the decode throughput compared to leading off-the-shell (OTS) Ethernet.

Evaluating scale-up technologies
Spec sheets often compare fabrics using simple bandwidth numbers: link rate, aggregate switch capacity, or headline bandwidth per device. Those numbers are useful, but they aren’t sufficient.
Evaluating scale-up capabilities for today’s AI workloads requires taking a factory-level view. Delivered, full-system performance determines how many tokens can be processed and produced per unit time, power, and factory footprint. It depends on all-to-all fabric bandwidth, end-to-end latency (that is, latency from every GPU’s HBM memory, through the fabric, to every other GPU’s HBM memory), the in-network compute for reductions and other collectives. It also requires fully integrated software at every level that can optimize routing, expose collectives, balance link traffic, pipeline data transfers, and has full support for the libraries and frameworks used by today’s AI workloads.
Delivered performance only matters to the extent that the factory is operational. The ability to run for extended periods without failure, continuously monitor system health, and support service and maintenance at the component level while the rest of the factory continues operations translates delivered performance to goodput — the capacity of a factory to produce over its full lifetime.
Achieving optimal delivered performance and goodput across many workloads through a complex technology is an incredibly difficult challenge, addressed through years of learnings with widely deployed scale-up infrastructure and a robust supply chain. Using an unproven technology stack or a solution not purpose-built for scale-up networking risks sub-optimal factory performance, disruptive downtime events, and undependable supply.
Together, these metrics drive the only result that matters: sustained, dependable ROI over the lifetime of the AI factory.
The Key Metrics for Scale-Up Networking in AI Factories
NVLink: Performance, resiliency, maturity
Now in its sixth generation, NVLink provides the delivered performance and resiliency needed to maximize factory goodput, and over ten years of scale-up technology investment and proven deployments including world-leading hyperscalers, Cloud Service Providers (CSPs), and supercomputing centers.
World-leading performance
With Vera Rubin NVL72, sixth generation NVLink provides 3.6 TB/s per GPU of bidirectional GPU-to-GPU bandwidth and 260 TB/s of rack-level GPU bandwidth in a 72-GPU domain. The end-to-end latency for GPU-to-GPU transfers is 3X lower than alternative solutions based on off-the-shelf Ethernet, and the packet rate is 10X higher.
3X
Lower latency
10X
Higher packet rate
130 TFLOPS
In-network compute
6th generation NVLink delivers lower latency and higher packet rates for GPU-to-GPU transfers compared to alternative solutions based on off-the-shelf Ethernet, as well as in-network compute capabilities for collective operations
NVLink Switch trays and NVLink spine of 5,000 cables form a single all-to-all topology so any GPU can communicate with any other GPU with uniform latency and bandwidth. Each tray includes four NVLink 6 switch chips, 28.8 TB/s of total tray bandwidth, and 14.4 TFLOPS of FP8 in-network compute. In a single Vera Rubin NVL72 rack, NVLink 6 provides 260 TB/s of aggregate bandwidth and 130 TFLOPS of in-network compute for accelerating reductions and other collective operations (e.g. all-reduce, reduce, broadcast, etc).
The NVLink technology roadmap includes support for scale-up domain sizes up to 1152 GPUs and connectivity through co-packaged optics.

Software is a critical piece of the scale-up networking stack and includes NVIDIA Dynamo, NVIDIA TensorRT-LLM, NVIDIA Collective Communications Library (NCCL) and more (discussed in detail in the section on platform maturity), all built on top of NVIDIA CUDA, the world-leading parallel computing platform, first released nearly 20 years ago
The hardware and software are developed with extreme co-design. This results in performance speedups that can’t be achieved through siloed stack elements. For today’s AI workloads, this translates directly into infrastructure economics.
As an example, in the transition from NVIDIA Hopper to NVIDIA Blackwell, which included doubling the NVLink bandwidth, expanding the size of the NVLink scale-up domain to 72 from 8 GPUs, and incorporating Dynamo for disaggregated inference, NVIDIA achieved a 50X improvement in MoE inference performance per watt.
The NVIDIA Vera Rubin platform further extends this by doubling both NVLink bandwidth and in-network compute.

As model architectures evolve, the hardware fabric and software communication stack evolve together. This extreme co-design approach to development requires tight collaboration across the entire stack. It’s not feasible to replicate with a siloed approach to development, especially under hyperscale deployment pressure, but it’s the difference between just adding accelerators versus scaling to useful delivered performance.
Built for speed and intelligent resiliency
AI factories must run continuously. As racks grow denser and more valuable, serviceability and fault management become part of the performance equation. A fabric that delivers high bandwidth but requires disruptive maintenance can reduce effective capacity and revenue.
NVLink 6 introduces management and intelligent resiliency features designed for AI factory operations. The features make rack maintenance easier, provide better insight into component performance and load balancing, and keep the factory running even if individual nodes require service or updates.
The sixth generation NVLink Switch supports intelligent resiliency features to maximize uptime and goodput:
- Control plane resilience
- Support for operation with partially populated racks
- Hot-swappable switch trays
- Software-defined routing with management controller fallback
- Dynamic traffic rerouting
- In-service software updates
- Fine-trained link telemetry for monitoring and fault attribution
These capabilities aren’t secondary. In a production AI factory, a failed link, switch, tray, or management controller shouldn’t force an entire rack out of service. Operators need fault isolation, telemetry, and service procedures that match the economic value of the infrastructure. NVLink 6 brings scale-up networking into the same operational discipline expected from the rest of the AI factory stack.
Mature technology on a fast cadence
The strongest technology strategies combine maturity with pace. NVLink has both.
NVLink is now in its sixth generation of purpose-built scale-up networking fabric. It has been production-deployed at scale for nearly a decade, with millions of NVIDIA chips deployed across NVLink-capable systems and an ecosystem of servers, racks, cables, switches, software, and operations tooling.
The NVIDIA annual platform cadence delivers new hardware capabilities at the pace of AI innovation, enabling the industry to rapidly deploy new models and workflows to meet the world’s AI needs.
In addition to hardware, NVIDIA and the community release software as a key component of the scale-up networking stack. These include:
- Dynamo: Open source framework for scaling generative AI and reasoning models in multi-node GPU environments, with disaggregated serving , dynamic GPU allocation, KV-cache-aware routing, and NIXL-based data movement.
- TensorRT-LLM: Open source NVIDIA library for high performance LLM inference on NVIDIA GPUs, using optimized kernels, compute/communication overlap, and lower-precision formats like NVFP4 and FP8 to boost throughput including for MoE models.
- NIXL: Open source NVIDIA library for fast data transfers across GPU memory, CPU memory, NVMe, and remote storage helping move KV-cache and inference state efficiently in distributed systems.
- NCCL: Open NVIDIA library for high-speed GPU communication, with more than 10 years of open source innovations. Includes topology-aware collectives, AI framework integration, and SHARP support for in-network reductions and other collectives.
- Fabric and Memory Management: Includes address space management APIs, extensible memory semantics, and built-in memory coherency for efficient HBM access.
Software innovations continue long after hardware release, enabling X factor speedups throughout the life of the product.
NVLink-C2C extends the fabric to CPUs
AI factories need more than GPU-to-GPU bandwidth. They also need high-bandwidth, coherent CPU-GPU connectivity for orchestration, data movement, memory management, storage services, and agentic workloads that mix compute phases.
NVIDIA NVLink-C2C provides that path. With Vera CPUs in the Vera Rubin NVL72 platform, NVLink-C2C delivers 1.8 TB/s of coherent bandwidth between CPUs and GPUs, 7x the bandwidth of PCIe Gen6. This enables high-speed data sharing and a unified coherent memory architecture across CPU and GPU memory. For infrastructure teams, that means fewer bottlenecks between control, memory, and compute. For software teams, it simplifies programming models and supports workloads such as KV-cache offload, multi-model execution, data processing, and agentic orchestration.
Vera itself is designed for this role, with 88 NVIDIA custom Olympus cores, high memory bandwidth, and power-efficient operation for AI factory workloads. In the NVIDIA platform, CPU, GPU, NVLink, HBM, system memory, networking, and software are co-designed to optimize performance.
NVLink Fusion: Semi-custom XPU infrastructure without starting over
Hyperscalers and AI natives build specialized silicon for targeted workloads and to offer options to their customers, but they face several challenges in deploying them. Integrating state-of-the-art scale-up networking, designing and deploying a complete rack-scale architecture, committing to data center design amid silicon supply risk, and supporting heterogeneous infrastructure each present unique difficulties.
NVIDIA NVLink Fusion provides that path. NVLink Fusion is the high-bandwidth, low-latency interconnect technology and IP that connects custom silicon to the NVIDIA world-leading AI infrastructure platform. With NVLink Fusion, they can leverage the proven NVLink scale-up stack and ecosystem to reduce development and deployment complexity, increase performance, and accelerate time-to-market for semi-custom AI factories. And by standardizing on a single unified architecture, NVLink Fusion simplifies operations across the data center, enables flexible reprovisioning of data center capacity, and allows custom AI XPUs to integrate seamlessly with GPUs for disaggregated compute.
This breaks a fundamental constraint by eliminating the need to choose between custom XPUs and a world-class AI platform. And because NVLink Fusion keeps pace with the NVIDIA technology roadmap, adopters can keep pace with the company’s annual cadence.
The path for production AI factories
The next AI factory won’t be won by the fastest standalone accelerator. It will be won by the infrastructure that can deliver the most intelligence, at the lowest cost per token, on the fastest cadence, with the least deployment risk.
NVLink delivers leading performance, with 3X lower latency and 10X higher packet rates as well as 130 TFLOPs of in-network compute, factory resiliency features, and a mature, proven platform. AI factories are becoming the defining infrastructure of the AI computing era. NVLink delivers the performance, operability, maturity, and continuous innovation that powers them.
Learn more about the NVIDIA Vera Rubin Platform, NVLink, and NVLink Fusion.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み