NVIDIA、アジェンティックAI時代を牽引するRubin GPUアーキテクチャを発表
NVIDIA は、エージェント型 AI ワークロードに対応する次世代 Rubin GPU アーキテクチャを発表し、Blackwell よりもエネルギー効率で 10 倍の処理能力を持つ新プラットフォームを提示した。
キーポイント
エージェンティック AI への対応強化
単発のプロンプト応答から、推論・計画・ツール使用・検証を行う持続的なエージェント型ワークロードへと重点が移行しており、Rubin アーキテクチャはこの要件に最適化されている。
Blackwell 比 10 倍のエネルギー効率
NVIDIA Rubin GPU は、単位エネルギーあたりのエージェンティックスループットで Blackwell よりも最大 10 倍の性能向上を実現し、大規模な AI ファクトリの運用コスト削減に寄与する。
次世代ハードウェアと NVFP4 の採用
拡張された Tensor Cores、新 HBM4 メモリサブシステム、第 3 世代 Transformer Engine を組み合わせ、NVFP4 形式で最大 50 ペタフロップスの性能を達成する。
データセンターの再定義
低遅延・高スループット・大規模 KV キャパシティを必要とするエージェント型推論に対応するため、データセンター全体を単一の計算ユニットとして統合する「NVIDIA Vera Rubin プラットフォーム」の構想が示された。
アジェンシー推論性能の劇的向上
Vera Rubinプラットフォームは、内部の2T MoEワークロードにおいて、世代間のパフォーマンスが10倍向上するパレートフロンティアを実現しています。
大規模エージェントとツール呼び出しのサポート
Rubin NVL72 plus Veraシステムは、従来のHopperやBlackwellと比較して約10倍のエージェント数と2倍のツール呼び出し処理能力を誇ります。
Agentic Inference のボトルネック解消
Rubin GPU とその協調設計されたスケールアップシステムは、データ移動、計算効率、長時間コンテキストの実行、ラック規模での展開に至るまで、エージェント推論の終端間(エンドツーエンド)のボトルネックを解決します。
重要な引用
Agentic workloads are not defined by a single prompt and response, but by sustained inference across many reasoning steps.
The data center must be reimagined as a single unit of compute
NVIDIA Rubin GPU, designed to deliver up to 10x more agentic throughput per unit of energy than NVIDIA Blackwell
Figure 1. Pareto frontiers illustrating the 10x generational uplift in agentic inference performance of the Vera Rubin platform (internal 2T MoE workload)
Rubin NVL72 plus Vera supports roughly 10 times more agents and twice as many tool calls.
The Rubin GPU and its co-designed scale-up system address the end-to-end bottlenecks of agentic inference, from data movement and compute efficiency to long-context execution and rack-scale deployment.
影響分析・編集コメントを表示
影響分析
この記事は、AI インフラが単なる計算資源から、自律的な意思決定と実行を担う「エージェント型 AI」の基盤へと進化していることを示す重要な転換点です。Rubin アーキテクチャの登場により、大規模な複雑タスク処理におけるエネルギー効率とスループットが飛躍的に向上し、企業レベルでの AI エージェント実装が現実的なものになると予想されます。
編集コメント
エージェンティック AI の台頭に伴い、従来の推論モデルの要件を超えた「持続的な推論」と「複雑なタスク実行」への対応が急務となっています。NVIDIA が提示する Rubin アーキテクチャは、この新たなパラダイムシフトに対する明確な技術的解答であり、次世代データセンターの設計思想そのものを書き換える可能性があります。
AI モデルの学習や人間向けのチャットインターフェースから始まった技術は、今や大規模な知能を生成し続ける常時稼働型の「AI ファクトリー」へと進化しました。これらのファクトリーの新たな使命は、推論を行い、計画を立て、ツールを活用し、中間結果を検証し、広大なコンテキストにわたる複雑な多段階タスクを実行する「エージェント型ワークフロー」を駆動することです。
エージェント型の負荷は、単一のプロンプトとレスポンスのやり取りで定義されるものではありません。それは多数の推論ステップにわたって持続的な推論処理を要するものです。これには、1 ステップあたりの低レイテンシ、高いデコードスループット、効率的な長文コンテキストへのアテンション機能、大容量の KV キャッシュ容量、そして密結合した GPU ドメイン間でモデルをスケールさせる能力が求められます。データセンターは計算資源の単一ユニットとして再設計されなければなりません。このビジョンを実現するのが、NVIDIA の「Vera Rubin プラットフォーム」です。
このプラットフォームの核心となるのは、NVIDIA Rubin GPU です。これは、エネルギー効率を単位あたりで比較した場合、NVIDIA Blackwell と比べてエージェント処理スループットを最大 10 倍に引き上げるよう設計されています(図 1)。拡張された精度柔軟性を備えたTensor Cores、新しい HBM4 メモリサブシステム、そして最大 50 ペタフロップスのNVFP4 パフォーマンスを実現する第 3 世代 Transformer Engine が連携し、エージェントワークロードの効率的な加速を可能にしています。
image*図 1. エージェント推論性能における世代ごとの 10 倍向上を示すパレートフロンティア(内部の 2T MoE ワークロード)*
本稿では、データ転送や計算効率、長文コンテキストの実行、そしてラック規模での展開に至るまで、エージェント推論のボトルネックをエンドツーエンドで解消する NVIDIA Rubin GPU と、それと協調設計されたスケールアップシステムについて解説します。
Rubin GPU アーキテクチャは、どのようにしてエージェントワークロードを支えるのか?
Rubin GPU(図 2)は、リトレイルサイズに制限された計算ダイを高密度かつ高効率で構成することで実現されています。これら 2 つのダイは、NVIDIA High-Bandwidth Interface(NV-HBI)と呼ばれる高速なダイ間リンクを通じて、1 つのパッケージ内に統合されています。
image*図 2. NVIDIA Rubin GPU チップアーキテクチャ*
このアーキテクチャの根幹にあるのは、推論、生成、検索、ツール利用といった作業が頻繁に切り替わるエージェントワークロードにおいて、膨大な計算リソースをいかに有効活用するかという課題です。3,360 億個のトランジスタ、224 のストリーミングマルチプロセッサ(SM)、そして 896 の Tensor Core が提供する圧倒的な計算密度と、数値形式に応じて精度を動的に切り替える第 3 世代 Transformer Engine がその解決策となります。この柔軟性により、Rubin GPU は高い精度を維持したまま、最大 50 ペタフロップスの NVFP4 推論性能を実現しています。
しかし、パフォーマンスは Tensor Core のスループットだけで決まるわけではありません。Rubin GPU では計算リソースが Graphics Processor Clusters(GPC)に整理され、大規模な集中型 L2 キャッシュによって支えられています。GigaThread Engine が作業を調整し、MIG Control が複数のワークロード向けに GPU を分割管理します。また NV-DEC がデコード処理を加速します。これらの機能が連携することで、Rubin GPU は計算密度を高い稼働率へと変換し、大規模なエージェントシステムを特徴づける多様で動的なワークロード全体で持続的なパフォーマンスを発揮できるようになります。
この利用率も、データが計算コアに到達する速度に依存します。Rubin は専用 HBM コントローラーと 12-Hi スタックを駆使し、最大 288 GB の HBM4 メモリを搭載することで、ピーク帯域幅として最大 22 TB/s を実現します。
高効率なデータ移動を複雑なデータレイアウト間で管理するのが、強化された Tensor Memory Accelerator (TMA) です。一方、NVIDIA NVLink 6 は、GPU から GPU への全対全通信のために NVLink スイッチへ最大 3,600 GB/s のスケールアップ帯域幅を提供します。また、CPU と GPU の整合性のある通信には NVLink-C2C が 1,800 GB/s を、ホスト接続には x16 PCIe Gen 6 が最大 256 GB/s をそれぞれ提供します。
最後に、大規模なエージェント型展開では、この実行ドメインをデータが通過する際の保護も不可欠です。TEE-I/O を備えた機密コンピューティングは、AI ファクトリー内での保存中・転送中・使用中のデータをすべて守るために設計されています。これら計算、メモリ、接続、セキュリティの各能力を合わせ持つことで、大規模化するモデルやスケールアップ領域にわたって効率的な実行を維持する必要があるエージェント型ワークロードに対する、GPU レベルの基盤が完成します。
Rubin GPU はどのようにして重要な推論経路を加速するのか?
ピーク計算能力だけでは、エージェント型の推論を加速するには不十分です。実世界の性能は、GPU がデータをいかに効率的に移動させ、行列演算を実行し、長いコンテキストの注意機構(アテンション)を処理し、依存関係のあるカーネル間で遷移するかにかかっています。
このセクションでは、これらの重要な実行経路におけるオーバーヘッドを削減するために設計された Rubin GPU の機能について解説します。
ラックスケールでの MoE 重みとトークンの移動を加速
Mixture-of-Experts (MoE) モデルは、多数の専門家ネットワーク間でトークンを動的にルーティングします。専門家の数が増えるほど、推論性能を高めるためには、効率的に専門家の重みを検索して移動させることが重要になります。
Rubin GPU は、Tensor Memory Accelerator を強化し、専門家モデルに多いデータ転送のオーバーヘッドを削減しました。改良された記述子(ディスクリプタ)の処理により、共通のレイアウトを持つがメモリの異なる場所に存在するテンソルを、ソフトウェア側でより効率的に扱えるようになります。
Rubin ではさらに、TMA に対するインライン記述子更新サポートを追加しました(図 3)。メモリ上の記述子を変更するのではなく、同じレイアウトを共有するテンソルに対しては単一の統一された記述子を維持し、メモリアドレスポインタやストライドなどのフィールドを、実行時に TMA 命令内で直接上書きできるようにしています。
image*図 3. Rubin は MoE の記述子共有を簡素化*
これにより、専門家の数が増加しても、MoE モデルの拡張性がより効率的に向上します。メタデータ管理やデータ転送のオーバーヘッドを減らすことで、Rubin は GPU の処理時間を有益な推論計算に集中させられます。これは、大規模な MoE モデルに依存するエージェントワークロードにおいて、スループットの向上を支えるものです。
ラックスケールでの行列演算の効率を倍増
Rubin は、1 クロックあたりの Tensor Core の処理スループットを 2 倍に引き上げます。これは K 次元で処理できるデータ量を倍増させることで実現されています。この最適化により、スループットがボトルネックとなるカーネルだけでなく、メモリアクセスやレイテンシがボトルネックとなるカーネルのパフォーマンスも向上します。
モデルの実行は多数の GPU に分散されるため、各 GPU が受け取る出力作業の断片は小さくなりますが、その一方でリダクション次元(K 次元)は依然として大きなままです。
K 次元を大きくする最大の利点は、ループ回数を減らせることです。図 4 に示すように、Blackwell では 4 回のイテレーションが必要だった GEMM 演算も、Rubin では 2 回で完了できます。ループ回数が減ればオーバーヘッドが削減され、Tensor Core の利用率が向上します。その結果、コンテキスト処理やデコード処理における GEMM 演算も、高いテンソル並列度でより効率的に実行できるようになります。
image*図 4. Rubin は K 次元の命令スループットを 2 倍に*
エージェント AI における長文脈処理の主要課題への対応
長文脈処理やエージェント AI のワークロードは、アテンション計算にますます大きな負荷をかけています。コンテキストウィンドウが拡大するにつれ、モデルはより多くのトークンを比較し、巨大なアテンションスコア行列を正規化し、その結果を次の層の出力生成に用いる値データに適用する必要があります。このため、アテンション計算は「ユーザーあたりの 1 秒間に生成できるトークン数」を向上させる上で、最も重要なパフォーマンスパスの一つとなっています。
Rubin は、活性化のスパース性と適応圧縮、そして改善されたソフトマックス処理能力を組み合わせることで、アテンション処理を加速します。Rubin の新しいスパース性機能をアテンションに安全かつ効果的に活用するシンプルな方法は以下の通りです。
まず、中間的なアテンションスコアを生成するために密な計算を行います。その後、Rubin はこの中間データを Tensor Memory から読み取り、構造化された 2:4 スパース圧縮形式に変換します。これにより、効率的に利用するための非ゼロ値とメタデータが生成されつつ、スコアの書き込みコストやストレージ要件を削減できます。その結果、後のアテンション処理段階ではデータを減らして動作できるようになりますが、モデルの他の部分から期待される密な出力フォーマットは維持されます。
image*図 5. Rubin のスパース性機能は、アテンションブロックと MLP ブロックの活性化に適用可能*
この圧縮された中間表現により、ソフトマックス処理と、2 つ目のアテンション GEMM(行列乗算)という 2 つの重要な場所で計算量を削減できます。ソフトマックスは非ゼロのアテンション値のみを処理し、その後の密な Q 行列との乗算では、ソフトマックスからの非ゼロ値と、最初の圧縮ステップから得たメタデータを活用したスパース MMA(行列演算)を使用します。
これにより、長文コンテキストアテンションにおいて最もコストのかかる部分での計算量とデータ転送量が削減され、トークンあたりの電力効率(tokens/watt)が向上します。周囲のモデルパイプラインのインターフェースを変更する必要はありません。
Rubin はソフトマックスのスループットも向上させます。Tensor Core のスループットが増大すると、アテンション行全体での指数関数計算とリダクションに依存するソフトマックスがボトルネックになる可能性があります。これに対処し、Rubin では指数関数のスループットを強化しています。具体的には、Blackwell ベースラインと比較して FP32 で 2 倍、BF16/FP16 で 4 倍の向上を実現し、高速化する行列演算にソフトマックスが追いつけるようにしました。
NVIDIA GPU プラットフォーム
FP32 指数関数スループット(SM あたりクロックあたり)
BF16/FP16 指数関数スループット(SM あたりクロックあたり)
Blackwell 1x 1x
Blackwell Ultra 2x 2x
Rubin 2x 4x
*表 1. Blackwell から Rubin へ移行すると指数関数スループットが向上*
これらの機能により、Rubin のアテンション加速は単一のカーネル改善に留まりません。活性化のスパース性が中間的なアテンション計算量を削減し、高速化した指数関数がソフトマックスのボトルネックを解消します。
カーネル実行効率の向上
推論がより大規模なモデルや多数の GPU にスケールするにつれ、純粋な Tensor Core のスループットだけでは性能全体を語ることはできません。GPU はカーネルから次のカーネルへも効率的に遷移する必要があります。特に推論においては重要で、活性化データがクリティカルパス上に存在するためです。あるカーネルが活性化データを生成してメモリに書き込み、次のカーネルがそのデータを読み込んで次のトークンを生成し続けるという流れにおいて、この移動効率が決定的な役割を果たします。
従来のプロデューサー・コンシューマー型の実行モデルでは、GPU のタイムラインに「バブル(アイドル状態)」が生じることがあります。あるプロデューサーカーネルが一部のタイルやスレッドブロックの処理を早く完了しても、依存関係が解消されるまでコンシューマーは有効な作業を開始できません。
Blackwell ではプログラムによる依存起動機能でこの課題を改善し、コンシューマー側の進行をより早期に可能にしましたが、必要な活性化データ(activation data)が利用可能になるまでは、依然として待機状態が続くケースがあります。
image*図 6. プロデューサー・コンシューマーの重なり:Blackwell の一括トリガー(上)と Rubin のタイルレベルでのトリガー(下)*
Rubin では、依存関係にあるカーネル間の調整をより細粒度で行えるようになりました。これにより、必要な入力データが利用可能になった時点で即座にコンシューマー側の作業を開始でき、プロデューサー側の処理全体が完了するのを待つ必要がなくなります。
その結果、GPU のタイムラインはより密に詰められ、アイドル状態の隙間が減り、依存関係にあるカーネル間の重なりを最大化できます。これは特に、アジェンティック AI 推論において重要です。ここでは活性化データがモデル内を順次移動するため、カーネル間のレイテンシがユーザーあたりのトークン生成速度(tokens per second)に直結します。
Rubin のメモリと通信機構は、どうやって高スループットな推論を支えているのか?
モデルの規模、コンテキストウィンドウ、GPU ドメインが拡大するにつれ、データ転送は計算と同等に重要な役割を担うようになります。Rubin は、GPU 内部およびスケールアップシステム全体における重み、活性化値、KV キャッシュデータ、通信トラフィックの流れを改善するために設計されています。以下のセクションでは、高いスループットでの推論を支えるメモリと通信の革新について解説します。
スケールアップ通信の高速化
推論が単一の GPU からラック全体のシステムへとスケールする際、通信はパフォーマンスのボトルネックとなる重要な経路の一部となります。通信を GPU カーネル内部に直接統合することで、カーネルは停止して CPU に制御を戻す必要がありません。計算が進行中であっても、NVLink を介して別の GPU へ直接データを書き込んだり、集約処理を行ったりできます。
従来の GPU から GPU への通信では、ペイロードデータの転送に加え、調整や同期処理が必要でした。これらのステップは遅延を引き起こし、特に分散推論ワークロードで通信が頻繁に発生する場合には、インターコネクトの帯域幅を圧迫します。
Rubin は、デバイス側から開始される NVLink 通信における「カウント書き込み(counted writes)」を導入しました。この機能により、受信側の GPU が転送完了をより効率的に追跡できるようになり、GPU 間のデータ転送における同期処理が簡素化されます。
image*図 7. カウント書き込みを活用し、Rubin は NVLink 通信を高速化します*
NVLink を活用した統合通信とカウント書き込み機能は、GPU 間のデータ転送を調整するための低遅延メカニズムを提供し、同期操作を待たずに計算処理を継続させることで、全体の処理効率を高めています。
最大限の電力・計算効率を実現するメモリサブシステムの共設計
推論におけるデコード(生成)フェーズは、本質的にメモリサブシステムにボトルネックを抱えています。重要なのはピーク帯域幅の数値そのものではなく、各カーネルがメモリサブシステム全体をいかに効率的に活用できるかです。
現代の推論やエージェント型ワークロードでは、長時間のコンテキスト、大規模な KV キャッシュ、対話形式でのトークン生成により、デコードフェーズにかかるエンドツーエンドの実行時間が長くなっています。このため、実際に達成されるメモリ帯域幅が、パフォーマンスを左右する重要なレバーとなっています。
原文を表示
What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale. These factories are now tasked with powering agentic workflows that reason, plan, use tools, verify intermediate results, and execute complex multistep tasks across vast contexts.
Agentic workloads are not defined by a single prompt and response, but by sustained inference across many reasoning steps. They demand low per-step latency, high decode throughput, efficient long-context attention, large KV cache capacity, and the ability to scale models across tightly coupled GPU domains. The data center must be reimagined as a single unit of compute, a vision realized with theNVIDIA Vera Rubin platform.
At the core of the platform is the NVIDIA Rubin GPU, designed to deliver up to 10x more agentic throughput per unit of energy than NVIDIA Blackwell (Figure 1). Enhanced Tensor Cores with expanded precision flexibility, a new HBM4 memory subsystem, and the third-generation Transformer Engine—delivering up to 50 petaflops of NVFP4 performance—work together to accelerate agentic workloads efficiently.

This post examines how the NVIDIA Rubin GPU and its co-designed scale-up system address the end-to-end bottlenecks of agentic inference, from data movement and compute efficiency to long-context execution and rack-scale deployment.
How does Rubin GPU architecture support agentic workloads?
The Rubin GPU (Figure 2) is constructed from reticle limited compute dies to achieve high density and efficiency. These two dies are unified on a single package through a high-speed inter-die link called the NVIDIA High-Bandwidth Interface (NV-HBI).

The architecture begins with the challenge of keeping an enormous amount of compute productive as agentic workloads shift between reasoning, generation, retrieval, and tool use. Its 336 billion transistors, 224 streaming multiprocessors (SMs), and 896 Tensor Cores provide the raw compute density, while the third-generation Transformer Engine adapts precision across numerical formats. That flexibility enables the Rubin GPU to deliver up to 50 petaflops of NVFP4 inference performance while preserving accuracy.
But performance depends on more than Tensor Core throughput. The Rubin GPU organizes compute resources into Graphics Processor Clusters (GPCs) with a large centralized L2 cache. The GigaThread Engine coordinates work, MIG Control partitions the GPU for multiple workloads, and NV-DEC accelerates decoding. Together, these capabilities help the Rubin GPU turn compute density into sustained utilization across the diverse and dynamic workloads that define large-scale agentic systems.
That utilization also depends on how quickly data can reach compute cores. Rubin integrates up to 288 GB of HBM4 memory, driven by dedicated HBM controllers and 12-Hi stacks, to deliver up to 22 TB/s of peak bandwidth. The enhanced Tensor Memory Accelerator (TMA) manages high-efficiency movement across complex data layouts, while NVIDIA NVLink 6 provides 3,600 GB/s scale-up bandwidth to the NVLink Switch for all-to-all GPU-GPU communication, NVLink-C2C delivers 1,800 GB/s for coherent CPU-GPU communication, and x16 PCIe Gen 6 provides up to 256 GB/s of host connectivity.
Finally, large-scale agentic deployments must protect data as it moves through this execution domain. Confidential Computing with TEE-I/O is designed to secure data at rest, in transit, and in use across the AI factory. Together, these compute, memory, connectivity, and security capabilities form the GPU-level foundation for agentic workloads that must sustain efficient execution across increasingly large models and scale-up domains.
How does the Rubin GPU accelerate critical inference paths?
Peak compute alone is not enough to accelerate agentic inference. Real-world performance also depends on how efficiently the GPU moves data, executes matrix operations, processes long-context attention, and transitions between dependent kernels. This section explores the Rubin GPU features designed to reduce overhead in these critical execution paths.
Accelerating rack-scale MoE weight and token movement
Mixture-of-experts (MoE) models dynamically route tokens across many expert networks. As the number of experts grows, efficiently locating and moving expert weights becomes increasingly important to inference performance.
The Rubin GPU enhances the Tensor Memory Accelerator to reduce data-movement overhead for expert-heavy models. Its improved descriptor handling allows software to work more efficiently with tensors that share common layouts but reside at different locations in memory.
Rubin improves this with inline descriptor update support for TMA (Figure 3). Instead of modifying the descriptor in memory, the kernel can keep one unified descriptor for tensors that share the same layout and override fields such as the memory pointer and stride directly in the TMA instruction at runtime.

This helps MoE models scale more efficiently as expert counts increase. By reducing metadata-management and data-movement overhead, Rubin enables more GPU time to be devoted to useful inference computation, supporting higher throughput for agentic workloads that rely on large MoE models.
Doubling the efficiency of rack-scale matrix operations
Rubin doubles the Tensor Core throughput per clock by doubling the amount of data it can process along the dimension. This optimization not only helps throughput bound kernels but also memory and latency bound kernels.
As model execution is split across many GPUs, each GPU often receives a smaller slice of the output work while the reduction dimension remains large.
The benefit of a larger dimension is fewer loop iterations. In Figure 4, a GEMM that requires four iterations on Blackwell can be completed with two iterations on Rubin. Fewer iterations reduce loop overhead, improve Tensor Core utilization, and help both context and decode GEMMs run more efficiently at high tensor-parallel scale.

Addressing key challenges in long-context processing for agentic AI
Long-context and agentic AI workloads put increasing pressure on attention. As context windows grow, the model must compare more tokens, normalize larger attention-score matrices, and apply those scores to the value data used to produce the next layer’s output. This makes attention one of the most important performance paths for improving tokens per second per user.
Rubin accelerates attention by combining activation sparsity with adaptive compression and improved softmax throughput. One simple, safe, and effective way to use Rubin’s new sparsity features in attention is as follows. The attention pipeline begins with a dense computation to generate the intermediate attention scores. Rubin can then load that intermediate data from Tensor Memory into a structured 2:4 sparse compressed form, generating both nonzero values and the metadata needed to use them efficiently while reducing the scores’ write cost and storage requirements. This allows the later attention stages to operate on less data while preserving the dense output format expected by the rest of the model.

This compressed intermediate representation reduces work in two key places: softmax and the second attention GEMM. Softmax can operate on the nonzero attention values, and the following multiplication with the dense matrix can use sparse MMA with the nonzeros from Softmax and the metadata from the original compression step. The result is less compute and data movement in the most expensive part of long-context attention improving tokens/watt, without requiring the surrounding model pipeline to change its interface.
Rubin also improves the softmax throughput. As Tensor Core throughput increases, softmax can become a bottleneck because it depends on exponential math and reductions across attention rows. Rubin increases exponential throughput, including 2x FP32 and 4x BF16/FP16 throughput versus the Blackwell baseline, helping softmax keep pace with faster matrix operations.
- *Together, these features make the Rubin attention acceleration more than a single-kernel improvement. Activation sparsity reduces the amount of intermediate attention work, while faster exponentials reduce the softmax bottlenecks.
Improving kernel execution efficiency
As inference scales across larger models and more GPUs, raw Tensor Core throughput is only part of the performance story. The GPU also needs to move efficiently from one kernel to the next. In inference, this is especially important because activations often sit on the critical path: one kernel produces activation data, writes it to memory, and the next kernel consumes that data to continue generating the next token.
Traditional producer-consumer execution can create bubbles in the GPU timeline. A producer kernel may complete work for some tiles or thread blocks early, but the consumer may not begin useful work until a broader dependency is resolved. The Blackwell programmatic dependent launch improves this by allowing earlier consumer-kernel progress, but dependent work can still wait for required activation data to become available.

Rubin enables more fine-grained coordination between dependent kernels. This allows consumer work to begin earlier as required input data becomes available, rather than waiting for a larger set of producer work to complete.
The result is a more tightly packed GPU timeline, reduced idle gaps, and improved overlap between dependent kernels. This is particularly valuable for agentic inference, where activations move sequentially through the model and kernel-to-kernel latency directly affects tokens per second per user.
How do Rubin memory and communication sustain high-throughput inference?
As models, context windows, and GPU domains grow, data movement becomes as important as computation. Rubin is designed to improve the flow of weights, activations, KV cache data, and communication traffic within the GPU and across scale-up systems. The following sections look at the memory and communication innovations that help sustain high-throughput inference.
Accelerated scale-up communications
As inference scales from a single GPU to full rack-level systems, communication becomes part of the critical performance path. When communication is fused directly inside a GPU kernel, the kernel does not stop and hand control back to the CPU; it directly writes data or performs reductions over NVLink to another GPU while computation is still in flight.
Traditional GPU-to-GPU communication requires coordination and synchronization work in addition to moving payload data. These steps can add latency and consume interconnect bandwidth, particularly when communication occurs frequently within distributed inference workloads.
Rubin introduces counted writes for device-initiated NVLink communication. This capability streamlines synchronization for GPU-to-GPU data transfers by allowing the receiving GPU to track transfer completion more efficiently.

NVLink fused communication with counted writes provides a lower-latency mechanism for coordinating GPU-to-GPU data movement, helping keep computation moving instead of waiting on synchronization operations.
Co-designing a memory subsystem for maximum power and compute efficiency
The decode, or generation, phase of inference is fundamentally memory subsystem bound. It is not about peak bandwidth specs but rather how efficiently each kernel can utilize the entire memory subsystem. Modern reasoning and agentic workloads amplify this constraint by spending more end-to-end runtime in decode: long contexts, large KV caches, and interactive token generation make achieved memory bandwidth a critical performance lever.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み