NVIDIA、AIファクトリーの電力効率最大化技術「DSX MaxLPS」を発表
本文の状態
日本語全文を表示中
詳細モードで約22分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
NVIDIA Developer Blog
NVIDIA は AI ファクトリの電力制約下での効率最大化を目指す「DSX MaxLPS」を発表し、動的な電力配分や温水冷却による熱効率向上により、1 メガワットあたりの AI 出力を最適化する技術体系を示した。
AI深層分析を開く2026年8月22日 01:02
AI深層分析
キーポイント
AI ファクトリの設計パラダイムシフト
従来の「データセンターにいくつの GPU を収められるか」という問いから、「利用可能なメガワットあたりどの程度の AI 出力が得られるか」へと焦点が移っている。
NVIDIA DSX MaxLPS の定義と目的
MaxLPS は Maximum Land Power Shell を指し、土地、電力供給、物理的シェルというサイトレベルの制約内で AI ファクトリのスループットを最大化する技術スイートである。
動的電力配分と熱効率の向上
同社は未使用の電力余力を GPU に動的に割り当てる機能や、温水液体冷却による 45°C の熱効率設計を導入し、PUE(電力使用効率)の改善を図っている。
静的ラックプロビジョニングの課題
ピーク需要に対応するために最大電力を割り当てる従来の静的な設計は、実際のワークロードでは電力が遊休化し、冷却やネットワークなどのオーバーヘッドが増大する問題がある。
静的ラック割当による電力の遊休
従来のデータセンター設計では全ラックが最大消費電力を同時に引き出す前提で電力を確保するため、各ラックは独立した電力島として扱われる。このため、あるラックに割り当てられた余剰電力は他のラックへ融通できず、結果として利用可能な計算リソースが遊休する。
重要な引用
The question is no longer how many GPUs fit in a data center, but how much AI output each available megawatt can deliver.
MaxLPS stands for Maximum Land Power Shell, the site-level constraints that define an AI factory: land, utility power, and the physical shell holding power, cooling, networking, and compute infrastructure.
Dynamic power allocation: Continuously monitors and allocates unused power headroom to GPUs
An isolated rack provisioned with excess power cannot lend that unused power to a neighbor that could turn it into tokens.
編集コメントを表示
編集コメント
電力制約が AI ファクトリの主要なボトルネックとなる中、NVIDIA が提示した MaxLPS の概念は、単なるハードウェアの性能向上を超え、インフラ全体の設計思想を変える重要な転換点である。特に温水冷却や動的電力配分といった具体的な技術への言及は、実運用におけるコスト削減と持続可能性を追求する企業にとって即座に検討すべき課題を示唆している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI ファクトリーは電力制約のある産業システムです。もはや「データセンターに何個の GPU を収められるか」ではなく、「利用可能な 1 メガワットあたり、どれだけの AI 出力を生み出せるか」が問われています。
AI インフェレンスワークロードにおいては、この観点から「アプリケーションレベルでの電力あたりのパフォーマンス」が、AI ファクトリーの効率を測る主要指標となります。
しかし、すべてのメガワットが収益に直結する計算リソースに変換されるわけではありません。GPU に到達する前に、配電、冷却、ネットワーク、ストレージ、バックアップ、施設維持費などが電力の一部を消費してしまいます。さらに、従来のラック割り当て方式はこの問題を悪化させます。これはデータセンター設計における古いアプローチで、最悪のピーク需要に対応するために、ラックあたりの最大消費電力を確保するものです。しかし実際にはワークロードごとに必要な電力は異なり、設定された最大電力の一部が未使用のままになることも珍しくありません。運用側では、障害への備えや柔軟な対応、将来の拡張のために、さらに余剰容量を確保しています。
NVIDIA が調査した代表的な電力予算の視点では、AI 出力に割り当てられる計算リソースが、サイト全体の供給電力のおよそ 60% を占めています。
NVIDIA DSX MaxLPS は、チップ、サーマル、システム、ソフトウェアの技術を統合したスイートで、固定された電力予算内で AI ファクトリーの処理能力を最大化します。MaxLPS とは「Maximum Land Power Shell(最大限の土地・電力・シェル)」の略称です。これは AI ファクトリーを定義するサイトレベルの制約、すなわち敷地、Utility 電源、そして電力・冷却・ネットワーク・計算インフラを支える物理的なシェルを指します。
MaxLPS は、以下の 3 つの層を最適化するために設計されています:
- 動的な電力割り当て: 未使用の電力余裕(ヘッドルーム)を常時監視し、GPU に自動的に割り当てます。
- 高度なパフォーマンス・パー・ワット技術: 固定された電力予算内でジョブレベルのパフォーマンスを向上させるソフトウェアによる電力最適化手法です。
- 45°C のサーマル効率とサイト設計: 温水液冷を採用して冷却オーバーヘッドを削減し、電力使用効率(PUE)を改善。これにより、同じ物理的枠組みの中で計算リソースを増やすことが直接可能になります。
なぜ静的なラック割り当てが電力の無駄を生むのか
従来のデータセンターの電力計画では、すべてのラックが同時に最大定格電力を消費すると仮定して十分な電力を確保しています。これはピーク需要に対する施設側の保護にはなりますが、各ラックを孤立した電力島として扱うことになります。過剰な電力で用意された孤立したラックは、その余剰電力を隣接するラックに貸し出すことができません。その隣のリソースこそが、電力をトークンに変換できる場所です。
AIファクトリー規模では、施設全体のオーバーヘッドやラック内の損失、さらに障害発生時や再起動、チェックポイント作成時の運用上の非効率性が原因で、AI負荷に割り当てられる電力が減少します。また、静的なラック割り当て方式では、ラックに割り当てられた電力の中に遊休分(ヘッドルーム)が生じ、それが固定化されてしまう問題があります。あるラックのピーク需要のために確保された電力がunusedのまま放置されている間に、別のラックはそれを活用できる可能性があります。
動的電力割当は、この回収可能なラックレベルの余剰電力を解消することを目指しています。
図 1 は、100MW の AIファクトリーにおける電力予算のウォータフォールを示しています。グリッドからの入力電力 100MWのうち、20MWが施設オーバーヘッドに割り当てられ、10MWがラック損失として消費され、さらに障害発生時や再起動、チェックポイント作成時の運用上の非効率性により10MWはAI負荷に利用できません。その結果、AI負荷に利用可能な電力は60MWとなります。各減算項目は、元々のグリッド入力に対する割合で表現されています。

図 1. 100 MW の AI ファクトリにおける電力予算のウォータフォールチャート(概略)
トレーニング、ポストトレーニング、推論は、計算バースト、メモリーバウンド実行、同期、チェックポイント作成、プリフェッチ、デコード、アイドルギャップ、ネットワークバウンド通信といった段階を通過します。各フェーズでは電力の消費パターンが異なります。システム安定性の観点から、ラックはワークロードのピーク電力に対応できるように設計されますが、実際のエネルギー消費量は、実質的な期間においてはそのピーク値を下回ります。
NVIDIA DSX MaxLPS の動的電力割当はどのように機能するのか?
NVIDIA DSX MaxLPS は、静的な電力割り当てに代わり動的な電力配分を実現します。これにより、データセンターの固定された電力予算内で、電力利用状況を監視・最適化する Dynamic Power Software (DPS)(現在は開発者プレビュー版)を活用できます。DPS は、電力会社レベルからラック、ノード、GPU に至るまでのデータセンターのトポロジーをモデル化し、包括的な電力管理システムとして機能します。
運用担当者は、リソースグループ、電力予算、および電力の割り当て・制限・強制に関するポリシーを定義します。これらの制約範囲内で、DPS は割り当てられた電力と実際の消費量を常時比較します。GPU やラックが予約されたレベルを下回って稼働している場合、DPS はその余剰分を同じ管理グループ内の他のリソースに割り当てます。サイト全体の電力上限は変更されませんが、利用可能な電力からより多くの生産性を引き出すことが可能になります。
図 2 に示すように、DPS は継続的な制御ループを実行します。システムは GPU、ラック、およびグループレベルの電力テレメトリデータを収集し、未使用の電力容量を特定します。その後、ポリシーに従って電力を再配分し、承認されたグループ電力予算への準拠を検証します。また、電力イベントや緊急時のポリシーが発生した場合は、ベストエフォート型で対応を行います。

この価値は、各ラックを個別の設定として扱うのではなく、ファーム全体にわたる多数の電力決定を調整することによって生まれます。サイト予算が変更された場合や、グリッド障害、メンテナンス、緊急事態などにより利用可能な電力が変動した場合でも、DPS はデータセンター全体の運用制限を自動的に適応させ、オペレーターが手動で各ラックの再計画を行う必要はありません。
図 3 では、静的プロビジョニングと MaxLPS の動的プロビジョニングにおける電力利用率を比較することで、これらの効率化による恩恵を示しています。静的プロビジョニングでは 170 kW の電力が遊休状態となっていますが、MaxLPS の動的プロビジョニングはこの余剰分を回収し、同じ 540 kW というサイト電力予算内で追加のラック展開を可能にします。

DSX Exchange は、AI ファクトリの運用向けに設計されたオープンソースのイベントバス(現在は開発者プレビュー版)です。これにより、DPS やその他の DSX サービスがビル管理システム、電力監視システム、冷却インフラ、グリッドインターフェース、計算スケジューラと連携する IT/OT 間のイベントバスが構築されます。MaxLPS は DSX Exchange を必須としませんが、この統合によって季節ごとの冷却余力や施設の電力イベントといった信号を DPS が検知し、対応できるようになります。
MaxLPS が実現する高効率なパフォーマンス・パー・ワットの実現
ラックレベルでの電源制御に加え、MaxLPS には各 GPU がすでに確保した電力をいかに効率的に活用するかを最適化するソフトウェア機能も含まれています。
DSX MaxLPS には、推論(inference)、トレーニング、メモリバウンド、計算バウンドといった一般的なデータセンター運用モードに対応した、最適化されたワークロードプロファイル電力ソリューション(WPPS)が用意されています。個々のジョブごとに電源、メモリ、周波数、およびノードの動作を調整するのではなく、オペレーターは検証済みのプロファイルを適用することで、コンピューティング動作をワークロードに適合させることができます。
Application Performance and Power Manager (APPM) は、選択された設定を参加する GPU に適用します。また、NVIDIA Dynamo などのソフトウェアを使用すれば、推論サービスにおけるラック間のパフォーマンスや電力動作をさらに最適化できます。このアプローチの核心は AI ファクトリの最適化です。GPU の構成、アプリケーションの動作、そしてサービングトポロジーを整列させることで、ファーム全体のパフォーマンスとワットあたりの出力向上を図ります。
推論はこれを明確に示す好例です。ワットあたりの 1 秒間あたりのトークン数を最大化することは、測定可能なワークロードのスループット増大として直接反映されるためです。同じプロファイルは、トレーニングやポストトレーニングの場面でも適用可能です。
NVIDIA の「Vera Rubin NVL72 AI ファクトリ」において、MaxLPS とデータセンターの電力計画を組み合わせることで、同じ電力予算内で Rubin GPU の容量を最大 40% 増強できると NVIDIA は予測しています。NVIDIA GB200 NVL72 で測定された結果と併せると、これらは動的電力管理がメガワットあたりの AI ファクトリ生産性をいかに向上させるかを示すものです。
図 4 は、代表的な推論ワークロードを用いて評価した MaxLPS の結果を示しています。Vera Rubin NVL72 では DeepSeek-R1 を、GB200 NVL72 では Kimi-K2.5 をそれぞれテストに使用しました。MaxLPS により、GB200 NVL72 でのラック電力は 125 kW から 90 kW に、Vera Rubin NVL72 では 136 kW から 101 kW に削減されます。これにより、ワークロードのスループットを維持したまま、同じ電力範囲内でそれぞれ 39%、35% のラック数を増やすことが可能になります。ワットあたりのパフォーマンスは、GB200 NVL72 で約 1.5 倍、Vera Rubin NVL72 で 1.3〜1.4 倍向上します。

図 4:NVIDIA GB200 NVL72 および GB300 NVL72 システムにおける代表的な推論ワークロードの検証
MaxLPS を活用した AI ファクトリーの規模設計
MaxLPS のインフラ設計は、施設全体の固定された電力容量を起点とし、そこから内側へ向かって計算を進めます。その目的は、MaxLPS 運転ポイントにおいてサイトが支えられる GPU ラックの最大数を決定することです。同時に、長期的な電力・冷却・スペース・ネットワーク容量の目標値もここで設定されます。
初期導入(Day 1)では、このインフラ上限を下回る規模で始めることも可能です。トレーニングやポストトレーニングを主軸とするワークロードからスタートするサイトの場合、平均ラック電力が MaxLPS の平均運転ポイントを上回って運用されることもあります。各ラックの消費電力が多くなるため、固定された施設容量内には初期段階では設置できるラック数が限られます。
したがって、MaxLPS は「最初からすべてのラックを埋めること」を義務付けるものではなく、ライフサイクル全体を通じたキャパシティ目標として定義されます。ハードウェアのライフサイクルを通じてワークロードが推論中心へとシフトしていくにつれ、平均ラック電力は MaxLPS の平均運転ポイントまで低下します。これにより、施設全体の電力容量を増やすことなく、追加のラック設置スペースを確保できる余地(ヘッドルーム)が生まれます。
図 5 に示す通り、MaxLPS の規模設計はこの進化を見据えて初期段階から計画されます。具体的には、GPU プロダクトファミリの選定、45°C DLC 入口運転におけるサイト PUE ターゲットの設定、MaxLPS GPU 数に合わせた東西ネットワークの設計、そして最終的に展開可能な GPU ラック位置数の算出を行います。

図 5:固定された電力枠内で MaxLPS を備えた AI ファクトリの規模を決定するプロセス。総サイト電力から PUE(エネルギー効率指標)で調整した IT 設備用電力へ、ネットワーク容量の減算を経て GPU ラック設置位置数を算出し、最終的に当日からの導入判断に至る流れを示しています。
想定される初日のワークロード構成比は、その設置位置のうち実際に初期に稼働させる数を決定します。MaxLPS の最大容量目標に合わせて、当初から必要なスペース、電力配分、冷却システム、ネットワーク容量をすべて設計しておくことで、運用者はワークロードの進化に応じて GPU やラックを段階的に追加できます。後から施設改修を行う必要はありません。
なぜ 45°C の液体冷却が重要なのか?
データセンターレベルでのソフトウェアによる電力制御は MaxLPS の物語における最大の要素ですが、それだけがすべてではありません。MaxLPS はさらに、チップレベル、熱管理レベル、システムレベルの技術革新にも依存しています。これらの進歩により、施設全体の電力予算のうち、より多くの割合を計算インフラへ直接配分することが可能になります。
「Vera Rubin NVL72 ラックは、45°C の液体冷却入口での運転を想定して設計されています。」(https://blogs.nvidia.com/blog/liquid-cooling-ai-factories/)。温水を使うのは一見すると逆説的に思えますが、その目的は性能、信頼性、寿命の要件を満たしつつ、効率的に熱を排出することです。冷却水の温度を高く保つことで、施設は機械的な冷却装置への依存を減らし、「フリークーリング(自然冷房)」を活用する機会を増やすことができます。これは外気や外部の水を利用して、最小限の機械的冷却で熱を放散させる手法です。
気候やサイト設計によっては、この方法によりエネルギー消費の大きいチラーや水資源を大量に使用する断熱式(蒸発式)クーラーへの依存度を下げることができます。高温時やシステム全体の信頼性を確保するためにはチラーが依然として重要ですが、必要な時にのみ使用することで、冷却に必要な電力全体と年間平均電力使用効率(PUE)を削減できます。
この冷却電力を計算リソースに振り向けることは、運用上の課題であると同時に施設設計の課題でもあります。熱放散システムは最悪の状況を想定して設計されますが、多くのサイトでは通常の運転条件下でチラーコンプレッサの消費電力は大幅に少なくて済みます。特にドライクーラーが負荷の大部分を処理できる場合などです。もし電気配電設備、ドライクーラー、制御システムがこの柔軟性に対応できるように設計されていれば、最悪の時間帯のために確保された冷却用の電力予算が、一年の大半において計算リソースへ回用可能になります。
技術的冷却システムのループを効率的に運用する
技術冷却システム(TCS)ループは、ラックと冷板の間、そして冷却剤分配ユニット(CDU)の間で熱を移動させる技術側の液体ループです。これは、チラーやドライクーラー、冷却塔、あるいはその他の施設設備を通じて熱を放出する広範な施設用冷却ループと共に動作します。
適切に運用された TCS ループは、以下の 3 つの制約条件をバランスよく満たす必要があります。
- GPU の熱負荷に応じて冷却剤の流量を調整し、施設のループ側で放熱の変動を受け入れやすくする
- 設計上の入口温度と熱的信頼性の限界内に留める
- ワークロードに起因する熱事象に対して即座に対応しつつ、システムの不安定化を防ぐ
多くの施設では、CDU の動作管理に比例・積分・微分(PID)制御を採用しています。PID 制御は一般的で有用な手法ですが、反応型であるという特徴があります。つまり、センサーが偏差を検知した後にのみ対応するのです。
大規模な AI ファクトリーでは、熱負荷が急速かつ同期して変化します。そのため、運用担当者は余裕を持たせるために、必要以上にループを低温で運転しているケースが多く見られます。
PID 制御による保守的なバッファは、計算処理に活用できる冷却電力を無駄にしてしまいます。これに対し、Agentic Control Systems(エージェント型制御システム)では、Phaidra の研究 including work by Phaidra に代表されるように、電力テレメトリと学習された制御ポリシーを活用して熱挙動を予測します。45 °C で動作可能なラックにおいて、Agentic Control は必須ではありません。ハードウェアや施設設計が定義する許容熱範囲(thermal envelope)の範囲内で動作し、Agentic Control はその上乗せとしての最適化機能です。
図 6 は、技術冷却ループと施設冷却ループを比較したものです。TCS ループは GPU のコールドプレートから熱を回収し、冷却剤分配装置(CDU)へ運搬します。その後、CDU で熱が施設冷却ループへ移され、施設のポンプやチラー、ドライクーラー、冷却塔などの設備機器によって建屋外へ排熱されます。オプションの監視機能と予測制御を導入することで、さらなる効率化が可能になります。

図 6. ファシリティ冷却ループとテクノロジー冷却システムループの比較
NVIDIA DSX MaxLPS の検証を開始する
一度構築された物理インフラは変更が困難です。既存施設のアップグレードでも、NVIDIA Vera Rubin NVL72 サイトの計画であっても、オペレーターは Day 1 にすべてのラックを配備しなくても、サイトレベルで NVIDIA DSX MaxLPS のリトロフィットまたは設計を検討すべきです。つまり、インフラに関する決定が固定される前に、電力トポロジの検証、冗長性、冷却挙動、テレメトリの利用可能性、ポリシー要件、ネットワーク容量、ラック配置の柔軟性、およびスペースやユーティリティの制約を確認する必要があります。
ソフトウェア導入を検討するチームには、NVIDIA Dynamic Power Software のドキュメント や NVL72 推論電力パイロット が出発点となります。スコープを定めた検証では、管理されていない静的なベースラインと、代表的なワークロードを使用した MaxLPS 管理の実行を比較し、スループット、レイテンシ、サービスエラー率、電力消費量、利用率、ポリシー準拠を追跡します。この結果により、オペレーターは Vera Rubin NVL72 の準備段階で、いつ、どの程度の速度でスケールすべきかを判断できます。
施設チームは、インフラが固定される前に、NVIDIA と早期に連携して 45°C の入口温度での動作検証、MaxLPS ラック配置の選択肢、そして電力配分の柔軟性を確認すべきです。サイトレベルでの設計判断は早く行い、ソフトウェアの検証結果に依存する必要はありません。ソフトウェアテストは並行して進めるか、後から実施しても問題ありません。
固定された電力をより多くの AI 出力に変えるには、AI ファクトリーには新しい運用モデルが必要です。静的なラック割り当てでは、実際のワークロードが有用な AI 出力として活用できたはずの電力が遊んでしまいます。
NVIDIA DSX MaxLPS はこの課題に対し、3 つの側面から解決策を提供します。まず、遊んでいるラック容量を活用する動的電力割当ソフトウェアです。次に、ワットあたりの出力を高めるパフォーマンス向上技術です。そして、冷却オーバーヘッドを削減しつつ、より多くの計算リソースを導入できる物理的な選択肢を維持する 45°C の熱設計と施設要件です。Vera Rubin NVL72 の展開を計画している運用者にとって、MaxLPS は固定された電力予算内で Rubin GPU の容量を最大 40% 増やすための道筋となります。
固定された電力予算内で、Rubin GPU の容量を最大 40% 増強できるよう、Vera Rubin NVL72 サイトの準備を進めましょう。まずは MaxLPS の概要 を読み込み、NVIDIA Dynamic Power Software のドキュメント や NVL72 推論電力パイロット を確認してください。また、45°C の熱設計やラック配置の選択肢、サイトレベルでの MaxLPS 対応状況を検証するためにも、早めに NVIDIA に相談することをお勧めします。
原文を表示
AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available megawatt can deliver. For AI inference workloads, this makes application-level performance per watt the key metric for measuring AI factory efficiency.
Not every megawatt translates to revenue-generating compute. Power distribution, cooling, networking, storage, backup, and facility overhead take a share of the power before it reaches a GPU. Static rack provisioning exacerbates this: an outdated approach to data center design allocates the maximum power draw per rack to meet worst-case peak demand, even though real workloads have different power needs and may leave some portion of that maximum power unused. Operators reserve additional capacity for failures, operational flexibility, and expansion.
In one representative power-budget view examined by NVIDIA, about 60% of delivered site power is allocated to compute for AI output.
NVIDIA DSX MaxLPS is a suite of chip, thermal, system, and software technologies that maximizes AI factory throughput within a fixed power budget. MaxLPS stands for Maximum Land Power Shell, the site-level constraints that define an AI factory: land, utility power, and the physical shell holding power, cooling, networking, and compute infrastructure.
MaxLPS is designed to optimize these three layers:
- Dynamic power allocation: Continuously monitors and allocates unused power headroom to GPUs
- Advanced performance per watt techniques: Software power optimization techniques that improve job-level performance at a fixed power budget
- 45° C thermal efficiency and site design: Cuts cooling overhead through warm-water liquid cooling, improving power usage effectiveness (PUE) to convert directly into more compute within the same fixed envelope
Why static rack provisioning strands power
Traditional data center power planning reserves enough power as though every rack could draw its specified maximum power simultaneously. That protects the facility against peak demand, but it treats each rack as an isolated power island. An isolated rack provisioned with excess power cannot lend that unused power to a neighbor that could turn it into tokens.
At AI factory scale, facility overhead, rack losses, and operational inefficiencies during failures, restarts, and checkpointing reduce the power available to the AI load. Separately, static rack provisioning can strand headroom within the power allocated to racks. Power reserved for one rack’s peak demand may sit unused while another rack could use it. Dynamic power allocation targets this reclaimable rack-level headroom.
Figure 1 shows a power-budget waterfall for a 100 MW AI factory. Of the 100 MW grid input, 20 MW is allocated to facility overhead, 10 MW to rack losses, and 10 MW is unavailable for AI load because of operational inefficiency during failures, restarts, and checkpointing. This leaves 60 MW available for AI load. Each deduction is expressed as a share of the original grid input.

Training, post-training, and inference move through compute bursts, memory-bound execution, synchronization, checkpointing, prefill, decode, idle gaps, and network-bound communication. Each phase draws power differently. For stability, a rack must be provisioned to handle the workload’s peak power draw, though actual energy consumption remains below that level for meaningful periods.
How does NVIDIA DSX MaxLPS dynamic power allocation work?
NVIDIA DSX MaxLPS replaces static power provisioning with dynamic power allocation, using Dynamic Power Software (DPS) (currently in Developer Preview) to monitor and optimize power utilization within the data center’s fixed power budget. DPS is a comprehensive power management system that models the data center topology from the utility level down to racks, nodes, and GPUs. Operators define resource groups, power budgets, and policies that govern how power can be allocated, constrained, and enforced.
Within those boundaries, DPS continuously compares allocated power against actual consumption. When GPUs or racks operate below their reserved level, DPS makes that headroom available to others in the same managed group. The site power envelope remains unchanged: DPS extracts more productivity from the available power.
DPS runs a continuous control loop, as shown in Figure 2. The system collects GPU-, rack-, and group-level power telemetry, identifies unused power capacity, reallocates power within policy, validates compliance with the approved group power budget, and responds to power events or emergency policies on a best-effort basis.

The value comes from coordinating many power decisions across the fleet rather than treating every rack as a one-off configuration. When the site budget changes, or when a grid, maintenance, or emergency event alters available power, DPS adapts operating limits across the data center without forcing operators to re-plan every rack by hand.
Figure 3 illustrates these efficiency gains by comparing the power utilization of static provisioning with that of MaxLPS dynamic provisioning. Static provisioning strands 170 kW of power, while MaxLPS dynamic provisioning reclaims this headroom, enabling an additional rack deployment within the same 540 kW site power budget.

DSX Exchange is an open source event bus for AI factory operations (currently in Developer Preview). It adds an IT/OT event bus connecting DPS and other DSX services to building management systems, electrical power monitoring systems, cooling infrastructure, grid interfaces, and compute schedulers. MaxLPS does not require DSX Exchange to function, but the integration exposes signals such as seasonal cooling headroom and facility power events that DPS can act on.
MaxLPS techniques for advanced performance per watt
Beyond rack-level power steering, MaxLPS also includes software features that optimize how each GPU uses the power it already has.
DSX MaxLPS includes optimized workload profile power solutions (WPPS) for common data center operating modes, including inference, training, memory-bound, and compute-bound configurations. Rather than tuning power, memory, frequency, and other node behavior for every job, operators apply validated profiles that align compute behavior with the workload.
The Application Performance and Power Manager (APPM) applies the selected configuration to participating GPUs, while software such asNVIDIA Dynamo can further optimize inter-rack performance and power behavior for inference services. The principle is AI factory optimization: align GPU configuration, application behavior, and serving topology to increase fleet-wide performance and output per watt.
Inference provides a clear example because maximizing tokens per second per watt translates the gain into measured workload throughput. The same profiles apply to training and post-training.
For NVIDIA Vera Rubin NVL72 AI factories, NVIDIA projects that MaxLPS, combined with data center power planning, can enable up to 40% more Rubin GPU capacity within the same power budget. Together with measured results on NVIDIA GB200 NVL72, these results demonstrate how dynamic power management can increase AI factory productivity per megawatt.
Figure 4 shows MaxLPS results evaluated using representative inference workloads: Vera Rubin NVL72 was tested with DeepSeek-R1, while GB200 NVL72 was tested with Kimi-K2.5. MaxLPS reduces provisioned rack power from 125 kW to 90 kW on GB200 NVL72 and from 136 kW to 101 kW on Vera Rubin NVL72, enabling 39% and 35% more racks, respectively, within the same power envelope while preserving workload throughput. Performance per watt improves approximately 1.5x on GB200 NVL72 and 1.3–1.4x on Vera Rubin NVL72.

How to size an AI factory site for MaxLPS
MaxLPS infrastructure design starts with a fixed gross facility power envelope and works inward to determine the maximum number of GPU rack positions the site can support at the MaxLPS operating point. It also establishes the long-term infrastructure target for power, cooling, space, and network capacity.
Day 1 deployment can be lower than this infrastructure limit. A site that begins with a training- or post-training-heavy workload may operate at an average rack power above the MaxLPS average operating point. Because each populated rack consumes more power, fewer racks can initially be deployed within the fixed facility envelope.
MaxLPS therefore defines the lifecycle capacity target, not a requirement to populate every rack from the start. As the workload mix shifts toward inference over the hardware lifecycle, average rack power can decline toward the MaxLPS average operating point, creating headroom to populate additional rack positions without increasing the facility power envelope.
As Figure 5 shows, MaxLPS sizing anticipates this evolution from the start by selecting the GPU product family, setting the site PUE target at 45° C DLC inlet operation, sizing the east-west network for the MaxLPS GPU count, and deriving the total deployable GPU rack-position count.

The expected day-one workload mix then determines how many of those positions are initially populated. With the associated space, power distribution, cooling, and network capacity configured up front for the full MaxLPS capacity target, operators can add GPUs and racks incrementally as workloads evolve without a later facility retrofit.
Why is 45° C liquid cooling important?
Software power steering at the data center level is the largest part of the MaxLPS story, but it is not the whole system. MaxLPS also depends on chip-, thermal-, and system-level advances that allow more of the facility power budget to reach the compute infrastructure.
Vera Rubin NVL72 racks are designed for 45° C liquid-cooling inlet operation. Warmer liquid may sound counterintuitive, but the goal is to remove heat efficiently while meeting performance, reliability, and lifetime requirements. Higher coolant temperatures can enable facilities to rely more often on “free cooling,” which uses outside air or water to reject heat with minimal mechanical chilling.
Depending on the climate and site design, this method can reduce reliance on energy-intensive chillers or water-intensive adiabatic (evaporative) coolers. Chillers remain important during hot conditions and for resilience, but using them only when needed reduces overall cooling power and average annual power usage effectiveness (PUE).
Recovering that cooling power for compute is a facilities design problem as much as an operating one. Heat rejection systems are sized for worst-case scenarios, yet many sites require far less chiller compressor power under typical operating conditions, especially when dry coolers can handle a larger portion of the load. If the electrical distribution, dry coolers, and controls are sized for that flexibility, power budgeted for cooling during the worst hour becomes available to compute during much of the year.
Operating the technology cooling system loop efficiently
The technology cooling system (TCS) loop is the technology-side liquid loop that moves heat between racks, cold plates, and the coolant distribution unit (CDU). It works alongside the broader facility cooling loop, which rejects heat through chillers, dry coolers, cooling towers, or other site equipment. A well-operated TCS loop balances three constraints:
- Adjusting the coolant flow rate to match GPU thermal demand, allowing the facility loop to absorb variation on the heat rejection side
- Staying within the designed inlet-temperature and thermal-reliability limits
- Responding to workload-driven thermal events without creating instability
Many facilities use proportional-integral-derivative (PID) control to manage CDU behavior. PID control is common and useful, but it is reactive—it responds only after sensors report a deviation. Large AI factories create fast, synchronized changes in thermal load, so operators often run the loop colder than necessary to preserve margin.
A conservative PID-controlled buffer, as described, wastes cooling power that could otherwise support compute. Agentic control systems, including work by Phaidra, use power telemetry and learned control policies to anticipate thermal behavior. A rack with 45 °C thermal efficiency does not require agentic control to operate: the hardware and facility design define the supported thermal envelope, and agentic control is an optimization on top of that.
Figure 6 shows the technology cooling loop versus the facility cooling loop. The TCS loop carries heat from GPU cold plates to the coolant distribution unit (CDU), which then transfers it to the facility cooling loop. Facility pumps, chillers, dry coolers, cooling towers, or other site equipment then reject the heat outside the building. Optional monitoring and predictive controls can help optimize efficiency.

Get started with NVIDIA DSX MaxLPS validation
Physical infrastructure is difficult to change once built. Whether upgrading an existing facility or planning NVIDIA Vera Rubin NVL72 sites, operators should retrofit or design for NVIDIA DSX MaxLPS at the site level, even if they do not deploy every rack on Day 1. That means validating power topology, redundancy, cooling behavior, telemetry availability, policy requirements, network capacity, rack-position optionality, and space or utility constraints before infrastructure decisions are fixed.
For teams evaluating the software path, the NVIDIA Dynamic Power Software documentation and NVL72 inference power pilot provide a starting point. A scoped validation compares an unmanaged static baseline against a MaxLPS-managed run using representative workloads, tracking throughput, latency, service error rate, power draw, utilization, and policy compliance. The results show operators when and how fast to scale as they prepare for Vera Rubin NVL72.
Facilities teams should engage NVIDIA early to validate site thermal design for 45° C inlet operation, MaxLPS rack-position optionality, and power distribution flexibility before infrastructure is fixed. The site-level design decision should be made early and does not depend on software validation. Software testing can proceed in parallel or follow at a later stage.
To turn fixed power into more AI output, power-limited AI factories need a new operating model. Static rack provisioning strands power that real workloads would otherwise turn into useful AI output.
NVIDIA DSX MaxLPS addresses this on three fronts: dynamic power allocation software that unlocks stranded rack capacity, performance-per-watt techniques that raise output per watt, and 45° C thermal and facility design that reduces cooling overhead while preserving the physical optionality to land more compute. For operators planning Vera Rubin NVL72 deployments, MaxLPS is a path to prepare sites for up to 40% more Rubin GPU capacity within fixed power budgets.
Prepare Vera Rubin NVL72 sites to provision up to 40% more Rubin GPU capacity within your fixed power budget. To get started, read the MaxLPS Overview, explore the NVIDIA Dynamic Power Software documentation, and the NVL72 inference power pilot. Engage NVIDIA early to validate the 45° C thermal design, rack-position optionality, and site-level MaxLPS readiness.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み