DeepSeek-V4-Pro の高速配信技術と H20 GPU での実装知見
本文の状態
日本語全文を表示中
詳細モードで約42分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
LMSYS Blog
LMSYS は DeepSeek-V4-Pro の大規模モデルを H20 GPU で効率的に運用する手法を発表し、ハードウェアの性能差をシステム最適化で縮小する成果を示した。
AI深層分析を開く2026年8月20日 01:48
AI深層分析
キーポイント
多様なサービスプロファイルの必要性
単一の設定では全てのワークロードに対応できないため、コンテキスト長やレイテンシ要件に応じてプリフェッチとデコードのプロファイルを切り替える戦略が示された。
ハードウェア制約下での最適化手法
NVIDIA Blackwell のような高性能GPUを備えていない H20 GPU においても、Attention-CP8 や MoE-TP8 の通信経路最適化により性能を引き出す技術が詳述されている。
H20 と B300 の性能比較と縮小
バッチサイズ1での単一ノード出力速度において H20 は B300 を下回るものの、最適化により両者のデコード性能比を 1.42 倍まで縮小することに成功した。
高スループットと長文コンテキストの達成
最適化されたプリフェッチではノードあたり 8.45k トークン/秒、100 万トークンのプロンプト処理を 43.7 秒で完了させるなど、スループット指向の設定でも高い数値を記録した。
最適化されたプロファイルごとのパフォーマンス
最適化されたプレフィルでは1ノードあたり8.45k入力トークン/秒、100万トークンのプロンプト処理に43.7秒を要する。スループット指向のデコードではDP16-EP16構成で1ノードあたり4.67k出力トークン/秒、平均TPOTは27.4msとなる。
重要な引用
One model needs multiple serving profiles.
Despite the substantial hardware gap, workload-specific system optimization narrows the observed decode performance ratio to 1.42x.
Optimized prefill reaches 8.45k input tokens/s per node and processes a 1M-token prompt in 43.7 seconds.
The contribution is a methodology, not a single benchmark.
編集コメントを表示
編集コメント
H20 GPU のような制約のある環境でも、システムレベルの最適化によって性能差を劇的に縮められることを示した実証研究である。大規模モデルの導入においてハードウェア選定だけでなく、ソフトウェアスタックの設計が成否を分ける重要な事例と言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
1. はじめに
DeepSeek-V4-Pro は、FP8 と FP4 の両方の重みに対応した、1.6 トリリオンパラメータの Mixture-of-Experts (MoE) モデルです。この規模のモデルは、HBM の容量が広く、計算スループットが高く、ネイティブな FP4 Tensor Core を備えた NVIDIA Blackwell GPU などのアクセラレータから恩恵を受けるのが自然です。しかし、それらの利点を欠いているにもかかわらず、H20 GPU は依然として広く展開されています。
ハードウェアの制約は、サービス要件を緩和するものではありません。長いコンテキストを持つプリフィルでは、最初のトークンまでの時間 (TTFT) を厳密に制御する必要があります。インタラクティブなデコードでは、各サービスティアが設定した 1 トークンあたりの出力時間 (TPOT) の目標値を満たす必要があります。持続的なトラフィックでは、集約スループットと KV キャパシティのバランスを取る必要があります。短い入力、長いコンテキスト、レイテンシに敏感なリクエスト、そして高い同時接続数は、それぞれ異なる形でシステムを圧迫します。これらすべてをうまく捌ける万能な設定は存在しません。
1 つのモデルに対して、複数のサービスプロファイルが必要です。 ワークロードの特徴、サービスレベル目標 (SLO)、そして測定されたハードウェアの挙動が組み合わさり、デプロイメントトポロジと実行パスを決定します。
- ワークロードに合わせたサービスプロファイルの選定。 ここで評価した構成では、プリフィルは測定されたコンテキスト長の範囲に応じて PP2 と PP4 の間で選択され、デコードは異なるレイテンシ、スループット、KV キャパシティの目標に合わせて最適化されたプロファイルを使用します。
- プリフィルパスの最適化。
Attention-CP8 → MoE-TP8とコンテキスト並列通信を最適化した上で、長いコンテキストと短いコンテキストのワークロードが実際に生み出すルーティング形状に合わせてチューニングを行います。
- デコードパスの最適化 DSpark の推論(speculative decoding)パスを最適化し、実行効率、エキスパートルーティング、通信と計算のオーバーラップを、異なるデコード SLO に合わせて微調整しました。
- レイテンシの限界突破 バッチサイズ 1 の条件下で、単一ノードの H20(141GB)リファレンス環境では 271 トークン/秒 の出力速度を達成しました。これは B300 で報告された 383.7 トークン/秒 LMSYS による 2026 年 7 月 6 日のブログ記事と比較した数値です。ハードウェアの性能差は依然として大きいものの、ワークロードに特化したシステム最適化により、観測されるデコード性能の比率は 1.42 倍 にまで縮小されました。詳細なベンチマーク設定と、ログベースのスループット抽出手法については 付録 B.3 をご覧ください。
- サービング能力の最大化 レイテンシの結果は、システムが到達できる性能の一つの極限に過ぎません。より広範なプロファイルファミリー全体では、最適化された事前処理(prefill)により、1 ノードあたり 8.45k トークン/秒 の入力速度を達成し、100 万トークンのプロンプト を 43.7 秒 で処理します。スループット重視のデコードにおいては、DP16-EP16 の構成で効率化されたリファレンス環境が、1 ノードあたり 4.67k トークン/秒 の出力速度を達成し、平均 TPOT は 27.4 ms となります。これらの結果は異なるプロファイルから得られたものであり、それぞれコンテキスト長、レイテンシ、スループット、キャパシティ制約の組み合わせに応じて選択・最適化されたものです。
この貢献は単一のベンチマークではなく、一つの手法論です。シナリオ固有のサービス(サービング)により、利用可能なハードウェア上で評価されたプロファイルの中から、各ワークロードがより最適な運用ポイントへ向かうことが可能になります。ここでは紹介したデプロイメントの選択肢、最適化手法、および測定結果が、計算資源、メモリ帯域幅、ネットワーク接続に制約のある環境で最先端モデルをサービスするチームにとっての実践的な参考になれば幸いです。
2. ハードウェア制約からサービスプロファイルへ
2.1 ハードウェア制約とサービスの役割
図 1. ハードウェアの格差:H20 と B300 の比較。
Blackwell は圧倒的な性能を提供しますが、H20 は展開規模において優れています。B300 はネイティブの FP4 Tensor Cores を備え、FP8 スループットは大幅に向上し、HBM 容量も格段に大きくなっています。一方、H20 は計算能力では B300 に及びませんが、大規模展開が可能であり、高いメモリ帯域と 900 GB/s の NVLink を提供します。
本研究の各ノードには、NVLink で接続された 8 つの GPU が搭載されています。Prefill(初期処理)はリクエストごとの永続的な状態を保持しないため、ハードウェア選定は主に TTFT(Time to First Token)、計算効率、通信効率が決定します。一方、Decode(生成処理)では、生成中にアクティブなすべてのリクエストの KV キャッシュを保持する必要があり、HBM 容量がコンテキスト長と同時実行性の直接的な制限となります。
本研究で検討したデプロイメントにおいては、この要件に基づき、Prefill には H20-96GB(容量がワークロードに十分であったため)を採用し、Decode には KV キャッシュ容量がボトルネックとなる H20-141GB を使用しました。
2.2 Capacity Choices
サービングの容量は、最終的には共有された HBM バジェットに依存します。モデル重みと各リクエストごとの KV 状態が同じメモリを競合するためです。ここでは「フルトークン容量」を、モデル重みとランタイムバッファの割り当て後に各ランクが保持できる最大数のフルアテンション KV トークンとして定義しています。これは直接バッチサイズの保証ではなく、メモリの上限を示すものです。
Humming MXFP4AFP8 で重みのフットプリントを削減する
まずは重みのフットプリントを減らすことです。 Humming MXFP4AFP8 は、ネイティブ FP4 Tensor Core を持たない H20 GPU において、MXFP4 エキスパート重みとオンライン FP8 アクティベーションを活用することで、重みのフットプリントと HBM 上のメモリトラフィックを削減します。SGLang との統合は sglang#23754 で利用可能です。Humming と SGLang の統合については、別の記事で詳しく取り上げます。モデルレベルでの精度結果と公開参照測定値は Appendix D.2 に記載されています。
オンライン C128 で KV 容量を拡大する
KV キャッシュに成長の余地を与えます。 オフライン C128 ベースラインでは、圧縮されたページごとにインデックスごとの状態を保持しますが、オンライン C128 はコンパクトな集約状態のみを維持し、より多くの HBM を KV キャッシュプールへ解放します。これにより追加の状態管理と推測検証の負荷が生じますが、テストでは TPOT の低下は見られませんでした。
容量向上効果の統合
容量向上効果は、重みと KV ステートの両方で複合的に現れます。 Humming MXFP4AFP8 は重みのフットプリントを削減することで、DP32-EP32 構成において Baseline FP8 + Offline C128 と比較してフルトークン容量を1.71 倍に拡大し、PP2-TP8 構成では4.47 倍に引き上げます。さらに Online C128 は C128 の補助状態領域を削減するため、Humming 導入後にさらに2.268 倍の増加をもたらします。これら 2 つの技術を組み合わせることで、DP32-EP32 ではベースラインに対して3.88 倍、PP2-TP8 では10.14 倍の容量向上を実現します。詳細なデータは付録 D.1 をご覧ください。
2.3 シナリオ別サービングプロファイル
Prefill プロフィール
最適なパイプライン深度は、パイプライン処理すべき作業量によって決まります。 PP2-CP8-TP8 と PP4-CP8-TP8 は、どちらも Attention-CP8 → MoE-TP8 という同じ実行パスを共有しています。トポロジーレベルでの主な違いはパイプライン深度です。PP2 はモデルを 2 つのステージに分散させるのに対し、PP4 では 4 つのステージを使用します。
短いコンテキストではパイプラインのオーバーヘッドを抑えることが優先され、長いコンテキストでは並列性をより多く引き出すことができます。入力データが短いとチャンク数が少なくなるため、深いパイプラインが十分に満たされず、充填や排気、ステージ間の転送にかかるコストが目立ってしまいます。一方、長いコンテキストでは十分な数のチャンクが存在するため、4 つのステージをすべて稼働させることが可能になります。各ステージのレイヤー数を減らすことで、追加されたノードはプレフィル処理における並列性を高めることに寄与します。
私たちのデプロイメントにおいては、これらの特性に基づき、短いコンテキストには PP2-CP8-TP8 を、長いコンテキスト向けワークロードには PP4-CP8-TP8 を採用しました。
低遅延デコードのプロファイル
低遅延を実現するには、実行パスを最短にすることが第一歩です。Single-node TP8 と PP2-TP8 は、どちらも Attention-TP8 → MoE-TP8 という同じ実行パスを採用しています。両者の違いは、モデルが単一のノード内に収められているか、複数のノード間に分割されているかという点にあります。
Single-node TP8 では、すべてのレイヤーを 1 つの H20-141GB ノード上に配置し、ステージ間の通信や同期を不要とします。一方、PP2-TP8 はモデルを 2 つのパイプラインステージに分割して実行します。
最も高速なトポロジーが、必ずしも最適なサービス形態とは限りません。単一ノードでの TP8 構成は実行パスが短く有利ですが、モデルの重みとサービング状態が同じノードの HBM を共有するため、KV キャッシュに割り当てられる領域に限界があります。そのため、長いコンテキストと大規模なバッチサイズを同時にサポートすることはできません。一方、PP2-TP8 はパイプラインオーバーヘッドが増加しますが、モデル重みを 2 つのノードに分散させることで、KV 状態のために利用可能な HBM を増やします。我々のレイテンシとキャパシティの目標を達成するため、バッチサイズ 1 のレイテンシ基準として単一ノード TP8 を採用し、低レイテンシサービングプロファイルには PP2-TP8 を使用しています。
高スループットデコードプロファイル
*図 6. 高スループットデコード:DP16-EP16 リファレンスと DP32-EP32 キャパシティプロファイル。
高スループットデコードでは、データ並列(DP)とエキスパート並列(EP)を同時に拡張します。 両方のプロファイルとも Attention-DP → MoE-EP の実行パスを使用しています。DP16-EP16 は最小の展開単位であり、DP32-EP32 は同じトポロジー内で DP と EP の両方を拡大した構成です。
スケールアウトは、1 GPU あたりの処理効率よりもリクエストの収容能力を優先します。より大きな EP(Expert Parallelism)グループを設定することで、エキスパート重みを複数の GPU に分散し、HBM を KV キャッシュ用に解放して、同時に処理できるリクエスト数を増やすことができます。一方で、各ノード内で完結する MoE トラフィックの割合は小さくなり、ノード間を跨ぐトラフィックの割合が大きくなるため、1 GPU あたりの効率が低下する可能性があります。
ここで評価したプロファイルでは、最小デプロイ単位および効率の基準として DP16-EP16 を使用し、リクエスト容量の拡大には DP32-EP32 を採用しています。
3. Prefill: コンピューティングと通信のバランス
Prefill(初期処理)のパフォーマンスはシステム全体の課題です。 エキスパートの不均衡、コンテキスト並列による通信、そして本番環境でのルーティングが複合的に TTFT(Time To First Token)を決定します。特定のカーネルだけを最適化しても不十分です。
3.1 なぜ MoE-TP が MoE-EP よりも優れているのか
通信量が少なくても、処理が遅くなる可能性があります。 MoE-EP ではルーティングされたトークンだけを交換しますが、実際の Prefill トラフィックではエキスパートに大きな偏り(スケー)が生じます。ホットなエキスパートを保持するランクは計算量が増え、ストレイガラー(遅延要因)となります。一方、他のすべてのランクは結合ステップで最も遅いパスを待たなければなりません。通信量が減っても、それがそのまま TTFT の短縮にはつながらないのです。
通信の最小化よりも、計算リソースのバランスを優先してください。ここで評価した H20 のプリフィルワークロードでは、PP2 と PP4 の両方で MoE-TP が採用されています。全シーケンスのオールギャザーとリデューススキャッターは通信量を増加させますが、そのトラフィックは高帯域幅の NVLink 上で行われるため、コストは安定しており予測可能です。すべての TP ランクが同じルーティングされたトークンに対してテンソル並列計算を実行するため、専門家の偏りがランクレベルでの長時間遅延(ロングテール)を引き起こすことはありません。このワークロードにおいては、予測可能な通信の方が、予測不可能な不均衡よりも安価です。
実装は sglang#24947 で公開されています。
3.2 プリフィル集合操作の高速化と融合
再利用可能な集合的高速パスを構築する。 MoE-TP は予測不可能な専門家の不均衡を、予測可能な集合通信に置き換えることで、通信効率を次のボトルネックとして浮き彫りにしました。私たちは TP と CP で対称的なメモリを再利用可能にし、AllReduce、AllGather、ReduceScatter が登録バッファによる高速パスと Hopper 向けのアクセラレーションを共有できるようにしました。この基盤となる上流の取り組みには、メモリープール所有権、コミュニケーター登録、MoE-TP 集合バッファ、そして CP Attention と KV-cache のバッファパスが含まれます。
次に、プリフィルのクリティカルパスを短縮する。 高速な集合通信だけでは、通信と計算の境界は消えません。32K の単一チャンクケースでは、コピーエンジン駆動の AllGather を FP8 量子化と共有専門家 GEMM と融合させ、その後に TopK 削減、共有専門家の加算、ReduceScatter を第 2 の Triton カーネルで結合する融合パスを構築しました。これにより 7 つの演算子を 3 つの実行グループに再編成し、PP4 の A/B テストで TTFT を約 3.5% 削減することに成功しました。
3.3 リアルなルーティング形状に向けた Humming のチューニング
汎用的なチューニングでは、重要な形状を見逃します。 Prefill ルーティングではトークンが 384 のエキスパートに不均等に分散されるため、実効的な M 次元は少数の離散値にクラスタリングされます。W13 と W2 も異なる形状で動作するため、単一の汎用的なヒューリスティックでは両方のパスを最適化できません。
本番環境のルーティングからチューニングを行います。 実際のルーティングヒストグラムから高頻度の形状を抽出し、W13 と W2 用に個別の正確な形状設定を構築します。その後、カーネルレベル、パイプラインステージレベル、マッチングされた A/B テストレベルで検証を行います。最適化の対象は M の合成範囲ではなく、実際にサービスしているルーティング分布そのものです。PP4 のマッチング A/B 実験(32K)では、選択された MoE カーネルのレイテンシが約 21% 短縮され、これがエンドツーエンドの TTFT を 11.35% 削減することにつながりました。
4. Decode: Optimizing Speculation and MoE Execution
デコード最適化は、実装においてプロファイル固有のものとなります。 PP2-TP8 では推論パイプラインステージ間の調整が必要ですが、DP32-EP32 は高並行下でのリファインメントステップとエキスパートルーティングの最適化に焦点を当てています。Humming の融合とオーバーラップは、これらのサービングトポロジの下にある共有 MoE ホットパスを改善します。
4.1 Low-Latency PP2-TP8: Extending DSpark Across Pipeline Stages
Pipeline parallelism splits the speculative loop. In PP2-TP8, target execution spans two pipeline stages, while the DSpark drafter resides only on the final stage. Stage 0 sends target hidden states to Stage 1, which performs verification, accepts tokens, and generates candidates for the next round.
Make two stages advance as one. Every speculative round crosses the pipeline boundary. We coordinate both stages and the required intermediate transfers under one execution protocol, preventing the stages from entering different rounds while avoiding redundant synchronization. The PP-specific DSpark integration is being upstreamed in sglang#32281.
4.2 High-Throughput DP32-EP32: Removing High-Concurrency Bottlenecks
本セクションの A/B 比較結果は、4K レゾリューションで DP32-EP32 の構成を用い、DP ランクあたり 32 件の同時リクエストを処理した条件下でのものです。
微調整には適切な実行形状を選択する。 微調整ステップでは、DSpark が生成した候補セットの再評価のために全語彙への投影が行われます。高並列環境では、行ごとのドット積計算で各アクティブな行に対して語彙重みを繰り返し読み込むため、デコードステップごとに持続的な遅延(テール)が発生します。そこで、アクティブな行を 1 つの転置された GEMM 演算に統合し、冗長なメモリアクセスを削減して微調整パスを短縮しました。その結果、GPU あたりのスループットは22.8% 向上しました。
測定されたルーティングに基づいてエキスパートを配置する。 DSpark のトラフィックには顕著なエキスポートの偏りが見られます。代表性のあるリクエストからルーティングの親和性を記録し、それを用いてエキスパート並列負荷分散(EPLB)や冗長エキスパートの設定を行います。これにより、一部のホットなエキスパートがクリティカルパスを繰り返し延長する問題を防止できます。GPU あたりのスループットは13.5% 向上しました。
4.3 Humming Decode Hot Path: Fusion and Overlap
これらの最適化はサービングトポロジーの下層に位置し、Humming ベースのデコードプロファイルで再利用可能です。以下のマッチング結果は、DP32-EP32 構成で 4K のコンテキスト長、各 DP ランクあたり 32 の並行リクエストを想定しています。
追加の量子化パスを削除する。 SwiGLU 活性化関数と量子化を融合させ、融合カーネルが W2 で必要なデータとスケールを直接生成するようにしました。これにより中間バッファへの繰り返しアクセスが不要となり、独立した量子化パスも排除されるため、W2 の処理を早期に開始できます。マッチングされた DSpark の A/B テストでは、GPU あたりのスループットが44.0%向上しました。
通信と W2 の実行を重畳させる。 以前の研究成果(Single-Batch Overlap (SBO)、sglang#9660)で提案した SBO メカニズムを、Humming-Aware SBOとして適応させました。タイルごとのシグナルにより、DeepEP は W2 の出力タイルが完成するのを待たずに、即座に対応する結合送信を開始できます。これは GEMM 全体の完了を待つ必要がないことを意味します。同じ運用ポイントでの初期のマッチングされた非スペキュレーション A/B テストでは、SBO は FP8 トランスポート層と比較してスループットを4.12%回復させました。
5. 評価:システム全体の向上とプロファイルのトレードオフ
5.1 プリフィル:累積的な向上とコンテキスト長のトレードオフ
PP2 は短文脈プロファイルの強化に寄与します。PP2 はすべての 9 つの入力長において性能が向上し、幾何平均スループットは36.5%増、ピーク総入力スループットは16,900 トークン/秒を達成しました。浅いパイプライン構造により短くリクエストの充填・排出オーバーヘッドが削減され、PP2 は少ないリソースで低い TTFT を維持できます。
PP4 はこの利点を長文脈へと引き継ぎます。PP4 も同様に 9 つのポイント全体で幾何平均スループット31.8%の向上を実現しました。文脈長が伸びるにつれて深いパイプラインには十分な作業量が確保され、固定コストを分散できるようになります。その結果、512K では総入力スループットが25,860 トークン/秒に達し、1M でも23,970 トークン/秒を維持します。
文脈長の変化が PP2 と PP4 のトレードオフを左右します。PP4 を基準とすると、PP2 は 4K で TTFT が16.7%低く、32K では19.5%低い値を示します。一方、8K、16K、64K では両者の差は2%以内にとどまります。しかし 128K から PP4 が決定的な優位性を示し始め、PP2 と比較して TTFT をそれぞれ26.2%、33.3%、42.1%、44.8%低下させます(128K、256K、512K、1M の順)。このため、ルーティングの境界点は普遍的な交差点ではなく、測定された文脈長の範囲に基づいた運用ポリシーとして扱います。
完全な TTFT と総入力スループットの結果は付録 A.1–A.2 に記載されています。
5.2 ローレイテンシデコード:パフォーマンスとキャパシティのトレードオフ
Optimized DSpark は、レイテンシの基準値を劇的に引き下げました。 図 15 に示す 4 つの入力長すべてにおいて、バッチサイズ 1 ではピーク TPOT が 74.8%–78.0% 削減されています。また、各測定ペアで共通する最大のバッチサイズにおいても、削減幅は 52.2%–60.0% を維持しています。この効果は短文脈や単一リクエストの実行に限定されるものではなく、8K から 1M の広範な入力長にわたって持続します。
観測されたサービング性能は、ピーク計算リソースの比率だけでは示唆できないほど高いレベルにあります。図 16 に示す 4 つの入力長において、PP2-TP8 の構成で最適化された DSpark はバッチサイズ 1 で 150–174 tokens/s を達成します。一方、単一ノードの TP8 リファレンスでは 183–271 tokens/s に達しています。
実際の実行パスで使用される精度を考慮すると、B300 は H20-141GB(B300 の FP4 対 H20 の FP8)と比較して Tensor Core のピーク計算能力が約 45.6 倍、メモリ帯域幅は 1.67 倍 と圧倒的に上回っています。しかし、観測された最大の生成レートは B300 で 383.7 tokens/s on B300、H20-141GB では 271 tokens/s であり、その比率は 1.42 倍 に過ぎません。このように強力なハードウェアリファレンスと比較しても、ワークロードに特化した最適化により、H20-141GB の観測サービング性能は大幅に向上し、B300 に迫る結果となっています。
本番環境の目標を達成するには、PP2-TP8構成が容量面で有利です。単一ノードでの TP8 は高速ですが、コンテキスト長 1M の場合、KV キャパシティはバッチサイズ 1 にしか対応できません。より大きなバッチや並行リクエストを受け付けることは不可能です。
一方、モデル重みを 2 つのパイプラインステージに分散する PP2-TP8 は、それぞれコンテキスト長 1M、512K、256K でバッチサイズ 4、8、16 をサポートします。Online C128 を採用した場合、トークンあたりの最大容量は11.04M tokens/rankに達します。
同様のコンテキスト長と並行処理能力を目標とする場合、レイテンシの基準としては単一ノード TP8 を維持しつつ、低レイテンシでのサービスプロファイルとしてPP2-TP8を採用することを推奨します。詳細なパフォーマンスおよび容量データは付録 B および付録 D.1 に記載されています。
5.3 高スループットデコード:最前線の成果とプロファイルのトレードオフ
図 17 は、システムが進化するにつれてスループットと対話性のフロンティアがどのように変化するかを示しています。横軸はユーザーあたりのトークン数/秒(tokens/s/user)で表される対話性、縦軸は GPU あたりのトークン数/秒(tokens/s/GPU)で表されるスループットです。DP/EP のプロファイルでは、各 DP ランクが 1 つの GPU にマッピングされます。対話性は「GPU あたりのスループット」を「DP ランクあたりの同時リクエスト数」で割った値として定義されます。右上方向に位置する点は、ユーザーが目にする生成速度と GPU の効率性のバランスが優れていることを意味します。4 つの曲線は、セクション 4 で紹介された個別の最適化による単なる改善ではなく、システム全体の累積的な進化を表しています。
MTP はマルチトークン予測(Multi-Token Prediction)を指し、(3, 1, 4)構成では、推測ステップが 3 回、top-k が 1、ドラフトトークン数が 4 つとなります。
システム最適化はフロンティア全体を押し上げます。 DP ランクあたり 32 の同時リクエストがある 4K 入力の場合、GPU あたりのスループットは 319.92 tokens/s/GPU から 703.15 tokens/s/GPU に向上し、2.20 倍の増加となります。一方、DP ランクあたり 1 のリクエストしかない 1M 入力では、27.05 tokens/s/GPUから66.82 tokens/s/GPUへと上昇します。最初の 3 つのシステムマイルストーンは、1M 入力において DP ランクあたり最大 1 リクエストしか処理できませんでしたが、最終的なシステムでは 4 リクエストに対応可能となり、スループットは 177.48 tokens/s/GPU に達しました。この運用範囲の拡大は、実行速度の向上とキャパシティの増大の両方によるものです。
図 18: GPU ごとのスループット(DP16-EP16 vs. DP32-EP32)
小規模なデプロイメント単位は、高い同時実行性を保つポイントで効率を維持します。
DeepSeek-V3/R1 の H20 でのサービングに関する以前の調査 [1] で、より小さな EP(Expert Parallelism)デプロイメント単位を使用すると、各ノード内の MoE(Mixture of Experts)トラフィックの大きな割合を保持できることがわかりました。DeepSeek-V4-Pro でも同様の利点が図 18 に示された動作ポイントで確認されています。DP ランクあたり 16 および 32 の同時リクエストの場合、DP16-EP16 は DP32-EP32 よりも GPU ごとのスループットが約 3.6%–20% 高い結果を示しました。
すべての同時実行レベルで完全なスウィープが単調であるわけではないため、DP16-EP16 を DP32-EP32 の普遍的な代替品とするのではなく、効率の基準として採用しています。
図 19: DP ランクあたりの長文コンテキストリクエスト容量
容量構成によって、高スループットを実現する最適なプロファイルは変化します。DP16-EP16 は GPU 単体あたりの効率は高いものの、DP32-EP32 は専門家の重みをより多くのランクに分散させ、KV キャッシュ用にさらに多くの HBM を確保できます。
コンテキスト長が 256K、512K、1M の場合、DP ランクごとの最大同時リクエスト数はそれぞれ 8, 4, 2 から 16, 8, 4 に増加し、一貫して 2 倍の拡張が見られます。我々のような長文コンテキストでの並列処理を目標とするデプロイメントでは、この追加容量により DP32-EP32 が容量指向の高スループットプロファイルとして最適となります。一方、DP16-EP16 は効率性の比較対象として依然として有用です。
完全なデータは付録 C および付録 D.1 に記載されています。
6. 結論
一つのモデルに、一つの妥協したプロファイルで対応する必要はありません。
我々は H20 上で DeepSeek-V4-Pro 専用のシナリオ別サービングスタックを構築しました。プリフィル(Prefill)処理はコンテキスト長に応じて PP2 と PP4 を切り替えます。デコード(Decode)では、低レイテンシには PP2-TP8 を、高スループットには DP32-EP32 を採用しています。
容量、デプロイメントトポロジ、実行パスを統合的に設計した結果、H20 は限られた計算資源とネイティブ FP4 Tensor Core の欠如という制約の中でも、1M トークンのコンテキストをサポートし、複数のサービング SLO(サービスレベル目標)を満たすことが可能となりました。
この転送可能な成果は、シナリオ駆動型の手法です。サービングプロファイルは、ハードウェア仕様や単独のベンチマーク結果だけで選定すべきではありません。まずはワークロード、SLO(サービスレベル目標)、コンテキスト長、並行処理数といった要件から始め、プロファイリングを通じてボトルネックとなるリソースを特定し、それを具体的なトポロジーと実行パスの決定に落とし込むことを推奨します。
この手法が、計算能力、メモリ容量、メモリ帯域幅、あるいは相互接続など、多様なリソース制約下で実用的なフロンティアモデル用サービングシステムを構築する AI インフラチームのお役に立ち、得られた知見をオープンソースコミュニティ全体と共有できることを願っています。
謝辞
SGLang フレームワークにおける優れた取り組みに対して、SGLang チームおよびコミュニティに感謝いたします。また、以下のチームおよび協力者の方々にも、ご支援と貢献に対し心より御礼申し上げます。
- Ant Group SCT チーム: Yongfei Xu, Qianyu Zhang, Zekai Gu, ZhiLin Huang, Fakang Wang, Jianhao Fu, Zhuoxuan Du, Xia Zhan, Chun Huang, Qi Liu, Xi Chen, Yuhan Mao, Peipeng Cheng, Hanlin Gao, Jinghua Yao
- Ant Group Venus チーム: Jinzhen Lin
- SGLang コミュニティ: Peng Zhang
付録 A. プリフィル結果
A.1 Humming PP2 プリフィル:ベースラインと最終プロファイルの比較
| 入力長 | ベースライン TTFT (ms) | ベースライン 総入力スループット (tokens/s) | 最終 TTFT (ms) | 最終 総入力スループット (tokens/s) |
|---|---|---|---|---|
| 4K | 775.8 | 5,280 | 573.3 | 7,140 |
| 8K | 1202.1 | 6,810 | 907.6 | 9,030 |
| 16K | 2059.8 | 7,950 | 1649.5 | 9,930 |
| 32K | 4137.5 | 7,920 | 2470.3 | 13,260 |
| 64K | 6195.7 | 10,580 | 4063.8 | 16,130 |
| 128K | 10744.4 | 12,200 | 7975.9 | 16,430 |
| 256K | 20542.2 | 12,760 | 15507.2 | 16,900 |
| 512K | 44544.6 | 11,770 | 34982.6 | 14,990 |
| 1M | 100304.2 | 10,450 | 79214.2 | 13,240 |
A.2 Humming PP4 プリフィル:ベースラインと最終プロファイルの比較
| 入力長 | ベースライン TTFT (ms) | ベースライン 総入力スループット (tokens/s) | 最終 TTFT (ms) | 最終 総入力スループット (tokens/s) |
|---|---|---|---|---|
| 4K | 924.6 | 4,430 | 687.9 | 5,950 |
| 8K | 1174.5 | 6,970 | 890.3 | 9,200 |
| 16K | 2202.0 | 7,440 | 1635.4 | 10,020 |
| 32K | 4185.6 | 7,830 | 3068.4 | 10,680 |
| 64K | 5252.4 | 12,480 | 3982.6 | 16,460 |
| 128K | 7793.4 | 16,820 | 5882.5 | 22,280 |
| 256K | 13210.7 | 19,840 | 10348.9 | 25,330 |
| 512K | 26350.1 | 19,900 | 20273.1 | 25,860 |
| 1M | 55532.3 | 18,880 | 43742.5 | 23,970 |
付録 B. ローレイテンシ・デコード結果
B.1 入力長とバッチサイズにわたるピーク TPOT
B.1.1 No-Spec PP2-TP8
| 入力長 / バッチサイズ (ピーク TPOT, ms) | 1 | 2 | 4 | 8 | 16 |
|---|---|---|---|---|---|
| 8K | 26.39 | 30.86 | 31.31 | 31.79 | 31.74 |
| 32K | 25.72 | 26.58 | 27.81 | 31.06 | 37.97 |
| 64K | 25.75 | 26.62 | 28.13 | 29.19 | 38.75 |
| 128K | 25.94 | 26.94 | 28.38 | 29.75 | 38.51 |
| 256K | 26.08 | 27.21 | 28.84 | 32.43 | 38.83 |
| 512K | 26.25 | 27.51 | 29.16 | 33.70 | - |
| 1M | 26.42 | 27.81 | 29.52 | - | - |
B.1.2 最適化された DSpark PP2-TP8
| 入力長 / バッチサイズ (ピーク TPOT, ms) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 5.91 | 6.76 | 7.97 | 10.00 | 14.55 | 19.23 |
| 8K | 5.80 | 6.87 | 8.85 | 10.48 | 15.18 | 19.60 |
| 32K | 6.14 | 7.04 | 8.39 | 10.83 | 14.86 | 20.46 |
| 64K | 6.15 | 7.13 | 8.73 | 10.39 | 15.49 | 21.65 |
| 128K | 6.77 | 7.02 | 8.91 | 11.59 | 16.17 | 24.78 |
| 256K | 5.76 | 6.98 | 8.61 | 11.98 | 17.72 | - |
| 512K | 6.35 | 7.95 | 9.87 | 14.30 | - | - |
| 1M | 6.65 | 8.92 | 12.43 | - | - | - |
B.2 バッチサイズ 1 の出力スループット
| 入力長 | No-Spec PP2-TP8 (tokens/s) | Optimized DSpark PP2-TP8 (tokens/s) | Single-Node TP8 (tokens/s) |
|---|---|---|---|
| 4K | - | 169 | 213 |
| 8K | 38 | 172 | 260 |
| 16K | - | - | 244 |
| 32K | 39 | 163 | 269 |
| 64K | 39 | 163 | 246 |
| 128K | 39 | 148 | 267 |
| 256K | 38 | 174 | 271 |
| 512K | 38 | 157 | 254 |
| 1M | 38 | 150 | 183 |
B.3 ベンチマーク設定
ハードウェア構成:H20(GPU 141GB)を 8 基搭載したデコードノード 1 ノード。
デコードサーバー
--tp-size 8 \
--mem-fraction-static 0.91 \
--max-running-requests 1 \
--cuda-graph-max-bs 1 \
--cuda-graph-bs 1 \
--moe-runner-backend humming \
--moe-a2a-backend none \
--speculative-algorithm DSPARK \
--speculative-num-draft-tokens 7 \
--speculative-dspark-block-size 7 \
--speculative-moe-runner-backend triton \
--speculative-moe-a2a-backend none
クライアントベンチマーク
python3 -m sglang.bench_serving \
--backend sglang-oai-chat \
--host 127.0.0.1 \
--port 8000 \
--model <MODEL_PATH> \
--dataset-name random \
--dataset-path <DATASET_PATH> \
--random-input-len 262144 \
--random-output-len 4096 \
--random-range-ratio 1.0 \
--num-prompts 10 \
--max-concurrency 1 \
--warmup-requests 0 \
--seed 1
出力スループットは、サーバーの TP0 Decode batch ログから抽出しました。サンプルのうち最も高い値と低い値を各 20% ずつ除外し、残りの平均値を採用しています。B300 の数値は、リンク先のソースで報告された設定に準拠しています。
付録 C. 高スループットデコード結果
C.1 DP32-EP32構成におけるFP8 + MTP (3, 1, 4) の結果
| 入力長 / DP ランクごとの同時リクエスト数 (トークン/秒/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 30.49 | 58.58 | 102.89 | 174.75 | 253.15 | 319.92 |
| 8K | 30.34 | 58.29 | 102.38 | 174.67 | 251.62 | 318.32 |
| 16K | 29.70 | 56.55 | 99.47 | 170.01 | 242.22 | 302.43 |
| 32K | 29.58 | 56.35 | 98.28 | 164.26 | 234.13 | - |
| 64K | 29.07 | 55.73 | 96.43 | 161.60 | - | - |
| 128K | 28.39 | 54.06 | 92.89 | 153.55 | - | - |
| 256K | 28.35 | 53.02 | 90.89 | - | - | - |
| 512K | 27.51 | 51.49 | - | - | - | - |
| 1M | 27.05 | - | - | - | - | - |
C.2 DP32-EP32 を用いた FP8 と最適化された MTP(3, 1, 4)
| 入力長 / DP ランクごとの同時リクエスト数 (トークン/秒/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 36.84 | 69.86 | 131.96 | 232.94 | 389.94 | 514.77 |
| 8K | 32.58 | 69.51 | 131.53 | 222.06 | 348.80 | 416.82 |
| 16K | 31.89 | 67.44 | 127.79 | 216.14 | 341.85 | 395.99 |
| 32K | 31.49 | 67.21 | 124.49 | 208.83 | 337.97 | - |
| 64K | 30.95 | 66.47 | 123.68 | 205.44 | - | - |
| 128K | 30.22 | 64.47 | 119.14 | - | - | - |
| 256K | 30.18 | 63.23 | - | - | - | - |
| 512K | 29.28 | 61.40 | - | - | - | - |
| 1M | 28.79 | - | - | - | - | - |
C.3 DP32-EP32 を用いた FP8 + DSpark の活用
| 入力長 / DP ランクごとの同時リクエスト数 (トークン/秒/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 53.1 | 94.8 | 181.2 | 338.1 | 495.8 | 591.8 |
| 8K | 44.5 | 88.4 | 170.1 | 317.3 | 495.5 | - |
| 16K | 43.6 | 88.3 | 165.3 | 308.8 | 455.5 | - |
| 32K | 43.0 | 87.3 | 161.0 | 298.4 | - | - |
| 64K | 42.3 | 86.3 | 158.0 | - | - | - |
| 128K | 41.3 | 83.8 | - | - | - | - |
| 256K | 41.2 | - | - | - | - | - |
| 512K | 40.0 | - | - | - | - | - |
| 1M | 39.3 | - | - | - | - | - |
C.4 DP32-EP32 と Humming MXFP4AFP8、オンライン C128、DSparkの組み合わせ
| 入力長 / DP ランクごとの同時リクエスト数 (トークン/秒/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 75.32 | 127.10 | 235.85 | 417.53 | 564.08 | 703.15 |
| 8K | 75.60 | 128.29 | 238.01 | 417.34 | 560.68 | 709.64 |
| 16K | 74.00 | 124.47 | 231.25 | 406.21 | 539.72 | 674.19 |
| 32K | 73.07 | 122.27 | 225.28 | 392.47 | 521.70 | 601.67 |
| 64K | 71.81 | 120.92 | 221.05 | 386.11 | 516.54 | 599.63 |
| 128K | 70.12 | 117.29 | 212.93 | 366.88 | 487.69 | - |
| 256K | 70.03 | 115.03 | 208.35 | 345.21 | 457.62 | - |
| 512K | 67.95 | 111.71 | 191.99 | 302.80 | - | - |
| 1M | 66.82 | 105.82 | 177.48 | - | - | - |
C.5 DP16-EP16 with Humming MXFP4AFP8 + Online C128 + DSpark
| 入力長 / DP ランクごとの同時リクエスト数 (トークン/秒/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 76.80 | 129.62 | 236.83 | 397.42 | 584.37 | 759.73 |
| 8K | 76.69 | 130.55 | 237.53 | 398.60 | 582.03 | 762.09 |
| 16K | 76.16 | 127.88 | 233.10 | 388.54 | 571.22 | 745.23 |
| 32K | 74.07 | 124.77 | 226.24 | 378.79 | 559.05 | 722.51 |
| 64K | 74.57 | 124.69 | 223.84 | 373.13 | 541.46 | 695.35 |
| 128K | 72.36 | 120.34 | 219.66 | 365.13 | 518.98 | - |
| 256K | 71.38 | 119.19 | 211.64 | 340.72 | - | - |
| 512K | 69.54 | 115.14 | 198.81 | - | - | - |
| 1M | 67.39 | 106.50 | - | - | - | - |
付録 D. キャパシティに関する結果
D.1 デコード時のキャパシティのスケーリング
| デコードプロファイル | 設定 | フルトークン容量 (tokens/rank) | 前ステージとの比較 | FP8 ベースラインとの比較 |
|---|---|---|---|---|
| DP32-EP32 | Baseline FP8 + Offline C128 | 1,475,328 | - | 1.00× |
| Humming MXFP4AFP8 + Offline C128 | 2,526,720 | 1.71× | 1.71× | |
| Humming MXFP4AFP8 + Online C128 | 5,731,328 | 2.268× | 3.88× | |
| PP2-TP8 | Baseline FP8 + Offline C128 | 1,089,024 | - | 1.00× |
| Humming MXFP4AFP8 + Offline C128 | 4,869,888 | 4.47× | 4.47× | |
| Humming MXFP4AFP8 + Online C128 | 11,044,906 | 2.268× | 10.14× |
D.2 ヒューミング精度検証
DP16-EP16 Humming MXFP4AFP8 + Online C128 + DSpark プロファイルの GSM8K1000 における評価を行いました。このプロファイルは、95.5% の完全一致精度を達成し、無効な回答が 1 つ、システムエラーはゼロでした。これは、私たちが設定した 95.0% の合格閾値を満たしています。
参考として、アップストリームの SGLang Humming インテグレーション では、DeepSeek-V4-Flash を用いた 200 例の GSM8K 評価において、以下の結果が報告されています。
| バックエンド | GSM8K 精度 |
|---|---|
| Marlin MXFP4A16 | 96.5%–97.0% |
| FlashInfer MXFP4 | 96.5%–97.0% |
| Humming MXFP4A16 | 96.5%–97.5% |
| Humming MXFP4AFP8 | 97.0% |
この公開比較において、Humming MXFP4AFP8 の精度低下は確認されませんでした。同モデルは DeepSeek-V4-Pro ではなく DeepSeek-V4-Flash を採用しているため、本稿のサービングプロファイルにおける精度損失を直接測定するデータではなく、外部リファレンスとして扱います。
原文を表示
1. Introduction
DeepSeek-V4-Pro is a 1.6-trillion-parameter Mixture-of-Experts (MoE) model released with both FP8 and FP4 weights. Models at this scale naturally benefit from accelerators such as NVIDIA Blackwell GPUs, which offer more HBM, higher compute throughput, and native FP4 Tensor Cores. Yet H20 GPUs remain widely deployed, despite lacking those advantages.
Hardware constraints do not relax serving requirements. Long-context prefill must still control time to first token (TTFT). Interactive decode must satisfy the time-per-output-token (TPOT) target of each service tier. Sustained traffic must balance aggregate throughput against KV-cache capacity. Short inputs, long contexts, latency-sensitive requests, and high concurrency stress the system in different ways; no universal configuration can serve all of them well.
One model needs multiple serving profiles. Workload characteristics, service-level objectives (SLOs), and measured hardware behavior jointly inform the deployment topology and execution path:
- Match serving profiles to the workload. In the configurations evaluated here, prefill selects between PP2 and PP4 based on the measured context-length range, while decode uses profiles optimized for different latency, throughput, and KV-capacity targets.
- Optimize the prefill path. We optimize Attention-CP8 → MoE-TP8 and context-parallel communication, then tune for the real routing shapes produced by long- and short-context workloads.
- Optimize the decode path. We optimize the DSpark speculative-decoding path, refine execution, expert routing, and communication–computation overlap for distinct decode SLOs.
Push the latency frontier. At batch size 1, the single-node H20-141GB reference reaches 271 output tokens/s, compared with the 383.7 tokens/s reported on B300. Despite the substantial hardware gap, workload-specific system optimization narrows the observed decode performance ratio to 1.42×. Detailed benchmark settings and the log-based throughput extraction methodology are provided in Appendix B.3.
Cover the serving envelope. The latency result represents only one edge of the system. Across the broader profile family, optimized prefill reaches 8.45k input tokens/s per node and processes a 1M-token prompt in 43.7 seconds. For throughput-oriented decode, the DP16-EP16 efficiency reference reaches 4.67k output tokens/s per node, corresponding to an average TPOT of 27.4 ms. These results intentionally come from different profiles, each selected and optimized for a different combination of context length, latency, throughput, and capacity constraints.
The contribution is a methodology, not a single benchmark. Scenario-specific serving allows each workload to move toward a better measured operating point among the profiles evaluated on the available hardware. We hope the deployment choices, optimization methods, and measurements presented here provide a practical reference for teams serving frontier models under compute, memory, bandwidth, or interconnect constraints.
2. From Hardware Constraints to Serving Profiles
2.1 Hardware Constraints and Serving Roles
Blackwell offers raw performance; H20 offers deployable scale. B300 provides native FP4 Tensor Cores, much higher FP8 throughput, and substantially more HBM. H20 cannot match its compute capability, but it remains available at scale and provides high memory bandwidth and 900 GB/s NVLink. Each node in this study contains eight GPUs connected by NVLink. Prefill does not retain long-lived per-request state, so its hardware choice is governed primarily by TTFT, compute, and communication efficiency. Decode must retain the KV cache of every active request throughout generation, making HBM capacity a direct limit on context length and concurrency. For the deployment studied here, this led us to use H20-141GB for decode and H20-96GB—whose capacity was sufficient for our prefill workloads—for prefill.
2.2 Capacity Choices
Serving capacity ultimately comes from a shared HBM budget: model weights and per-request KV state compete for the same memory. We define full-token capacity as the maximum number of full-attention KV tokens that each rank can hold after model weights and runtime buffers have been allocated. It is a memory ceiling rather than a direct guarantee of admissible batch size.
Reducing Weight Footprint with Humming MXFP4AFP8
Reduce the weight footprint first. Humming MXFP4AFP8 uses MXFP4 expert weights with online FP8 activations to reduce weight footprint and memory traffic on H20 GPUs, which lack native FP4 Tensor Cores. The SGLang integration is available in sglang#23754. We will cover the Humming/SGLang integration in a dedicated follow-up post. Model-level accuracy results and public reference measurements are provided in Appendix D.2.
Expanding KV Capacity with Online C128
Give the KV cache room to grow. The Offline C128 baseline retains per-index state for each compressed page. Online C128 instead maintains a compact aggregate state, releasing more HBM to the KV-cache pool. It introduces additional state maintenance and speculative-verification work, but we observed no TPOT regression in our tests.
Combined Capacity Gains
Capacity gains compound across weights and KV state. By reducing the weight footprint, Humming MXFP4AFP8 expands full-token capacity to 1.71× the Baseline FP8 + Offline C128 configuration for DP32-EP32 and 4.47× for PP2-TP8. Online C128 then reduces the C128 auxiliary-state footprint, providing another 2.268× increase on top of Humming. Combined, the two techniques raise capacity to 3.88× the baseline for DP32-EP32 and 10.14× for PP2-TP8. Appendix D.1 provides the complete data.
2.3 Scenario-Specific Serving Profiles
Prefill Profiles
The right pipeline depth depends on how much work there is to pipeline. PP2-CP8-TP8 and PP4-CP8-TP8 share the same Attention-CP8 → MoE-TP8 execution path. At the topology level, their primary difference is pipeline depth: PP2 distributes the model across two stages, while PP4 uses four.
Short contexts favor lower pipeline overhead; long contexts expose more parallelism. Short inputs produce fewer chunks, leaving a deeper pipeline underfilled and making fill, drain, and cross-stage transfer costs more prominent. Long contexts provide enough chunks to keep four stages busy; with fewer layers per stage, the additional nodes translate into more prefill parallelism. In our deployment, these characteristics led us to use PP2-CP8-TP8 for shorter contexts and PP4-CP8-TP8 for long-context workloads.
Low-Latency Decode Profiles
Low latency starts with the shortest execution path. Single-node TP8 and PP2-TP8 share the same Attention-TP8 → MoE-TP8 execution path; the difference is whether the model is partitioned across nodes. Single-node TP8 places all layers on one H20-141GB node and avoids cross-stage communication and synchronization. PP2-TP8 partitions the model across two pipeline stages.
The fastest topology is not always the most serviceable one. Single-node TP8 has the shorter execution path, but model weights and serving state share the HBM of one node, leaving limited room for the KV cache. It cannot simultaneously support long contexts and larger batch sizes. PP2-TP8 pays additional pipeline overhead but distributes the model weights across two nodes, releasing more HBM for KV state. For our latency and capacity targets, we use single-node TP8 as the batch-size-1 latency reference and PP2-TP8 as the low-latency serving profile.
High-Throughput Decode Profiles
High-throughput decode scales data and expert parallelism together. Both profiles use the Attention-DP → MoE-EP execution path. DP16-EP16 is the smallest deployment unit; DP32-EP32 expands both DP and EP within the same topology.
Scale-out prioritizes request capacity over per-GPU throughput. A larger EP group distributes expert weights across more GPUs, releasing HBM for the KV cache and admitting more concurrent requests. At the same time, a smaller fraction of MoE traffic remains within each node, while a larger fraction crosses nodes, which can reduce per-GPU efficiency. In the profiles evaluated here, we use DP16-EP16 as the smallest deployment unit and efficiency reference, and DP32-EP32 to expand request capacity.
3. Prefill: Balancing Compute and Communication
Prefill performance is a system problem. Expert imbalance, context-parallel communication, and production routing shapes jointly determine TTFT; optimizing an isolated kernel is not enough.
3.1 Why MoE-TP Instead of MoE-EP
Less traffic can still take longer. MoE-EP exchanges only routed tokens, but real prefill traffic exhibits significant expert skew. Ranks that own hot experts perform more computation and become stragglers; all other ranks wait for the slowest path at the combine step. Lower communication volume does not translate into lower TTFT.
Balance compute before minimizing traffic. For the H20 prefill workloads evaluated here, both PP2 and PP4 use MoE-TP. Full-sequence all-gather and reduce-scatter introduce more communication, but the traffic remains on high-bandwidth NVLink and has stable, predictable cost. All TP ranks execute tensor-parallel computation over the same routed tokens, preventing expert skew from becoming a rank-level long tail. For this workload, predictable communication is cheaper than unpredictable imbalance. The implementation is available in sglang#24947.
3.2 Accelerating and Fusing Prefill Collectives
Build a reusable collective fast path. MoE-TP replaces unpredictable expert imbalance with predictable collective traffic, making communication efficiency the next bottleneck. We made symmetric memory reusable across TP and CP, allowing AllReduce, AllGather, and ReduceScatter to share registered-buffer fast paths and applicable Hopper acceleration. The supporting upstream work spans memory-pool ownership, communicator registration, MoE-TP collective buffers, and the CP Attention and KV-cache buffer paths.
Then shorten the Prefill critical path. Faster collectives alone do not remove the boundaries between communication and computation. For the 32K single-chunk case, we built a fused path that overlaps a copy-engine-driven AllGather with fused FP8 quantization and shared-expert GEMM, then combines TopK reduction, shared-expert addition, and ReduceScatter in a second Triton kernel. This reorganizes seven operators into three execution groups and reduces TTFT by approximately 3.5% in a matched PP4 A/B.
3.3 Tuning Humming for Real Routing Shapes
Generic tuning misses the shapes that matter. Prefill routing distributes tokens unevenly across 384 experts, so the effective M dimension clusters into a small set of discrete values. W13 and W2 also operate on different shapes, so a single generic heuristic cannot optimize both paths.
Tune from production routing. We extract high-frequency shapes from real routing histograms, build separate exact-shape configurations for W13 and W2, and validate them at the kernel, pipeline-stage, and matched A/B levels. The optimization target is not a synthetic range of M, but the routing distribution we actually serve. In a matched PP4 A/B at 32K, selected MoE kernel latency falls by approximately 21%, translating into an 11.35% end-to-end TTFT reduction.
4. Decode: Optimizing Speculation and MoE Execution
Decode optimization is profile-specific in our implementation. PP2-TP8 requires coordination across speculative pipeline stages, while DP32-EP32 focuses on optimizing the refinement step and expert routing at high concurrency. Humming fusion and overlap improve the shared MoE hot path beneath these serving topologies.
4.1 Low-Latency PP2-TP8: Extending DSpark Across Pipeline Stages
Pipeline parallelism splits the speculative loop. In PP2-TP8, target execution spans two pipeline stages, while the DSpark drafter resides only on the final stage. Stage 0 sends target hidden states to Stage 1, which performs verification, accepts tokens, and generates candidates for the next round.
Make two stages advance as one. Every speculative round crosses the pipeline boundary. We coordinate both stages and the required intermediate transfers under one execution protocol, preventing the stages from entering different rounds while avoiding redundant synchronization. The PP-specific DSpark integration is being upstreamed in sglang#32281.
4.2 High-Throughput DP32-EP32: Removing High-Concurrency Bottlenecks
The matched A/B results in this subsection use DP32-EP32 at 4K with 32 concurrent requests per DP rank.
Choose the right execution shape for refinement. The refinement step applies a full-vocabulary projection to rescore DSpark's candidate set. At high concurrency, the row-wise dot-reduce repeatedly reads the vocabulary weights for every active row, creating a persistent tail in each decode step. We combine active rows into one transposed GEMM, reducing redundant memory traffic and shortening the refinement path. Per-GPU throughput improves by 22.8%.
Place experts from measured routing. DSpark traffic also exhibits significant expert skew. We record routing affinity from representative requests and use it to configure expert-parallel load balancing (EPLB) and redundant experts, preventing a small number of hot experts from repeatedly extending the critical path. Per-GPU throughput improves by 13.5%.
4.3 Humming Decode Hot Path: Fusion and Overlap
These optimizations sit below the serving topology and can be reused by Humming-based decode profiles. The matched results below use DP32-EP32 at 4K with 32 concurrent requests per DP rank.
Remove the extra quantization pass. We fuse the SwiGLU activation with quantization so that the fused kernel directly produces the data and scale required by W2. This eliminates repeated access to an intermediate buffer and removes the standalone quantization pass, allowing W2 to start earlier. In the matched DSpark A/B, per-GPU throughput improves by 44.0%.
Overlap communication with W2. We adapt the Single-Batch Overlap (SBO) mechanism from our previous work (sglang#9660) into Humming-Aware SBO. Per-tile signals allow DeepEP to begin the corresponding combine send as soon as a W2 output tile completes, without waiting for the entire GEMM. In an earlier matched non-spec A/B at the same operating point, SBO recovers 4.12% throughput relative to the FP8-transport tier.
5. Evaluation: System Gains and Profile Trade-offs
5.1 Prefill: Cumulative Gains and Context-Length Trade-offs
PP2 strengthens the short-context profile. PP2 improves at all nine input lengths, with a geometric-mean throughput gain of 36.5% and a peak total input throughput of 16,900 tokens/s. Its shallower pipeline reduces fill-and-drain overhead for short requests, allowing PP2 to maintain lower TTFT with fewer resources.
PP4 carries the gains into long context. PP4 delivers a geometric-mean throughput gain of 31.8% across the same nine points. As context length grows, the deeper pipeline has enough work to amortize its fixed cost: total input throughput reaches 25,860 tokens/s at 512K and remains 23,970 tokens/s at 1M.
Context length shifts the PP2/PP4 trade-off. Relative to PP4, PP2 lowers TTFT by 16.7% at 4K and 19.5% at 32K. The two profiles remain within 2% at 8K, 16K, and 64K. PP4 establishes a decisive advantage from 128K onward, reducing TTFT relative to PP2 by 26.2%, 33.3%, 42.1%, and 44.8% at 128K, 256K, 512K, and 1M, respectively. We therefore treat the routing boundary as an operating policy derived from the measured context-length range rather than a universal crossover point.
Appendix A.1–A.2 provide the complete TTFT and total-input-throughput results.
5.2 Low-Latency Decode: Performance and Capacity Trade-offs
Optimized DSpark resets the latency baseline. Across the four input lengths shown in Figure 15, Optimized DSpark reduces peak TPOT by 74.8%–78.0% at batch size 1. At the largest batch size shared by each pair of measurements, the reduction remains 52.2%–60.0%. The gain holds from 8K through 1M rather than being confined to short contexts or single-request execution.
Observed serving performance is much closer than peak-compute ratios alone suggest. Across the four input lengths shown in Figure 16, Optimized DSpark on PP2-TP8 reaches 150–174 tokens/s at batch size 1. The single-node TP8 reference reaches 183–271 tokens/s. For the precisions used by the actual execution paths, B300 has approximately 45.6× the peak Tensor Core compute of H20-141GB (B300 FP4 versus H20 FP8) and 1.67× its memory bandwidth. Yet the highest observed generation rates are 383.7 tokens/s on B300 and 271 tokens/s on H20-141GB, respectively—a ratio of 1.42×. Even against this much stronger hardware reference, workload-specific optimization brings the H20-141GB reference substantially closer in observed serving performance.
Capacity favors PP2-TP8 for our production targets. Single-node TP8 is faster, but at a 1M context it has enough KV-cache capacity only for batch size 1. It cannot admit a larger batch or more concurrent requests. By distributing model weights across two pipeline stages, PP2-TP8 supports batch sizes 4, 8, and 16 at 1M, 512K, and 256K, respectively. With Online C128, its full-token capacity reaches 11.04M tokens/rank. For context-length and concurrency targets similar to ours, we recommend retaining single-node TP8 as the latency reference and using PP2-TP8 as the low-latency serving profile. Appendix B and Appendix D.1 provide the complete performance and capacity data.
5.3 High-Throughput Decode: Frontier Gains and Profile Trade-offs
Figure 17 shows how the throughput–interactivity frontier evolves with the system. The horizontal axis is interactivity in tokens/s/user, and the vertical axis is throughput in tokens/s/GPU. In these DP/EP profiles, each DP rank maps to one GPU; interactivity is per-GPU throughput divided by the number of concurrent requests per DP rank. Points farther toward the upper right provide a better combination of user-visible generation speed and GPU efficiency. The four curves represent cumulative system evolution rather than the isolated gain of any optimization in Section 4.
MTP denotes multi-token prediction; the (3, 1, 4) configuration uses three speculative steps, top-k 1, and four draft tokens.
System optimization moves the entire frontier. At 4K with 32 concurrent requests per DP rank, per-GPU throughput rises from 319.92 tokens/s/GPU to 703.15 tokens/s/GPU, a 2.20× increase. At 1M with one request per DP rank, it rises from 27.05 tokens/s/GPU to 66.82 tokens/s/GPU. The first three system milestones can each process only one request per DP rank at 1M; the final system supports four and reaches 177.48 tokens/s/GPU. The expanded operating envelope comes from both faster execution and greater capacity.
Smaller deployment units preserve efficiency at selected high-concurrency operating points. In our earlier work on serving DeepSeek-V3/R1 on H20, we found that a smaller EP deployment unit can keep a larger fraction of MoE traffic within each node. DeepSeek-V4-Pro shows the same advantage at the operating points plotted in Figure 18: with 16 and 32 concurrent requests per DP rank, DP16-EP16 delivers approximately 3.6%–20% higher per-GPU throughput than DP32-EP32. The full sweep is not monotonic across every concurrency level, so we use DP16-EP16 as an efficiency reference rather than a universal replacement for DP32-EP32.
Capacity shifts the preferred high-throughput profile. DP16-EP16 is more efficient per GPU, but DP32-EP32 distributes expert weights across more ranks and releases additional HBM for the KV cache. At 256K, 512K, and 1M, the maximum concurrent requests per DP rank increase from 8, 4, and 2 to 16, 8, and 4, respectively—a consistent 2× expansion. For deployments with long-context concurrency targets similar to ours, this additional capacity favors DP32-EP32 as the capacity-oriented high-throughput profile, while DP16-EP16 remains useful as an efficiency reference. Appendix C and Appendix D.1 provide the complete data.
6. Conclusion
One model does not require one compromise profile. We built a scenario-specific serving stack for DeepSeek-V4-Pro on H20. Prefill switches between PP2 and PP4 according to context length. Decode uses PP2-TP8 for low latency and DP32-EP32 for high throughput. By co-designing capacity, deployment topology, and execution path, H20 can sustain 1M-token contexts and meet multiple serving SLOs despite limited compute and the absence of native FP4 Tensor Cores.
The transferable result is a scenario-driven methodology. Serving profiles should not be selected from hardware specifications or isolated benchmarks alone. We recommend starting from the workload, SLO, context length, and concurrency, then using profiling to identify the binding resource and translate it into concrete topology and execution-path decisions. We hope this methodology helps AI infrastructure teams build practical frontier-model serving systems under diverse resource constraints—whether the bottleneck is compute, memory capacity, memory bandwidth, or interconnect—and share those lessons with the broader open-source ecosystem.
Acknowledgements
We would like to thank the SGLang Team and Community for their outstanding work on the SGLang framework. We also thank the following teams and collaborators for their support and contributions:
- Ant Group SCT Team: Yongfei Xu, Qianyu Zhang, Zekai Gu, ZhiLin Huang, Fakang Wang, Jianhao Fu, Zhuoxuan Du, Xia Zhan, Chun Huang, Qi Liu, Xi Chen, Yuhan Mao, Peipeng Cheng, Hanlin Gao, Jinghua Yao
- Ant Group Venus Team: Jinzhen Lin
- SGLang Community: Peng Zhang
Appendix A. Prefill Results
A.1 Humming PP2 Prefill: Baseline vs. Final Profile
| Input Length | Baseline TTFT (ms) | Baseline Total Input Throughput (tokens/s) | Final TTFT (ms) | Final Total Input Throughput (tokens/s) |
|---|---|---|---|---|
| 4K | 775.8 | 5,280 | 573.3 | 7,140 |
| 8K | 1202.1 | 6,810 | 907.6 | 9,030 |
| 16K | 2059.8 | 7,950 | 1649.5 | 9,930 |
| 32K | 4137.5 | 7,920 | 2470.3 | 13,260 |
| 64K | 6195.7 | 10,580 | 4063.8 | 16,130 |
| 128K | 10744.4 | 12,200 | 7975.9 | 16,430 |
| 256K | 20542.2 | 12,760 | 15507.2 | 16,900 |
| 512K | 44544.6 | 11,770 | 34982.6 | 14,990 |
| 1M | 100304.2 | 10,450 | 79214.2 | 13,240 |
A.2 Humming PP4 Prefill: Baseline vs. Final Profile
| Input Length | Baseline TTFT (ms) | Baseline Total Input Throughput (tokens/s) | Final TTFT (ms) | Final Total Input Throughput (tokens/s) |
|---|---|---|---|---|
| 4K | 924.6 | 4,430 | 687.9 | 5,950 |
| 8K | 1174.5 | 6,970 | 890.3 | 9,200 |
| 16K | 2202.0 | 7,440 | 1635.4 | 10,020 |
| 32K | 4185.6 | 7,830 | 3068.4 | 10,680 |
| 64K | 5252.4 | 12,480 | 3982.6 | 16,460 |
| 128K | 7793.4 | 16,820 | 5882.5 | 22,280 |
| 256K | 13210.7 | 19,840 | 10348.9 | 25,330 |
| 512K | 26350.1 | 19,900 | 20273.1 | 25,860 |
| 1M | 55532.3 | 18,880 | 43742.5 | 23,970 |
Appendix B. Low-Latency Decode Results
B.1 Peak TPOT Across Input Lengths and Batch Sizes
B.1.1 No-Spec PP2-TP8
| Input Length / Batch Size (Peak TPOT, ms) | 1 | 2 | 4 | 8 | 16 |
|---|---|---|---|---|---|
| 8K | 26.39 | 30.86 | 31.31 | 31.79 | 31.74 |
| 32K | 25.72 | 26.58 | 27.81 | 31.06 | 37.97 |
| 64K | 25.75 | 26.62 | 28.13 | 29.19 | 38.75 |
| 128K | 25.94 | 26.94 | 28.38 | 29.75 | 38.51 |
| 256K | 26.08 | 27.21 | 28.84 | 32.43 | 38.83 |
| 512K | 26.25 | 27.51 | 29.16 | 33.70 | - |
| 1M | 26.42 | 27.81 | 29.52 | - | - |
B.1.2 Optimized DSpark PP2-TP8
| Input Length / Batch Size (Peak TPOT, ms) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 5.91 | 6.76 | 7.97 | 10.00 | 14.55 | 19.23 |
| 8K | 5.80 | 6.87 | 8.85 | 10.48 | 15.18 | 19.60 |
| 32K | 6.14 | 7.04 | 8.39 | 10.83 | 14.86 | 20.46 |
| 64K | 6.15 | 7.13 | 8.73 | 10.39 | 15.49 | 21.65 |
| 128K | 6.77 | 7.02 | 8.91 | 11.59 | 16.17 | 24.78 |
| 256K | 5.76 | 6.98 | 8.61 | 11.98 | 17.72 | - |
| 512K | 6.35 | 7.95 | 9.87 | 14.30 | - | - |
| 1M | 6.65 | 8.92 | 12.43 | - | - | - |
B.2 Batch-Size-1 Output Throughput
| Input Length | No-Spec PP2-TP8 (tokens/s) | Optimized DSpark PP2-TP8 (tokens/s) | Single-Node TP8 (tokens/s) |
|---|---|---|---|
| 4K | - | 169 | 213 |
| 8K | 38 | 172 | 260 |
| 16K | - | - | 244 |
| 32K | 39 | 163 | 269 |
| 64K | 39 | 163 | 246 |
| 128K | 39 | 148 | 267 |
| 256K | 38 | 174 | 271 |
| 512K | 38 | 157 | 254 |
| 1M | 38 | 150 | 183 |
B.3 Benchmark Settings
Hardware: one 8× H20-141GB decode node.
Decode server
--tp-size 8 \
--mem-fraction-static 0.91 \
--max-running-requests 1 \
--cuda-graph-max-bs 1 \
--cuda-graph-bs 1 \
--moe-runner-backend humming \
--moe-a2a-backend none \
--speculative-algorithm DSPARK \
--speculative-num-draft-tokens 7 \
--speculative-dspark-block-size 7 \
--speculative-moe-runner-backend triton \
--speculative-moe-a2a-backend none
Client benchmark
python3 -m sglang.bench_serving \
--backend sglang-oai-chat \
--host 127.0.0.1 \
--port 8000 \
--model <MODEL_PATH> \
--dataset-name random \
--dataset-path <DATASET_PATH> \
--random-input-len 262144 \
--random-output-len 4096 \
--random-range-ratio 1.0 \
--num-prompts 10 \
--max-concurrency 1 \
--warmup-requests 0 \
--seed 1
Output throughput is extracted from the server's TP0 Decode batch log lines; we discard the highest and lowest 20% of samples and average the remainder. The B300 number follows the setup reported in the linked source.
Appendix C. High-Throughput Decode Results
C.1 DP32-EP32 with FP8 + MTP (3, 1, 4)
| Input Length / Concurrent Requests per DP Rank (tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 30.49 | 58.58 | 102.89 | 174.75 | 253.15 | 319.92 |
| 8K | 30.34 | 58.29 | 102.38 | 174.67 | 251.62 | 318.32 |
| 16K | 29.70 | 56.55 | 99.47 | 170.01 | 242.22 | 302.43 |
| 32K | 29.58 | 56.35 | 98.28 | 164.26 | 234.13 | - |
| 64K | 29.07 | 55.73 | 96.43 | 161.60 | - | - |
| 128K | 28.39 | 54.06 | 92.89 | 153.55 | - | - |
| 256K | 28.35 | 53.02 | 90.89 | - | - | - |
| 512K | 27.51 | 51.49 | - | - | - | - |
| 1M | 27.05 | - | - | - | - | - |
C.2 DP32-EP32 with FP8 + Optimized MTP (3, 1, 4)
| Input Length / Concurrent Requests per DP Rank (tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 36.84 | 69.86 | 131.96 | 232.94 | 389.94 | 514.77 |
| 8K | 32.58 | 69.51 | 131.53 | 222.06 | 348.80 | 416.82 |
| 16K | 31.89 | 67.44 | 127.79 | 216.14 | 341.85 | 395.99 |
| 32K | 31.49 | 67.21 | 124.49 | 208.83 | 337.97 | - |
| 64K | 30.95 | 66.47 | 123.68 | 205.44 | - | - |
| 128K | 30.22 | 64.47 | 119.14 | - | - | - |
| 256K | 30.18 | 63.23 | - | - | - | - |
| 512K | 29.28 | 61.40 | - | - | - | - |
| 1M | 28.79 | - | - | - | - | - |
C.3 DP32-EP32 with FP8 + DSpark
| Input Length / Concurrent Requests per DP Rank (tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 53.1 | 94.8 | 181.2 | 338.1 | 495.8 | 591.8 |
| 8K | 44.5 | 88.4 | 170.1 | 317.3 | 495.5 | - |
| 16K | 43.6 | 88.3 | 165.3 | 308.8 | 455.5 | - |
| 32K | 43.0 | 87.3 | 161.0 | 298.4 | - | - |
| 64K | 42.3 | 86.3 | 158.0 | - | - | - |
| 128K | 41.3 | 83.8 | - | - | - | - |
| 256K | 41.2 | - | - | - | - | - |
| 512K | 40.0 | - | - | - | - | - |
| 1M | 39.3 | - | - | - | - | - |
C.4 DP32-EP32 with Humming MXFP4AFP8 + Online C128 + DSpark
| Input Length / Concurrent Requests per DP Rank (tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 75.32 | 127.10 | 235.85 | 417.53 | 564.08 | 703.15 |
| 8K | 75.60 | 128.29 | 238.01 | 417.34 | 560.68 | 709.64 |
| 16K | 74.00 | 124.47 | 231.25 | 406.21 | 539.72 | 674.19 |
| 32K | 73.07 | 122.27 | 225.28 | 392.47 | 521.70 | 601.67 |
| 64K | 71.81 | 120.92 | 221.05 | 386.11 | 516.54 | 599.63 |
| 128K | 70.12 | 117.29 | 212.93 | 366.88 | 487.69 | - |
| 256K | 70.03 | 115.03 | 208.35 | 345.21 | 457.62 | - |
| 512K | 67.95 | 111.71 | 191.99 | 302.80 | - | - |
| 1M | 66.82 | 105.82 | 177.48 | - | - | - |
C.5 DP16-EP16 with Humming MXFP4AFP8 + Online C128 + DSpark
| Input Length / Concurrent Requests per DP Rank (tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 76.80 | 129.62 | 236.83 | 397.42 | 584.37 | 759.73 |
| 8K | 76.69 | 130.55 | 237.53 | 398.60 | 582.03 | 762.09 |
| 16K | 76.16 | 127.88 | 233.10 | 388.54 | 571.22 | 745.23 |
| 32K | 74.07 | 124.77 | 226.24 | 378.79 | 559.05 | 722.51 |
| 64K | 74.57 | 124.69 | 223.84 | 373.13 | 541.46 | 695.35 |
| 128K | 72.36 | 120.34 | 219.66 | 365.13 | 518.98 | - |
| 256K | 71.38 | 119.19 | 211.64 | 340.72 | - | - |
| 512K | 69.54 | 115.14 | 198.81 | - | - | - |
| 1M | 67.39 | 106.50 | - | - | - | - |
Appendix D. Capacity Results
D.1 Decode Capacity Scaling
| Decode Profile | Configuration | Full-Token Capacity (tokens/rank) | Vs. Previous Stage | Vs. FP8 Baseline |
|---|---|---|---|---|
| DP32-EP32 | Baseline FP8 + Offline C128 | 1,475,328 | - | 1.00× |
| Humming MXFP4AFP8 + Offline C128 | 2,526,720 | 1.71× | 1.71× | |
| Humming MXFP4AFP8 + Online C128 | 5,731,328 | 2.268× | 3.88× | |
| PP2-TP8 | Baseline FP8 + Offline C128 | 1,089,024 | - | 1.00× |
| Humming MXFP4AFP8 + Offline C128 | 4,869,888 | 4.47× | 4.47× | |
| Humming MXFP4AFP8 + Online C128 | 11,044,906 | 2.268× | 10.14× |
D.2 Humming Accuracy Validation
We evaluated the DP16-EP16 Humming MXFP4AFP8 + Online C128 + DSpark profile on GSM8K1000. The profile achieves 95.5% exact-match accuracy, with one invalid response and zero system errors, passing our 95.0% acceptance threshold.
As a public reference, the upstream SGLang Humming integration reports the following results on a 200-example GSM8K evaluation of DeepSeek-V4-Flash:
| Backend | GSM8K Accuracy |
|---|---|
| Marlin MXFP4A16 | 96.5%–97.0% |
| FlashInfer MXFP4 | 96.5%–97.0% |
| Humming MXFP4A16 | 96.5%–97.5% |
| Humming MXFP4AFP8 | 97.0% |
No accuracy degradation is visible for Humming MXFP4AFP8 in this public comparison. Because it uses DeepSeek-V4-Flash rather than DeepSeek-V4-Pro, we treat it as an external reference rather than a matched accuracy-loss measurement for our serving profile.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み