GPU4台必要なら時間単価は商品ではない
Vast.ai のデータ分析により、GPU 単体の時価と、複数 GPU を同一マシンで必要とする実ワークロードの価格には大きな乖離があり、特に大規模な構成では供給不足や価格高騰が顕著であることが示された。
AI算出
市場分析ainew評価高い
AI インフラの供給制約(ベースリスク)を定量的に示した独自分析であり、単なるニュース報知ではなく市場構造の洞察を提供している。ただし、対象はグローバルなクラウド市場であり、日本固有の事例や規制に関する言及はない。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 50
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
この記事の3ポイント
- 1
ヘッドライン価格と実需価格の乖離
- 2
構成要件による供給の急減
- 3
ワークロード要件の厳格性
なぜ重要か・誰に関係するか
この分析は、AI 開発者や企業が GPU クラウドリソースを調達する際の現実的な課題を浮き彫りにしており、単純な単価比較に基づく予算策定がリスクを伴うことを示唆しています。特に大規模モデルのトレーニング需要が高まる中、供給側の物理的制約(同一マシンの構成数)が市場価格に直結するため、リソース計画にはより高度な予測と柔軟性が求められます。
AI深層分析を開く2026年7月25日 16:20
キーポイント
ヘッドライン価格と実需価格の乖離
GPU の単体時価(例:H200 で約$3.93/時間)は、複数 GPU を同一マシンで必要とするワークロードの実質コストとは一致しない。
構成要件による供給の急減
4 枚の H200 が同一マシンに配置された環境を求めると、利用可能な供給量が半分になり、価格も約 4% 上昇する。8 枚構成ではサンプル内で完全な供給が見当たらない。
ワークロード要件の厳格性
トレーニングやファインチューニングには、散在した GPU 時間ではなく、十分な相互接続と長期利用可能な同一クラスタ内の複数 GPU が不可欠である。
重要な引用
An H200 rented for about $3.93 per GPU-hour. Without knowing anything else, it looks pretty straightforward.
But if you want four identical H200s in the same machine, half the qualifying supply disappears, and the remaining supply costs 4% more.
Eight scattered GPU-hours don't substitute for one hour on eight co-located GPUs.
編集コメントを表示
編集コメント
GPU クラウド市場の「見かけの安さ」と「実需の高騰」をデータで明確に示した貴重な分析です。大規模 AI モデルの開発においては、単体スペックだけでなく、相互接続や構成数の制約がコストと納期に直結するため、この視点は極めて重要です。
関連する開示についてはこちらをご覧ください。
最近、計算リソースのレンタル市場である Vast.ai 上で GPU の賃貸価格を調査しました。H200 の単価は、1 GPU あたり約 3.93 ドル/時間でした。これだけ見ると非常にシンプルに思えます。必要なチップを選び、時間あたりの料金を必要な GPU 数と利用時間に掛ければ、請求額が算出されるはずです。
しかし、同じマシン内に H200 を 4 枚積んだ構成を求めると、条件を満たす供給量の半分が消え、残りの供給も 4% ほど高くなります。8 枚必要であれば、残念ながら手が出せません。サンプル内にはどんな価格でも見つかりませんでした。レンタルレートと、実際に使える計算リソースの価格は異なるのです。
つまり、計算リソースの価格表示は「 headline rate(表面上のレート)」に基づいていますが、多くのワークロードがそのように消費されるわけではありません。見出しでは H100 と B200 はそれぞれ異なる価格で表示されていますが、学習やファインチューニング、高スループット推論には、同じマシンまたはクラスター内に複数の同一 GPU を配置し、十分な相互接続と十分な利用可能時間を確保する必要があります。8 台の分散した GPU を 1 時間ずつ使うことと、8 台の隣接する GPU で 1 時間使うことは、決して同等ではありません。
表面上の数字と実務的なワークロードにおける価格の差を定量化するため、Vast.ai からオンデマンドオファーを取得し、以下の条件を満たすものだけを抽出しました。検証済みであること、現在レンタル可能であること、少なくとも 7 日間利用可能であること、そして信頼性スコアが 99% 以上のホストに紐付いていることです。
このフィルタリング処理の結果、対象となったのは5種類のアクセラレーターモデルでした。具体的には A100 SXM4、H100 SXM、H200、B200、そして L40S です。
同じ物理マシン上で異なるバンドルサイズが複数のオファーとして表示されているケースが多いため、ここでは物理マシン単位で重複を除外しました。その結果、37台のマシンと 93 個の GPU がリストアップされました。
次に、1 個、2 個、4 個、8 個という同じ仕様の GPU を必要とする場合、どれだけの供給量があれば対応可能か、また各サイズで最も安い利用料はいくらになるかを調査しました。価格指標としては、条件を満たす機器のうち最安値の上位3台までの中央値を採用し、極端に安い単価が結果を歪めるのを防ぎました。
*Diagram generated by GPT 5.6 Sol*
この分析から興味深い事実が浮かび上がりました。それは、必ずしも大規模なクラスターの方が高価になるとは限らないという点です。必要な GPU の数が増えるほど利用可能な在庫は減少しますが、その割に価格が連動して上昇するわけではありません。H200 のデータがこの傾向を最も明確に示しています。
1 個から 2 個へ増やしても、ほぼ同じ価格で利用可能な物理マシンの割合は 92% を維持しました。しかし、4 個になると利用可能な在庫は半減し、GPU 時間あたりの価格は 4% 上昇して 3.93 ドルから 4.08 ドルとなりました。そして 8 個を必要とするケースでは、条件を満たす機器は一つも存在しませんでした。
H100 と B200 の供給はさらに逼迫していました。H100 は 2 GPU リクエストには対応可能でしたが、4 GPU を必要とするリクエストには対応できませんでした。B200 のサンプルでは物理マシンが 2 台確認できましたが、どちらも最大 2 GPU までしかサポートしておらず、4 GPU や 8 GPU の利用可能性はゼロでした。
A100 は予想通り、供給が減ると価格が上昇しました。8 GPU マシンも存在しましたが、1 GPU あたりの価格がベースラインの 0.60 ドルから 1.29 ドルに跳ね上がり、連続性プレミアム(コンティニュイティ・プレミアム)は 114% に達しています。
L40S は両刃の剣でした。リストされた在庫のうち、8 GPU リクエストに対応できる大規模マシンにあるのはわずか 35% だけですが、その唯一の候補機は安価でした。これは滑らかな供給曲線ではなく、薄く断片的な注文簿の状態です。
通常、希少性は価格に影響します。石油を例に挙げれば、需要が増加して供給が一定であれば価格は上昇します。しかし、コンピューティングリソースでは必ずしもそうはなりません。計算資源は、構成と可用性によって割り当てられるのです。市場全体で数百 GPU 時間の在庫があっても、購入者がそれらを組み合わせて稼働可能なクラスターを構築できないケースがあります。
ボトルネックとなるのは、ハードウェアの世代、共設置(コロケーション)、ネットワーク、信頼性、予約期間、CPU、ストレージ、そして地理的条件です。
しかし、適切な構成が存在しない場合、希少性は取引価格の上昇として現れません。取引自体が成立しないからです。そのため、1 GPU や 2 GPU のレンタル価格だけで算出された指数では、コンピューティングリソースは豊富で安定しているように見えてしまいます。一方で、大規模クラスターを必死に見つけようとする企業は、どんなに高額な価格であっても見つからないというジレンマを抱えています。
この理由から、大規模な購入者は二国間容量契約に頼るのです。これらの契約は、特定の時刻における特定の構成を確保します。余剰な容量は Vast.ai などの市場で取引されることになります。
最初の計算先物契約は、標準化された GPU アワー基準に対して決済されると考えられます。デリバティブには単純な定義と信頼できるインデックスを支える十分な観測データが必要であるため、これは理にかなっています。しかし、将来の 4 枚または 8 枚の H100 GPU 需要に対するヘッジとして H100 の先物を買う企業にとって、実質的なベースリスクが存在します。先物は広範な H100 価格の上昇にはヘッジ機能を持ちますが、共置された容量が不足していることへの対策にはなりません。ヘッジは成立しても、ワークロードは実行されません。
エネルギー市場では、この残余リスクを「場所による価格差(ロケーション・ベース)」や、供給の確約性と中断可能性の間のスプレッドとして管理しています。同様に、計算市場もトポロジーに特化したバージョンを構築するでしょう。中心には汎用的な GPU アワーを置き、クラスターサイズ、相互接続の品質、地域、可用性保証などに対してはプレミアムまたは別契約を設ける形です。そして、これらのクラスターアワーや容量オプションが、基礎となる GPU アワーよりも活発に取引されるようになる可能性も十分にあります。
Vast.ai は単一の市場プレイスに過ぎません。そこに掲載されているのは広告としてのオファーであり、完了した取引ではありません。また、そこに表示される供給量は、世界の設置済み総容量を反映しているわけではありません。大規模なクラスターの推定値のいくつかは少数の機械に基づいているため、詳細なプレミアムについては注意深く扱う必要があります。
この利用パターンの分析は、単に「サンプルが少ないから」と片付けられるものではありません。信頼性閾値を 98% から 99.5% の間で変えても結果は変わりませんでした。B200 に 4 GPU 構成の提供はなく、H200 に 8 GPU 構成の提供もありませんでした。A100 も大規模クラスター利用には依然としてプレミアム価格がついています。
このサンプル数が少ないこと自体が重要なポイントです。 headline となるような GPU の平均価格は、一見活発な市場があるように見せているかもしれませんが、実際に必要な特定の構成(例えば 4 GPU や 8 GPU)の市場は、ほとんど存在していないのです。
本ニュースレターがお役に立てたなら、ぜひ同僚にも共有してください。
Share Buy the Rumor; Sell the News
コメントや質問、ご意見はいつでも歓迎します。私と直接つながりたい場合は、以下のいずれかをご利用ください。
- Twitter でフォローする
- LinkedIn で接続する
- メールを送る(dave [at] davefriedman dot co へ。ドメインは .com ではありません!)
データ取得用のスクリプトは GPT 5.6 Sol が作成しました。データソース:Vast.ai offer-search API。スナップショット撮影日:2026 年 7 月 23 日。
原文を表示
*Please see relevant disclosures here.*
I recently surveyed1 GPU rental prices on Vast.ai, the compute rental marketplace. An H200 rented for about $3.93 per GPU-hour. Without knowing anything else, it looks pretty straightforward. You pick the chip, you multiply the hourly rate by the number of GPUs you need and the number of hours you need, and you get an invoice.
But if you want four identical H200s in the same machine, half the qualifying supply disappears, and the remaining supply costs 4% more. If you want eight, you’re out of luck: there’s nothing in the sample at any price. The rental rate and the price of usable compute are different numbers.
In other words, compute pricing coverage runs on headline rates, but many workloads don’t consume them that way. The headlines say that an H100 costs one amount, and a B200 a different amount. But training, fine-tuning, and high-throughput inference need multiple identical GPUs in the same machine or cluster, with adequate interconnect and a long enough availability window. Eight scattered GPU-hours don’t substitute for one hour on eight co-located GPUs.
To size the gap between headline numbers and prices for real workloads, I pulled on-demand offers from Vast.ai and kept the ones that were verified, currently rentable, available for at least seven days, and attached to a host with a reliability score of 99% or better.
The sample that resulted from this filtering covered five accelerator models: A100 SXM4, H100 SXM, H200, B200 and L40S. Multiple offers often represent different bundle sizes on the same box, so I deduplicated by physical machine. The result was 37 machines and 93 listed GPUs.
Then I asked how much supply could satisfy a requirement for one, two, four, or eight identical GPUs, and what the best available price was at each size. The price measure is the median of up to the three cheapest eligible machines. This limits the pull of a single unusually cheap listing.
An interesting result of this analysis is that larger clusters don’t necessarily cost more. Supply of large clusters thins as the requested cluster grows, usually with no matching move in the quoted price. H200 shows this most cleanly. Going from one GPU to two left 92% of listed physical inventory eligible at essentially the same price. Going to four cut eligible inventory to 50% while the price measure rose 4%, from $3.93 to $4.08 per GPU-hour. At eight GPUs, nothing qualified.
Supply of H100 and B200 were thinner still. Qualifying H100 supply supported a two-GPU request but not a four-GPU one. The B200 sample held two physical machines, both topping out at two GPUs, so four- and eight-GPU availability was zero. A100 behaved like you would expect: price increased as supply declined. Eight-GPU machines existed, and the price measure climbed from $0.60 per GPU-hour at the one-GPU baseline to $1.29, a contiguity premium of 114%. L40S cut both ways. Only 35% of its listed inventory sat in machines big enough for an eight-GPU request, but the one machine that qualified was cheap. This is a thin cross-sectional order book, not a smooth supply curve.
Usually, scarcity affects prices. Think about oil: if demand rises while supply stays constant, prices rise. But this is not what seems to happen with compute. Compute rations through configuration and availability instead. A marketplace can hold hundreds of GPU-hours in aggregate while a buyer still can’t assemble them into a working cluster. The binding constraints are hardware generation, co-location, networking, reliability, reservation duration, CPUs, storage and geography.
But when there are no suitable configurations, scarcity never shows up as a higher transaction price. The trade doesn’t happen. So an index built from one- and two-GPU rental prices makes compute look abundant and stable. Yet at the same time a company scrambling to find a large cluster can’t find one at any price.
This explains why large buyers rely on bilateral capacity agreements. Those contracts secure a specific configuration at a specific time. The excess capacity ends up on marketplaces like Vast.ai.
We can assume that the first compute futures contracts will settle against standardized GPU-hour benchmarks. This makes sense, since derivatives need simple definitions and enough observations to support a credible index. However, a company that buys H100 futures to hedge a future four- or eight-GPU requirement carries real basis risk. The futures will hedge a rise in the broad H100 price and do nothing about the lack of co-located capacity. The hedge pays but the workload doesn’t run.
Energy markets manage this residual risk as location basis and the spread between firm and interruptible supply. Similarly, the compute markets likely will build a version around topology. Think generic GPU-hours at the center, with premiums or separate contracts for cluster size, interconnect quality, region and availability guarantees. And these cluster-hours and capacity options may well end up trading more actively than the base GPU-hour.
Vast.ai is a single marketplace. Its listings are advertised offers, not completed transactions, and its visible supply is not global installed capacity. Several of the larger cluster estimates rest on a handful of machines, so treat the precise premiums with caution.
The availability pattern is harder to dismiss. It held when I varied the minimum reliability threshold between 98% and 99.5%. B200 still had no four-GPU offer. H200 still had no eight-GPU offer. A100 still carried a large-cluster premium. The small sample is partly the point: a headline GPU price can rest on a seemingly active market while the market for one specific usable configuration is almost empty.
If you enjoy this newsletter, consider sharing it with a colleague.
Share Buy the Rumor; Sell the News
I’m always happy to receive comments, questions, and pushback. If you want to connect with me directly, you can:
- follow me on Twitter,
- connect with me on LinkedIn, or
- send an email to dave [at] davefriedman dot co. (Not .com!)
Scripts to retrieve the data were written by GPT 5.6 Sol. Data source: Vast.ai offer-search API. Snapshot taken July 23, 2026.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み