AI ラボの収益目標と需要増により計算コストが 10 倍に上昇する可能性
AI ラボの収益が急拡大する未来において、計算資源の使用効率化や推論コストへの依存が高まる中、計算リソース自体の価格が10倍に跳ね上がる可能性が示唆されている。
AI深層分析を開く2026年7月30日 23:13
AI深層分析
キーポイント
収益と計算資源の乖離構造
ラボの収益が年率10倍になる一方で計算資源の使用量が3倍程度しか増えない場合、利益率の向上、計算価格の高騰、または推論への支出比率拡大のいずれかが必須となる。
市場における3つの要因の進行
Anthropic の推論での利益率が40%から80%超へ上昇し、スポット価格も底値から40%以上高騰しており、OpenAI などの計算支出に占める推論の割合も25%から50%近くへと急増している。
推論依存への抵抗と戦略
ラボは計算資源を推論に多く割くことを好まず、これはAI 進歩が停滞したことの表明と同義となるため、現状のモデルは1年後には陳腐化すると見なしてビジネスケースを構築している。
価格高騰の必然性
主要ラボ間の競争優位性が明確でない限り、計算資源の使用効率化による利益率拡大に限界が来るため、計算リソース自体のコスト上昇が避けられないと分析されている。
AI ラボ向け計算リソースの価格高騰
Google と Anthropic は SpaceX から GPU を月額 9 億ドルで借りており、これはスポット価格の約 2 倍である。
重要な引用
Why compute might get 10x more expensive in coming years
Lab margins have to increase, The price of compute has to increase, Labs have to spend a greater fraction of their compute on inference.
If you're spending most of your compute on inference, you're kind of declaring that AI progress has stalled.
Google is reportedly paying $900 million a month for 110K GPUs that are a blend of GB200s and GB300s.
編集コメントを表示
編集コメント
本記事は、AI ラボの収益構造と計算資源コストの関係性について鋭い洞察を提供しており、業界の先行きを読む上で重要な示唆を含んでいる。特に推論コストへの依存度上昇が技術進歩の停滞を示す指標となり得るという指摘は、今後の投資判断や戦略策定において注視すべき点である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
私は、記事の執筆時間を2時間に制限した超短編ブログを試しに書いてみたいと思います。重要なサブ質問のすべてを網羅することはできないかもしれませんが、これこそが私が興味を持つさまざまなトピックについて進歩を遂げる唯一の方法です。
今日は、今後数年間の研究機関における計算資源(コンピュート)の状況についてお話しします。
Anthropic の収益は前年比で10倍になりました。同社は今年末までに約1,000億〜1,500億ドルの収益を達成する可能性があります。この傾向が続くためには、来年末までに1兆ドルの収益を上げる必要があります。なぜこの傾向が継続しなければならないのかという深い理由はなく、むしろ継続しない可能性も十分にあります。究極的には、これはAI の能力に関する問題です。しかし、仮にこの傾向が続いたとしましょう。その世界では何が真実でなければならないのでしょうか。
研究機関の計算資源は前年比3倍になっています。収益を10倍にしつつ、計算資源の使用量を3倍に抑え続けるためには、以下の3つの要素のうちいずれか、あるいは複数の組み合わせが実現されなければなりません。1. 研究機関の利益率が高まること、2. 計算資源の価格が上昇すること、3. 研究機関が計算資源を推論(インファレンス)に割く割合を増やすことです。
私の理解では、基本的に以下の3つの現象が同時に進行しています。
第一に、AnthropicのFable推論における利益率は2025年の40%から、今年にはおそらく80%を超えると見られています[1](ただし、限界計算リソースについては後述する通り、必ずしもこの数字に反映されているわけではありません)。
第二に、スポット価格での計算リソースは2月の安値から40%以上上昇しており、実際にはラボが支払う金額の方がさらに多い可能性を秘めています(これも後述します)。
第三に、Epochのデータによると、OpenAIの2024年の計算リソース支出のうち約4分の1が推論に充てられていました。しかし現在では、その割合は50%以上に達している可能性が高いです。
ラボ側としては、この3番目の現象(計算リソース支出のより大きな割合を推論に回すこと)にはなりたくないと考えています。冗談めかして言われるように、推論からの収益を得る目的は、投資家により多くの資金を提供してもらい、より大きなモデルを訓練するための計算リソースを購入させるためです。もし計算リソースの大部分を推論に使ってしまえば、それは実質的に「AIの進歩が停滞した」と宣言しているようなものです。つまり、さらに訓練に投資する価値がないと判断され、ビジネスの本質は単なるクラウドプロバイダーと同じものになってしまうからです。
しかし、ラボ側はこの見解には同意していません。彼らは、1年後には非常に使い物にならなくなるモデルをサービス提供することで、より賢いモデルの継続的な訓練というビジネスケースを構築しようとしていると考えています。
残る選択肢は、1 つ目の実験室の利益率が高まるか、2 つ目の計算リソースのコストが跳ね上がるかのどちらかです。もし先頭を走る 1〜2 の実験室が他社と比べて圧倒的な差をつけていれば、前者のケースが優勢になるでしょう。市場において利益率は、自社の優位性が次に良い代替案よりもどれだけ大きいかによって決まります。
効果 1 が支配的となるためには、来年末までに利益率が 95% 前後に達している必要があります。私にはその数字は非常に不自然に思えます。しかし、AI 実験室の収益が今後数年間で驚異的なペースで成長し続けるというシナリオであれば、十分にあり得る話だと言えます。
では、来年末までに 1 兆ドル規模という異常な収益世界が実現する背景を説明するもう一つの要素、つまり計算リソースのコストが大幅に上昇するという点について解説しましょう。すでにこの動きは始まっています。特に実験室が必要とする計算リソースの価格高騰は顕著です。彼らは当然ながらスポットインスタンス(需要に応じて価格が変動する一時的なサーバー)には頼れません。重み付けデータや顧客情報のセキュリティを確保し、かつ高い利用率と柔軟性を得るためには十分な規模のインフラが必要です。
この特定のセグメントにおける計算市場がいかに異常な状態にあるかを示す例として、Google と Anthropic が SpaceX から計算リソースを借りている価格に注目してください。報道によると、Google は GB200 と GB300 を組み合わせた 11 万個の GPU を月額約 9 億ドルでリースしているそうです。これは同種の GPU のスポット価格の時間単価と比較して約 2 倍に相当します。さらに、現在のスポット価格自体も 2 月の時点からすでに 40% 上昇しています。
ここで強調したいのは、AI モデルが賢くなるほど、同じ計算リソースからより多くの収益を生み出すようになるという結論です。もし H100 相当のハードウェア上で動作し、現在のソフトウェアエンジニア市場価格に匹敵する能力を持つ真の人並みの AI エンジニアが登場したとすれば、その H100 の年間レンタル料は 25 万ドルを超えるはずです。これは現在のスポット価格の約 15 倍です。
もちろん、ソフトウェアエンジニアが 1,000 万人も増えれば、一人あたりの付加価値は低下し、結果として H100 が現在よりも 15 倍の収益を生み出すことはできないだろうと考えるかもしれません。しかし、私はそれが本当に正しいとは限りません。もしこの議論を AI ではなく人間に対して適用すれば、それは古典的な「労働の塊の誤謬(lump of labor fallacy)」に陥ることになります。経済学者の多くは、専門化やイノベーションが労働の価値を高めるため、高度なスキルを持つ移民の流入が長期的には賃金を下げないと考えています。もしかすると、今回の労働供給の急増はあまりにも規模が大きくスピードも速いため、この一般的な知見が適用できなくなるのかもしれません。しかし、もし標準的な経済学の原則に従うなら、計算リソースの限界価値(ひいてはその限界価格)は驚くほど高くなるはずです。
そのような世界では何が起きるのでしょうか?
最先端モデルが計算リソースの収益化を巧みに行うようになるほど、後追いは困難になっていきます。もし 2028 年までにソフトウェアエンジニアリングが自動化され、計算コストが現在の 15 倍に跳ね上がっていたとすれば、収益源を持たないあなたが、最先端ラボとの間で計算リソースを巡る競争で勝つのは極めて難しくなるでしょう。
最も優秀で効率的なモデルを訓練できれば、今日よりもはるかに高いマージンを設定できるようになります。これはアルチアン=アレン効果の一例です。異なる品質を持つ 2 つの商品に固定費が加わると、高品質な商品の方が相対的に安くなり、需要が高品質側へシフトします。H100 を 1 時間あたり 20 ドルで使用する状況では、同じ結果を得るために計算リソースをより多く消費する、劣った非効率なモデルを使うのは極めてコストが高く、愚かな行為となります。つまり、計算資源をより効率的に活用できるモデルを開発できれば、ラボは非常に大きなプレミアム価格を設定できるようになるのです。元々、基盤となる計算リソースに対して(直接的であれ間接的であれ)莫大な費用を支払う必要があるなら、同じ量の計算リソースで最も効率よく動作する最高峰のモデルを使うために追加料金を支払う方が合理的です。
現在、多くの人気のある AI アプリケーションは利用コストが高騰してしまいます。AI が現時点で人間労働と比較すれば比較的安価である理由は、トップレベルの人間がこなせるタスクをまだすべて実行できないからです。しかしそのうち、この状況も変わります。そうなれば、GPU を使って短尺動画の質の低いコンテンツ(スロップ)を生産する行為は、コスト面で淘汰されることになるでしょう。
同時に、こうした予測は過去の「希少性に関する人々の誤解」というパターンに当てはまります。例えばサイモン=アールリッヒの賭けがそうです。ポール・アールリッヒは、1990 年までの 10 年間に商品バスケットのコストが増加するのではなく減少すると予想し、この賭けを行いました。この賭けは人気経済議論で有名ですが、これはアールリッヒのマルサス的世界観が誤りであったことを示す例として語られることが多いです。つまり彼は、市場シグナルや人間の創意工夫がいかに希少な投入資源をより効率的に利用する道を見つけ出すかを過小評価していたのです。(ただし別の分析では、もしこの賭けが異なる年代に行われていれば、アールリッヒが勝っていた可能性も示唆されています。)
私は、シモン=アールリッヒの商品バスケットが計算資源(コンピュート)の適切な参照対象ではないと推測しています。なぜなら、金属の採掘に比べて計算資源の供給ははるかに非弾力的であり、需要の急増を吸収する能力も著しく低いからです。
年間の計算容量が 3 倍になるという数字は、以下の要素を掛け合わせた結果です。1.4 倍はムーアの法則によるもの、1.2 倍は新ファブ(工場)の建設によるものですが、これは少なくとも 2030 年まで EUV ツールの供給不足によってボトルネックになります。そして 1.8 倍は、AI が他のデバイスから最先端ウェーハの割当を奪い取ることで生まれるものです。
この AI の需要が 2027 年末に壁にぶつかるのは確実で、その頃には N3 プロセスにおける AI の占める割合が現在の 60% から 86% に達します。これらの入力要素のいずれかを組み合わせ直して計算スケーリングを劇的に加速させるのは困難であり、特に最後の要因については、ウェーハ割当が飽和状態に達したことで、1〜2 年以内に成長が止まるでしょう。
将来、計算資源は再び安くなる時期が必ず訪れると強調しておきたいです。その時、ロボットがシリカ砂の海岸や銅鉱山からコンピュータを製造できるようになるはずです。そうなれば、計算資源の価格は原材料費と工具のコストに近づき、低下するはずです。
私が今話しているのは、現在の状況です。AI の計算能力は年間で 3 倍になっていますが、これは AI が年々どれほど有用になっているかによる価格上昇効果を相殺するには十分ではありません。
ところで、アントの収益が年間で10倍に拡大している一方で、計算リソースは3倍程度しか伸びていないという事実は、モデル事業において極めて強い規模の経済(スケールメリット)が存在することを示唆しています。これは論理的にも納得できる話です。モデルを学習させる際、一度きりのコストで多様なスキルを獲得し、それをすべてのユーザーに按分して利用できます。一方、人間の労働では、各ケースごとにゼロから再トレーニングが必要となるため、この点とは異なります。
私は知能の集中によるパワーの偏在が懸念されるため、このような強い規模の経済が存在する世界で暮らしていることは望ましくないと考えています。しかし、現実にはそうであるようです。
推論のブレンデッドマージンは70%を超えるとされています。これに基づけば、Fable APIのマージンも80%を超える可能性が高いと推測されます。これは直感的な判断に基づく総括的な見解です。
原文を表示
I want to experiment with very quick blog post where I time-box writing for 2 hours. I won’t be able to nail down a lot of important sub-questions, but the alternative is just not making any progress on a lot of different topics I’m curious about.
Today I want to talk about the compute situation of the labs over the coming years.
Anthropic revenue has 10xed year over year. Anthropic likely ends the year with ~$100–150B of revenue. For this trend to continue, Anthropic would have to make $1T in revenue by the end of next year. There’s no deep reason why the trend needs to continue, and it very well might not - it’s ultimately a question about AI capabilities. But suppose it does. What would have to be true about that world?
Lab compute 3x-es year over year. For a lab to 10x revenue while continuing to only 3x compute, some combination of the following 3 things has to happen: 1. Lab margins have to increase, 2. The price of compute has to increase, 3. Labs have to spend a greater fraction of their compute on inference.
My understanding is that basically all 3 of these things have been happening: 1. Anthropic went from 40% margins in 2025 to probably >80% this year for Fable inference1 (though perhaps not on the marginal compute - see more below). 2. Spot prices for compute are up 40%+ from the February trough, and that likely understates how much more the labs have to pay (again, more below). 3. Roughly a quarter of OpenAI’s 2024 compute spend went to inference according to Epoch, and it’s certainly closer to 50% if not higher now.
Labs would prefer *not* to do 3 (spend greater and greater shares of compute on inference). As has been said jokingly, the point of inference revenue is to convince investors to give you more money to buy more compute to train bigger models. If you’re spending most of your compute on inference, you’re kind of declaring that AI progress has stalled, because it’s not worth investing more in training, and your business is basically that of a cloud provider. The labs do not think this is true - they think they are serving models that will look extremely shitty within a year in order to build up the business case to continue training smarter models.
So that leaves 1. lab margins will increase, or 2. compute gets more expensive. It will be more of the former if the leading 1 to 2 labs are significantly ahead of the competition. In a market, your margins are set by how much better you are than the next best alternative. For effect 1 to dominate, margins would have to be in the mid-90s percent by the end of next year. That sounds quite crazy to me. But it does sound plausible to me that AI lab revenues will keep growing astonishingly fast.
So that leaves one more effect to explain how this world of crazy $1T revenue by end of next year might come to be: compute gets a lot more expensive. As I mentioned, this is already starting to happen. The price increase is even stronger when we look at the tranche of compute that the labs need - they obviously can’t rely on spot instances - they need security for their weights and customer info, and enough scale to get good utilization and flexibility. To look at how crazy the compute market is in that tranche, consider the price at which Google and Anthropic are renting compute from SpaceX. Google is reportedly paying $900 million a month for 110K GPUs that are a blend of GB200s and GB300s. That’s roughly 2x the spot price per hour for those GPUs. And the current spot price is itself 40% higher than it was in February.
I want to emphasize the key conclusion here: as AI models become smarter, they’ll better monetize the same amount of compute. If a true human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That’s 15x today’s spot prices.
Of course you might expect that if we have 10 million extra software engineers, the marginal value of a software engineer would decrease and so that H100 wouldn’t necessarily be able to produce 15x more revenue than it currently does. But I actually don’t know if that’s true. If we applied this argument to people instead of AI, then this would be the classic lump of labor fallacy. Economists generally believe that high-skilled immigration does not decrease wages in the long run because of how specialization and innovation increase the value of labor. Maybe this labor supply shock is so big and so fast that this general heuristic no longer applies. But if we believe what standard economics says about labor, then the marginal value of compute (and thus the marginal price of compute) should become astonishingly high.
What would happen in such a world?
- As the top models get better and better at monetizing compute, it becomes harder and harder to catch up. If by 2028 we’ve automated software engineering and the price of compute is 15x higher than it is right now, then it’s going to be much more difficult for you, with no revenue, to compete for compute against the frontier labs.
- If you can train the best, most efficient model, then you’ll be able to charge MUCH higher margins than you can today. This is the Alchian–Allen effect: when a fixed cost gets added to two goods of different quality, the premium good becomes relatively cheaper, so demand shifts toward it. At $20 per H100-hour, it’s going to be extremely costly and stupid to use a weaker, less efficient model, because it’s going to burn more tokens running your expensive compute to get the same result. So labs will be able to charge a very large premium if they can train a model that better economizes this scarce resource. If you’re gonna have to pay so much for the underlying compute in the first place (directly or indirectly), you might as well pay extra to use the very best, most efficient model that runs on that same amount of compute.
- A lot of current popular applications of AI get priced out. The reason AI is relatively cheap right now, at least in comparison to human labor, is partly that it can’t do a lot of things that top humans can do. At some point that will no longer be the case. And so using GPUs to make short-form video slop will just get priced out.At the same time, this kind of prediction does pattern match onto the ways people have been wrong about scarcity in the past. I’m thinking of the Simon–Ehrlich wager, where Paul Ehrlich bet that the cost of a basket of commodities would increase rather than decrease in the decade leading up to 1990. This wager is famous in popular economics discussion, because it’s supposed to illustrate how Ehrlich’s Malthusian worldview was falsified — he underestimated the way in which market signals and human ingenuity would find ways to better economize scarce inputs. (Though other analysis shows that if this bet had been made in a different decade, Ehrlich would have won).
- I’m guessing the Simon-Ehrlich basket of commodities is not the correct reference class for compute, because compute supply is much less elastic, and much less capable of absorbing large demand shocks, than the extraction of different metals is. That 3x in compute capacity per year comes from multiplying the following numbers together;: 1.4x from Moore’s Law, 1.2x from new fab construction (bottlenecked through at least 2030 by EUV tool supply), 1.8x from AI taking leading edge wafer allocation from other devices (which will start to hit a wall by end of 2027, when AI will go from 60% of N3 to 86%). Seems hard to shift any 3 of those inputs into compute scaling much more, and the final one will actually hit a will within a year or two as wafer allocation gets saturated.
- I want to clarify that at some point in the future compute becomes cheap again. At some point robots will be able to turn shores of silica sand and mines of copper into computers. At which point the price of compute should fall closer to the cost of raw inputs and tools. I’m just talking about this current regime where AI compute merely 3x-es year over year, which is not fast enough to offset the price effect of how much more useful AI is becoming year over year.
- By the way, the fact that Ant revenue has been 10x-ing year over year, while compute has only been 3x-ing, suggests that there are very strong economies of scale in the model business.Logically this makes sense - when you train a model, you pay this one time cost of learning all these different skills that you can then amortize across all your users. (Unlike with human labor, where each instance has to be retrained from scratch).I wish we didn’t live in a world with such strong economies of scale of intelligence (because I’m worried about power concentration). But it seems we do.
Blended inference margins are >70%, so I’m guessing Fable API margins are >80% - total vibe claim.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み