170 社調査:AI インフラ導入は加速、コスト管理は遅れ
本文の状態
日本語全文を表示中
詳細モードで約25分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
VentureBeat AI
ベンチャーベイトの調査によると、企業の AI インフラ導入は生産環境で定着したが、コスト管理の遅れにより GPU 利用率が低く、次期投資先として専門クラウドへの移行が進んでいる。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 17:08
AI深層分析
キーポイント
購買判断基準の変化
企業は総所有コスト(TCO)よりもパフォーマンスや GPU の入手可能性を優先し、成功指標としてもコストではなく稼働率と信頼性を重視する傾向が明確になっている。
コスト管理の遅れと非効率
自社で GPU を運用する企業の 69% が利用率 50% 以下であり、AI コストを厳密に追跡できる企業は半数未満にとどまり、コスト対効果への満足度が最も低い。
次期投資先の方向性
現在のスタックからの脱却として、44% の企業が AI 専門クラウドの評価を最優先計画に掲げており、実際に利用している企業は 5% 未満というギャップが存在する。
AI 専用クラウドへの移行と非 NVIDIA アクセラレータの台頭
44% の企業が AI 専用クラウドを次回の支出計画で最優先しており、そのネット動向は最も強い。同時に、非 NVIDIA アクセラレータへの関心も 39% に達している。
プロバイダーの切り替え意向と現状の使用状況
企業の 62% が今後 12 ヶ月以内にプロバイダーの切り替えまたは追加を検討しているが、現在の使用量は CoreWeave や Lambda を含め新参クラウドは依然として低い。
重要な引用
performance and GPU availability now outrank total cost of ownership
among the 155 enterprises that operate their own GPUs, 69% report utilization of 50% or less
AI-specialized clouds are the top planned evaluation area at 44%
AI-specialized clouds are the top planned evaluation area at 44% and carry the strongest net momentum of any infrastructure approach (+36)
編集コメントを表示
編集コメント
企業の AI 戦略が「速度と可用性」に偏りすぎた結果、コスト管理という基礎的な課題が深刻化している現状は警戒すべき信号である。導入の成熟度が高まるほど、リソース効率の可視化と専門クラウドの適切な活用が競争優位性の鍵となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
170 社にわたる調査で、AI インフラは決定的に本番環境へ移行したことが明らかになりました。現在、3 分の 2 の企業が AI ワークロードを実稼働させており、3 割がスケーリングされた運用を行っています。しかし、そのインフラにかかるコストを把握する能力は、この急速な普及に追いついていません。
企業は購買判断において、コストの優先順位を静かに下げました。今や総所有コスト(TCO)よりもパフォーマンスと GPU の入手可能性が重視され、価格よりも信頼性が成功の指標となっています。本番環境のプレッシャーにさらされているチームにとっては、この順序変更は合理的な選択です。しかし、それは不快な事実へと突き当たります。AI コンピュートのコストを厳密に追跡できている企業は半数に満たず、多くの GPU が稼働容量の半分以下で運用されています。さらに、次なる投資先として注目されているのは特殊化されたクラウドですが、実際にそれを利用している企業は 20 社に 1 社未満です。
今回の VentureBeat Pulse Research は、企業の AI インフラとコンピュート動向を調査しました。各組織が導入プロセスのどの段階にあるか、現在どのような基盤上で AI を稼働させているか、どのように購入・測定しているのか、そして次なる投資先はどこに向かっているのか。最も示唆に富むのは、その根底にあるコンピュートの経済性を、どれほど明確に見えているかという点です。
これは運用段階にある企業群の現状です。企業の 2 分の 3(66%)が AI ワークロードを実環境で稼働させており、そのうち 29% は大規模な AI の実運用を報告しています。まだ AI ワークロードを一切導入していないのはわずか 4% に過ぎません。
この成熟度はインフラストラクチャの構成にも表れています。平均的な企業は 3 つのプラットフォームを併用しており、OpenAI(49%)、Google Gemini(48%)、Microsoft Azure(47%)、Google Cloud(42%)がほぼ半数の環境で採用されています。主要なプラットフォームとして一つだけ選ぶよう問われた場合、Azure が 26% で首位に立ちます。
最も重要な変化は、企業がどのように判断を下すかという点にあります。既存のクラウドやデータスタックとの統合性が依然として最大の選定基準(40%)ですが、パフォーマンス(レイテンシとスループット)が 35% で第 2 位に浮上し、GPU の入手可能性も 24% で第 3 位となりました。これらは総所有コスト(TCO)の 22% を上回っています。
評価基準においても同じ傾向が見られます。稼働率と信頼性が主要な成功指標であるという回答が企業の 51% に、開発者生産性が 39% に達し、どちらも「100 万トークンあたりのコスト」(31%)を上回っています。実運用のプレッシャーにさらされている企業は、速度と可用性を求めて導入を進め、評価基準としても優先度を高めています。その結果、コストに関する項目は相対的に順位を下げています。
経済状況がコントロールできていれば、これは特筆すべきことではないでしょう。しかし実際には、そうではありません。自社で GPU を運用している 155 の企業のうち、69% が稼働率を 50% 以下と報告しており、半分を超える利用率を達成できているのはわずか 23% に過ぎません。さらに 12% は稼働率自体を測定していません。
AI コストとそのリターンを厳密に追跡している企業は半数未満(47%)で、大規模な AI 運用を行っている企業に限っても、この数字は 56% にしか達しません。「コスト対効果」に関する満足度は、全体的な満足度(4.14 点)と比較して最も低く 3.87 点でした。これは、測定を行わないと評価が難しい項目において、特に不満が目立っていることを示しています。
次の支出の方向性は、現在のスタックから離れるものへと向かっています。AI 特化型クラウドは、44% の企業が最も重点的に検討する領域であり、インフラストラクチャのアプローチの中で最も強いネットモメンタム(+36)を記録しています。しかし、CoreWeave や Lambda のような特定のプレイヤーの現在の利用率はそれぞれ 3.5% に過ぎず、その他のニュークラウド企業の利用率も 3% を下回っています。
一方、Nvidia 以外のアクセラレータへの関心は 39% に達しています。また、62% の企業が今後 12 ヶ月以内にプロバイダの切り替えや追加を検討しているものの、検討対象となっているのはすでに運用中の大手企業に集中しています。
調査方法
VentureBeat は、企業向け AI インフラストラクチャ、計算リソース、推論コストに関する経済性を調査対象とした「Pulse Research」シリーズの一環として本アンケートを実施しました。回答は従業員数 100 人以上の組織に限定されており(サンプル数 n=170)、最も小規模な 1~100 人のセグメントは除外されています。データは 2026 年 7 月の単一波から抽出されたもので、複数月にわたる集計ではありません。そのため、本レポートは横断的な分析に留まり、月次間の推移を推測するものではありません。すべての数値は 2026 年 7 月の調査結果のみに基づいています。また、一部の質問では複数の選択肢を選ぶことが可能だったため、回答割合の合計が 100% を超える場合があります。
組織規模別に見ると、本調査は同シリーズで通常見られる中堅企業中心の構成から、より大規模な市場層へと広がっています。従業員数 251~1,000 人(28%)と 1,001~5,000 人(25%)が主要な回答者層を占め、そのほかに 10,001 人以上(19%)、101~250 人(15%)、5,001~10,000 人(12%)が含まれています。つまり、回答者の 57% が従業員数 1,000 人以上の組織に所属しています。
役職別には、管理職が 48% を占め、次いで個別貢献者(Individual Contributors)が 27%、C レベル役員が 12%、VP やディレクターが 9% です。購買権限については信頼性の高い回答が集まっており、AI ソリューションの最終決定権を持つ層が 39%、推奨や影響力を行使する層が 43% を占めています。
業界別では、テクノロジー・ソフトウェア分野が 35% と最大規模を占め、次いで製造業(14%)、金融サービス(12%)、ヘルスケア・ライフサイエンス(9%)となっています。
回答者数は 170 名で、方向性を示す指標として十分な信頼性がありますが、これは精密な測定値というよりは傾向を示すシグナルと捉えるべきです。このデータは確率サンプリングではなく自己選抜によるものであり、大規模なハイパースケール事業者の視点ではなく、実際に AI インフラを構築・運用している組織の声を反映したものと解釈するのが適切です。
知見 1:3 分の 2 がパイロット段階を超えている
現在、企業の 3 割が AI を本番環境で大規模に稼働させています
各企業が AI 導入のどの段階にいるかを尋ねたところ、この調査対象グループはもはや実験段階にはなく、すでに次のフェーズに進んでいることがわかりました。
企業の 66% が AI ワークロードを生産環境で運用しており、29% は「大規模な本番運用」を行っているとしています。一方で、概念実証(PoC)の段階にあるのは 30% に過ぎず、まだ着手していないのはわずか 4% です。これは同シリーズの過去の調査と比較して、はるかに運用フェーズに踏み込んだサンプル構成となっています。この傾向は、回答者の 57% が従業員数 1,000 人以上の大企業に属しているという、上位層中心の構成と一致しています。
この成熟度が、その後のすべての議論の土台となっています。本レポートで取り上げられるインフラストラクチャに関する意思決定は、まだパイロット段階にあるチームではなく、実際に稼働しているワークロードを抱え、現実的なコストに直面している組織によって行われています。これが、第 5 の発見で示された購入基準の再編成を説明する背景です。ここでは、コストよりもパフォーマンスと可用性が優先されます。これは、ライブシステムを運用するチームにとっての当然の帰結です。
また、第 6 と第 7 の発見における重要性も高まります。実験段階でリソース利用率やコストを計測できない企業には計画上の課題があり、大規模な本番環境でそれらを把握できない企業には運用上の課題があるのです。
発見 2:スタックは「ハイパースケーラーと API」の組み合わせであり、プラットフォームは平均して 3 つある
専門的な GPU クラウドはまだ存在感を示していない
企業が現在 AI の実行にどのプロバイダーやプラットフォームを利用しており、どれを主軸としているかを尋ねました。答えは依然として既存の大手です。複数の大手を同時に利用しています。
現在のスタックは「ハイパースケーラーと API」の組み合わせであり、単一ではありません。企業は平均して 3 つのプラットフォームを名挙げています。汎用クラウドと主要なモデル API が、現在の実装のほぼすべてを占めています。具体的には、OpenAI、Gemini、Azure、Google Cloud の 4 つのプラットフォームが、企業の 10 社中 4 社以上で利用されています。
その中で「最も主軸とするプラットフォーム」を一つ選ぶよう求められた場合、Microsoft Azure が 26% で首位に立ち、Google Cloud が 19% で続きます。一方、モデルプロバイダー全体では、OpenAI(14%)、Gemini(14%)、Anthropic(8%)を合わせると、35% の企業がこれを主軸として選んでいます。
AI インフラの話題を独占している専門的な「ネオクラウド」GPU プロバイダーは、実際にはまだ限定的な存在に過ぎません。CoreWeave と Lambda はそれぞれスタックの 3.5% に登場し、Baseten は 3% です。Crusoe、Nebius、Fireworks、Together、Anyscale はいずれも 2% 以下です。これらを合わせると、主要プラットフォームとして名指しされるのは企業のわずか 1% に過ぎません。
一方、13% の企業が独自でオープンソースのセルフマネージド型スタックを運用しており、9% が自社の GPU クラスタを稼働させています。どちらも、専門クラウド全体のカテゴリーよりも規模が大きいのです。この対比こそが、「発見 3」における評価意図を注意深く読むべき理由となります。
これらの数値を読む際の注意点:方法論セクションで説明した通り、本調査のサンプルは自己申告型であり、回答者は利用するすべてのプロバイダーをカウントしました(平均 3.0 件選択)。したがって、この数値は支出額や主要プラットフォームとしての地位ではなく、スタック内での存在度を表しています。どこに重心が置かれているかをよりよく示すのは、「主要プラットフォーム」に関する別の質問です。
このような方法で構築されたサンプルでは、市場全体の支出を重み付けした調査とは異なるプロバイダー構成が現れます。これらの数値は、AI 活用層が現在運用している環境の姿として読み解き、業界全体での市場シェア推計との乖差は、サンプル固有の性質と捉えてください。
発見 3:次の投資先は、まだ導入していないインフラへ
AI 専門クラウドが評価リストで首位に立ち、最も強い勢いを示しています
今後12ヶ月にわたり、企業がどのAIインフラストラクチャを評価する予定か、また各カテゴリで活用を増やすか減らすかを想定しているかを尋ねたところ、両方の回答は現在のスタックから離れる方向を示しています。
このレポートが示す最も鋭い緊張関係は、本シリーズがこれまでに連続する波の中で記録してきたものと同じです。最も多く挙げられた評価予定領域である「AI 特化型クラウド」は44%に達しますが、実際にこれらの企業のうち3.5%しか利用していないという事実(ファインディング2)と対照的です。ほぼ4割(39%)がNvidia製以外のアクセラレータの評価を検討しており、その1/4が次世代のNvidiaシリコンを、さらに分散型コンピューティングネットワークでさえも18%が評価対象としています。
方向性を問う質問は単なる反復ではなく、この傾向を裏付けています。各アプローチについて今後活用を増やすか減らすか、あるいは現状維持かを尋ねたところ、企業はAI特化型クラウドを最も高いネットモメンタム(+36)として位置づけました。これは「より多く行う」と回答した42%に対して「少なくする」が6%という結果で、推論API(+34)やハイパースケイラー(+30)を上回っています。一方、オンプレミスおよびコロケーションインフラはネットモメンタムが僅か+5と後れを取り、22%という相当な割合の企業が利用を縮小すると回答した唯一のカテゴリとなっています。オンプレミス以外のすべてのアプローチでネットでの拡大が見られ、特化型クラウドは最も小さなベースから急速に拡大しています。
現在の利用状況と照らし合わせると、これは単なる漸進的な調整ではありません。これは、企業が数回の波にわたって示唆し続けてきたがまだ実行に移していない基盤再構築の最前線です。
評価率が 44% である一方、実際の利用率は 3.5% に過ぎないというギャップは、このデータセットにおける「意図から行動へ」の転換幅として最大級です。この格差がどう解消されるか——新クラウド企業が評価を本番導入へと転換できるか、それとも大手クラウド事業者が自社の AI インフラで需要を取り込むか——が、このカテゴリにおける最大の未解決課題となっています。
【発見 4】6 割の企業が移行を検討中。特に既存ベンダーの間でその傾向が顕著
高い離脱意向を持つ一方で、検討対象はすでに運用しているスタックに限定されている
企業はいつ、どのインフラプロバイダへ切り替え、あるいは追加する予定なのか、また現在どのようなプロバイダを検討しているのかを尋ねました。
計算リソースのような基盤的なカテゴリにおいて、これは相当な規模の移行意向です。企業の 62% が今後 12 ヶ月以内にプロバイダの切り替えまたは追加を検討しており、そのうち 29% は来四半期だけでも実施する計画を立てています。現状維持を計画しているのはわずか 39% です。
関心が向かう先こそが、より有用なシグナルです。最も移行を検討されているのは、すでに導入済みのプロバイダーたちです。OpenAI と Google Cloud はそれぞれ 29%、Microsoft Azure が 28%、Gemini が 25%、Anthropic が 16%、Oracle Cloud が 14%、AWS が 13% です。
一方、調査 3 で評価リストの上位に挙がった専門クラウドへの移行検討は、これらに比べてはるかに限定的です。CoreWeave は 4%、Lambda は 3.5% に過ぎず、残りは 2% 以下となっています。さらに 8% の企業は現在評価中ですが、候補リストはまだ固まっていません。
この 2 つの知見に矛盾はありません。それぞれ異なる時間軸で動いているからです。調査 3 で示されたニュークラウドへの関心は、「AI コンピュートが最終的にどこで稼働すべきか」という 12 ヶ月規模の評価シナリオです。一方、直近の四半期における移行は、既存プレイヤー同士のシェア争いと、すでに契約を結んでいるプロバイダー間での支出の集約が主因です。
44% という評価率を近未来のピプラインと捉えるベンダーには、4% の検討率という事実も併せて考慮すべきでしょう。
調査 5:コストよりもパフォーマンスが優先される
総所有コストは、レイテンシや GPU の入手可能性に後れをとる
企業が AI インフラストラクチャプロバイダーを選定する際、何を最も重要視するか。また、稼働開始後に成功の主要指標として何を測るかについて尋ねました。その答えはいずれも「価格」から離れています。
既存のシステムスタックとの統合は、40% の企業が最も重視する要因であり、すでに3つのプラットフォームを運用中で、適合しない第4のプラットフォームを追加したくないという層の傾向と一致しています。しかし、それより下の項目については状況が変化しています。
パフォーマンスが2位(35%)、GPUへのアクセスと可用性が3位(24%)となり、どちらも総所有コスト(TCO:22%)を上回っています。微細なオートスケーリングは18%、百万トークンあたりのコストは16%で、以前はこのシリーズでは突出した低い値でしたが、今でも最下位です。
測定基準についても同様の傾向が見られます。稼働率と信頼性が51%の企業にとって主要な成功指標であり、開発者の生産性やデプロイ速度(39%)、百万トークンあたりのコスト(31%)、レイテンシ(27%)、スループット(25%)を大きく引き離しています。つまり、運用上の指標が経済的な指標を圧倒的に上回っているのです。
これはファインディング1で示された本番環境のチームにとって一貫した姿勢です。稼働中のワークロードを扱うチームは、まずシステムが安定して稼働するか、そしてどれだけ迅速にリリースできるかを最優先します。しかし、この姿勢はファインディング7と矛盾する点があります。
総所有コスト(TCO)が購入基準として4位に後退したまさにその時、53% の企業が計算リソースのコストを厳密に追跡できていないという事実があります。不都合な読み方をすれば、コストがランキングで下位になったのは、スタックの中で最も見通しが悪い要素だからであり、測定できない基準は測定可能な基準に敗れる傾向があるためです。
ファインディング6:GPU はより高温になるが、大半はまだ冷ややかだ
GPU オペレーターのおよそ 7 割が、稼働率を 50%以下と報告
企業が実際に GPU 容量のどの程度を活用しているかについて調査を行いました。このデータは、自社で GPU を運用している 155 社を対象としたものです。なお、API のみを利用して自社運用を行っていない 15 社は対象から除外されています。
現在稼働中の計算リソースは、以前に比べてもまだ冷ややか(非効率)な状態にあります。GPU を運用する企業のうち約 7 割(69%)が、稼働率を容量の半分以下と回答しました。特に 26〜50%の範囲にある企業が 46%を占めています。さらに、25%以下の稼働率で動いている企業も全体の 4 分の 1(26%)に達しています。
一方で、50%を超える稼働率を実現している企業は 23%です。これは単なる誤差ではなく、効率化に成功した少数派と言えるでしょう。
さらに気になるのは、稼働率を全く測定していない残りの 12%です。彼らは「非効率だ」と主張することも、「無駄を検出する」こともできないため、両方向から隠れてしまっています。また、稼働率が成熟度とともに向上するという一般的な予想とは異なり、大規模な AI を本番環境で運用している企業でも、50%を超えるのは 22%に過ぎません。これは他のすべての企業における 24%と統計的に有意差がありません。つまり、スケール(規模)を拡大しただけでは、自動的に稼働率の高い fleet が生まれるわけではないのです。
稼働していないアクセラレーターは、コストのかかる資産です。このレポートにおける最も明確な指標の一つがこれです。企業は専門クラウドや次世代半導体の評価を検討している(調査結果 3)一方で、すでに所有するリソースの多くが未使用のまま放置されています。現在の設備には大きな効率化の余地があり、その一部ではコストすら測定できていない状況です。
調査結果 7:支出の内訳を把握できる企業は半数に満たない
大規模運用企業でも厳格なコスト追跡は 56% に留まる
企業が AI インフラへの投資額とそのリターンを定量化できているか、また現状の運用に対する満足度はどうなのかを尋ねました。会計上の信頼性は、支出の増加にはまだ追いついていません。
コスト管理は資金の流れに遅れをとっています。企業の 47% が AI コンピュートの費用とリターンを厳格に追跡しているだけで、大半は部分的な把握にとどまっています(39%)。定量化できていない企業も 15%、優先順位が低いという回答も 6% に上ります。成熟度が高まるにつれて改善は見られますが、根本的な解決にはなっていません。大規模に AI を本番環境で運用している企業でも厳格な追跡は 56% に達するのみで、それ以外の企業では 43% です。このサンプルの中で最も運用面で先進的なセグメントであっても、AI コンピュートの費用やリターンを正確に把握できない企業が 10 社中 4 社以上も存在します。
現在のインフラに対する満足度は、やや肯定的ではあるものの、明らかな偏りがある。5 段階評価で全体としての満足度は平均 4.14、導入の容易さは 4.04 だが、「費用対効果」は 3.87 と最も低く、これが課題となっている。これは、実際には企業がまだ定量化できない経済的な関係性に対して不満を抱いていることを示している。
「発見事項 5」と併せて読むと、この状況は単なる矛盾ではなく、自己強化されていることがわかる。コストは購入基準の中で第 4 位に後退したものの、依然としてスタックの中で最も見えない要素であり、企業もまたこの最も見えない要素を最低評価している。より優れた計測ツールが導入されたとしても、企業が何を購入するかが変わるとは限らない。しかし、パフォーマンスや可用性のために支払っている代償が妥当なものかどうかを知ることは可能になるだろう。
発見事項 8:メモリのフロンティアはまだ未開拓
Dell と Nvidia が先行しているが、市場は散在しており、5 社に 1 社は現状を把握できていない。
大規模推論における新たな制約、すなわち GPU 計算からメモリ(特に KV キャッシュ容量)へのシフトに対応するため、企業はどうするべきか尋ねた。この分野はまだ初期段階であり、分断されている状態だ。
メモリ制約という現実が存在するものの、その管理は依然として不十分です。市場をリードするのは Dell(24%)と Nvidia(21%)で、残りはオープンソースのツール群(12%)、MLA や量子化といったモデルレベルでの効率化技術(11%)、そしてそれぞれが少数派に留まる多数のストレージベンダーへと散らばっています。いずれのアプローチも過半数を占めるに至っておらず、トップ2社の合計でも市場の半分には達していません。
最も示唆的なのは、企業の約 5 社に 1 社(19%)が、この制約を認識していない(7%)か、まだ対策を始めていない(12%)という事実です。推論コストやアーキテクチャを根本から変えるはずの転換点において、これはまだ初期段階で不安定な市場であることを物語っています。また、これは「発見 7」における測定ギャップとも整合しています。現在の計算リソースのコストを定量化できていない企業は、次にどの制約がコスト増を招くかを予測する立場にありません。メモリボトルネックが迫っている一方で、この層の多くはまだ眼前の問題すら把握できていないのです。
結論:速度のために購入し、コストについては目隠し状態
従業員 100 人以上の組織は AI インフラを実運用に移しており、3 分の 2 が稼働中のワークロードを処理し、そのうち 3 割がスケーリングされた環境で動作しています。こうした企業の購買行動も成熟しており、平均して 3 つのプラットフォームを併用し、選定基準として統合性とパフォーマンスを重視します。また、成功の指標としてはシステムのアップタイムと開発者の生産性を挙げています。稼働中のシステムを運用するチームにとって、この優先順位付けは正しいものです。
成熟していないのは、コスト管理の仕組みだ。総所有コスト(TCO)は選定基準の中で4位に転落し、100 万トークンあたりのコストに至っては最下位となっている。その一方で、企業の 53% が計算リソースのコストを厳密に追跡できておらず、GPU オペレーター事業者の 69% が半分の容量以下で稼働している。さらに 12% は利用率を測定することさえしていない。
彼らが運用するインフラにおける「コストパフォーマンス」は、最も評価が低い項目となっている。しかし、その判断を下す際、多くの企業がそれを裏付ける計測ツールを持っていないのが実情だ。コストが重要でなくなったわけではない。ただ、見えない存在になってしまったのだ。そして、購入基準は静かに、「実際に目に見えるもの」を中心に再編成されつつある。
一方、次の支出の行先は現在のスタックを超えている。専門化された AI クラウドが 44% で最上位の評価対象となっており、あらゆるアプローチの中で最も強いネットモメンタムを有している。現状の使用率は 3.5% に過ぎず、近未来での切り替えを検討する割合も 4% だが、これは「意図から行動へ」の転換率における最大のギャップを示すデータだ。
Nvidia 以外のアクセラレーターにも注目が集まり、39% を占めている。そして、その次に来る制約要因——大規模推論における計算リソースからメモリへのシフト——については、5 社に 1 社が認識しておらず、あるいは対応していないのが現状だ。
7 月の調査では 170 社が回答し、このシリーズで通常よりも上位市場層にまでリーチしました。これは方向性を示すデータですが、その傾向は明確です。企業は AI インフラの運用には長けていますが、コスト管理についてはまだ不十分です。今後の調査で問われるのは、再構築の波が来る前にインフラの可視化が追いつくのか、それとも企業が次のレイヤーを購入し続けるのかという点です。
原文を表示
Across 170 enterprises, AI infrastructure has moved decisively into production — two-thirds now run AI workloads live and three in 10 run them at scale — while the ability to account for what that infrastructure costs has not kept pace. Enterprises have quietly demoted cost in the buying decision: performance and GPU availability now outrank total cost of ownership, and reliability outranks price as the measure of success. That reordering is rational for teams under production pressure, but it lands on an uncomfortable fact — fewer than half can rigorously track what their AI compute costs, most GPUs still run at half capacity or less, and the next dollar is aimed at specialized clouds that fewer than one in twenty of them actually use.
This wave of VentureBeat Pulse Research examines enterprise AI infrastructure and compute: where organizations are in their deployment journey, what they run AI on today, how they buy and measure it, where the next investment is aimed, and — most revealingly — how well they can see the economics of the compute underneath it all.
This is an operational cohort. Two-thirds of enterprises (66%) have AI workloads running in production, and 29% describe AI in production at scale, with only 4% not yet running AI workloads at all. That maturity shows in the stack: the average enterprise runs three infrastructure platforms, with OpenAI (49%), Google Gemini (48%), Microsoft Azure (47%), and Google Cloud (42%) all present in roughly half of them. Asked to name one primary platform, Azure leads at 26%.
The most consequential shift is in how enterprises decide. Integration with the existing cloud and data stack remains the top selection factor at 40%, but performance — latency and throughput — has climbed to second at 35%, and access to GPU availability to third at 24%, both ahead of total cost of ownership at 22%. The same ordering governs measurement: uptime and reliability is the primary success metric for 51% of enterprises and developer productivity for 39%, ahead of cost per million tokens at 31%. Enterprises under production pressure are buying and measuring for speed and availability, and have moved cost down the list.
That would be unremarkable if the economics were under control, but they're not. Among the 155 enterprises that operate their own GPUs, 69% report utilization of 50% or less and only 23% clear the halfway mark; 12% do not measure utilization at all. Fewer than half (47%) rigorously track what their AI compute costs and returns, and even among enterprises running AI in production at scale that figure only reaches 56%. Value for money is the weakest of three satisfaction scores at 3.87, against 4.14 for overall satisfaction — the softness landing precisely on the dimension hardest to judge without measurement.
The next round of spending points away from the current stack. AI-specialized clouds are the top planned evaluation area at 44% and carry the strongest net momentum of any infrastructure approach (+36), yet CoreWeave and Lambda each registers at 3.5% of current usage and the rest of the neocloud field sits below 3%. Non-Nvidia accelerators draw 39%. And 62% of enterprises intend to switch or add a provider within 12 months — though the consideration set is dominated by the same incumbents they already run.
Methodology
VentureBeat fielded this survey as part of its ongoing Pulse Research series, this one focused on enterprise AI infrastructure, compute, and inference economics. Responses are filtered to organizations with more than 100 employees (n=170; the survey’s smallest size band, 1–100 employees, is excluded), drawn from a single July 2026 wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends; all figures are drawn from the July fielding only. Several questions were multiple-select, so those shares can sum to more than 100%.
By organization size this wave reaches further up-market than the mid-market skew this series usually carries: 251–1,000 employees (28%) and 1,001–5,000 (25%) lead, with 10,001+ (19%), 101–250 (15%), and 5,001–10,000 (12%) filling out the rest — meaning 57% of respondents sit above 1,000 employees. By role it spans managers (48%), individual contributors (27%), the C-suite (12%), and VPs and directors (9%); on purchasing authority it is buyer-credible, with 39% final decision-makers and another 43% recommenders or influencers for AI solutions. Technology/Software is the largest industry at 35%, followed by Manufacturing (14%), Financial Services (12%), and Healthcare/Life Sciences (9%).
At 170 respondents the sample is large enough to read directionally with reasonable confidence, but it should still be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It is best read as the view from organizations actively building and operating AI infrastructure rather than from the largest hyperscale operators.
Finding 1: Two-thirds are past the pilot
Three in 10 now run AI in production at scale
We asked where organizations sit in their AI deployment journey. This cohort has largely moved beyond experimentation.
Two-thirds of enterprises (66%) have AI workloads running in production, and 29% describe AI in production at scale. Only 30% remain in proofs of concept and just 4% have not started. This is a materially more operational sample than this series has typically drawn, consistent with its up-market composition — 57% of respondents sit above 1,000 employees.
That maturity is the frame for everything that follows. The infrastructure decisions in this report are being made largely by organizations with production workloads and real bills, not by teams still sizing a pilot. It explains the reordering of buying criteria in Finding 5, where performance and availability displace cost — the priorities of teams running live systems. It also raises the stakes on Findings 6 and 7: an enterprise that cannot measure utilization or cost during experimentation has a planning problem, while one that cannot measure them in production at scale has an operating one.
Finding 2: The stack is hyperscaler-and-API, three platforms deep
The specialized GPU clouds still barely register
We asked which providers and platforms enterprises currently use to run their AI, and which one they treat as primary. The answer remains the incumbents — several of them at once.
The current stack is hyperscaler-and-API, and it is plural: enterprises name three platforms on average. The general-purpose clouds and the major model APIs account for essentially all current deployment, with four platforms — OpenAI, Gemini, Azure, and Google Cloud — each presents in more than four of every 10 enterprises. Asked to pick one primary platform, Microsoft Azure leads at 26%, with Google Cloud second at 19%; the model providers together take 35% of primary status when OpenAI (14%), Gemini (14%), and Anthropic (8%) are combined.
The specialized “neocloud” GPU providers that dominate AI-infrastructure headlines remain marginal in practice. CoreWeave and Lambda each appear in 3.5% of stacks, Baseten in 3%, and Crusoe, Nebius, Fireworks, Together, and Anyscale each at or below 2%. Combined, they are named as the primary platform by 1% of enterprises. Meanwhile 13% run a custom open-source self-managed stack and 9% operate their own GPU clusters — both larger footprints than the entire specialized-cloud category. That contrast is what makes the evaluation intentions in Finding 3 worth reading closely.
A note on reading these shares: As described in the methodology section, this sample is self-selected and this question counted every provider a respondent uses — an average of 3.0 selections each — so the figures measure presence in the stack rather than spending or primary status. The separate primary-platform question is the better guide to where the center of gravity sits. A sample built this way will show a different provider mix than a spend-weighted census of the broader market; read these shares as a portrait of what this AI-active cohort runs today, and treat gaps against industry-wide market share estimates as a property of the sample rather than a contradiction of either.
Finding 3: The next dollar goes to infrastructure they don't yet run
AI-specialized clouds top the evaluations list and carry the strongest momentum
We asked where enterprises plan to evaluate AI infrastructure over the next 12 months, and whether they expect to do more or less with each category of infrastructure. Both answers point away from the stack they run today.
Here is the report’s sharpest tension, and it is the same one this series has now recorded across successive waves. The single most-cited planned evaluation area — AI-specialized clouds, at 44% — is the category that 3.5% of these enterprises actually use (Finding 2). Nearly four in 10 (39%) intend to evaluate non-Nvidia accelerators, a quarter next-generation Nvidia silicon, and even decentralized compute networks draw 18%.
The direction-of-travel question corroborates it rather than merely repeating it. Asked whether they expect to do more, less, or about the same with each approach, enterprises put specialized AI clouds at the highest net momentum (+36, with 42% doing more against 6% doing less), ahead of inference APIs (+34) and hyperscalers (+30). On-prem and co-located infrastructure is the laggard at +5, the only category where a substantial share — 22% — report pulling back. Every off-premises approach is net-expanding; the specialized clouds are expanding fastest from the smallest base.
Read against current usage, this is not incremental adjustment. It is the leading edge of a re-platforming that enterprises have been signaling for several waves and have not yet executed. The gap between a 44% evaluation rate and a 3.5% usage rate is the single widest intent-to-action spread in this dataset, and how it resolves — whether the neoclouds convert evaluation into deployment, or whether the hyperscalers absorb the demand with their own AI infrastructure — is the open question of the category.
Finding 4: Six in 10 plan to move, mostly among the incumbents
High churn intent, but the consideration set is the stack they already run
We asked whether and when enterprises plan to switch or add an infrastructure provider, and which providers they are considering.
For a category as foundational as compute, this is a substantial amount of intended movement: 62% of enterprises intend to switch or add a provider within 12 months, and 29% within the next quarter alone. Only 39% plan to stand still.
Where that interest points is the more useful signal. The providers drawing the most switching consideration are the ones enterprises already run — OpenAI and Google Cloud (29% each), Microsoft Azure (28%), Gemini (25%), Anthropic (16%), Oracle Cloud (14%), and AWS (13%). The specialized clouds that top the evaluation list in Finding 3 draw far less concrete switching consideration: CoreWeave 4%, Lambda 3.5%, and the remainder at or below 2%. A further 8% are evaluating with no shortlist yet.
The two findings are not in conflict; they operate on different clocks. The neocloud interest in Finding 3 is a 12-month evaluation thesis about where AI compute should eventually run. The switching in the next quarter is mostly incumbents trading share and enterprises consolidating spend among providers they already hold contracts with. Vendors reading the 44% evaluation figure as near-term pipeline should weigh it against a 4% consideration rate.
Finding 5: Performance overtakes cost, in buying and in measurement
Total cost of ownership falls below latency and GPU availability
We asked what matters most when enterprises select an AI infrastructure provider, and what they treat as the primary measure of success once it is running. Both answers have moved away from price.
Integration with the existing stack remains the top selection factor at 40%, which is consistent with a cohort running three platforms and unwilling to add a fourth that does not fit. What has changed is everything below it. Performance sits second at 35% and GPU access and availability third at 24%, both ahead of total cost of ownership at 22%. Fine-grained autoscaling draws 18% and cost per million tokens 16% — no longer the outlier it once was in this series, but still last.
Measurement follows the same logic. Uptime and reliability is the primary success metric for 51% of enterprises, well ahead of developer productivity and deployment speed (39%), cost per million tokens (31%), latency (27%), and throughput (25%). Taken together, the operational metrics dominate the economic one by a wide margin.
This is a coherent posture for the production cohort in Finding 1 — teams running live workloads care first about whether the system stays up and how fast they can ship on it. But it sits uneasily beside Finding 7. Total cost of ownership has been demoted to fourth as a buying criterion at exactly the moment when 53% of enterprises still cannot rigorously track what their compute costs. The uncomfortable reading is that cost has fallen down the list partly because it remains the hardest thing in the stack to see, and criteria that cannot be measured tend to lose to criteria that can.
Finding 6: The GPUs run warmer, but most still run cold
Roughly seven in 10 GPU operators report 50% utilization or less
We asked what share of their GPU capacity enterprises actually utilize. Figures here are reported on the 155 enterprises that operate their own GPUs; 15 consume exclusively via API and run none.
The compute already in place runs cold, though less so than this series has recorded before. Roughly seven in ten GPU-operating enterprises (69%) report utilization at or below half capacity, with the 26–50% band alone accounting for 46%. About a quarter (26%) run at 25% or below. Against that, 23% now clear the 50% mark — a meaningful efficient minority rather than a rounding error.
The remaining 12% who do not measure utilization at all are the more troubling number, because they are invisible in both directions: they cannot claim efficiency and cannot detect waste. And utilization does not improve with maturity in the way one might expect — among enterprises running AI in production at scale, 22% clear the 50% mark, statistically indistinguishable from the 24% among everyone else. Scale is not, by itself, producing better-utilized fleets.
Idle accelerators are expensive accelerators, and this remains the clearest single measure of the gap in this report: enterprises are planning to evaluate specialized clouds and next-generation silicon (Finding 3) while the capacity they already own sits substantially unused. The efficiency headroom in the current fleet is large, and for one in eight enterprises, entirely unmeasured.
Finding 7: Fewer than half can account for what they spend
Rigorous cost tracking reaches only 56%, even among at-scale operators
We asked whether enterprises can quantify the cost and return of their AI infrastructure spend, and how satisfied they are with what they run. Confidence in the ledger still lags the spending.
Measurement trails money. Fewer than half of enterprises (47%) rigorously track the cost and return of their AI compute; the majority track only partially (39%), cannot quantify it yet (15%), or have not prioritized it (6%). Maturity helps but does not solve it: among enterprises running AI in production at scale, rigorous tracking reaches 56%, against 43% for everyone else. Even in the most operationally advanced segment of this sample, more than four in ten cannot account precisely for what their AI compute costs or returns.
Satisfaction with current infrastructure is moderately positive and tellingly uneven. On a five-point scale, overall satisfaction averages 4.14 and ease of implementation 4.04, while value for money trails at 3.87 — the softness landing on the one dimension that requires measurement to assess. Enterprises are, in effect, expressing dissatisfaction with an economic relationship most of them cannot yet quantify.
Read with Finding 5, the picture is self-reinforcing rather than merely inconsistent. Cost has slipped to fourth among buying criteria while remaining the least visible property of the stack, and the least visible property is the one enterprises rate lowest. Better instrumentation would not necessarily change what enterprises buy — but it would let them know whether the trade they are making for performance and availability is a good one.
Finding 8: The memory frontier is still unclaimed
Dell and Nvidia lead a scattered field, and one in five has no view
We asked how enterprises would address the emerging constraint in large-scale inference — the shift from GPU compute to memory, specifically KV-cache capacity. The field remains early and fragmented.
The memory frontier is real but barely governed. Dell leads at 24% and Nvidia follows at 21%, with the remainder scattering across open-source tooling (12%), model-level efficiency techniques such as MLA and quantization (11%), and a long tail of storage vendors each in low single digits. No approach commands anything close to a majority, and the two leaders together account for less than half the field.
Most telling is that roughly one in five enterprises (19%) either do not recognize the constraint (7%) or have not begun to address it (12%). For a shift that will reshape inference cost and architecture, this is an early and unsettled market. It is also consistent with the measurement gap in Finding 7 — enterprises that cannot yet quantify what their current compute costs are in a poor position to anticipate which constraint will drive that cost next. The memory bottleneck is arriving while most of this cohort is still working to see the one in front of it.
The bottom line: Buying for speed, blind on cost
Organizations with more than 100 employees have moved AI infrastructure into production — two-thirds run live workloads, three in ten at scale — and their buying behavior has matured accordingly. They run three platforms on average, select on integration and performance, and measure success on uptime and developer velocity. For teams operating live systems, that is the right set of priorities.
What has not matured is the accounting. Total cost of ownership has fallen to fourth among selection criteria and cost per million tokens sits last, at the same moment that 53% of enterprises cannot rigorously track what their compute costs, 69% of GPU operators run at half capacity or less, and 12% do not measure utilization at all. Value for money is the lowest-rated attribute of the infrastructure they run — a judgment most of them are making without the instrumentation to support it. Cost has not become unimportant; it has become invisible, and the buying criteria have quietly reorganized around what can actually be seen.
Meanwhile the next round of spending points past the current stack. Specialized AI clouds are the top evaluation target at 44% and carry the strongest net momentum of any approach, against a 3.5% usage rate and a 4% near-term switching consideration — the widest intent-to-action spread in the data. Non-Nvidia accelerators draw 39%. And the constraint after this one, the shift from compute to memory in large-scale inference, is unrecognized or unaddressed by one enterprise in five.
At 170 respondents in a single July wave, reaching further up-market than this series typically does, this is a directional read — but the direction is consistent. Enterprises have become good operators of AI infrastructure and have not yet become good accountants of it. The open question for later waves is whether the instrumentation catches up before the re-platforming arrives, or whether enterprises buy the next layer of
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み