トークン請求が到来:AI の暴走するコストを管理するための業界の駆け引き
AI コストの急増により企業経営が脅かされる中、Linux Foundation がクラウドコスト管理の FinOps に倣った「Tokenomics Foundation」を設立し、業界全体でトークン使用量の可視化と統制への転換が始まった。
キーポイント
AI コストの予算崩壊と経営危機
Uber や Microsoft など大手企業が AI コーディングやツール利用で予算を短期間で超過し、CEO の「コストを度外視せよ」という指示が現実的なコスト管理の必要性に転換している。
トークン消費の爆発的増加要因
単価は低下しているにもかかわらず、Claude Opus 4.5 や GPT-5.1 などの高性能モデルによる自律型エージェント(AI Agents)の普及が、トークン消費量を劇的に増大させている。
業界標準「Tokenomics Foundation」の設立
Linux Foundation が新機関を立ち上げ、クラウド分野の FinOps のように AI トークン利用におけるコスト管理の基準と可視化ツールを提供し始める。
AIコストの爆発的増加と管理の遅れ
CEOの「コストを気にせず最良モデルを使う」指示や、エージェント機能の普及により消費量が急増し、一部の企業では数億ドル規模の使用料請求が問題となっている。
高コストと生産性の不確実性
トークン使用量が多いエンジニアは生産性が高い傾向にあるものの、バグや書き直しの増加も伴っており、支出に対する明確な収益効果の測定が困難である。
従来型ツールの限界と管理手法の見直し
トークンコストは従来のクラウドコストとは桁違いのデータ規模となるため、スプレッドシートや既存ツールでは追跡できず、会計システムや監視ツールの根本的な再構築が必要である。
AI コスト管理市場の拡大と多様なプレイヤー
Pay-i や Paid といった専門ベンダーに加え、Ramp、Datadog、New Relic などの既存大手も AI 支出管理やトークンレベルの可視化機能を追加し、FinOps Foundation の多くのベンダーがこの領域に参入している。
重要な引用
Our conversations are never about that now. Now the conversations are about, 'hey, we're spending so much. What visibility do you have? What auditability do you have?'
We started hearing existential crises, and the whole conversation shifted from tokenmaxxing and 'go fast' to 'we need guardrails, how do we control this?'
A Priceline employee told TechCrunch that a routine Cursor contract renewal came back 4-5x more expensive.
"It's like the crack-cocaine epidemic... They let you try it to get you hooked on it, and now you're kind of beholden to it."
"Tracking token costs is a trillions-of-rows-a-month data problem. You can't just stick that into whatever spreadsheet or even basic tool."
"The financial report for how much you spend on Anthropic, even if you call the Opus model, some of the spend will be on Sonnet or Haiku, because they are smart enough to do it."
影響分析・編集コメントを表示
影響分析
この記事は、AI 導入が初期の「速度と機能」から「持続可能性とコスト管理」へと転換点にあることを示す重要な指標である。Linux Foundation による標準化の動きは、今後数年間で AI ツールの調達・運用プロセスに根本的な変化をもたらし、企業における AI ROI の算定基準を再定義する可能性がある。
編集コメント
AI の普及が加速する中で、コスト管理の重要性が技術的な機能性を超えて経営課題として浮上した象徴的な出来事です。今後は「使い放題」から「最適化・制御」へのパラダイムシフトが業界全体で本格化するでしょう。
業界全体で、企業は AI の価格に頭を悩ませ始めました。Uber は 2026 年の AI コーディング予算を 4 月までに使い果たし、Microsoft も開発者向けの Claude Code ライセンスを数ヶ月後に撤回しました。Priceline の従業員は TechCrunch に対し、Cursor の契約更新が通常時の 4〜5 倍の費用になると報告しています。
トークンあたりの価格は低下しているにもかかわらず、AI 採用の推進と、より自律的なエージェント(autonomous agents)の台頭により、トークンの消費量はますます増加しています。2025 年初めに「食べ放題」サブスクリプションで大量に利用した企業たちは今、資金がどこへ消えたのかを把握し、支出を抑制し、予算の破綻から ROI(投資対効果)の一部を救済できるかどうかを検討しようと必死です。
一方、これに応える市場も形成されつつあります。スタートアップ、既存ベンダー、そして新たな標準化団体はすべて、企業が支出を追跡するためのツールと言語を提供するために競い合っています。
「6 ヶ月前なら、顧客との会話では『何ができるのか?十分なのか?』という話ばかりだった」と、今週ニューヨーク市で開催されたイベントで、OpenAI のエンタープライズ担当責任者である Alexander Embiricos 氏は TechCrunch に語った。「しかし現在、その種の会話はもはやない。現在の対話は、『ああ、こんなに支出している。可視化はできているのか?監査可能性はあるのか?トークン制御機能は何か?モデルの効率性はどうか』という内容だ」
この背景を踏まえ、Linux Foundation は今週、Tokenomics Foundation の設立計画を発表した。これはクラウド利用コストに対して FinOps が果たした役割と同様に、AI トークンに関するコスト管理の規律を確立することを目的とした新たな基準策定機関である。
「4 月と 5 月に、企業から『なんてことだ、2026 年通年のトークン予算の 3 倍も使っているのに、まだ 4 月なのに』という声が聞こえ始めた」と、Linux Foundation のプロジェクトである FinOps Foundation の執行ディレクター J.R. Storment 氏は TechCrunch に語った。「存続に関わる危機感の声が上がり、議論は tokenmaxxing や『スピード重視』から、『ガバナンスが必要だ。どう制御すべきか』へとシフトした」
テック業界全体に響き渡る叫び声は、CEO たちがチームに対して最良のモデルを使用し、コストなど構わず迅速に進めるよう熱心に要求した結果として生じました。11 月にリリースされた新しいモデル、例えば Anthropic の Claude Opus 4.5、OpenAI の GPT-5.1、Google の Gemini 3 Pro は、エージェント型ツール(agentic tools)に大きな改善をもたらしましたが、その結果消費量が急増しました。ある企業は、従業員の使用制限を設定しなかったために、なんと 5 億ドルの Claude 請求書を抱えることになったと報じられています。
「これはコカイン中毒の流行のようです」と Priceline の IT 財務シニアディレクターである Chris Reed は語り、同社が特定のグループにトークン制限を設け始めたことを指摘しました。「最初は試させて中毒者にさせるのです。今ではあなたはその依存から逃れられなくなっているようなものです。」
エンジニアリング運用プラットフォーム Faros AI の CEO である Vitaly Gordon は、最近ある CTO と話し合い、その CTO がこう話したと伝えました。「私のエンジニアの一人が先月トークンに 4 万ドルを費やしました。私は彼を止めるべきか、それとも他の全員に彼のようにさせるよう伝えるべきか、本当に判断がつきません。」
Faros AI が 4 月に発表した 20,000 人の開発者に対する 2 年間の調査では、出力は増加している一方でバグや書き直しも増えていることが分かりました。同様にエンジニアリング管理プラットフォームの Jellyfish も、最も多くのトークンを使用するエンジニアの方が、AI をあまり使用しないエンジニアよりも約 2 倍生産性が高い一方で、その成果を達成するために使用するトークンの数は 10 倍に達していることを発見しました。
Jellyfish の研究責任者である Nicholas Arcolano は、TechCrunch 宛てのメールで、AI への支出が爆発的に増加している主な要因はエージェント機能によるものであり、開発者あたりの消費量は 9 ヶ月間で約 18.6 倍に上昇していると語りました。全体として、これらの統計データは、支出額が示唆するほどには生産性に関する根拠が明確ではないことを意味しています。
「極端な支出が報われるかどうかは、最終的に出荷されたコードのビジネス価値(例えば収益)にかかっていますが、多くの企業はまだそれを測定できていません」と Arcolano は述べています。
その測定上の問題の一部は、今日 AI が利用されている規模の大きさによるものです。
「クラウドコストを追跡することは、月間数億行規模のデータの問題ですが、トークンコストを追跡するのは月間数兆行規模のデータの問題です。それを単にスプレッドシートや基本的なツールに押し込むことはできません。それを行うには、ツール、仕様、会計システムを根本的に見直す必要があります」と Storment は言います。
Priceline では、Reed 氏はすでに不一致を目撃しています。彼はベンダーが報告した利用状況と Priceline の内部データとの間に問題があることを指摘しました。
「私はキャリアの初期に通信費管理で働き始めましたが、通信からクラウド、そして AI に至るまで、同じような類似点が見えています」と彼は言います。「何か新しいものを導入するたびに、請求エラーや監査、最適化の機会が生じやすくなります」
この問題を取り巻く市場が形成され始めています。GenAI 投資の費用とパフォーマンスを追跡・測定・最適化する純粋なプレイヤー企業として、Pay-i が挙げられます。一方、Paid は開発者がコストを追跡し、利用状況を測定し、サブスクリプション料金ではなく実際の価値に基づいてユーザーに請求できる仕組みを提供しています。
また、Jellyfish、Waydev、Faros AI といった企業は、開発者ツールの ROI(投資対効果)を証明するための AI エージェント監視機能を提供しています。Storment 氏によると、FinOps ファウンデーション内の 180 のベンダーのほとんどがこの分野に傾倒しているとのことです。
既存の流通経路を持つ企業も、この新たな市場を活用するために新機能を追加しています。Ramp は最近 AI 支出管理 に参入しました。Datadog と New Relic は、クラウドコスト管理、トークンレベルの観測性(observability)、GPU 監視などのサービスを追加しました。来週の FinOps X コンファレンスでは、AWS がエンタープライズ向け AI 支出に焦点を当てた新たな財務管理機能を発表すると予想されています。
NEA のパートナーであるティファニー・ラックは、トークンの効率性と観測可能性は、「ハーン層またはアプリ層」に追加される可能性が高いと考えています。彼女は、企業向け AI エージェントを開発するスタートアップの Factory を例に挙げました。Factory は今週、各タスクに対して最適なモデルを自動的に選択するモデルルーターを ローンチ しました。
ゴードンは、フロンティア研究所やその他のモデルプロバイダーが、OpenRouter スタイルの最適化を採用し、クエリを最も安価なモデルへ誘導するようになるだろうと予測しています。これはすでに企業向けの Claude の請求書において現れ始めている傾向です。
「Anthropic への支出額に関する財務報告では、Opus モデルを利用しているとしても、その一部は Sonnet や Haiku に充当されることになります。なぜなら、それらはそれを賢く実行できるからです」とゴードンは述べています。「これはますます一般的な現象になっていくと思います」
しかし、これらのツールはいずれも、トークンがいくらで、何を生成し、ベンダー間で支出をどう比較すべきかという共通言語や共有定義なしに構築されています。まさにそこが、Tokenomics Foundation が有用性を証明しようとしている場所です。
財団は「トークノミクス」の標準的な定義と枠組み、AI トークンの利用と請求に関するオープンスタンダード・仕様・指標、そして知能あたりのコストやワットあたりのトークン数といった AI 経済のための新指標を構築しています。また、トークンファクトリの有効性と消費効率にわたる指標も定義する計画です。同グループは7月に正式発足を予定しており、来週の FinOps X コンファレンスでさらに多くのメンバーを発表しようとしています。
「トークン経済学は、これまでこの規模で管理してきたどのものよりも本質的に抽象的で不透明です」とセールフォースの首席可用性責任者であるニシャン・グプタ氏は声明で述べています。「これは、業界がクラウドのために築いてきた運用上の筋肉とは異なるものを必要とします。」
とはいえ、ゴールドマン・サックスは2030年までにグローバルなトークン利用量が24倍に増加すると予測しています。すでに予算超過となっている企業には今すぐの解決策が必要ですが、財団による最初の成果物はまだ数ヶ月先です。
「おそらく私たちは蒸気機関車を作りましたが、まだ組み立てラインを確立できていないのです」とゴードンは語りました。
アルコラーノ氏によると、賢明なアプローチは広範かつ中程度の採用です。
「最も高い投資対効果(ROI)を得るには、重度ユーザーをさらに引き上げるのではなく、広範な中間層を低利用から中利用へと移行させることです」と彼は述べました。
*ラッセル・ブランドムとティム・フェルンホルツが本報道に貢献しました。*
当記事内のリンクを通じてご購入いただいた場合、私たちは少額のコミッションを獲得する可能性があります。これは私たちの編集の独立性には影響しません。
AI の突進するコストを管理するための業界内での激しい駆け引きの内部事情:トークン請求書が期限を迎える(続き 8/8)
生成 AI モデルの利用が増加するにつれ、企業は「トークン」と呼ばれる計算リソースの使用量に対して課金される新たな現実と向き合わなければなりません。これは従来のソフトウェアライセンスモデルとは根本的に異なる構造であり、クラウドプロバイダーや大規模言語モデル(LLM: Large Language Model)の開発者、そしてそれらを利用するエンドユーザーの間で複雑なコスト管理の課題を生んでいます。
主要な AI プラットフォームを提供する企業は、トークンベースの課金体系を拡大し続けています。これは、入力されたテキストや生成された回答の長さ(トークン数)に応じて料金が変動する仕組みです。このモデルにより、ユーザーは必要な分だけ支払うという柔軟性を得る一方で、予測不能な使用パターンが突発的な高額請求につながるリスクも孕んでいます。
業界では、コストを抑制するための多様な戦略が模索されています。例えば、トークン効率の高い軽量モデルの採用や、キャッシュ(記憶)機能を活用して重複する計算を回避する技術、あるいはバッチ処理による最適化などが挙げられます。また、内部ツールとして独自の AI エージェントを開発し、外部 API への依存度を下げる動きも加速しています。
しかし、これらの対策は万能ではありません。特に大規模なデータ処理や複雑な推論が必要なケースでは、依然として膨大なトークン消費を避けられません。そのため、多くの企業がリアルタイムでの使用量監視システムを導入し、予算超過の警報を発信する仕組みを整備しつつあります。
専門家の間では、「トークン管理」がこれからの AI 運用における必須スキルになるとの見方が強まっています。財務部門と技術部門の連携を強化し、AI の投資対効果(ROI)を明確に評価できるフレームワークの構築が急務となっています。
今後の展望として、より効率的なアルゴリズムの開発や、ハードウェアレベルでの最適化が進めば、トークンあたりのコスト低下が期待されます。しかし、その間も企業は慎重なリソース配分と継続的なコスト分析を余儀なくされるでしょう。AI の爆発的成長の裏側で、静かに進行する「トークン請求書」への対応は、業界全体の持続可能性を左右する重要な課題となっています。
この状況を踏まえ、各社は自社の AI 戦略を見直し、短期的な利益追求だけでなく、長期的なコスト構造の安定化を図る必要性に迫られています。成功の鍵は、技術的な最適化と経営的な視点の融合にあると言えるでしょう。
原文を表示
Across the industry, companies are starting to balk at the price of AI. Uber blew through its entire 2026 AI coding budget by April. Microsoft revoked its developers’ Claude Code licenses months after enabling them. A Priceline employee told TechCrunch that a routine Cursor contract renewal came back 4-5x more expensive.
Even though per-token prices have fallen, the push for more AI adoption and increasingly autonomous agents have driven token consumption higher and higher. Companies that gorged themselves in early 2025 on all-you-can-eat subscriptions are now scrambling to understand where their money is going, pull back spending, and figure out whether they can salvage some ROI from the wreckage of their budgets.
Meanwhile, a market is forming to meet them there. Startups, established vendors, and a new standards body are all racing to give companies the tools and language to track what they spend.
“Six months ago, I would have a conversation with a customer and it would be all about ‘What can it do? Is it good enough?’” Alexander Embiricos, OpenAI’s head of enterprise, told TechCrunch at an event in New York City this week. “Our conversations are never about that now. Now the conversations are about, ‘hey, we’re spending so much. What visibility do you have? What auditability do you have? What token controls do you have? What is the efficiency of your models?’”
It’s against this backdrop that the Linux Foundation this week unveiled plans for the Tokenomics Foundation, a new standards body that aims to instill the same cost discipline around AI tokens that FinOps did for cloud spend.
“In April and May, I started hearing from companies: ‘Oh my god, we are 3x over our entire 2026 token budget and it’s only April,’” J.R. Storment, executive director of the FinOps Foundation, a project under the Linux Foundation, told TechCrunch. “We started hearing existential crises, and the whole conversation shifted from tokenmaxxing and ‘go fast’ to ‘we need guardrails, how do we control this?’”
The cries heard round the tech world followed fervent demands from CEOs pushing their teams to use the best models and move fast, costs be damned. New models released in November like Anthropic’s Claude Opus 4.5, OpenAI’s GPT-5.1, and Google’s Gemini 3 Pro brought significant improvements to agentic tools, which have multiplied consumption. It’s how one company reportedly found itself with a $500 million Claude bill after forgetting to set usage limits for employees.
“It’s like the crack-cocaine epidemic,” said Chris Reed, senior director of IT finance at Priceline, noting the company had begun placing token limits on certain groups. “They let you try it to get you hooked on it, and now you’re kind of beholden to it.”
Vitaly Gordon, CEO of engineering operations platform Faros AI, said he recently spoke to a CTO who told him: “One of my engineers spent $40,000 on tokens last month, and I genuinely don’t know whether I should stop him or should I go and tell everyone else to be like him.“
A two-year study of 20,000 developers that Faros released in April found that output was rising, but so were bugs and rewrites. Jellyfish, an engineering management platform, similarly found engineers who used the most tokens were about twice as productive as those who used AI less, but they spent 10x the number of tokens to get there.
Nicholas Arcolano, head of research at Jellyfish, told TechCrunch via email that expenditure on AI is exploding in large part due to agentic features, with per-developer consumption rising about 18.6x in nine months. All in all, these stats make the productivity case murkier than the spending suggests.
“Whether extreme spend pays off comes down to the ultimate business value of shipped code (e.g. revenue), which most companies still can’t measure,” Arcolano said.
At least some of that measurement issue is the sheer scale at which AI is being used today.
“Tracking cloud costs is a hundreds-of-millions-of-rows-a-month data problem,” Storment said. “Tracking token costs is a trillions-of-rows-a-month data problem. You can’t just stick that into whatever spreadsheet or even basic tool. You’ve got to fundamentally rethink your tooling, your specs and your accounting systems to do that.”
At Priceline, Reed is already seeing discrepancies. He noted issues between a vendor’s reported usage and Priceline’s internal data.
“I started my career in telecom expense management, and I’m seeing all the same parallels, from telecom to cloud to AI,” he said. “Anytime you introduce something new, it’s ripe for billing errors and audit and optimization opportunities.”
A market is beginning to form around this problem. There are the pure-play companies, like Pay-i, which tracks, measures, and optimizes the costs and performance of GenAI investments. Paid, meanwhile, lets developers track costs, measure usage, and bill users based on actual value rather than subscription fees.
Then there are companies like Jellyfish, Waydev, and Faros AI, which all provide AI agent monitoring to prove the ROI of developer tools. Storment says most of the 180 vendors within the FinOps Foundation are leaning toward this space.
Companies with existing distribution are also adding new features to capitalize on this new market. Ramp has recently moved into AI spend management; Datadog and New Relic have tacked on services like cloud cost management, token-level observability, and GPU monitoring. At the FinOps X conference next week, AWS is expected to introduce new financial management features geared toward enterprise AI spending.
Tiffany Luck, a partner at NEA, thinks token efficiency and observability will likely be added in at the “harness or app layer.” She pointed to Factory, a startup that makes AI agents for enterprises, which this week launched a model router that automatically picks the right model for every task.
Gordon expects frontier labs and other model providers to adopt OpenRouter-style optimization to drive queries to the cheapest models — a trend already showing up on enterprise Claude bills.
“The financial report for how much you spend on Anthropic, even if you call the Opus model, some of the spend will be on Sonnet or Haiku, because they are smart enough to do it,” Gordon said. “I think this will become more and more of a thing.”
But all these tools are being built without a common language or shared definitions for how much a token costs, what it produces, and how to compare spend across vendors. That’s where the Tokenomics Foundation hopes to prove useful.
The Foundation is building a canonical definition and framework for “tokenomics;” open standards, specifications and metrics for AI token usage and billing; as well as new metrics for AI economics, like cost-per-intelligence or tokens-per-watt. It also plans to define metrics across token factory effectiveness and consumption efficiency. The group is planning a formal launch in July, and is about to announce more members at the FinOps X conference next week.
“Token economics is fundamentally more abstract and opaque than anything we’ve managed at this scale before,” Nishant Gupta, chief availability officer at Salesforce, said in a statement. “It requires a different operational muscle than the one the industry built for cloud.”
That said, Goldman Sachs projects global token usage to multiply by 24 times by 2030. The companies already over budget need solutions now, and the foundation’s first deliverable is still months away.
“Maybe we created a steam engine, but we still haven’t figured out the assembly line,” said Gordon.
According to Arcolano, the smart move is broad, moderate adoption.
“The best ROI comes from moving the broad middle from low to moderate usage, not pushing heavy users higher,” he said.
*Russell Brandom and Tim Fernholz contributed to this reporting.*
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み