AI の請求額が制御不能に。Cloudflare が今すぐ解決します
Cloudflare は、AI 利用コストの管理不能化という課題に対し、AI Gateway に支出制限機能と ID ベースの予算・ルーティング機能を追加することで、企業における AI 投資対効果(ROI)の可視化と制御を可能にする新機能を発表した。
キーポイント
AI 支出管理の課題と背景
多くの企業が「後で請求書を確認する」という方針で AI を推進したが、その結果、誰がどのモデルを使用したか不明瞭なまま巨額の請求額が発生するという問題が生じている。
Cloudflare AI Gateway の新機能
AI Gateway に支出コントロール機能を追加し、既存の ID プロバイダー(Identity Provider)と連携した ID ベースの予算管理やモデル選択ルーティングをクローズドベータで提供開始する。
コスト最適化と可視性の重要性
すべてのタスクに最上位モデルを使用する非効率を防ぎ、適切なツールを選択させることで、各チームごとの支出内訳を明確にし ROI の計算を可能にする。
ドル単位の予算管理機能
AI Gateway がトークン数ではなくドル単位での累積使用料をリアルタイムで追跡し、モデル、プロバイダー、ユーザー、チームなど任意の次元に予算制限を設定できるようになりました。
柔軟なリセットとフォールバック戦略
予算は日次・週次・月次などの固定またはローリング期間で設定可能であり、上限到達時にはデフォルトでリクエストをブロックするか、Dynamic Routes を経由して代替モデルへ自動振り替えることが可能です。
ID 連携による詳細なコスト可視化
Cloudflare Access を介した認証情報(JWT)からユーザー ID を抽出しメタデータとして付与することで、組織全体での個人別・チーム別のトークン消費量とコスト帰属を一元管理できます。
IDP グループに基づく詳細な予算管理
Cloudflare Access と連携することで、従業員やサービスごとの認証情報を基に、個人やチーム単位でモデル利用の予算制限とポリシーを自動適用できます。
重要な引用
For fear of falling behind, many companies have pushed their employees to use AI as aggressively as possible. The edict was clear: 'Move fast, we'll figure out the bill later.'
You can't calculate ROI on your AI spend without visibility on what you're spending, and you can't protect that ROI without controls.
The problem is that most tasks don't need a frontier model. A code review summary doesn't need the same model as a complex architecture refactor.
These are true cost control measures in the form of budgets set in dollars, not tokens, that track cumulative spend across all requests.
We solved this by enabling AI Gateway to add identity to every request... This makes per-user token consumption, team-level usage breakdowns, and cost attribution across the organization all visible in one place.
When a user hits their limit, requests can be downgraded to a cheaper model or blocked.
影響分析・編集コメントを表示
影響分析
本発表は、AI 導入が急拡大する中で浮上した「コスト管理のブラックボックス化」という普遍的な課題に対し、インフラ層(Gateway)から解決策を提供する画期的な動きです。企業にとっては、AI の利用を止めずに、かつ財務リスクを管理しながら ROI を最大化するための具体的な手段を得たことになります。
編集コメント
AI の利用拡大に伴うコスト管理は、今や CIO や CFO が直面する最優先課題の一つです。Cloudflare がインフラレベルで ID ベースの制御を提供することは、現場の開発者が自由に AI を使いながら財務規律を維持するための重要なステップと言えます。
今や、AI の支出について心配していない CIO は地球上に一人もいません。CFO たちも次第に不安を強めています。
取り残されることを恐れて、多くの企業が従業員に対して AI を可能な限り積極的に利用するよう迫っています。その指示は明確でした。「速く動け、請求書は後で考える」と。そして大部分において、この方針は機能しました:AI は積極的に関わったチームにとって真に変革的なものとなりました。
しかし、コストは現実のものです:トークン支出における巨額の請求書や痛みを伴う超過利用に関する数え切れないほどの恐怖譚を耳にしてきました。
本日、Cloudflare AI Gateway における支出管理機能を発表するとともに、Cloudflare Access と既存のアイデンティティプロバイダーを活用した、アイデンティティ駆動型の予算設定とルーティングのためのクローズドベータ版も開始します。
数百社に AI ストラテジーについて話を聞いてきた中で、私たちは共通するストーリーを目にしてきました:ある企業がすべてのエンジニアに対して共有 API キーを通じて最先端モデルへのアクセスを許可します。利用が急増し、月末になると財務部門が請求書を取り出し、誰も資金の行方説明できません。機械学習チームが新しいパイプラインをトレーニングしたのか?それともインターンが Claude Opus をメールのトリアージに使用していたのか?あるいは週末に 5000 万トークンを消費して暴走した継続的インテグレーションジョブだったのか?API キーには誰が利用したかが表示されないため、誰も知りません。
ガイドラインがない場合、スタッフは一般的に利用可能な最も大きなモデルを求めます。なぜそうしないでしょうか?予算もなく、可視性もなく、ルーティングロジックもないのであれば、すべてのタスクに対して最も強力なモデルを使用するのが合理的な選択です。問題は、ほとんどのタスクが最先端モデルを必要としないことです。コードレビューの要約には、複雑なアーキテクチャのリファクタリングと同じモデルは不要です。ログパーサー(log parser)にも、顧客向けコンテンツジェネレーターと同じモデルは不要です。最も強力かつ高価なものをデフォルトとするのではなく、タスクに適切なツールを選択できるようにすべきです。また、支出がどこに向かっているかを簡単に確認できる仕組みも必要です。
何に支出しているかの可視性がなければ AI への投資対効果(ROI)を計算できず、コントロールがなければその ROI を保護することもできません。ビジネスの他のすべての項目には予算とチームごとの配分があり、AI への支出も例外であってはなりません。
What AI Gateway is
AI Gateway は、アプリケーションと AI プロバイダーの間に位置します。OpenAI、Anthropic、Google、または他のプロバイダーに直接呼び出すのではなく、リクエストはまず AI Gateway を経由してルーティングされます。
これにより、すぐにいくつかの有用なツールが利用可能になります:
異なるプロバイダーやモデル間で簡単に切り替え可能な統一された請求書(billing)
すべてのプロバイダーにわたるログ記録 — すべてのリクエスト、トークン数、コストを一つの場所で管理
レスポンスキャッシュ(caching)
レート制限(rate limiting)
コンテンツガードレールと、個人識別情報(PII)や機密情報がモデルに到達する前にブロックする機能
しかし、AI Gateway には、誰がいくら使っているかを確認したり、AI 利用費に制限を設けたりするための簡単な方法がありませんでした。
アカウント全体の集計使用状況は確認できました。しかし、エンジニアリング部門のジェーンさんが今月 Claude に 2,000 ドルも使い果たした一方で、データサイエンスチーム全体ではわずか 400 ドルしか使っていなかったといった詳細までは見ることができませんでした。「エンジニアリング部門には Frontier モデルで月額 5,000 ドル、インターン生には Kimi K2.6 で月額 200 ドル」といった予算設定もできませんでした。
それが今日から変わります。
支出制限:AI 利用のための予算
AI Gateway は、支出制限をコア機能としてサポートするようになりました。これはトークンではなくドル単位で設定される予算という形をとった、真の費用管理機能です。すべてのリクエストにわたる累積支出を追跡し、従来のレート制限とは独立して動作します。
制限は、モデル、プロバイダー、またはユーザー、チーム、アプリケーションといった管理者が定義したカスタム属性など、任意の組み合わせの次元に対して適用できます。期間設定は、月初、月曜日の午前中、あるいは深夜にリセットされる固定ウィンドウと、継続的に累積するローリングウィンドウの両方に対応し、日次、週次、月次のいずれにも設定可能です。
image
AI Gateway は、モデルの価格に基づいてリクエストごとのコストを計算し、リアルタイムで設定された制限に対する累積支出を追跡します。分析ダッシュボード上でモデルごとの支出を簡単に追跡でき、モデル、プロバイダー、または任意のカスタム属性でフィルタリングすることも可能です。
予算制限に達した際、どのような対応をするか選択肢があります。AI Gateway はデフォルトでそれ以上のリクエストをブロックします。あるいは、Dynamic Routes を通じてルールを設定し、支出制限に達した後、リクエストをフォールバックモデルへルーティングすることも可能です。これにより、厳格な支出上限がエンジニアのワークフローを妨げることを防ぎます。また、制限に達した際にアラートを送信する機能も追加する作業を進めています。
支出制限は本日、すべてのプランにおける AI Gateway ユーザー向けにオープンベータとして利用可能です。ダッシュボードのゲートウェイ設定または API を経由して設定できます。
私たち自身でもこの仕組みを利用しています
Cloudflare 内ではすでにトークンコストを追跡しています。Cloudflare の全従業員が毎日 AI ツールを使用しており、AI Gateway を通じて月間数百万件のリクエストと数十億のトークンを処理しています。この規模で直面するすべての企業が抱える同じ問いに直面しました:誰が何をどの程度使用しているのか、そしてそれをどのように予算化すべきかです。
私たちは、AI Gateway によって各リクエストにアイデンティティ(ID)を追加できるようにすることで、この課題を解決しました。従業員が Cloudflare Access を介して認証を行うと、JSON Web Token (JWT) からそのアイデンティティを抽出し、AI Gateway のリクエストのメタデータとして付加します。これにより、ユーザーごとのトークン消費量、チームレベルでの使用状況の内訳、組織全体のコスト帰属が、一つの場所で可視化されます。
アイデンティティ駆動型の予算とポリシー(クローズドベータ)
支出制限に加えて、本日、アイデンティティ駆動型の予算とポリシーをクローズドベータとして発表します。
AI Gateway の支出制限機能を使えば、モデルやプロバイダー、カスタム属性ごとに予算を設定できます。ただし、アプリケーション側がそのメタデータを渡す必要があり、AI Gateway は受け取った内容をそのまま信頼します。検証された自動的な帰属管理を実現するには、アイデンティティが必要です。
Cloudflare Access と組み合わせることで、AI Gateway は各リクエストを発行した主体を特定できます。単なるアカウント名だけでなく、どの従業員か、どのアイデンティティプロバイダー(IdP)グループに所属するか、どのサービスかなどまで把握可能です。
実際の運用では以下のようになります。


ユーザーごとの予算設定も可能です。例えば、一般社員には月額 500 ドル、シニアエンジニアには月額 2,000 ドルといった具合です。ユーザーが予算上限に達した場合は、リクエストをより安価なモデルへ降格させるか、ブロックすることもできます。
チームごとのモデルポリシーも設定可能です。例えば、ML チームには Claude Opus と GPT-4o を利用させ、ブランドデザインチームには生成画像・動画モデルへのアクセス権限を与えます。インターンは Workers AI 上のオープンソースモデルを利用します。これらのポリシーは、すでに管理している既存の IdP グループ(アイデンティティプロバイダーグループ)に直接マッピングされます。
CI/CD パイプラインや自律型エージェントにおいて、Access サービストークンを使用すれば、各エージェントに名前付きのアイデンティティを付与できます。今週、コードレビューボットが 500 万個のトークンを消費した一方で、ドキュメント生成ツールは 50 万個しか使用していないことが確認できるでしょう。もしあるエージェントが制御不能な状態になりつつあるなら、他のエージェントに影響を与えることなく予算ポリシーを適用できます。
すべての AI Gateway ログエントリには、認証されたアイデンティティ(メールアドレス、IdP グループ、サービストークン名)が含まれます。これらを分析プラットフォームへエクスポートすれば、カスタム機能を構築することなく、ユーザー別・チーム別のコスト内訳を取得できます。
内部では、AI Gateway エンドポイントに対して Cloudflare Access アプリケーションを作成し、IdP グループに基づいてポリシーを構成します。開発者やエージェントがリクエストを実行すると、通常の CLI デバイスコードフロー(device-code flow)を通じて OAuth により認証が行われます。AI Gateway がトークンを検証し、アイデンティティを抽出します。独自の Worker を記述したり、JWT を自分で解析したり、信頼ベースのメタデータヘッダーに依存する必要はありません。
私たちは最近、社内 AI エンジニアリングスタックの構築方法について記事を書きました。今日公開するのはまさにその仕組みです。あなたもこれを利用でき、自ら構築する必要はありません。
クローズドベータへのアクセスをご希望の場合は、こちらからサインアップしてください。
次なるステップ:コスト管理からコスト最適化へ
予算を設定することは必要不可欠です。しかし、一度予算を設けたら、それをいかに最大限に活用するかという課題が残ります。
現実には、すべてのリクエストが最先端モデルを必要とするわけではありません。要約タスクは、品質に大きな低下を伴わずに、より小さく安価なモデルで実行できますが、大規模なコードのリファクタリングには最先端の技術が必要になることもあります。しかし、制御手段がない場合、人々はほぼ例外なく最も高度なモデルを選択してしまいます。
そのための解決策は次期リリースで提供されます:AI Gateway でインテリジェントかつタスクベースのルーティングを構築中です。各リクエストに対して分析を行い、最低コストで最高の結果をもたらすモデルへ自動的にルーティングします。現在も活発に開発中ですので、開発者向けドキュメントや変更履歴をご確認ください。
Get started
AI Gateway の利用は無料です。すべてのユーザーが今すぐ使用可能な支出制限機能があります。
まだ行っていない場合は、ゲートウェイを作成し、アプリケーションをそのゲートウェイへ指向させてください。その後、ダッシュボードまたは API を通じて支出制限を設定します。まずは監視モードで高い制限値を設定し、適用を開始する前に現在の利用パターンを理解することから始めましょう。
ユーザーごとの追跡やチームベースのポリシーが必要な場合は、ID ドライブ型予算クローズドベータ版にご登録ください。Access 統合の設定を支援いたします。
現在、AI コストをどのように管理されているか、ぜひお聞かせください。Cloudflare Community で議論に参加するか、より広範な AI セキュリティ戦略についてご相談いただくためにご連絡ください。
原文を表示
There isn't a CIO on the planet not worried about AI spend right now. CFOs are increasingly nervous, too.
For fear of falling behind, many companies have pushed their employees to use AI as aggressively as possible. The edict was clear: "Move fast, we'll figure out the bill later." And for the most part, it worked: AI has been genuinely transformational for the teams that leaned in.
But the costs are real: we’ve heard countless horror stories of huge bills and painful overages on token spend.
Today, we're announcing spend controls in Cloudflare AI Gateway, and a closed beta for identity-driven budgets and routing using Cloudflare Access and your existing identity provider.
As we’ve spoken with hundreds of companies about their AI strategy, we’ve seen a common story: The company gives every engineer access to frontier models through a shared API key. Usage takes off. At the end of the month, finance pulls the invoice and nobody can explain where the money went. Was it the machine learning team training a new pipeline? Was it an intern running Claude Opus on email triage? Was it a runaway continuous integration job that burned through 50 million tokens in a weekend? Nobody knows, because the API key doesn't tell you who used it.
Without guidelines, staff will generally reach for the biggest model available. And why wouldn't they? If there's no budget, no visibility, and no routing logic, the rational move is to use the most powerful model for everything. The problem is that most tasks don't need a frontier model. A code review summary doesn't need the same model as a complex architecture refactor. A log parser doesn't need the same model as a customer-facing content generator. It should be easy to select the right tool for the job, rather than defaulting to the most powerful and expensive one. And it should be simple to see where the spend is going.
You can't calculate ROI on your AI spend without visibility on what you're spending, and you can't protect that ROI without controls. Every other line item in a business has a budget and per-team attribution and AI spend should be no different.
What AI Gateway is
AI Gateway sits between your applications and AI providers. Instead of calling OpenAI, Anthropic, Google, or any other provider directly, your requests route through AI Gateway first.
This immediately gives you several useful tools:
Unified billing to easily switch between different providers and models
Logging across all providers — every request, token count, and cost in one place
Response caching
Rate limiting
Content guardrails and the ability to block Personally Identifiable Information (PII) and secrets before they reach the model
However, AI Gateway didn’t have an easy way to answer who is spending what or how you might set limits on AI spend.
You could see aggregate usage across your account. But you couldn't see that Jane from engineering burned through \$2,000 on Claude this month while the entire data science team only used \$400. You couldn't set a budget that said "engineering gets \$5,000/month on frontier models, interns get \$200/month on Kimi K2.6."
That changes today.
Spend limits: budgets for AI usage
AI Gateway now supports spend limits as a core feature. These are true cost control measures in the form of budgets set in dollars, not tokens, that track cumulative spend across all requests, operating independently of traditional rate limiting.
You can scope limits to any combination of dimensions: model, provider, or admin-defined custom attributes like user, team, or application. Windows can be fixed (resets on the first of the month, Monday, or midnight) or rolling, and set to daily, weekly, or monthly.
image
AI Gateway calculates cost per request based on the model's pricing, and tracks cumulative spend against your limit in real time. You can easily track your model spend on our analytics dashboard and filter by model, provider, or any custom attribute.
You have options for what happens when the budget limit is reached. AI Gateway will block further requests by default. Or you can set up rules through Dynamic Routes to route requests to a fallback model after you’ve hit a spend limit, so that a hard spending cap won’t kill your engineers’ workflow. We’re working to add the capability for you to also send alerts when a limit is reached.
Spend limits are available in open beta today for all AI Gateway users across all plans. Configure them in your gateway settings in the dashboard or via the API.
We use this ourselves
We're tracking token costs inside Cloudflare already. Every Cloudflare employee uses AI tools daily, routing millions of requests and billions of tokens per month through AI Gateway. We faced the same question every company faces at this scale: who's using what, and how do we budget for it?
We solved this by enabling AI Gateway to add identity to every request. When an employee authenticates via Cloudflare Access, we extract their identity from the JSON Web Token (JWT) and attach it as metadata on the AI Gateway request. This makes per-user token consumption, team-level usage breakdowns, and cost attribution across the organization all visible in one place.
Identity-driven budgets and policies (closed beta)
In addition to spend limits, today we’re also announcing identity-driven budgets and policies as a closed beta.
Spend limits in AI Gateway let you set budgets by model, provider, or custom attributes. But your application has to pass that metadata, and AI Gateway trusts whatever it receives. For verified, automatic attribution, you need identity.
When combined with Cloudflare Access, AI Gateway can see who is making each request — not just which account, but which employee, which identity provider (IdP) group, which service, etc.
Here's what that looks like in practice.
image
image
You can set per-user budgets, say \$500/month for individual contributors and \$2,000 for senior engineers. When a user hits their limit, requests can be downgraded to a cheaper model or blocked.
You can set per-team model policies. For instance, your ML team gets Claude Opus and GPT-4o. The brand design team can access generative image and video models. Interns use open-source models on Workers AI. These policies map directly to your existing IdP groups, the same identity provider groups you already manage.
For CI/CD pipelines and autonomous agents, Access service tokens allow you to give each agent a named identity. You can see that your code review bot used 5 million tokens this week while your documentation generator used 500,000. If one agent is running out of control, apply a budget policy without affecting any others.
Every AI Gateway log entry will include the authenticated identity: email, IdP group, service token name. Export these to your analytics platform, and you've got a cost-by-user-by-team breakdown without building anything custom.
Under the hood, you create a Cloudflare Access application for your AI Gateway endpoint and configure policies based on your IdP groups. When a developer or agent makes a request, they authenticate via OAuth, using the typical CLI device-code flow. AI Gateway validates the token and extracts the identity. You don't need to write a custom Worker, parse JWTs yourself, or rely on honor-system metadata headers.
We recently wrote about how we built our internal AI engineering stack. This is what we are making available today — so you can use it, too, and you don't have to build it yourself.
If you would like access to the closed beta, sign up here.
What's next: from cost control to cost optimization
Setting a budget is necessary. But once you’ve got a budget, how do you make the most of it?
The reality is that not every request needs a frontier model: a summarization task can run on a smaller, cheaper model without meaningful quality loss, while a large-scale code refactor might require the bleeding edge. But without controls, people will almost always opt for the most advanced model.
A solution for that is coming next: We're building intelligent, task-based routing in AI Gateway. For each request, we can analyze and automatically route it to the model that will give you the best result at the lowest cost. This is in active development, so follow our developer docs and changelog.
Get started
It’s free to get started with AI Gateway. Spend limits are available now for all users.
If you haven't already, create a gateway and point your applications at it. From there, set up spend limits in the dashboard or via API. Start with a high limit in monitoring mode to understand your current usage patterns before you start enforcing.
If you want per-user attribution and team-based policies, sign up for the identity-driven budgets closed beta, and we'll get you set up with the Access integration.
We want to hear how you're managing AI costs today. Join the conversation on Cloudflare Community or reach out to discuss your broader AI security strategy.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み