Databricks、大規模AIコーディングツールのコスト管理手法を公開
本文の状態
日本語全文を表示中
詳細モードで約18分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Databricks AI Engineering
Databricksは、大規模展開におけるAIコーディングツールの利用が出力速度を劇的に向上させる一方、コストの指数関数的増加という課題に直面し、収益を上回るリスクがあるとして、その解決策となるコスト管理手法を報告した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月8日 02:45
AI深層分析
キーポイント
AI コーディングツールのコスト爆発リスク
大規模な AI ツール導入において、出力向上と引き換えにコストが指数関数的に増加し、収益を脅かすパラドックスが生じている。
効率フロンティアの概念定義
単なる低価格モデルではなく、特定の知能レベルに対して最適な価格帯を提供する「効率フロンティア」が重要であると指摘している。
実証済みの管理アプローチとツール公開
Databricks は Stripe や Uber などの事例に基づき、メタハネスや AI ゲートウェイといったインフラをオープンソース化し、アクセスの容易さとコスト抑制を両立させる手法を提示した。
効率フロンティアの定義と重要性
効率フロンティアは、与えられた知能レベルに対して最良の価格点を持つモデルの集合によって定義される。このフロンティアは知能のフロンティアよりも急速に進展しており、週単位でより高い知能対価格比を持つ新モデルがリリースされている。
ベンチマーク評価とモデル移行の実践
実際のコーディングタスクにおけるパフォーマンスを反映するため、多くの企業が社内開発ミックスに即した自動評価システムを構築している。評価結果によっては新モデルが品質向上をもたらさない場合やコスト増となるため、導入を見送る判断も下される。
重要な引用
nearly every company deploying AI tools at scale has hit the same wall: exponentially growing costs
The efficiency frontier is defined by the set of models that have the best price point for a given level of intelligence
achieving a 'dual mandate': (a) providing broad access to AI tooling, with minimal friction, and (b) keeping aggregate costs inside of a roughly fixed envelope per user
"The efficiency frontier is defined by the set of models that have the best price point for a given level of intelligence."
編集コメントを表示
編集コメント
本稿は、AI ツールの導入拡大がもたらす経済的リスクに対する現実的な対策を提示しており、技術選定におけるコスト意識の重要性を再認識させる内容である。特に、他社との協働で得られた知見をオープンソース化して共有する姿勢は、業界全体の成熟度を高める上で意義深い動きと言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI コーディングツールは莫大な価値をもたらします。Databricks において、エージェントによるコーディングは追跡しているすべての速度指標を測定可能な形で改善し、一部のチームでは出力が桁違いに向上しました。しかし、大規模に AI ツールを導入する企業はほぼ例外なく、同じ壁に直面しています。それは*指数関数的に増大するコスト*です。この曲線は持続不可能であり、放置すれば最終的に収益を凌駕してしまいます。
支出の爆発的な増加により、企業は奇妙なジレンマに陥っています。一方では AI 変革を最大限に進め、強力なツールを従業員の手元に届けたいと願う一方で、他方では AI がもたらす効率化そのものを損ない、場合によっては逆転させる恐れのある集計コストの現実と向き合わなければなりません。
幸いにも、大規模導入の先駆者たちはこの課題を解決する一連のアプローチに収束し、「二重の使命」を達成しています。それは (a) 摩擦を最小限に抑えつつ AI ツールへの広範なアクセスを提供すること、そして (b) ユーザーあたりの集計コストをおよそ一定の範囲内に抑えることです。
本稿では、Databricks の経験や Stripe、Coinbase、Uber、Ramp といったデジタルネイティブ企業との対話に基づき、実証済みのコスト管理手法を解説します。以下の表は現在の手法と関連する節約効果を要約したもので、数値は開発チームへの非公式な調査に基づく方向性の目安です:
これらの手法の中には、多くの企業がすでに利用しているソフトウェアで簡単に実装できるものもあります。一方で、エンドユーザーのクライアントを改変したり、モデル間でトラフィックをシフトさせたりする手法には、新たなインフラストラクチャが必要です。Databricks では、主要なインフラコンポーネントをオープンソース化するか、無料で提供しています。具体的には、エンドユーザー向けのメタハルネス(Omnigent)と AI Gateway(Unity AI Gateway)です。
なお、本記事では、他の企業との対話で言及されたソフトウェアについても触れています。
コーディングモデルにおける「効率フロンティア」
コーディング関連の支出を、リリースされるより効率的なモデルへ移行する上で最も大きなコスト削減効果をもたらすのがこのポイントです。この点については議論の余地があります。「より安価なモデル」という単純な説明は、実はモデルのコストと品質の間に存在する微妙な関係を隠蔽しているからです。
一般的に「フロンティアモデル」という言葉は、「最高知能を持つモデル」を指し、フロンティアラボも主にその知能の向上に注力しています。現在では、これらのモデルが数学やサイバーセキュリティにおける新たな問題も解決できるようになっています。
しかし、AI を大規模に展開する際により重要になるのは別の種類のフロンティアです。それが「効率性のフロンティア」です。この効率性のフロンティアは、「特定の知能レベルに対して最もコストパフォーマンスが良いモデルの集合体」として定義されます。日常のコーディング作業の多くが数学的証明や新たなセキュリティ洞察を必要としないため、実務において重要なのは、一般的なソフトウェアエンジニアリングの品質基準を満たすモデルのコストです。この「効率性のフロンティア」は知能のフロンティアよりもはるかに速く進化しており、以前より単位価格あたりの性能が優れた新しいモデルがほぼ毎週のように登場しています。
コスト削減のレバー #1: オープンソースおよび低コストモデルへの移行
| ステップ1。 モデルの効率フロンティアを測定する | ステップ2。 ユーザーを最適なモデルへ移行する |
|---|---|
| Databricks 推奨モデル(2026年8月6日時点)GLM 5.2Opus 4.8GPT 5.6-Sol*注:データおよび知識処理タスクには、Databricks は *Genie One* を使用します。上記はコアソフトウェア開発向けです。 |
より新しく、効率的なモデルを迅速に採用することは、コスト削減において最も大きな成果をもたらす手法です。しかし、その恩恵を得るためには、自社の既存モデルに対してどの新しいモデルが実際に優位性を示すのかを把握する必要があります。
これは容易ではありません。なぜなら、公開されているベンチマークは、コーディングタスクにおける実世界でのパフォーマンスを示すのにあまり役立たないからです。そのため、多くの企業が、自社の開発環境に即したより代表的な評価を行うために、独自の自動評価システムを構築しています。
Databricks は最近、そのようなベンチマークの例を公開しました(例)。これによると、GLM モデルは非常に競争力のある価格パフォーマンス比を示していました。この評価結果を受け、私たちは社内の開発者向けに GLM の展開を開始しました。
しかし、新しいモデルが常に効率のフロンティアを押し上げるわけではありません。むしろ、評価によってネガティブな結果が出ることも珍しくありません。Stripe は Opus 4.7 が Opus 4.6 と比べて品質に有意な改善をもたらさず、コストが増加したという調査結果を発表し、社内で Opus 4.7 の利用を見送りました。Databricks も同様に、Opus 5.0 を Opus 4.8 と比較した際に、コストの悪化を確認しています。
ハーネスとモデルの柔軟性を活用する
コスト削減における最大の成果は、新しいモデルへ切り替えることから生まれます。そのため、モデルの柔軟性に対応できるエンドユーザー向けツールの採用が、コスト抑制の重要な要素となっています。
特定のモデルに関連して最も一般的に使用されるツールは「ハーネス(harness)」と呼ばれています。独自開発の最先端モデルでは、特定のハーネスと連携するように設計されるケースが増えており、あるモデルには特定のハーネスの方が「より良く機能する」という状況が生まれています。企業がモデルへの依存を避けたい場合、主に以下の 2 つのアプローチがあります。
ユーザーにハーネスの切り替えを依頼する
一つ目のアプローチは、開発者に Claude Code、Codex、Cursor など複数のハーネスを提供し、コスト効率の高いモデルへ移行したい際に、開発者自身がハーネスを切り替えてもらう方法です。これにより、ユーザーは可能な限り好みの環境で作業できますが、この手法の欠点は、個々の開発者にとっての切り替えコストが高くなることです。
もし切り替えコストが高すぎると、ハーネス自体が事実上のモデルファミリーへのロックイン(囲い込み)となり、より競争力のあるモデルへ支出を移す能力が制限されてしまいます。
メタハarnessを活用する。 新たな、そして急速に普及しているアプローチとして、開発者に共通のユーザー体験を提供しつつ、リクエストを基盤となる harness(プロプライエタリおよびオープンソース双方)へ配信する「メタハarness」を使用する方法があります。この手法により、モデルや harness の独立性を維持しながら、開発者の切り替えコストも削減できます。Databricks では、Omnigent を活用する開発者にとってこれがデフォルトモードとなっています。また、一部の企業では、自社の開発ツールチェーンと連携するカスタム内部メタハarness を構築しています。
コスト削減のレバー #2: 動的なリクエストおよびタスクルーティング
ユーザーに自らタスクに適したモデルを選ばせるのではなく、自動的なモデルおよびツールの選択が、エージェント型コーディングワークフローからさらに効率を絞り出す可能性があるという研究が増えています。ルーティングのアプローチは概ね以下の3 つのカテゴリーに分けられます:
要求レベルのルーティング
クライアント(コーディングハッチなど)と基盤モデルの間に、状態を保持するプロキシが配置されます。このプロキシは、各推論リクエストに対して回答可能な最も低コストのモデルへリクエストをルーティングしようとします。
エージェントユースケースにおけるルーティングでは、サーバーサイドキャッシングの影響も考慮する必要があります。大規模なコンテキスト処理において、キャッシュミス(コールドヒット)が発生すると非常に高額なコストがかかるためです。
ルーティングに関する新製品が次々と登場し、初期段階ながら有望な結果を示しています。具体例としては、Cursor の Router、OpenRouter の AutoRouter、Ramps の Router 機能、そして Databricks が提供する Unity AI Gateway 内の Smart Routing 機能などが挙げられます。
タスクレベルのルーティング(メタハッチ)
クライアント側のプロセスが、タスクの複雑さに応じてユーザーの要求を異なるハッチに振り分けます。例えば、「コンポーネント X を Y に名前変更する」といった単純なタスクもあれば、「レイテンシを削減するための設計検討事項を探る」ようなオープンエンドで複雑なタスクもあります。
ディスパッチャ(メタハッチとも呼ばれます)は、各タスクにどのレベルの基盤モデルが必要かを判断し、そのタスク全体をモデルに一任します。このパターンをサポートするメタハッチの例として、Omnigent が挙げられます。
エスカレーションと委譲のパターン:単一のハネスで、高コストかつ高度な知能を持つモデルと低コストのワーカーモデルをペアリングします。いくつかのアプローチでは、Claude のアドバイザーツールのように、安価なモデルが主導権を握り、「タスクにさらに多くの計算資源が必要だと判断した際にエスカレーション」を行います。逆のパターンも存在します。Cognition の Devin Fusion では、高コストのモデルがメインループとなり、必要な作業のみを低コストのモデルへ選択的にアウトソースします。Databricks 内部の結果によると、AI ゲートウェイのスマートルーティング機能は、最も高価なモデルとほぼ同等の品質を維持しつつ、タスクあたりの平均コストを常に 30% 以上削減できることが示されています。他の企業との対話でも、同様の結果が報告されています。
コスト管理のレバー #3:開発者に可視性、トリップワイヤー、予算枠を提供する
「ユーザーに月額予算を与えて完了」というシンプルな結論で記事全体を始めるべきだったのではないかと思うと意外かもしれません。しかし、特定の支出額に達した瞬間に利用が完全に停止されるような「厳格な予算制限」は、話を聞いたどの企業でも最終手段としてしか使われていません。
AI 関連の支出管理においてトークン単位の厳格な予算枠が必ずしも効果的ではないのには、主に 2 つの理由があります。第一に、開発者が予算上限に達した際に AI ツールへのアクセスを完全に遮断すれば、生産性が著しく低下してしまいます。企業も従業員も、そのような結果は望んでいません。第二に、「支出が多い」ユーザーの中には、AI を活用することで劇的な効率化を実現し、莫大な成果を生み出している人々が含まれています。こうしたユーザーを萎縮させることは、本末転倒です。
そのため、多くの企業が採用しているのは、厳格な利用制限ではなく、エンドユーザーへの可視性を高め、支出が増えるにつれて段階的に摩擦(確認プロセスなど)を導入する、よりニュアンスに富んだ漸進的なアプローチです。
「可視化」:取材に応じた企業はすべて、ユーザーに対して利用状況の即時フィードバックを提供する仕組みを持っていました。多くの企業がさらに、より安価なモデルを活用してコストを削減するための具体的なアドバイスや洞察も提供しています。ユーザーがどのツールで最大の ROI(投資対効果)を得られるかを判断できるよう、すべてのツールの利用状況を把握できることが重要です。*Databricks の開発者ダッシュボード*
「支出のゲート」:開発者は、支出額に応じて段階的なアクションの実行や承認を求める対象となります。最もシンプルな支出ゲートの形式は、ユーザー自身で解除可能なもので、支出率が特定の閾値を超えたことを警告するものです。Databricks では、誤ってまたは意図せず支出が増大することを防ぐために、自己解除型のゲートが有効な手段であることを発見しました。さらに、管理フローを通じた明示的な予算承認を必要とする段階も導入可能です。
「モデルのダウンシフト」:開発者が支出ゲートに到達した場合、トークンアクセス権限を完全に停止するのではなく、より低コストなモデルへ切り替える(ダウンシフト)ことができます。最安値のモデルは最先端の知能モデルと比べて圧倒的に安価であるため、この手法により、巨額の継続的な支出を発生させることなく開発作業を継続することが可能になります。
停止機能の限界
究極のケースでは、ほとんどのシステムがユーザーをすべてのトークンアクセスから完全に停止する能力を保持しています。前述した通り、これは往々にして一時的な措置に過ぎず、AI をいかに効率的に活用するかについての対話の出発点となります。
コスト削減の手法 #4:トークンのオーバーヘッド低減
ユーザーが AI コーディングエージェントに対して比較的シンプルな要求(例:「このバグを調査して修正してください」)を入力すると、そのエージェントは膨大な量の関連コンテキストを集め、多数のツールを呼び出し、コードベースを検索し、企業が提供するスキルやシステム情報を統合します。高価な LLM 推論が行われる頃には、ユーザーが最初に記述した内容は AI システムに投入されるデータのごく一部に過ぎず、コストの大部分はユーザーが明示的に含めなかったコンテキストによって支配されています。
文脈の肥大化を防ぐための技術はまだ新しい分野ですが、いくつか有望なアプローチが検討され始めています。具体的には以下のような手法です。
- アクティブなコンテキストをより頻繁に圧縮(コンパクト化)するよう強制する
- 会話量が少なくトークン効率の高いハーンネスを使用するか、既存のハーンネスを調整してトークンのオーバーヘッドを減らす
- 人気のあるツールを検証し、その冗長性を削減する
- 開発者にタスクをより小さな個別の作業単位に分割させ、コンテキストのスコープを狭めるよう促す
コンテキストサイズが大きくなると、プロンプトキャッシュも全体のパフォーマンスに大きな役割を果たします。プロプライエタリおよびオープンソースの LLM には、プロンプトキャッシュを有効化し、キャッシュ保持期間を調整できる設定が用意されています。
キャッシュへの書き込みにはコストがかかりますが、キャッシュからの読み取りは推論あたりのコストを劇的に削減できます。このトレードオフは各社の具体的なワークロードに依存するため、デフォルトのキャッシュ設定を手動でチューニングして全体のヒット率を上げれば、コスト面で劇的な改善が見込めます。
Databricks では、ハネスとキャッシュ設定を比較的シンプルな方法で調整した結果、開発者の品質低下を確認することなく、生成トークン数と関連コストを約 50% 削減できました。私たちはこの分野の技術について引き続き探求しており、さらに有意義な最適化の可能性が残っていると考えています。
*余分な推論呼び出しを排除しキャッシュ書き込みを減らすことで、セッションあたりのトークンを劇的に削減しました。*
AI ゲートウェイの設計パターン
上記の技術手法には、多くの暗黙的な要件が伴います。新しいモデルを迅速に活用するためには、「モデルメニュー」を一元的に管理できる場所が必要であり、エンドユーザーは複数のモデルを柔軟に組み合わせられるツールチェーンを備えている必要があります。また、複数の AI ツール全体で予算の可視化を実現するには、統一されたコスト観測機能が存在しなければなりません。さらに、コンテキストの肥大化に対処するためには、一般的なツール呼び出し出力を観察し、圧縮や凝縮を強制する仕組みが求められます。
これらの課題は、「AI ゲートウェイ」と呼ばれる新しいインフラストラクチャソフトウェアによって包括的に解決されています。AI ゲートウェイとは、以下のすべての機能を一元的に提供する場所です。
- 基盤となるモデル(プロプライエタリモデルおよびオープンソースモデルの両方)へのアクセスを管理・仲介するキャパシティ管理。
- モデルのダウンシフトといった複雑な予算ポリシーを含む、予算の追跡と強制執行。
- エンドユーザー向けツールの設定管理。これにより、モデルのホワイトリスト、圧縮設定、およびローカルで調整されるその他の側面を強制します。
- 下流での効率分析やベンチマークに活用できるよう、コーディングセッションのトレースをログ記録すること。
Databricks では、これらの機能すべてに Unity AI Gateway を積極的に活用しています。
すべてを統合する
AI コーディングコストの指数関数的な増加は避けられない運命ではなく、解決可能なエンジニアリングおよびガバナンスの問題です。この課題を克服した企業には共通の戦略があります。
それは、知能の向上よりも効率性の限界追求に執着し、モデルの柔軟性を損なわないツールの採用、最も安価で能力のあるモデルへ知的にリクエストをルーティングする、堅苦しい予算管理から可視化と段階的な摩擦への転換、そして実際の支出を支配するトークンオーバーヘッドの削減です。これらの手法は、AI 導入の意義となった生産性向上を犠牲にする必要はありません。むしろ、これらを組み合わせることで、組織は予測可能なコスト枠組みの中で、低摩擦な広範なアクセスという二つの要請を満たすことが可能になります。
企業のコスト管理を支援する新しいインフラ抽象化が出現しています。Databricks では、コスト管理スタックの主要コンポーネントをオープンソースまたは無料ソフトウェアとして公開しました。中央管理用の Unity AI Gateway と、開発者向けツールの Omnigent です。これらコンポーネントは、毎日数千の企業によって利用されています。この技術環境が急速に進化する中、より多くの企業が知見を共有し、手法を比較することを歓迎します。
本稿では、Uber、Stripe、Coinbase、Ramp のインフラリーダーの方々にご協力いただき、コメントやレビューをいただきました。また、初期草案に対するフィードバックを提供いただいた Thrive Capital にも感謝いたします。
原文を表示
AI coding tools deliver immense value: at Databricks, agentic coding has measurably improved every velocity metric we track and, in some teams, driven an order-of-magnitude gains in output. But nearly every company deploying AI tools at scale has hit the same wall: *exponentially growing costs*. That curve is unsustainable - left unchecked it will eventually overtake revenue. The spend explosion has left enterprises in a paradoxical situation: on the one hand, desiring to maximally push AI transformation and put powerful tools in the hands of employees, and on the other hand, having to reconcile with an aggregate cost profile that threatens to undermine or even reverse the very efficiency gains AI provides.
Fortunately, several of the earliest large-scale adopters have converged on a set of approaches that solve this puzzle, achieving a “dual mandate”: (a) providing broad access to AI tooling, with minimal friction, and (b) keeping aggregate costs inside of a roughly fixed envelope per user. This post outlines proven cost management techniques, based on our experience at Databricks and conversations with several other digital-native companies, including Stripe, Coinbase, Uber, and Ramp. The table below summarizes current techniques and associated savings; the numbers are directional, based on an informal survey of development teams:
Some of these techniques can be easily implemented with software many companies already use. Others require new infrastructure, particularly techniques that modify end-user clients or shift traffic across models. At Databricks, we’ve open sourced or made freely available our key infrastructure components: an end user meta-harness (Omnigent) and our AI Gateway (Unity AI Gateway). For completeness, this post also covers software used by other companies we spoke with.
The “Efficiency Frontier” for Coding Models
The single greatest cost lever in moving coding spend to more efficient models as they are released. This point bears some discussion, as the simple explanation of "cheaper models” in fact hides a nuanced relationship between model cost and quality.
Colloquially, the term *frontier model* means “the highest intelligence model,” and *frontier labs *largely focus on advancing peak intelligence. Frontier models can now solve novel problems in math or cybersecurity. But when AI is deployed at scale, a different type of frontier matters more: the *efficiency frontier*. The efficiency frontier is defined by the set of models that have the best price point *for a given level of intelligence*. Most day-to-day coding doesn't require mathematical proofs or novel security insights, so what matters in aggregate is the cost of models that meet the quality bar for typical software engineering work. This "efficiency frontier” is advancing far faster than the intelligence frontier, with new models being released almost weekly that present better intelligence-per-unit-price than prior models.
Cost Lever #1: Moving to open source and lower cost models
| Step 1. Measure model efficiency frontier | Step 2. Migrate users to best models |
|---|---|
| Databricks Recommended Models (as of August 6, 2026)GLM 5.2Opus 4.8GPT 5.6-Sol*Note: For data and knowledge work tasks, Databricks uses *Genie One*; the above is for core software development.* |
Rapidly adopting newer, more efficient models delivers the largest cost wins of any technique. But to capture those gains, a company first needs to know which models actually beat its incumbents. This can be difficult because public benchmarks do a poor job of indicating real-world performance on coding tasks. To size up new models, many companies have built automated evaluations that they believe are more representative of their internal development mix. Databricks recently published an example of such a benchmark, in which we observed highly competitive price/performance for GLM models. That benchmark led us to roll GLM out to developers internally. Often, new models do not advance the efficiency frontier,and evaluations frequently produce negative results: Stripe found that Opus 4.7 did not meaningfully improve quality over Opus 4.6, while increasing cost. They therefore declined to make Opus 4.7 available internally. Databricks saw similar cost regressions when comparing Opus 5.0 to 4.8.
Harness and Model Flexibility
Since the biggest wins come from switching to new models, adopting end user tooling that allows for model flexibility is becoming a critical component of keeping costs down. The tool most commonly used in concern with a particular model is called* harness*. Proprietary frontier models are increasingly co-designed to work well with specific harnesses, meaning certain harnesses “work better” with certain models. If a company wants to preserve model independence there are roughly two approaches:
Ask users to switch harnesses. One approach is to provide developers with a set of harnesses (Claude Code, Codex, or Cursor) and then ask them to switch between harnesses when a company wants to migrate spend to lower cost models. This lets users work in their preferred harness when possible, but the downside of this approach is that switching costs for an individual developer can be high. If switching costs become too high, the harness itself becomes a de facto lock-in to a model family, limiting the ability to move spend to more competitive models.
Use a meta-harness. A new and increasingly popular approach is to use a *meta-harness* that surfaces a common user experience to developers while dispatching requests to underlying harnesses (both proprietary and open source). This approach allows both model/harness independence while also reducing developer switching costs. At Databricks, this is the default mode for developers who leverage Omnigent. Some companies we talked to have built custom internal meta-harnesses that integrate with their development toolchain.
Cost Lever #2: Dynamic Request and Task Routing
Instead of asking users to choose task-appropriate models themselves, a growing body of research suggests that automatic model and tool selection may further squeeze efficiency out of agentic coding workflows. Routing approaches roughly fall into three categories:
- Request Level Routing: A stateful proxy sits in between a client (such as a coding harness) and the underlying foundation models. The proxy attempts to route requests to the lowest-cost model capable of answering each inference request. Routing for agentic use cases also needs to account for server-side caching, since a cold cache hit has a very high cost for large context workloads. A new wave of products is showing early, promising results for routing. Examples are: Cursor Router, OpenRouter’s AutoRouter, Ramps Router feature and Databricks own Smart Routing feature in Unity AI Gateway.
- Task Level Routing (Meta Harness): A client-side process dispatches user tasks to different harnesses based on the complexity of the task. A user task might be “rename this component from X to Y” (a simple task) or an open-ended task like “Explore design considerations that would reduce latency” (a complex task). The dispatcher, often called a Meta Harness, examines which level of underlying model is required for a task and then delegates that entire end-to-end task to the model. Omnigent is an example of a Meta Harness that supports this pattern.
- Escalation/Delegation Patterns: A single harness pairs two models (an expensive, high-intelligence model and a cheap worker model). In some approaches, such as Claude’s Advisor Tool, the cheaper model runs the show and escalates when it thinks a task requires more horsepower. The inverse pattern also exists: In Cognition’s Devin Fusion, the higher cost model is the main loop, and it selectively outsources work to a cheaper model.Internal results at Databricks suggest that our AI Gateway Smart Router is able to consistently reduce average task cost by more than 30%, while roughly matching the quality of the most expensive model in the working set. Other companies we spoke with have seen similar results.
Cost Lever #3: Giving developers visibility, tripwires, and budgets
It may be surprising that this entire article did not start and end with “Give users a monthly budget and be done with it.” *Hard budgets*, where usage is entirely cut off at a specific spend threshold, are often used only as a last resort option in every company we spoke with. There are two reasons that hard token budgets are not particularly effective for AI spend management: First, if a developer hits their budget ceiling, cutting off further access to AI tools would be debilitating to productivity. Neither the company or employee actually wants that outcome. Second, at least *some* of the “high spending” users are in fact those who have achieved monumental efficiency gains with AI and are producing immense output. Discouraging those users is self-defeating.
Instead of a hard user spending cap, most companies are adopting a more nuanced and progressive approach that focuses on visibility for end users and increased degrees of friction as spend increases.
- Visibility: Every company we spoke with had a mechanism to provide near-instantaneous feedback to users on their ongoing spend, with many also offering specific tips or insights on how to reduce spend by using less expensive models. It is important that users be able to see their spend across all tools, since they may want to influence their choice of tool where they get the highest ROI.
A developer dashboard at Databricks
- Spend Gates: Developers can be asked to take actions or seek approvals at increasing levels of spend. The simplest form of spend gate is one that can be self-cleared and serves as a warning that the spend rate is increasing above some threshold. At Databricks, we’ve found self-clearing gates a useful mechanism for preventing accidental or unintentional spend. Further gates can be introduced that require explicit budget approval (often through a management chain).
- Downshifting: If a developer has hit a spend gate, they can be downshifted to a lower-cost model rather than being entirely suspended from token access. Since the lowest-cost models are drastically less expensive than frontier-intelligence models, this technique allows developers to continue getting work done without incurring massive ongoing spend.
- Suspension: In the limit case, most systems do retain the ability to fully suspend users from all token access. As stated above, this is often a temporary measure only and the starting point for a conversation about how to efficiently leverage AI.
Cost Lever #4: Reducing Token Overhead
When a user types a relatively simple request into an AI coding agent (such as “Please investigate and fix this bug.”), that agent subsequently gathers massive amounts of relevant context, invokes a large number of tools, searches through the codebase, and integrates skills or system information provided by the company. By the time costly LLM inference occurs, the user's initial statement accounts for only a negligible fraction of the data fed into the AI system, meaning costs are dominated by context the user did not explicitly include. Techniques in reducing context bloat are still new, but several promising approaches are being explored, such as:
- Coercing more frequent compaction (compression) of the active context.
- Using harnesses that are “less chatty” (more token efficient), or tuning existing harnesses to generate less token overhead.
- Auditing popular tools and decreasing their verbosity.
- Encouraging developers to break tasks into smaller individual units of work, decreasing context scope.
When contexts get large, prompt caching also plays a meaningful role in overall performance. Both proprietary and open source LLMs have settings that allow you to enable prompt caching and tune how long the cache is stored. Cache writes cost money, but cached reads can drastically reduce per-inference cost. This trade-off is dependent on a company’s specific workload, so hand-tuning of default cache settings to increase overall cache hit rate can have drastic improvements to overall cost.
At Databricks, relatively simple tuning of our harness and caching settings led to an almost 50% reduction in the number of generated tokens and associated costs, with no observed quality degradation for developers. We continue to explore techniques in this area and think meaningful additional optimization remains possible.
*A drastic reduction in tokens per session by eliminating extraneous inference calls and reducing cache writes.*
The AI Gateway design pattern
The techniques above had many implicit technical requirements: To rapidly take advantage of new models, companies must have a central location where the “model menu” is managed, and end-users must have a toolchain that supports model mixing. To provide budget visibility across multiple AI tools, a unified cost observability capability must exist. To manage context bloat, companies need a way to observe typical toolcall outputs and enforce compression or compaction. These needs are collectively being solved by a new class of infrastructure software, best described as an AI Gateway. An AI gateway is a central location where all of the following occur:
- Capacity management and proxying of access to underlying models (both proprietary and OSS models).
- Budget tracking and enforcement, including complex budget policies such as model downshifting.
- Configuration management for end-user tools, to enforce model allow-lists, compaction settings, and other locally mediated aspects.
- Logging of coding session traces for downstream efficiency analysis and benchmark.
At Databricks, we rely heavily on Unity AI Gateway for all of these capabilities.
Putting it all together
The exponential growth of AI coding costs is not an inevitability, it's a solvable engineering and governance problem. Companies that have tamed it share a common playbook: relentlessly chase the efficiency frontier rather than the intelligence frontier, adopt tooling that preserves model flexibility, route work intelligently to the cheapest capable model, replace hard budgets with visibility and progressive friction, and cut the token overhead that dominates real-world spend. None of these techniques requires sacrificing the productivity gains that made AI adoption worthwhile in the first place; together, they let organizations satisfy the dual mandate of broad, low-friction access within a predictable cost envelope.
A set of new infrastructure abstractions is emerging to give companies the tools to manage their costs. At Databricks, we’ve released the key components in our cost management stack as open source or free software products: Our Unity AI Gateway for central management and Omnigent for developer tooling. Thousands of companies use these components every day. We invite more companies to share findings and compare techniques as this technology landscape rapidly evolves.
*Acknowledgements: Thank you to infrastructure leaders at Uber, Stripe, Coinbase, and Ramp who provided commentary and reviews of this article. Thank you to Thrive Capital for feedback on an early draft of this article.*
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み