Databricks、Unity AI Gateway でスマートルーティング導入しコスト削減
本文の状態
日本語全文を表示中
詳細モードで約16分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Databricks AI Engineering
Databricks は Unity AI Gateway にスマートルーティング機能を追加し、コーディングタスクにおいて最も高価なモデルを不要な場合に使用しないことで、1 件あたりのコストを 30% 以上削減できると報告した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月14日 04:27
AI深層分析
キーポイント
Smart Routing の実装と機能
Databricks は Unity AI Gateway に「Smart Routing」機能をベータ版として追加し、タスクの複雑度に基づいて自動的に最適なモデルやコーディングハーンネスを割り当てる仕組みを実装した。
コスト削減と性能維持の実証
同社によると、内部コードワークロードでは主要モデル(Opus 5)の65%のコストで同等以上の成果を達成し、公開ベンチマークでも半額未満のコストで同等の性能を発揮したという。
Omnigent との統合による最適化
コーディングエージェント用のメタハーンネス「Omnigent」を併用することで、モデル選択だけでなくツール(ハーンネス)の組み合わせまで自動最適化し、開発者の手動選定負担を排除する。
既存ツールの拡張性
この機能は Claude Code や Codex などの既存の開発者向けツール上で直接動作するため、ワークフローの変更や学習コストを最小限に抑えて導入が可能である。
単純タスクには安価なモデルを使用
タスクの複雑さに応じて安価なモデルを選択し、高度な性能が必要な場合にのみ上位モデルへエスカレートする。
重要な引用
Smart Routing works directly in Claude Code and Codex, allowing you to optimize the tools developers already use.
On internal coding workloads, Smart Routing outperformed any single model at just 65% of the cost per task of a leading model like Opus 5.
Just leveraging lower cost models can save you 50%+, but it's incredibly daunting for users.
Most of the win comes from using cheaper models for simpler tasks.
編集コメントを表示
編集コメント
Databricks は、モデル選定の複雑さを解消する「Smart Routing」により、コスト削減と生産性維持の両立を現実的なレベルで実現した。特に既存ツールとの親和性を重視した設計は、現場での導入障壁を下げる重要な要素となる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
コーディングタスクにおける価格と性能のフロンティアには、非常に多様なモデルやハルネスが存在します。2026 年だけで 33 の新モデルがリリースされるほどです。
以前、Databricks のコードベースに対するベンチマーク調査についてまとめた記事で、モデルは能力レベルごとにクラスター化されており、フラグの切り替えや単一ファイルの編集、範囲を限定したバグ修正など、日常的な作業には最も高価なモデルが必ずしも必要ではないことを示しました。
では、開発者の生産性を損なうことなく AI コーディングコストを削減するにはどうすればよいでしょうか。最大の機会の一つは、すべてのタスクを最も高性能(かつ高価格)のオプションにデフォルト設定するのではなく、各タスクに適したモデルを選択することです。安価なモデルを活用するだけでコストを 50% 以上削減できる可能性がありますが、ユーザーにとっては非常にハードルが高い現実でもあります。
優れたモデルやハルネスが乱立する中、コーディングエージェントの利用者は常に「選択の過多」に直面しています。すべてのタスクに対して最適なモデルを選ぼうとして時間を浪費するのではなく、多くの人が最も高性能なモデルを最高難易度の設定で固定し、そのまま作業を進めています。ユーザーに選択を迫ったり、生産性を阻害する厳しい制限をかけたりするのではなく、私たちは新たなアプローチが必要だと確信しました。
そのため、Unity AI Gateway の新たなコスト制御機能として「Smart Routing」をベータ版としてリリースします。Unity AI Gateway は、企業全体で AI へのアクセス管理や支出統制、ポリシー適用を一元的に行えるプラットフォームですが、Smart Routing を追加することで、タスクの複雑さに応じて最適なモデルを自動選定するインテリジェントな最適化が可能になります。この機能は Claude Code や Codex の中で直接動作するため、開発者がすでに使用しているツールをそのままにしながらコスト効率を最大化できます。
さらに、モデルのルーティングだけでなく、コードエージェント向けのメタ・ハーネス Omnigent を活用することで、モデルとハーネスの両面から最適化を実現します。これにより、開発者は手動で組み合わせを選ぶ必要なく、タスクに最適な構成を自動的に得ることができます。
その効果は数字が物語っています。社内でのコーディングワークロードでは、Smart Routing は主要なモデルである Opus 5 のコストのわずか 65% で、単一モデルよりも優れたパフォーマンスを発揮しました。また、公開ベンチマークにおいても、Opus 5 と同等の性能を維持しながら、コストは半分以下に抑えられています。
ここからは、私たちが得た知見を紹介します。
単純なタスクにはより安価なモデルを活用することで、コスト削減の大部分を実現しています。社内では多様なタスクが発生しますが、その多くは高機能なモデルを必要としません。そこで、ルーターを設定し、簡単なタスクには低コストモデルを選択させつつ、最先端モデルのパフォーマンスが求められる複雑な作業については「エスカレーション(上位モデルへの引き上げ)」を行うようにしています。
タスク開始時に利用可能な情報(説明とメタデータ)のみを用いても、良好な結果が得られました。回答やテストケース、リポジトリに関する詳細などは提供していませんでしたが、Smart Routing は依然として効果的にモデルを選択できました。
タスクの複雑さを正確に評価し、必要に応じてエスカレーションする点については、まだ改善の余地が大きく残っています。完璧な先見性を持つルーターであれば、現在の費用のほんの一部で全てのモデルを上回る性能を発揮できるでしょう。また、実際のセッションでは、作業が進行中に複雑化していないかを再評価することも有効です。このギャップを埋めることは、研究課題であると同時にハッチネス設計上の課題でもあります。改善のためには、実際のユーザーフィードバックから学ぶ必要があります。
それでは、どのようにしてこれを構築したのかを見ていきましょう。
インテリジェントなモデルルーティングはどのように動作するのか?
インテリジェントなモデルルーティングは、複雑さ、能力、コストなどの要因に基づいて、タスクに最も適したモデルを選択します。コーディングエージェントにおいては、このルーティングをいつ実行するかが重要な判断ポイントとなります。
ルーティングには一般的に 2 つのアプローチがあります。
リクエストごとのルーティング:特定のセッション内では、一部のチームがメッセージの各リクエストをプロンプムの複雑さのみに基づいてルーティングするアプローチを検討しました。しかし、スケールした運用において課題となるのはコスト構造です。実際にはコストは「キャッシュヒット率」によって支配されます。高いキャッシュヒット率を維持するには、連続するターン(対話)を同じモデルに、そして現在では人気のあるモデルについては同じリソースレベルへルーティングする必要があります。
タスク認識型ルーティング:セッション開始時にはまずタスクの複雑さを評価し、それに適したモデルとハーンネス(実行環境)を提案します。この選択はセッション全体を通じて維持されます。これによりキャッシュヒット率が保たれ、将来的な最適化の余地も生まれます。具体的には、キャッシュが古くなった際(例えばコンパクションイベントが発生した時など)に必要に応じてモデルやリソースレベルのアップグレード・ダウングレードが可能になります。
スマートルーターの仕組みについて
私たちはキャッシュ効率を維持しつつ、各コーディングタスクに適したモデルと実行環境をマッチングさせるため、「タスク認識型ルーティング」を採用しました。最も興味深い課題は、タスクを開始する前にその難易度を判断することです。まずはシンプルに始めたいと考えたため、現在のルーターでは単一のポリシーを用い、すべてのタスクに対して同じ基準で適用しています。
まず、タスクを分類します。そのために、低コストかつ低遅延のモデルを使用し、タスクの説明を読み込んで、いくつかのセマンティックなフィールドにラベル付けします。具体的には、「システムのどの部分が変更されるか」「プロンプトに含まれるコード証拠(スニペット、トレースバック、あるいは明示的な情報なし)」「失敗の様相」「修正が局所的に見えるかどうか」「どのようなプロジェクトに属するか」です。
これらの情報から、ルーターはタスクタイプファミリーと言語ファミリーを導き出します。フロンティアモデルを使用すると、すべてのリクエスト(コストを抑えたい単純なものも含めて)に負荷がかかるため、この抽出器は意図的に小さく高速に設計されています。
次に、どのモデルクラスが最適かを三角測量によって特定します。ルーターはデフォルトで中規模のモデルを選択し、ラベル情報に基づいて方向を調整します。タスクがフロンティアレベルの能力と知識を必要とする場合は高価なモデルへエスカレートし、そうでない場合はより安価なモデルへ委譲します。これにより、単一のポリシーですべてのモデルスイートを活用することが可能になります。
初期結果は有望です。他社がアクセスできない独自の内部ベンチマークでは 35% のコスト削減を達成しました。一般化できることを示す公開コーディングベンチマークでも、56% のコスト削減を実現しています。自社のユースケースに関する知見の深化や、設計パートナーとの連携が進むにつれて、この効果はさらに拡大すると予想されます。
コーディングタスクをどのようにしてモデルやハッチェス間でルーティングするか?
Smart Routing はルーティングの決定を担当しますが、その決定に基づいて実際に行動を起こす仕組みも必要です。この機能は Claude Code や Codex 内でネイティブに動作しますが、コーディングエージェントにおいては、適切なモデルを選ぶだけでなく、最適なコーディングハッチェスを選択することでパフォーマンスをさらに向上させることができます。エンジニアが選択したモデルとハッチェスを最大限に活用できるよう支援するには、個々のコーディングセッションの上に位置し、それらを統括するレイヤーが必要です。そのため私たちは Omnigent を構築しました。
Smart Routing は Omnigent において2つのレベルで実装されています。
第一に、Omnigent を利用する開発者は、特定のコーディングハッチェスを手動で選択する代わりに Smart Routing を選定できます。これにより、Omnigent が各タスクに対して自動的にハッチェスとモデルを選択します。モデルのルーティングは Unity AI Gateway の Smart Routing によって支えられています。この設計により、開発者や管理者は、クライアントを毎回変更することなく、組織レベルの方針の提供や過去の会話履歴の利用オプションといったカスタマイズを柔軟に行うことができます。
これにより、すべてのサブエージェントの起動は Smart Routing API を経由して行われ、サブエージェントは異なるハーンチスとモデルを活用できるようになります。ユーザーからの初期プロンプトは往々にして不十分で、その複雑さを判断するのが難しいため、サブエージェントを使用すれば、新しい情報に基づいて作業を調整し、キャッシュをリセットして明確な指示を出すことが可能になります。
1 つのタスクでも、計画段階と並列実行されるサブエージェント作業の間で、微妙なルーティング判断が行われることがあります(例:大規模なコードベースの要約には低コストモデルを割り当て、アーキテクチャ設計には高価なモデルを使用するなど)。これにより、さらに大きなコスト削減効果が得られます。
モデルルーティングが機能しているかどうかをどう評価するか?
効果的なモデルルーティングは、コストだけでなく開発者の生産性も最適化の対象とする必要があります。コストだけを追求してはいけません。ルーティング技術はまだ初期段階であり、大幅な改善が必要であるため、フィードバック信号を持つことが極めて重要です。
その第一歩として、私たちはすべてのコーディングセッションのトレースを記録し、後で評価に活用できるようにしました。コストと開発者の体験の両方を考慮する必要があります。生産性を犠牲にしてまでコストだけを最適化してはなりません。
Unity AI Gateway では、コーディングエージェントのトレースを Unity Catalog に記録できます。これは極めて機微なデータであり、成熟した企業では高度なタグ付けとアクセスポリシーによって管理される必要があります。
ルーターの変更を評価するために、AI モデルと人間のレビューの両方を用いてトレースを分析しました。ルーティング前の自社セッションでこの分析を実行した際、デフォルトモデルが最も高価であるという理由だけで、多くのセッションがその必要性のない作業に「フロンティア・モデル」の予算を費やしていることがわかりました。
実際には、ルーターの有効性を検証する最善の方法は、以下の指標を継続的に監視することです:
- モデル別のセッション内訳
- ルートされたモデルによってエンドツーエンドで完了したセッション数
- ルーティングによるドル単位の節約額
次世代のインテリジェント・モデルルーティングへの展望
この分野には大きな可能性があり、引き続き本格的な研究を継続する計画です。まだ初期段階のため、学ぶべきことが多くあります。
最初の課題は、実際のユーザー行動と一致しない信頼性の低いベンチマークデータでした。ベンチマークタスクは非常に整然としており、それぞれが独立した作業指示として到着します。ルーターはこうしたタスクでは良好に動作しますが、実際のセッションは全く異なるケースが多いのです。
- 最初のプロンプトは精度が低いことがほとんどです。開発者が最初に入力するのは仕様書ではなく、症状や大まかな意図を示すものだからです。しかし、ルーターはその最初のメッセージを読み込んで判断を確定してしまいます。
- セッションが再利用されるため、最初の要求には正しかった判断も、4 回目のリクエストでは間違っている可能性があります。その際、再度ルーターに問い直す仕組みがありません。
そのため、より多くの情報を収集し、新しい手法を試すために、以下の新たな方向性で研究を進めています:
タスクのスコーピングが無料の段階から始めよう。PR 審査、サブエージェントの起動、バッチ移行、 scheduled jobs は、すべて機械によって作成されたタスク記述に基づいて事前に完全に指定されています。この種のタスクには、現在のルーティング手法をそのまま適用でき、誰かの作業習慣を変える必要もありません。そのため、まずはここから展開します。
数回やり取りした後でルーティングしよう。対話型の業務では、最初のプロンプトが判断を下すのに最も不適切なタイミングです。無理にそこで決める必要はありません。むしろ、安価なモデルで初期のやり取りを処理し、不明点を明確にした上でタスクが形になってからルーティングするのが有効です。速く小さなモデルで探索する方がコストと体験の両面で改善されるため、このアプローチを採用すべきです。
セッションを小さく保とう。一つのタスクに集中したセッションほどルーティングが効きやすく、コストも抑えられます。ツール側で、話題が変わった際に新しいセッションを開始するのが自然な選択肢となるよう設計することで、これを促進できます。
切り替えのコストを下げる。上記のすべてを実現するには、セッション中にモデルを切り替えることが低コストである必要があります。現在の状況ではキャッシュヒット率がコストの大部分を占めるため、大規模運用でセッション中のモデル切り替えは現実的ではありません。文脈圧縮(context compaction)が自然な接合点となります。なぜなら、そこではすでにキャッシュミスが発生しているからです(これは Cognition の Devin Fusion が採用している手法です)。将来的には、ルーティング層がキャッシュミスを明示的に価格設定するようになり、現在のルーティング利用方法にそのコストを埋め込む必要がなくなることを目指しています。
ルーティングは通常、コスト削減のための手段として語られます。私たちの成果の多くは簡単なタスクに対してより低い価格を支払うことから生まれますが、同じ仕組みを使って、より良い結果を得るためにいつ支出を増やすべきかを判断することも可能です。
価値最大化には二つの側面があります:必要であれば安価なモデルを使い、その価値に見合う場合は迷わず支出を増やすことです。
コーディングツールはユーザーにより多くのトークン消費を促す傾向がありますが、私たちが本当に目指すべきは、ドルあたりの生産的な「アウトプット」の最適化であり、トークンの量ではありません。必要であればより安価で高速なモデルを選ぶことは、単にお金を節約するだけでなく、時間を節約し、真に必要なタスクのために貴重な最先端リソースを確保することにもつながります。
Smart Routing の利用開始
Smart Routing は現在、Unity AI Gateway を通じてベータ版として提供されています。複雑度に基づいてコーディングタスクを適切なモデルに自動的にルーティングするため、チームは各タスクに最適なモデルを選択することで、最先端レベルのパフォーマンスを実現しながらコストを 30% 以上削減できます。
開発者の選択肢や生産性を制限することなく AI コーディングコストを削減したいチームにとって、Smart Routing は手動でモデルを選定したり、単純な支出上限に頼ったりする従来の手法に対する代替案となります。また Omnigent を利用すれば、複数のモデルやコーディングハーンレスにわたってインテリジェントなルーティングを拡張することも可能です。
利用を開始するには、以下のドキュメントページをご覧ください:
- モデルのルーティング: Unity AI Gateway でスマートルーティングを有効にする
- モデルとハッチの組み合わせによるルーティング: Omnigent (v0.9.0+) と組み合わせてスマートルーティングを使用する
Unity AI Gateway については、公式サイト をご覧ください。
原文を表示
The price and performance frontier for coding tasks features a huge diversity of models and harnesses: in 2026 alone, we’ve seen 33 new models released. In our prior post about benchmarking against the Databricks codebase, we found that models cluster into capability tiers and that much everyday work (e.g., flipping a flag, a single-file edit, a well-scoped bug fix) did not require the most expensive models.
So how do you reduce AI coding costs without sacrificing developer productivity? One of the biggest opportunities is matching each task to the right model instead of defaulting every task to the most capable (and most expensive) option. Just leveraging lower cost models can save you 50%+, but it’s incredibly daunting for users. With a proliferation of great models and capable harnesses, coding agent users are constantly faced with choice overload. Instead of wasting time trying to select the best model for every task, many are setting the most capable at the highest effort and moving on. Instead of asking users to choose or stunting productivity with hard caps, we knew we needed to innovate.
That’s why we’re launching the next major cost control in Unity AI Gateway: Smart Routing, now available in Beta. Unity AI Gateway provides a central place to get access to AI, manage spend, and enforce controls across your entire enterprise, and Smart Routing adds intelligent optimization by automatically matching tasks to the right model based on complexity. Smart Routing works directly in Claude Code and Codex, allowing you to optimize the tools developers already use.
And we’re going beyond model routing. With Omnigent, our meta-harness for coding agents, teams can leverage the full power of Smart Routing by optimizing across both models and coding harnesses, giving developers the right combination for the task without having to choose it themselves.
The results speak for themselves: On internal coding workloads, Smart Routing outperformed any single model at just 65% of the cost per task of a leading model like Opus 5. On public benchmarks, Smart Routing matched Opus 5 on performance at *less than half the cost*.
Here’s what we learned:
- Most of the win comes from using cheaper models for simpler tasks. We see a wide diversity of tasks internally, and most of them don’t need a premium model. We configured the router to select cheaper models for simpler tasks, but to “escalate” complex work that requires frontier-model performance.
- We saw good results using only the information available at the start of the task (description and metadata). We didn’t provide the answer, tests, or anything about the repository, and Smart Routing was still able to choose effectively.
- There’s still substantial room for improvement in effectively sizing work complexity and escalating when needed. A router with perfect foresight would beat every single model at a fraction of what ours spends, and in real-life sessions, it can also be helpful to reassess midway through whether the task has gotten more complicated. Closing this gap is both a research and harness design problem. We need to learn from real user feedback to improve.
Let’s walk through how we built this.
How does intelligent model routing work?
Intelligent model routing selects the model best suited for a task based on factors like complexity, capability and cost. For coding agents, an important decision is when that routing should happen.
There are generally two approaches to routing:
- Per-request routing: In a given session, some teams have explored routing each request purely based on the complexity of the prompt for that message. The challenge is that, at scale, costs are dominated by cache hit rate. Having a high cache hit rate requires routing consecutive turns to the same model (and, currently, to the same effort level for the popular models).
- Task-aware routing: At the beginning of the session, you start by assessing the complexity of the task and then suggest a model and harness for it that you stick to for the duration of the session. This preserves the cache-hit rate and provides an opportunity for future optimization, such as upgrading or downgrading as needed when the cache becomes stale (e.g. when there is a compaction event).
How our Smart Router works
We opted for task-aware routing to preserve cache efficiency while matching each coding task to the appropriate model and harness. The most interesting problem is judging how difficult a task is before starting it. We wanted to start simple, so our router currently uses a single policy and applies it to every task the same way.
First, we classify the task. For this, we use a cheaper, low-latency model that reads the task description and labels it with a handful of semantic fields: what part of the system changes, what code evidence the prompt carries (a snippet, a traceback, or nothing explicit), how it appears to be failing, how localized the fix looks, and what kind of project it belongs to. From these, the router derives a task-type family and a language family. Using a frontier model would tax every request (even the simple ones we want to save on), so the extractor is intentionally small and fast.
Then, we triangulate on which model class is best. The router defaults to a medium-sized model and uses the labels to move in either direction, escalating to a more expensive model when the task demands frontier-level capability and knowledge or delegating down to a cheaper one when it does not. This means the single policy can leverage a whole suite of models.
The early results are promising. Against our own internal benchmark, which no labs have had access to, we saw 35% savings. Against public coding benchmarks that demonstrate our results generalize, we achieved 56% cost savings. We expect to see this grow as we learn more about our own use cases and with our design partners.
How do you route coding tasks across models and harnesses?
Smart Routing handles the routing decision, but then you need to be able to act on it. It works natively inside of Claude Code and Codex, but for coding agents, we see better performance by choosing not only the right model, but also the right coding harness. Helping engineers take advantage of the chosen model and harness requires a layer sitting above the individual coding sessions to orchestrate across them. This is why we built Omnigent.
Smart Routing is implemented in Omnigent at two levels:
First, developers using Omnigent can select Smart Routing instead of manually choosing a specific coding harness. Omnigent then automatically selects both the harness and model for each task, with model routing powered by Smart Routing in Unity AI Gateway. This design gives developers and admins the freedom to provide customizations, such as org-level guidance or the option to use previous conversation history, without changing the client every time.
This also means all sub-agent launches go through the Smart Routing API, allowing sub-agents to leverage a different harness and model. The user's initial prompt is often underspecified and hard to judge in terms of complexity, so sub-agents allow you to adjust new work based on new information, with a fresh cache and clear instructions. A single task can experience nuanced routing decisions across planning and parallel sub-agent work (e.g., you can route large codebase summarization tasks to cheaper models while designing the architecture with more expensive ones), leading to even more substantial savings.
How do you evaluate whether model routing is working?
Effective model routing needs to optimize for both cost and developer productivity, and not only cost alone. It’s critical to have feedback signals since routers are still early technologies that will need substantial iteration. Our first step here was to log all coding session traces for later evaluation. We want to consider both cost and developer experience – we don’t want to optimize cost at the expense of productivity.
With Unity AI Gateway, traces for coding agents can be recorded into Unity Catalog. This is extremely sensitive data, and it needs to stay governed by sophisticated tagging and access policies in most mature enterprises.
We used both AI models and human review to analyze traces to evaluate changes to the router. When we ran this analysis on our own sessions before routing, we observed that a large share of sessions were spending frontier-model money on work that did not need it, simply because the default model was the most expensive. In practice, the way to validate that the router is useful is to continuously monitor the following metrics:
- Breakdown of sessions by models
- # sessions that are completed by the routed models end-to-end
- Amount of dollar savings from routing
Where we are taking intelligent model routing next
We believe this is an area with significant opportunity and plan to continue conducting substantial research here. It’s still early, so we have a lot to learn.
Our first challenge has been unreliable benchmarking data that doesn’t match real user behavior. Benchmark tasks are unusually well-behaved, with each arriving as a self-contained statement of work. While routers perform well on such tasks, real sessions often are not like that at all.
- Opening prompts are rarely precise, since what a developer types first is a symptom or a rough intention rather than a specification, and our router reads that first message and commits.
- Sessions get reused, so the decision that was right for the first request can be wrong by the fourth, and nothing asks the router again.
So we’re researching a few new directions to gather more information and try new techniques:
- Start where task scoping is free. PR reviews, sub-agent launches, batch migrations and scheduled jobs are fully specified at the outset because a machine wrote the task statement. Routing works on this class today without changing anyone's habits, which is why we are deploying here first.
- Route after a few turns, not on the first one. For interactive work, the opening prompt is the worst available moment to decide, and nothing forces us to decide there. Instead, it can be useful to let a cheap model handle the initial exchange and ask clarifying questions, then route the task once it has taken shape. It is better to explore with a fast and small model, so this should improve cost and experience together.
- Make sessions smaller. Sessions that stay on one task route better and cost less, and tooling can encourage that by making a fresh session the obvious move when the subject changes.
- Make switching cheap. All of the above need mid-session model changes to be affordable. In today’s world, with costs dominated by cache-hit-rate, switching mid-session is untenable at scale. Context compaction is the natural seam, since a cache miss is already happening there (this is how Cognition’s Devin Fusion does it). Over time, we want the routing layer to explicitly price cache misses rather than bake that into how we use the router.
Routing is usually pitched as a way to spend less. While most of our wins come from paying lower prices for easy work, the same machinery can help us decide when to spend more to get a better outcome. Valuemaxxing cuts both ways: take the cheap model when it suffices, and be confident in spending more when the value justifies it.
Coding tools often incentivize users to consume more and more tokens, but what we really want is to optimize productive *output* per dollar, not tokens. Picking a cheaper, faster model when it suffices doesn’t just save money. It also saves time, and it keeps scarce frontier capacity available for the tasks that genuinely need it.
Try Smart Routing today
Smart Routing is now available in Beta through Unity AI Gateway. It automatically routes coding tasks to the right model based on complexity, helping teams achieve frontier-level performance with 30%+ cost savings by selecting the best model for every task. For teams looking to reduce AI coding costs without limiting developer choice or productivity, Smart Routing provides an alternative to manually selecting models or relying on blunt spending caps. And with Omnigent, you can extend intelligent routing across models and coding harnesses.
To get started, visit our docs pages:
- For model routing: Enable Smart Routing through Unity AI Gateway
- For model + harness routing: Use Smart Routing with Omnigent (v0.9.0+)
Learn more about Unity AI Gateway by visiting our website.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み