The Pulse:新たなトレンド、スマートモデルルーティング
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Pragmatic Engineer
ゲルゲーが、エンジニアリング部門における AI 支出削減の動向を解説するニュースレターで、スマートなモデルルーティングという新トレンドを取り上げている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
こんにちは、Pragmatic Engineer ニュースレターの特別無料号をお届けするゲルゲーです。毎号、私はシニアエンジニアやエンジニアリングリーダーの視点から、ビッグテックとスタートアップを取り上げています。今回は、以前の『The Pulse』記事で取り上げた 4 つのトピックのうちの一つを扱います。フルサブスクライバーには、3 週間前に以下の記事を配信しました。このメールを転送された方は、こちらから購読できます。
2 週間前、私はエンジニアリング部門における AI への支出削減を試みる企業のトレンドについて取り上げました。その際、大手企業のエンジニアリング責任者の一人に話を聞いたところ、「適切なタスクに対して最適なモデルを選別する『インテリジェント』なルーターがあれば」と願っているとのことでした。
このような願いが生まれる理由は明白です。トークンあたりの価格はモデルによって大きく異なり、安価で平均的なモデルと最先端のモデルの間では、簡単に 10〜20 倍もの差が生じます。
このようにメリットが明らかなため、現在こうしたソリューションが存在するかどうかを調べてみました。その結果は以下の通りです。いつもの免責事項:私はこれらのベンダーとは一切関係がなく、紹介料を受け取ったこともありません!
ベンダー:
Factory Router: セッションごとに最適なモデルを自動選定し、コストを 20〜25% 削減できると主張しています。詳細はこちら。
Not Diamond: コーディングモデルの自動選定を行い、約 30% のコスト削減を実現すると主張しています。OpenRouter で採用されており、その裏側で動作しています。詳細はこちら。
Augment Code 社の Prism: コーディングタスクに対して自動的に「最良」のモデルを選択します。詳細はこちら。
Morph による Model Router。プロンプトに基づき、モデルのリストからモデル選択を提案する API です。詳細はこちら
Weave router:Codex、Claude Code、Cursor の内部で動作するトークンルーターです。「ハード」なリクエストはフロンティアモデルに留まり、「イージー」なものはオープンソースモデルへ転送されます。詳細はこちら
組み込みのルーティング機能を備えた AI ゲートウェイ。API ゲートウェイは職場で LLM を利用するための一般的な方法となっています。
OpenRouter:プロンプトを分析した後に最適なモデルを選択する「自動ルーター」機能を搭載しています。内部では Not Diamond を使用しています。詳細はこちら
Kilo Gateway:価格対性能比が最も優れていると判断されたモデルへリクエストをルーティングします。独自のモデルキーの使用をサポートし、サービスはルーティング専用としても利用可能です。詳細はこちら
Requestly.ai:コスト、レイテンシ、可用性、および多数の構成設定に基づき、自動的にリクエストを適切なモデルにルーティングします。詳細はこちら
LiteLLM:「自動ルーター」機能により、入力コンテンツに基づいて最適なモデルを自動的に選択するルーティングルールを定義できます。セットアップは手動で行う必要がありますが、多くの他の AI ゲートウェイと比較してより高い制御権を得られます。詳細はこちら
Envoy AI Gateway:一部のルーティング設定を提供するオープンソースのゲートウェイですが、ルーティングエンジンがコスト最適化やスマートなモデル選択よりも可用性に重点を置いているように感じられます。詳細はこちら
Cursor と GitHub Copilot もどちらも「Auto」モデル選択機能を持っており、自動的にモデルを選択します。Cursor の場合、これは固定価格のモデルであり、節約されたコストは Cursor 側に留まり、顧客には還元されませんが、他の多くのモデルよりも安価です。Copilot の Auto モードではインテリジェントなモデル選択が行われますが、このモードについて私が聞いた開発者からの肯定的なフィードバックはほとんどありませんでした。Pro プランの場合、Copilot はかなり古いモデルをサポートしています:GPT-5.5 や Opus 4.8 は利用できません。ただし、これらは Pro+ 以上のプランでは利用可能です。
インテリジェントルーティングに対する需要は極めて高いようです。私は Factory AI の共同創設者兼 CEO である Matan Grinberg に尋ねました。彼は次のように語りました:
「需要は計り知れないほど高く、特に企業(大企業)からの要望が顕著です。この提供を開始して以来、ほぼすべての銀行の CEO と面談してきました。彼らは支出をコントロールするレイヤーを求めつつ、高品質なコードを生成したいと考えているからです。
テック業界のほとんど全員が、オープンモデルは往々にして十分であることを認識し始めています。過去 6 ヶ月間、オープンモデルの使用は厳密に増加しています。私の推測では、トークン支出の観点から、ホストされたオープンモデルのパフォーマンスはコーディング関連業務の約 60% に十分であると考えられます。」
私には、「インテリジェントルーティング」が業界標準(テーブルステークス)となるように感じられ、そのためほぼすべての AI ベンダーが何らかのバージョンを構築し、多くの新規ベンダーがこの種の機能を提供するようになるものと予想されます。
もしここにリストされていないベンダーをご存知であれば、元の The Pulse アーティクルにコメントを追加して、そこでさらに多くの選択肢を確認してください。
この抜粋が掲載された The Pulse の完全な号をお読みいただくか、あるいはすべての The Pulse 号をチェックしてみてください。
原文を表示
Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from a previous The Pulse issue. Full subscribers received the article below three weeks ago. If you’ve been forwarded this email, you can subscribe here.
Two weeks ago, I covered a trend of companies trying to reduce spending on AI within their engineering departments. While talking to my sources about this, one head of engineering at a larger company told me that they wished there was an ‘intelligent’ router that picks the right model for the right task.
The reason for such a wish is clear; prices for tokens vary greatly per model, and there can easily be a 10-20x difference between a cheap, average model, and a state-of-the-art one.
I did some digging into whether any solutions like this currently exist because the benefits look obvious, and what I found is listed below. Usual disclaimer: I have no affiliation with these vendors, and have not been paid to mention any of them!
Vendors:
Factory Router: automatically selecting the right model per session, claiming 20-25% cost savings. More details.
Not Diamond: auto-selection of coding models, claiming around 30% cost savings. Used by OpenRouter, under the hood. More details.
Prism by Augment Code. Choosing the “best” model automatically for coding tasks. More details.
Model Router by Morph. An API to suggest model selection for a prompt, based on a list of models. More details
Weave router: a token router that works inside Codex, Claude Code and Cursor. “Hard” requests stay on frontier models, while “easy” ones go to open source ones. More details
AI gateways with routing built in. API gateways are popular ways to use LLMs in workplaces.
OpenRouter: comes with “auto router” functionality where, after analyzing the prompt, the best one is selected. Uses Not Diamond under the hood. More details
Kilo Gateway: route requests the model considered the best price-per-value. Supports using your own model keys, and using the service only as a router. More details
Requestly.ai: automatically route requests to the right model based on cost, latency, and availability, and tons of configuration. More details
LiteLLM: define routing rules that automatically select the best model, based on input content with the “auto routing” functionality. The setup is more manual, but you get more control than with many other AI gateways. More details
Envoy AI Gateway: an open source gateway that offers some routing configuration, though it feels that the routing engine focuses more on availability, not cost optimization and smart model routing. More details
Cursor and GitHub Copilot also have an “Auto” model selection that does automatic model selection. For Cursor, it’s a fixed-price model where any savings made are for Cursor: they are not passed on to customers, but the model is cheaper than most others. For Copilot, the Auto mode results in intelligent model selection – but I’ve not heard much positive feedback about this mode from the few devs I asked about it. For Pro plans, Copilot supports pretty old models: GPT-5.5 and Opus 4.8 are not available. These are, however, available on the Pro+ and above plans.
Demand seems to be extremely high for intelligent routing. I asked Matan Grinberg, cofounder and CEO at Factory AI, who told me:
“Demand has been off the charts, especially from the enterprise [from large companies.] I’ve met with practically every bank CEO since we launched this offering, because they want a layer to control spend, while still generating high-quality code.
Pretty much everyone in tech is starting to see that open models are often sufficient. We’re seeing open model usage strictly increasing the last six months. My guess is that hosted open models are sufficient in performance for around 60% of coding-related work, in terms of token spend.”
It feels to me that “intelligent routing” will become table stakes, and so we can expect pretty much all AI vendors to build some version of it, and many new vendors to offer this kind of functionality.
If you know of any additional vendors not listed, you can add a comment on the original The Pulse article, and see more options there.
Read the full issue The Pulse that this excerpt was from, or check out all The Pulse issues.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み