フロンティアモデル価格競争とオープンウェイト人気、ルーティング需要を牽引
本文の状態
日本語全文を表示中
詳細モードで約12分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Latent Space
Frontier モデルのコスト高騰とオープンウェイトモデルの台頭により、Stripe が OpenRouter を買収したように企業におけるモデルルーティングが AI デプロイの重要要素となっている。
AI深層分析を開く2026年8月19日 07:06
AI深層分析
キーポイント
モデルルーティングの重要性増大
Frontier モデルのコスト競争と Kimi K3 や Qwen3.8-Max などのオープンウェイトモデルの能力向上により、タスクに応じた最適なモデル選択が AI デプロイの鍵となっている。
Glean の市場での成功と戦略
元 Google エンジニアである Arvind Jain が率いる Glean は直近で年間収益 3 億ドルを達成し、LLM の必要性がないタスク(計算など)の回避や、複数の AI を統合した「メタ・ハーネス」としての役割を果たしている。
Glean のモデルルーティング機能
Glean は従業員による手動選択、管理者による制限、およびタスクに応じた動的な自動モードという 3 つのレベルでモデル選定を提供しており、顧客は経済的理由から自動モードを主に利用している。
コスト削減とモデルルーティングの必要性
最新の高機能AIモデルはトークンあたりの価格が倍増しており、複雑なタスクへの使用によりユーザーあたりのコストが大幅に上昇している。Gleanのようなルーティングシステムは、適切なモデルを割り当てることでClaude Codeと比較して4倍のコスト効率を実現し、企業規模での費用対効果を向上させている。
大規模な人間フィードバックループによる最適化
全社的な導入により得られる膨大なユーザー行動データから、タスクの種類やモデル選択の傾向、不満足時のアップグレードパターンを把握できる。このスケーラブルなフィードバックループが、より正確で効率的なモデルルーティングシステムの改善に貢献している。
重要な引用
"A big goal of Glean is to avoid using LLMs for tasks where we don't need them," Jain told Latent Space.
"You can think of Glean today as a superset of ChatGPT, Claude, Gemini, Grok... Glean combines the power of all of them into one experience."
"Ultimately our business is to deeply understand your data, knowledge, and information, but also how work happens inside your company."
"AI models have been getting expensive... the costs have gone up a lot."
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

最先端モデル開発企業間の熾烈な競争と、Kimi K3 や Qwen3.8-Max といったオープンウェイトモデルの性能向上に伴い、モデルルーティングは AI デプロイにおいて重要な要素となっています。Stripe が OpenRouter を 70 億ドル以上で買収したというニュースも記憶に新しいですが、このトレンドは企業側でも同様に活発です。
元 Google のディステングイッシュド・エンジニアである Arvind Jain が共同創設し率いる Glean は、大規模組織への AI 導入を専門としています。昨年 6 月の 1.5 億ドルのシリーズ F ラウンドを経て時価総額は 72 億ドルに達しましたが、今年に入ってからは年間経常収益(ARR)が 3 億ドルに到達し、15 ヶ月間で 3 倍の成長を遂げました。
Glean のミッションの一部は、各タスクにどのモデルを使用するか、あるいはそもそも LLM を使用する必要があるのかを判断することです。
「Glean の大きな目標の一つは、LLM が必要ないタスクで無理に使わないことです」と Jain は Latent Space に語っています。「Glean 内では、単純な足し算や掛け算を行うようなクエリが時折見られます。そんな場合は電卓を使えば十分です。」
しかし、Glean が主に目指しているのは、Jain 氏が「非常に強力なパーソナル・コワーカー」と呼ぶものを企業従業員に提供することです。つまり、主要な LLM を統括するメタ・ハーネスのような役割を果たすことが求められています。

Glean は昨年 9 月に第 3 世代の Glean Assistant を発表しましたが、現在ではエージェントが同社のシステムの大きな部分を占めています。
Jain 氏は「現在の Glean は、ChatGPT、Claude、Gemini、Grok のすべての機能を包含する超集合体と捉えてください」と語りました。「私たちが日常的に利用しているさまざまな AI プロダクトの力を、Glean は一つの体験として統合しています。」
企業にとって、AI 技術を組織内に導入することは課題の半分です。もう半分の課題は、組織内のナレッジを AI システムに取り込むことです。
「究極的に私たちのビジネスは、お客様のデータや知識、情報を深く理解することですが、同時に貴社内でどのように業務が行われているかを把握することも含まれます」と Jain 氏は述べています。
Glean におけるモデルルーティングの実践
では、実際の運用においてモデルルーティングとは何を意味するのでしょうか。基本的に Glean では、3 つのレベルでモデル選択が可能です。
- 従業員が明示的にモデルを選択できる
- 管理者が利用可能なモデルを制限したり、使用制限を設定したりできる
- Glean の自動モードは、タスクごとに動的に最適なモデルを選択する

特定のタスク向けにモデルを構成する。
実は、Glean の顧客が自動モードを主に選択する理由は経済的なものです。
「なぜ人々はモデルルーティングについて話し、それに対して興奮しているのでしょうか?その主な理由はコストです」と Jain は語りました。
Glean の共同創設者でエンジニアリードの Tony Gentilcore 氏は最近、「Glean は Claude Code よりもコストパフォーマンスが 4 倍高い。タスクあたりの平均費用は Glean が 0.45 ドルであるのに対し、Claude Cowork は 1.84 ドルだ」と主張しました。その要因として同氏は、Glean の「ハネスとルーティング機能」を挙げています。
個人レベルでは、LLM プロバイダーへの月額 20 ドル、100 ドル、あるいは 200 ドルのサブスクリプションから大きな価値を得ている人も少なくありません。しかし企業にとっては、ユーザーあたりのコストが簡単に制御不能な水準まで膨れ上がってしまいます。
「AI モデルは高騰しています」と Jain は指摘します。「Opus や最新の GPT など、最も高度なモデルを見てみましょう。これらは非常に強力であり、以前のモデルよりも複雑なタスクを処理できます。しかしトークンあたりのコストは高く、以前のモデルの 2 倍から 4 倍になることもあります。さらにユーザーは、これらのモデルを使ってより長いタスクを実行します。その結果、昨年に比べてユーザーあたりの支出が 10 倍、20 倍にも膨れ上がってしまうのです。つまり、コストは大幅に上昇しているのです。」
人間のフィードバックループ
Glean の台頭におけるもう一つの重要な要因は、同社が一般のビジネスユーザーが AI をどのように利用しているかを把握できる点にあります。このプロダクトは「コワーカー」として全従業員に展開される可能性があり、またすべての部署や機能においてエージェントを構築・展開するためにも使用されています。
顧客事例を見ると、Zillow では7,000人の従業員の8割が導入を完了し、Booking.com では「Glean が社内で初めて採用された AI プラットフォームとなりました」と報告されています。この浸透度の高さは、企業が実際にどのように AI を活用しているかを把握する上で、Glean にとって極めて有利な立場をもたらしています。
Jain はこう語ります。「広範な視点から、人々が実際には何をしているのかを観察できるようになりました。AI をどのようなタスクに用いているか、最初にどのモデルを選択するか、また満足できない場合に、より適切な結果をもたらす別のモデルへ切り替えるタイミングを把握できるのです」。
この大規模な人間によるフィードバックループが、モデルルーティングシステムの改善に寄与しています。
Here's Waldo, gathering raw materials
Glean のアーキテクチャのもう一つの柱が、Waldo という名前のモデルです。Jain によれば、これは大規模言語モデルの上に位置する存在で、今年4月に「Glean が初めて発表したエージェント型検索モデル」として発表されました。

Glean によると、このエージェント型検索モデルである Waldo は「レイテンシを50%削減し、トークン使用量を25%減らすことで、高度な処理が必要な作業にのみ最先端のモデルを割り当てることが可能になります」。
技術ブログ記事において、Waldo はユーザーの問い合わせに対するフィルタリングプロセスとして描かれています。具体的には、「質問をどのように分解するか、どのツールを使用すべきか、次に何を参照すべきか、そして高品質な回答を提供するためにフロンティアモデルに引き渡すのに十分な証拠が揃ったかどうか」を判断する役割です。
これはつまり、Glean 側で Jain が「原材料」と呼ぶタスクに必要な要素が特定された後に、モデルのルーティングが行われることを意味します。
「LLM のトークンを消費することなく、作業に必要な原材料を組み立てることができます」と彼は付け加えました。
この仕組みから導き出される帰結として、関連性の低いデータで埋め尽くされたフロンティアモデルよりも、文脈理解に優れコストの安いモデルの方が性能を発揮するケースがあることが挙げられます。
オープンウェイトモデルの急激な台頭
Jain 氏は現在、企業からのオープンウェイトモデルへの関心が著しく高まっていると確認しました。その主な理由はコスト懸念です。ただし、この動きが顕在化したのはここ数ヶ月のことです。
「昨年までは、(オープンソース LLM の)利用はごくわずかで、誰も真剣にオープンソースを検討していませんでした」と彼は語りました。背景には、これらのオープンソースモデルの多くが米国以外で開発されていることへの「偏見」があったことも一因です。
しかし突如として、企業顧客の間での関心が高まりました。

Jain 氏が 2026 年 7 月 27 日に投稿した、オープンウェイトモデルを支持するツイート。
「ここ 3 ヶ月で AI の利用コストが高騰したため、企業は AI 投資を維持することが困難になりつつあります」と Jain は語ります。「タスク実行においてオープンソースの方が桁違いに安価であることが判明し、大きな関心を集めています。現在では、多くの企業がオープンソースモデルを AI 戦略の核として位置づけることを検討しています。」
さらに、組織はもはや特定の 1〜2 のプロバイダーに依存する傾向にありません。この背景には、オープンウェイトモデルの台頭という潮流があります。
「もはや、単一のモデルプロバイダーや、せいぜい 2 つ程度に頼る気はありませんし、オープンソースなしでは生き残れないと考える企業が増えています」と Jain は指摘します。
Evals(評価)
2026 年の AI について真剣に議論するには、LLM の出力品質を評価する「Evals」の話題を外せません。Glean がどのように評価を実施し、その結果をモデルルーティングシステムにフィードバックしているのかを尋ねてみました。
Jain 氏によると、同社には「内部テストシステム」があり、異なるクエリクラスにおける実世界のワークロードと代替オプションを比較検証しています。具体的には、モデル自身に経路選択を行わせると同時に、「より安価なモデル」や「やや高価なモデル」など他のモデルで同じタスクを実行し、並列して結果を比較します。

Glean が品質を監視する方法。
Glean はその後、「AI による評価者」を活用し、モデルルーターがどれほど正確に機能したかを判定します。
「つまり、新しい実世界のトラフィックデータに基づいて継続的に学習が更新される仕組みです」と Jain は説明しました。基本的には、ユーザーの代わりにモデルルーターが処理を行う一方で、裏側では同じタスクを並行して実行し、その結果をフィードバックとして活用しています。
これは実世界の利用のうちごく一部(ごくわずかな割合)に対して行われるものですが、Glean のような大規模なスケールであれば、モデルルーターの訓練と改善には十分すぎるほどのデータ量となります。
企業向け検索からエンドツーエンド AI プラットフォームへ
今後、Latent Space で注目していくトレンドの一つは、AI システムがどのように企業内で導入されているか、そして一部の組織がいかにして「AI ネイティブ」へと完全移行しているかという点です。
これらの動向を監視する上で、Glean は特に興味深い対象と言えます。同社は企業向け AI 分野において最も初期のプレイヤーの一つであり、2019 年初頭に設立されました。当初は企業向け検索ソリューションの開発に注力していました。Jain 氏によれば、Glean は「ビジネス向けにトランスフォーマーや言語モデルを活用した最初の企業」です。

Glean の「AI Answers」は、組織内のドキュメントを直接参照して回答を生成します。
2023 年 4 月、swyx が Glean の創設エンジニアだった Deedy Das にインタビューを行いました。現在ではベンチャーキャピタル企業 Menlo Ventures のパートナーとなった Das ですが、当時(Glean 設立から約 4 年後の 2023 年)でも、同社の焦点はまだ主にエンタープライズ検索にありました。
しかし 2026 年 now、企業の AI 活用は検索にとどまりません。AI は社員のワークフローに不可欠な要素へと進化しています。
これが Glean を「より魅力的な」AI 企業へと変えました。Das 自身が昨年 11 月の『Latent Space』ポッドキャストへの復帰時にこう語っています。「私が Glean で最も気に入っている点は、かつては地味で魅力のない会社だったのに、後になってその魅力が際立つようになったことです」と。
この流れは、モデルルーティングの議論へと自然につながります。議論の最後を飾った Arvind Jain は、Glean を「エンドツーエンドの AI プラットフォーム」と呼び、顧客企業において「非常に重宝されている」と強調しました。これにより Glean は、「効果的なモデルルーティングを行うために必要なデータ」を保有できる立場にあるのです。
原文を表示

With the intense competition among frontier model companies, together with ever-increasing power of open-weight models like Kimi K3 and Qwen3.8-Max, model routing has become a key part of AI deployment. We’ve just seen Stripe buy OpenRouter for over $7B, but the trend is equally hot in enterprises.
Glean, co-founded and led by ex-Google Distinguished Engineer Arvind Jain, specializes in bringing AI to large organizations. It was last valued at $7.2B after a $150M Series F fund raise last June. This year, it reached $300 million in annual recurring revenue (ARR) — a three-fold increase over 15 months.
Part of Glean’s mission is to select which model to use for each task — or indeed if an LLM is even required.
“A big goal of Glean is to avoid using LLMs for tasks where we don’t need them,” Jain told Latent Space. “Sometimes you’ll see queries in Glean where people are adding two numbers or multiplying two numbers. They could have used a calculator to do that.”
But what Glean is mostly trying to do is bring what Jain calls “one really powerful personal co-worker” to enterprise employees. And that means being a kind of meta-harness for leading LLMs.

Glean announced its third-generation Glean Assistant last September; these days, agents are a big part of Glean’s system.
“You can think of Glean today as a superset of ChatGPT, Claude, Gemini, Grok,” Jain said. “All these different AI products that we’ve been using day to day, Glean combines the power of all of them into one experience.”
With enterprises, bringing AI technology into an organization is just half the challenge. The other half is bringing organizational knowledge into the AI systems.
“Ultimately our business is to deeply understand your data, knowledge, and information, but also how work happens inside your company,” Jain said.
How model routing is done in Glean
So what does model routing mean in practice? Basically, Glean offers three levels of model selection:
Employees can explicitly choose a model.
Administrators can restrict models or impose usage limits.
Glean’s automatic mode selects a model dynamically for each task.

Configuring models for certain tasks.
It turns out automatic mode is mostly chosen by Glean’s customers for economic reasons.
“Why are people talking about model routing? Why are they excited about it? It’s mostly because of cost,” Jain told us.
Another co-founder of Glean, engineering lead Tony Gentilcore, recently claimed that Glean “is 4x more cost-effective” than Claude Code, “averaging $0.45 per task versus $1.84 for Claude Cowork.” He put that down to Glean’s “harness and routing capabilities.”
Individually, many of us are getting great value out of our $20, $100 or $200 monthly subscription to an LLM provider. But for an enterprise, the per-user costs can easily spiral out of control.
“AI models have been getting expensive,” Jain said. “Like, if you look at Opus or the latest models of GPT, the most advanced models. Not only are they very powerful, they can run much more complex tasks than the previous models. But on a per token basis, they’re more expensive — sometimes double or quadruple the rates of the previous models. And then users actually use them to run much longer tasks. So you’re spending, like, 10 times, 20 times, more, on a per user basis, than what you were doing last year. So the costs have gone up a lot.”
The human feedback loop
Another key factor in Glean’s rise is that it gets to see how ordinary business users are using AI. The product is potentially deployed to every employee as a “coworker,” and it’s also used to build and deploy agents across all departments and functions.
Among its customers, Zillow reports 80% adoption across 7,000 employees, while at Booking.com, “Glean became the first AI platform adopted company-wide.” That kind of penetration gives Glean an enviable view into how AI is being used in enterprises.
“So we are getting to observe what people are actually doing with AI on a very broad basis,” said Jain. “We are getting to see when they’re on different types of tasks with AI, what models do they select first, and when they are not satisfied, when they actually upgrade to some other model [that] actually gives them the right results.”
This human feedback loop, at scale, helps improve the model routing system.
Here’s Waldo, gathering raw materials
Another part of Glean’s architecture is a model called Waldo, which Jain described as sitting on top of the large language models. Waldo was introduced in April as “Glean’s first agentic search model.”

Glean claims that Waldo, its agentic search model, “reduces latency by 50% and tokens by 25%, reserving advanced models for work that needs them.”
In a technical blog post, Waldo was portrayed as a kind of filtering process for user queries: it “decides how to break down the question, which tools to use, what to read next, and when it has enough evidence to hand off to a frontier model for a high-quality answer.”
This means the model routing is happening after Glean has determined what Jain calls the “raw materials” that are needed for the task.
“We’re able to assemble the raw materials needed to do the work without burning LLM tokens,” he added.
A corollary of this is that a cheaper model with better context may outperform a frontier model loaded with irrelevant data.
The rapid rise of open-weight models
Jain confirmed there is now significant interest from enterprises in open-weight models, primarily due to cost concerns. But this has only happened over the past few months.
“Last year, the usage [of open source LLMs] was minuscule and nobody was really seriously considering open source,” he said. Partly that was because of the “stigma” of many of these open source models being developed outside the US.
But suddenly, interest among enterprise customers has risen.

Jain’s tweet on July 27, 2026, in support of open-weight models.
“So in the last three months, because AI got so expensive, businesses have started to find it untenable to maintain these AI investments,” Jain said. “Given that open source is an order of magnitude cheaper to do tasks, it has created a lot of interest. Today, I can say that in most enterprises, they are considering open source models to be a key part of their AI strategy.”
More than that, organizations tend not to rely on just one or two providers anymore — and the rise of open-weight models is driving this trend.
“Nobody is willing anymore to rely on only one model provider, or two, and nobody thinks that they can survive without open source,” Jain said.
Evals
You can’t have a serious conversation about AI in 2026 without discussing evals — assessing the quality of results from LLMs. I asked how Glean goes about doing evals and how that is fed back into the model routing system.
Jain said they have “internal testing systems” where they compare real-world workloads, across different query classes, with alternative options. So they let the model choose a route and in parallel they try to complete the same task with “some other models which are maybe a little bit less expensive and a little bit more expensive.”

How Glean monitors quality.
Glean then uses “AI-based judges” to determine “how spot-on the model router was.”
“So there’s this continuous learning that gets updated with new real-world traffic, where basically what is happening is that you let the model router do the work for the user, but behind the scenes you run the same task,” Jain explained.
He added that this is done for only “a small fraction” of the real-world usage, but at Glean’s scale that’s more than enough to help train and improve the model router.
From enterprise search to end-to-end AI platform
One of the trends we’ll be monitoring going forward on Latent Space is how AI systems are being implemented within enterprises — and how some of these organizations are going full-on AI-native.
Glean is an especially interesting company to monitor for these trends, since it was one of the very first enterprise-facing AI companies. It was founded in early 2019, initially to tackle enterprise search. As Jain put it, Glean was “the first player to work with transformers and language models for businesses.”

Glean’s AI Answers draws “directly from your organization’s documentation.”
In April 2023, swyx interviewed Deedy Das of Glean. Das, who is now a partner at venture firm Menlo Ventures, was a founding engineer at Glean. But even at that point, in 2023 — about four years into Glean — the focus was still mostly on enterprise search.
Now, in 2026, enterprises aren’t just using AI for search. AI is becoming an integral part of every employee’s workflow.
That makes Glean a much ‘sexier’ AI company, as Das himself said on his return to the Latent Space podcast last November. “Broadly, one of the things that I love about Glean is it’s such a boring unsexy company that became sexy later,” he said.
This brings us full circle back to model routing. Arvind Jain ended our discussion by calling Glean an “end-to-end AI platform” that gets “used very heavily” by its enterprise customers. This, he added, allows Glean to “have that data that is required to do effective model routing.”
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み