Mozilla、AI の次はモデルよりインフラが重要と主張
本文の状態
日本語全文を表示中
詳細モードで約14分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Mozilla AI
Mozilla は、AI モデルの性能向上自体よりも、組織レベルでの管理インフラ不足がボトルネックとなっている現状を指摘し、生産環境への移行に伴うコスト・ガバナンス・所有権課題に対応する新プラットフォーム「Otari」の開発を発表した。
AI深層分析を開く2026年8月4日 12:09
AI深層分析
キーポイント
AI導入のフェーズ転換と課題
2022年から2026年にかけて、企業のAI活用は実験段階から生産環境への移行期へと急速に進み、現在はコスト管理やデータ所有権といったインフラ上の深刻な課題が顕在化している。
モデル性能よりインフラ管理の重要性
多くのタスクで十分な性能を持つモデルが増えている一方で、組織規模での運用に必要な信頼性、監査可能性、コスト制御を実現する管理基盤が不足していることが最大のボトルネックとなっている。
新プラットフォーム「Otari」の登場
Mozilla は既存のモデル性能の向上を否定するものではなく、むしろその多様性を管理するためのインフラとして、「Otari」という新たなプラットフォームの開発に着手したと発表した。
採用速度の加速と断片化のリスク
AI の生産環境への導入が少数の大手企業からあらゆるセクターの数千チームへと拡大する中で、複数のモデルを併用することによるシステム断片化という予期せぬ問題が発生している。
モデルの断片化による運用混乱
チームは複数のプロバイダーやバージョンが混在する環境で、API や価格、レイテンシの違いにより柔軟性がむしろ運用上の混沌を生んでいる。
重要な引用
The Model's the Easy Part - How to Get, and Keep, Value
Not because models aren't good enough... But, the infrastructure to manage them at an organizational level doesn't exist yet.
The real change is adoption velocity.
What looked like flexibility quickly became operational chaos.
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
imageモデルは簡単だ。価値をどう獲得し、維持するか
ここ数年の企業における AI の進化について、私の見解をお伝えします。
2022 年秋、世界は景気後入に突入すると予測されていました。IT バジェットは 2023 年度に向けて凍結状態でした。
2022 年 11 月 30 日、ChatGPT が登場します。これにより、経営陣の非技術系メンバーが AI と対話できる具体的な窓口を得ました。それは単なるオンライン上のチャットインターフェースです。
CEO、CFO、CRO などが休暇で帰った際、初期段階の生成 AI に驚かされます。テイラー・スウィフトをエミネム風にラップさせたり、メールを要約したり、旅行計画について会話したりするのです。感銘を受けた彼らは、2023 年に生成 AI 向けの予算を解禁し、社内での実験プロジェクトに割り当てました。
2023 年:実験と試行錯誤の時代です。新規予算は生成 AI 向けのみだったため、IT チームの注目はすべてそこに集中しました。
2024 年:有望な成果が出なかった生成 AI プロジェクトの約 90% が淘汰され、残りの 10% が GRC(ガバナンス・リスク管理・コンプライアンス)プロセスを経て本番導入へと進み始めました。
2025 年:アプリケーションが本番環境へ移行します。その際、ガードレールの強度やコスト管理、ROI(投資対効果)の測定レベルにはばらつきがあります。(また、2025 年後半にエージェント型コーディングが現実味を帯び始め、製品展開の速度が向上しました。)
2026 年、生成 AI の本番環境での利用が社内・社外ともに爆発的に増加します。年間予算は数ヶ月で使い切られてしまうでしょう。主権 AI(ソブリン AI)の議論や、最先端ラボへの政府関与の高まりに伴い、システムとデータの所有権に関する懸念も高まっています。

端的に言えば、議論の焦点は「AI を実験すべきか?」から「なぜまだ本番環境で使えていないのか?」へ、そして「ROI(投資対効果)はどうだ?一体いくら掛かったんだ!私のデータはどこへ行ったんだ!」へと移り変わりました。
全く異なる議論になっていますね。
実験なら安価です。API キーを取得して様子を見て、次に進むだけです。しかし本番運用は別問題です。スケールした信頼性、監査可能性、コスト管理が求められます。また、地理的な制約の厳守、オンプレミスでの計算リソース利用、自社のモデルを所有する権利など、厳しい要件が発生することもあります。
だからこそ私たちは Otari を構築しました。モデル自体が不十分だからではありません。むしろ逆で、多くのタスクにおいて十分な性能を持つモデルはすでに存在します。問題は、組織レベルでそれらを管理するためのインフラストラクチャがまだ存在しないことです。その基盤を私たちが作っています。
過去 2 年で何が変わったのか
議論を呼ぶ見解:AI の次の時代は、モデルの性能そのものよりもインフラストラクチャが鍵となる。
確かにモデルの能力は劇的に向上しました。特にオープンソースやオープンウェイトモデルの進化は目覚ましいものです。しかし、真の変化は採用の速度にあります。AI はもともと限られたリソースを持つ一部のテック企業だけのものでしたが、今ではあらゆる業界に数千ものチームが導入しています。その普及に伴い、誰も十分に予測していなかった問題が浮き彫りになりました。
まず挙げられるのは「分断化」です。多くのチームは単一のモデルしか使っていません。むしろ、数十ものモデルを使いこなしているのが実情です。高レベルで見れば、テキスト要約には GPT リリース版を、コーディング支援には Claude を使い、機密性の高い処理にはオンプレミス型のオープンウェイトモデルを採用するといった具合です。たとえ「OpenAI 一筋」の企業だとしても、チームによっては GPT-5.5、GPT-5.4、GPT-5.4-mini、GPT-5.4-nano、そしてデプロイ時に良好な動作を示したまま更新されていないレガシーモデルなどを使い分けています。それぞれに独自の API、価格体系、レイテンシ特性、レート制限が存在します。当初は柔軟性として歓迎されていたものが、いつしか運用上の混沌へと変貌してしまったのです。
次に「コストの不可視化」です。AI の推論コストは非線形的に増大します。テスト環境では月額 200 ドルで済んでいた機能も、利用状況が変化すれば本番環境では月額 2 万ドルに跳ね上がることもあります。この問題は、近い将来のフロンティア・ラボによる IPO を機にトークンに対するベンチャーキャピタル(VC)からの補助金が撤廃され、トークンの真のコストがより明確になるにつれて、いっそう重要度を増しています。多くのチームは請求書を受け取るまで、自分が負担すべきコストの実態に気づきません。また、問題が悪化する前にその兆候を可視化できるプロバイダー横断型のネイティブツールも存在しません。
3 つ目はガバナンスの欠落です。AI が金融、医療、法務、教育といった規制産業へ進出するにつれ、「どのモデルが、いつ、誰に、何を言ったのか」という追跡可能性がコンプライアンス上の必須要件となっています。さらに世界各地で議論されている「主権 AI」の動向は、こうした要件をさらに複雑化させています。現在のインフラでは、この課題に対応できる仕組みがありません。
複数のプロバイダー管理における課題
実際には、マルチプロバイダー環境がどのような複雑さを生むのかを見てみましょう。製品チームは 3〜4 の異なるモデルプロバイダーへルーティングを行っており、各プロバイダー内にも複数のモデルが存在します。さらに、その場で用意されたローカルソリューションへの接続も行われています。障害発生時のフォールオーバーロジックも独自に実装済みです。コスト管理にはスプレッドシートが使われており、エンジニアたちは直感と不完全なデータに基づいて手動でどのリクエストをどのモデルが処理するかを調整しています。
これは持続可能なアーキテクチャではありません。
問題なのはチームが間違ったことをしているからではなく、ツールの進化が遅れているのです。クラウドコンピューティングが成熟した際、組織はサーバーの手動管理から脱却し、複雑さを抽象化するプラットフォームを採用しました。AI においても同様の転換点に達しています。モデル自体が計算資源(compute)であり、その上に位置する制御層こそが欠落している部分です。
コストの可視化は最優先課題
コストは戦略的な課題として過小評価されがちですが、2026 年には状況が変わりつつあります。組織が「トークン最大化(tokenmaxxing)」と呼ばれる行為がいかに不確かなリターンに対して資本を燃やしているかを理解し始めたからです。
しかしながら、AI インフラストラクチャを経費センターとして捉えている多くの組織は、その考え方を根本的に誤っています。真に問うべきは「いくら使っているか」ではなく、「必要な成果を可能な限り最低のコストで得られているか、そして私たちはそれを把握できているか」という点です。これらは全く異なる問いかけです。
現在、ほとんどのチームはこの二つ目の問いに答えられません。プロバイダー間で成果あたりのコストを比較することも、価値に見合わない予算の浪費が生じている経路をリアルタイムで可視化することもできません。また、「このユースケースでは X 以上は使わない(ただし Y が発生した場合は例外)」といったポリシーを設定し、それを自動的に執行する仕組みも持てていません。
財務チームは、こうした可視性を最終的に要求します——実際にはすでに要求を始めています。これに先手を打つエンジニアリングチームは、対応が遅れたチームに対して構造的な優位性を得ることになるでしょう。
なぜ制御がますます重要になるのか
私が繰り返し述べている主張、そして私たちが構築しようとしているものの核心はこうです。制御こそが、新たな堀(モート)となるのです。
ここ数年、各チームは「どのモデルを使うか」で競い合ってきました。しかしその優位性は失われつつあります。モデルがコモディティ化し、最上位モデル間の性能差も縮小しています。さらに、ラズパイで動作するレベルに至るまで、あらゆるパフォーマンス段階に複数の選択肢が存在します。
これからの競争の焦点は、運用層に移っています。重要なのは、大規模かつ信頼性高く、コスト効率よく、安全に AI を展開できるかという点です。
これはインフラストラクチャの問題です。
ここで言う「コントロール」とは何を意味するのでしょうか?それは、コスト、能力、レイテンシ、コンプライアンスに応じてリクエストを賢くルーティングすることを指します。また、AI が何を行い、なぜそうしているのかをリアルタイムで可視化できることも含まれます。組織レベルの方針を設定し、各チームが毎回ゼロから仕組みを作り直すことなく、一貫してそれを適用できるようにすることも意味します。さらに、アプリケーション層を書き換えることなくプロバイダーを切り替えられるようにすることです。
コントロールとは、AI を成熟したエンジニアリング分野として運用することを意味します。
なぜ Mozilla は Otari の構築を決めたのか
Mozilla は常に、特定のインターネットのあり方を信じてきました。オープンで分散型、公共の利益のために統治され、市民によって運営されるインターネットです。AI インフラストラクチャが向かっている方向を見たとき、私たちはよく知っているパターンを目にしました。
少数のプロバイダーによる権力の集中。そして、検査や修正、制御ができない不透明なシステムに依存する多くの組織。
この物語は過去にも何度も繰り返されてきました。誰かが介入しなければ、結末は目に見えています。
Otari は、その課題に対する私たちの答えです。LLM 向けのコントロールプレーンであり、コアはオープンソースで設計されています。組織が自社の AI インフラに対して真の自律権を持つことを可能にするためのものです。単なるルーターでも、コストダッシュボードでもありません。アプリケーションとモデルの間に位置する完全なコントロールレイヤーとして、AI を自らの条件で運用するための可視性、ガバナンス、柔軟性を提供します。
私たちは意図的にオープンソースで構築しました。医療、教育、防衛、金融、そして市民技術といった分野で、この必要性が最も高い組織にとって、ベンダーロックインは許容できません。確かに、オープンソースは一つの流通チャネルであり、私たちのコアバリューにも合致していますが、それ以上の意味を持っています。世界経済を動かす最重要産業や機関、組織における採用には、不可欠な要件なのです。
Otari の機会:新カテゴリの定義
エージェント時代は到来するのではなく、すでに始まっています。ソフトウェアにおける次のアーキテクチャシフトは、新たなモデルのリリースではありません。それは「エージェントハネス」です。数十、あるいは数百もの AI エージェントを並列で調整するシステムであり、それぞれがモデル呼び出しを行い、コストを発生させ、管理が必要なデータにアクセスします。これをスケールして管理する複雑さは、現在のインフラが処理できる範囲を桁違いに超えています。これが解決すべき課題ですが、まだ適切な答えは見つかっていません。
このギャップこそが、新たなカテゴリです。単なる機能やニッチ市場ではありません。真剣な AI 導入を行うすべての組織に必要となる基盤層なのです。
今まさに制御のための計測(インストゥルメンテーション)に取り組む組織は、その運用上の優位性を複利のように積み上げていきます。一方、後回しにする組織は、もともとそれを想定して設計されていないシステムに、後付けでガバナンスを適用せざるを得なくなります。
私たちは Otari を、まさにそのような制御プレーンとして構築しています。コミュニティがこのカテゴリのあり方を形作れるようオープンソース化しており、最も必要としている組織が実際に導入できることを目指しています。医療機関、公共機関、金融機関、そして社会インフラ——これらはブラックボックス化するベンダーに依存することはできません。
ServiceNow の Amit Zavery 氏はこれを率直に指摘しました。「AI やエージェントの導入を検討する際、すべての顧客が懸念するのは制御です」と。Dell Inc. のマイケル・デルは、インフラストラクチャの重要性を強調しました。クラウドが提供したのは弾力的なスケーラビリティですが、「コスト予測可能な大規模なエンタープライズデータ上でのエージェント AI」については、約束も、おそらく実現もできないと述べています。
さらに驚くべきことに、Palantir の Alex Karp 氏さえも、選択肢とオープンソース AI の重要性について言及しています。「技術的な顧客が求めているのは、計算資源、モデル、データスタック、そしてアルファ(利益)に対する制御権です。彼らは生産手段を自らが所有していることを確認したいのです」と。
最も重要な機関や大企業は、自らのスタックを所有する必要があります。Otari は、その必要性に向けた一歩なのです。
AI の未来は、どのモデルが勝つかという話だけではありません。重要なのは、そのモデルの上層を誰が制御するかです。私たちは、この制御権は適切に運用されれば、すべての人に帰属すべきだと考えています。
その証明として、Otari をオープンソースで構築しています。Otari.ai のホストインスタンスをご利用いただくか、完全なオープンソースの GitHub リポジトリからセルフホスト型のゲートウェイをセットアップしてください。推論も AI スタックも、すべてがあなたの手にあります。
原文を表示
imageThe Model’s the Easy Part - How to Get, and Keep, Value
Here’s how I see the evolution of AI in enterprises over the last few years:
Autumn of 2022, the world thinks it’s going into a recession. IT budgets are frozen for 2023.
November 30, 2022: ChatGPT launches, and the non-technical parts of the C-Suite have a tangible interaction point with AI - a simple chat interface available online.
CEOs, CFOs, CROs go home for the holidays and are wowed by early-stage GenAI - making Taylor Swift rap like Eminem, summarizing emails, chatting about travel plans. Impressed, they unlock IT budget only for GenAI pet projects in 2023.
2023: pet projects, experimentation. The only new budget was in GenAI, so that’s where all IT teams focused.
2024: the great culling of 90% of GenAI pet projects not being promising, and the 10% that were starting to go through GRC for deployment.
2025: applications go into production, with varying levels of guardrailing, cost tracking, and ROI measurement. (Also, agentic coding becomes real in late 2025 - so product deployment velocity increases.)
2026: internal and external usage of GenAI in production explodes. Annual budgets are blown away in months or less. Concerns around system and data ownership increase with sovereign AI discussions and increased government involvement with frontier labs.
imagePut simply, we’ve moved from "should we experiment with AI?" to "why isn't this in production yet?" to “what’s the ROI, and my lord how much did that cost?!, and where did my data go?!.”
Very different discussions!
Experimenting is cheap: spin up an API key, see what happens, move on. Production is different. You need reliability, auditability, cost control at scale. You may need hard constraints over geography, on-premises compute, the ability to own your own models. That's why we built Otari. Not because models aren't good enough - in fact, the opposite, so many models are good enough for so many tasks. But, the infrastructure to manage them at an organizational level doesn't exist yet. We're building it.
What Changed in the Last Two Years
A contentious take: the most important shift hasn't been model capability. Models have improved dramatically, sure. Open source and open weight models especially. But the real change is adoption velocity. AI in production has gone from a handful of well-resourced tech companies to thousands of teams across every sector. With that came problems nobody fully anticipated.
First: fragmentation. Most teams aren't using one model - they’re using dozens. At a high-level, that might be a GPT release for text summarization, Claude for coding, an on-prem open-weight model for something sensitive. But even if they’re an “OpenAI shop”, they’ll still have teams using GPT-5.5, -5.4, -5.4-mini, -5.4-nano, and legacy models that worked well at deployment and haven’t been touched since. Each with its own API, pricing, latency profile, rate limits. What looked like flexibility quickly became operational chaos.
Second: cost opacity. AI inference scales non-linearly. A feature that costs $200/month in testing can cost $20,000/month in production if usage shifts. This is only getting more important as the “VC subsidies” on tokens lift with the upcoming frontier lab IPOs, and the true cost of a token becomes less opaque. Most teams don't find out what they’re on the hook for until they get the invoice. There's no native tooling across providers to surface this before it's too late.
Third: governance gaps. As AI moves into regulated industries - finance, healthcare, legal, education - "which model said what, when, to whom, and why" becomes a compliance requirement. And the sovereign AI discussions happening worldwide add complexity to these requirements. Current infrastructure has no answer for this.
The Challenge of Managing Multiple Providers
Here's what multi-provider complexity actually looks like in practice. A product team is routing to three or four different model providers, with multiple models per provider - and they’re routing to ad hoc local solutions. They've built custom failover logic for outages. They've got spreadsheets tracking costs. Engineers are manually tuning which model handles which request type based on gut feel and incomplete data.
This isn't a sustainable architecture.
The problem isn't that teams are doing something wrong. The tooling just hasn't caught up. When cloud computing matured, organizations stopped managing servers manually and adopted platforms that abstracted the complexity away. We're at the same inflection point with AI. Models are the compute. The control layer above them is what's missing.
Cost Visibility Is a First-Class Problem
Cost is underappreciated as a strategic issue, although that’s changing in 2026 as organizations start to realize how much “tokenmaxxing” is burning capital for questionable return. That said, most organizations treating AI infrastructure as a cost center are thinking about it wrong. The real question isn't "how much are we spending?" It's "are we getting the outcome we need at the lowest cost possible, and do we even know?" Those are very different questions.
Right now, most teams can't answer the second one. They can't compare cost-per-outcome across providers. They can't see in real time which routes are burning budget without proportionate value. They can't set policy ("never spend more than X on this use case, unless Y happens") and have it enforced automatically.
Finance teams will eventually demand - frankly, are now demanding - this kind of visibility. Engineering teams that get ahead of it will have a structural advantage over those that don't.
Why Control Becomes Increasingly Important
Here's the thesis I keep coming back to, and what we’re building for: control is the new moat.
For the past few years, teams competed on which model they used. That advantage is eroding. Models are commoditizing. The marginal difference between top-tier models is shrinking, and there are multiple competitive options at every level of model performance down to what can be run on a Raspberry Pi. The next competitive layer is operational: who can deploy AI reliably, cost-effectively, and safely at scale?
That's a question of infrastructure.
What does "control" actually mean here? It means routing requests intelligently by cost, capability, latency, and compliance. It means real-time observability into what your AI is doing and why. It means setting policies at the org level and having them enforced consistently, without asking every team to reinvent the wheel. It means swapping providers without rewriting your application layer.
Control means operating AI like a mature engineering discipline.
Why Mozilla Decided to Build Otari
Mozilla has always believed in a particular kind of internet: open, decentralized, governed in the public interest, and governed by the public. When we looked at where AI infrastructure was heading, we saw a familiar pattern. Consolidation of power in a small number of providers. Most organizations dependent on opaque systems they couldn't inspect, modify, or control.
We've seen this story before. We know how it ends if nobody intervenes.
Otari is our answer to that. It's a control plane for LLMs, open-source at its core, designed to give organizations genuine agency over their AI infrastructure. Not just a router. Not just a cost dashboard. A full control layer that sits between your applications and your models, giving you the visibility, governance, and flexibility to operate AI on your own terms.
We built it open source deliberately. The organizations that need this most, in healthcare, education, defense, finance, and civic tech, can't afford the lock-in. Yes, open source is a distribution channel - and one that aligns with our core values - but we see it as more than that: an absolute requirement for adoption in the most important industries, agencies, and organizations that drive the world economy.
Otari’s Opportunity: Defining a New Category
The agentic era isn't coming, it's here. The next architectural shift in software isn't another model release. It's agent harnesses: systems that coordinate dozens or hundreds of AI agents in parallel, each making model calls, each generating cost, each touching data that needs to be governed. The complexity of managing that at scale is orders of magnitude beyond what current infrastructure handles. This is the problem that needs to be solved, and it doesn't have a good answer yet.
That gap is the category. Not a feature, not a niche. A foundational layer that every serious AI deployment needs. The organizations that instrument for control now will have an operational advantage that compounds. The ones that wait will be retrofitting governance onto systems that were never built for it.
We're building Otari to be that control plane, open-source so the community can shape what this category becomes, and so the organizations that need it most can actually adopt it. Healthcare systems, public agencies, financial institutions, civic infrastructure: they can't depend on black-box vendors. ServiceNow's Amit Zavery said it directly: "Every customer, when they're thinking of AI adoption and agentic, they're worried about control." Michael Dell of Dell Inc. made the infrastructure case - what cloud delivered was elastic scale: "what it didn't promise, and cannot perhaps deliver, is cost-predictable agentic AI at scale on sensitive enterprise data." And heck — even Palantir's Alex Karp is weighing in on choice and open source AI: "What the technical customers want is control over their compute, their models, their data stack and their alpha. They want to know they own the means of production." The most important institutions and enterprises need to own their stack, and Otari is a step toward that necessity.
The future of AI isn't just about which model wins. It's about who controls the layer above the models. We think that control, done right, should belong to everyone. We're building Otari in the open to prove it. Pop on over to our hosted instances at Otari.ai or set up your self-hosted gateway from our fully open-source GitHub repository - own your inference, own your AI stack.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み