オープンウェイトモデルが AI ゲートウェイ利用の 29% に到達
TLDR AI の分析によると、AI Gateway におけるオープンウェイトモデルの利用シェアが 29% に達し、業界の多様化とコスト最適化の傾向を明確に示している。
キーポイント
オープンウェイトモデルの市場浸透率
AI Gateway の利用状況において、オープンウェイトモデルのシェアが 29% に到達し、プロプライエタリモデルと拮抗する地位を確立した。
インフラストラクチャの多様化
企業や開発者が单一の大手プロバイダーに依存せず、複数のオープンソースモデルをゲートウェイ経由で柔軟に利用する動きが加速している。
コストと制御の最適化
この傾向は、ランニングコストの削減やデータプライバシーの確保といった実務的なニーズに応えるための戦略的選択であることを示唆している。
重要な引用
Open-Weight Models Reached 29% of AI Gateway Usage
AI Gateway の利用状況において、オープンウェイトモデルのシェアが 29% に達したことが示された
影響分析・編集コメントを表示
影響分析
このデータは、AI エコシステムにおけるパワーバランスの変化を象徴しており、大手プロプライエタリモデルへの依存度低下と、カスタマイズ性やコスト効率を重視するオープンウェイトモデルの台頭を裏付けています。企業にとっては、ベンダーロックインからの脱却と、より柔軟な AI 戦略の構築が可能になる重要な指標となります。
編集コメント
29% という数値は、オープンソースモデルが単なる実験段階を超え、本格的な業務利用の基盤として定着したことを示す画期的な指標です。特に AI Gateway を介した利用という文脈は、企業レベルでの実装が進んでいることを意味しており、今後のアーキテクチャ設計において無視できないトレンドと言えます。
AI Gateway 生産インデックス — 2026 年 7 月
毎月、AI Gateway は、本番環境のアプリケーションと AI ラボの間で数十兆トークンのやり取りを仲介しています。これにより、現在の企業が実際にどのように AI を活用しているのかを把握することが可能になります。その実態をここでは公開します。5 月と 6 月に発表した生産インデックスレポートは、それぞれ こちら、こちら でご覧いただけます。
リンクをコピーする見出し 2026 年 7 月サマリー
7 月のインデックスは、2026 年 6 月に収集された AI Gateway のデータに基づいています。
- AI ゲートウェイのトークン利用量は前月比 29% 増、支出も 27% 増加しました。一方、1 トークンあたりの単価は 5 月に約 20% 上昇した後、この月は横ばいとなりました。
- オープンウェイトモデルが占めるゲートウェイのトークン利用量は 29% に達し、4 月の 11% から大幅に伸びました。支出は全体の 4% を下回る水準です。DeepSeek は利用量で 22.6% を記録し、3 位につけました。Google に僅か 2 ポイント及ばず、GLM 5.2 もリリースから 2 週間足らずでゲートウェイの主要モデル(利用量ベース)に名を連ねています。
- Anthropic はトークン利用量の 32% を占める一方で、支出の 61% を握りました。特にリスクの高いユースケースでは、支出シェアが 72% 以上を占めています。
- モダリティごとにリーダーは異なります。画像生成では OpenAI の GPT Image が 53% を占め、Google の Nano Banana(39%)を上回りました。動画分野の支出では中国勢が 3 分の 2 を握り、xAI の Grok Imagine は動画生成量の 42% を担い、全モデルの中で最も高いシェアを記録しました。
- Claude Fable 5 は 6 月 9 日にリリースされ、わずか 4 日で Opus 4.8 のリクエストボリュームの 22% に達しました。しかし、米国の輸出管理指令により、当月残りの期間アクセスが停止されました。
AI 投資は継続して増加、平均トークン単価は横ばい
6 月のトークン利用量と支出はほぼ同率で成長し(それぞれ 29% と 27%)、1 トークンあたりの平均価格は安定しました。先月までは支出の伸び率が利用量の約 2 倍に達しており、これが平均トークンコストを約 20% 押し上げる要因となっていました。
市場全体の支出額は増えたものの、トークン単価は上がっていません。4 月以来、オープンウェイトモデルのトークン利用量は全体的なボリュームの約 9 分の 1 からほぼ 3 分の 1 にまで拡大しました。ただし、ゲートウェイ上の平均トークン価格はその約 10 分の 1 の水準に留まっています。これだけでもトークン単価は下がるはずでしたが、6 月に安価な利用量が増えた一方で、クローズドウェイトの最先端モデルの価格も 1 トークンあたり約 12% 上昇しました。この二つの動きが互いに相殺された結果、平均的なトークン単価は横ばいとなりました。
これは、6 月のレポートで指摘されたルーティングの規律が、集計データとして明確に現れている証拠です。つまり、大量処理には低コストモデルを割り当て、リスクの高い処理には最先端モデルを維持しています。AI への投資はまだ急速に拡大しており、各企業は異なるモデル間で支出をどう配分するか、より戦略的なアプローチを取り始めています。
オープンウェイトが生産環境のトークンの約 3 分の 1 を担当
6 月のゲートウェイ利用データによると、オープンウェイトモデルは全トークンの 29% を処理しましたが、支出額は全体のわずか 4% 未満でした。つまり、費用対効果で言えば、ドル換算では約 25 分の 1 の出費で、トークン量のほぼ 3 分の 1 を賄っている計算になります。
オープンウェイトモデルのトークンシェアは 4 月以来、約 3 倍に拡大しました。この増加分の大部分は、DeepSeek の継続的な成長によるものです。現在、DeepSeek はゲートウェイ上のトークン供給源として Anthropic と Google に次ぐ第 3 位の規模となり、シェアは 22.6% を記録しています。一方、Google のシェアは 4 月の急伸が反動で解消され、24% に低下しました。これにより、DeepSeek は第 2 位との差をわずか 2 ポイントまで詰め寄っています。
現在、企業の顧客の約 8 人に 1 人が、本番環境でオープンウェイトモデルを採用しています。現在のペースが続けば、AI Gateway における利用量では、オープンウェイト専門のラボがすぐに第 2 位の規模に達するでしょう。
最新の参入事例がこの加速傾向を裏付けています。Z.ai は、Opus 4.8 の約 5 分の 1 の価格で、長期にわたる自律型タスク(アジェンシー)向けに設計された MIT ライセンスのオープンウェイトモデル「GLM 5.2」をリリースしました。6 月 16 日の API 公開から月末までのわずか 2 週間で、1 日あたりのトークン利用量は約 50 倍に急増。最終週には AI Gateway のトークン数ランキングで 11 位、単独の日では最高 7 位を記録しました。このモデルは、そのファミリー全体が 6 月に消費したトークンの 76% を、わずか 2 週間弱のうちに占めてしまったのです。
以前、最も急速な移行を示したと分析した「Gemini 3.1 Pro」でさえ、同様のシェアに達するには第 2 月を要しました。
コピーリンク見出しへ
フロンティア(最先端)ラボは、支出のほとんどを独占し続けています
6 月の AI Gateway における支出の 95% を、米国のトップ 4 フロンティアラボが占めました。Anthropic は、トークン数の 32% を使用しながら、支出全体の 61% を獲得しています(5 月は 65% でしたが、4 月と同水準です)。また、コーディングエージェント、バックオフィス業務の自動化、アプリ生成など、ミスのコストが極めて高い重要なユースケースにおける支出の 72% 以上も独占しました。
OpenAI のトークンシェアは 12.5% から 10.3% に減少しましたが、支出シェアは 13.3% から 16.1% に増加。その結果、市場平均と比較して 1 ヶ月でトークンあたりのコストが約 50% 上昇しました。DeepSeek と対照的に、OpenAI への利用量は減ったものの、より高価な処理を依頼する顧客が増えたためです。この傾向は、本番環境のデータにのみ明確に表れています。
各モダリティには異なるリーダーがいる
AI Gateway における画像生成は、実質的に 2 つのラボによる争いになっています。6 月のデータを見ると、OpenAI の「GPT Image」シリーズが生成された画像の 53% を占め、支出の 52% を握っていました。一方、Google の「Nano Banana」は約 39% の画像を生成し、支出では 43% を獲得しています。これら以外のファミリーが 5% を突破したケースはありません。
動画モダリティにおいては、現在中国のラボがプレミアム層を支配しています。ByteDance の「Seedance」は、生成された動画数の 3 分の 1 しか占めていないにもかかわらず、支出では 49% をリードしました。xAI の「Grok Imagine」は、支出の 19% で動画の 42% を生成しており、これはテキスト分野で DeepSeek が占めているボリュームディスカウントのポジションと全く同じです。中国のラボ 3 社(Seedance、Kling、Alibaba の Wan)を合わせると、動画関連の支出全体の約 3 分の 2 を握っています。このカテゴリはまだ初期段階にあるため、これらのシェアは今後変動していくでしょう。
要約すると、テキストは Anthropic が、画像は OpenAI が、動画の生成量は xAI が、そして動画の支出は ByteDance がそれぞれリードしています。各モダリティで最上位モデルを運用するには、少なくとも 3 つのラボにルーティングする必要があります。
US の輸出管理規制により、Claude Fable 5 は発売からわずか数日で利用停止へ
Anthropic は 6 月 9 日に「Claude Fable 5」をリリースし、急速な採用が進みました。なんと 4 日後には、Opus 4.8 への 100 リクエストに対して Fable の利用が 22 リクエストに達しました。
しかし、6 月 12 日に Fable 5 を含む米国の輸出管理指令が発効し、Anthropic はこれに対応してアクセスを停止せざるを得ませんでした。その結果、同モデルは月末までオフライン状態が続きました。規制解除は 6 月 30 日に行われ、7 月 1 日に利用が再開されています。
6月のデータから見える傾向
- バックオフィス向けエージェントは、1 トークンあたりのコストが最も高いワークロードです。全体のトークンの 5% を占めるのみですが、利用料金の 14% を占めています。
- B2B のユースケースは、6月のトークン数の 46% を担いましたが、利用料金の 60% に達しました。一方、B2C はその逆で、トークン数は全体の 43% ですが、利用料金は 26% です。
- Google の利用量は消費者向けのワークロードに集中しています。6月にはパーソナルアシスタント関連のトークンの 57% と教育関連の 54% を処理しましたが、コーディングエージェント関連は 2% に満たない程度でした。
このレポートについて
本分析は、2026 年 6 月までの Vercel AI Gateway から収集した匿名化された集計ルーティングデータに基づいています。
測定に関する補足:
Spend は、市場価格(公開リスト価格)を採用し、各自の API キーを持つチーム間での比較を可能にする標準化された指標を提供します。
Volume は、AI Gateway を経由してルーティングされたトークンの総数をカウントしたものです。
B2C、B2B、およびユースケース分類は集計値であり、特定のチームやワークロードが個別に特定されることはありません。
画像と動画の数値は、生成されたメディア(成功したリクエストのみ)の数を示しており、トークン数ではありません。画像・動画モデルはトークン課金ではないため、これらのシェアは実際に生成された画像や動画の割合を反映しています。
オープンウェイトモデルの評価は 2 つのアプローチで行われます。まず、Volume と Spend のシェア(ラボ別シェアチャート)ではプロバイダー別に分類し、Gateway で自社モデルを提供する 4 つのラボ(DeepSeek、MiniMax、Moonshot、Z.ai)をカウントします。これは保守的な見積もりです。一方、エンタープライズ採用率の数値はモデル単位で分類され、どのプロバイダーが提供したかにかかわらず、Qwen、Gemma、GPT-OSS、Kimi などのオープンウェイトモデル全体をカウント対象とします。
Copy link to headingPrevious reports
原文を表示
*AI Gateway Production Index — July 2026 *
Every month, AI Gateway routes tens of trillions of tokens between production applications and AI labs, giving us a view of what AI usage actually looks like in today’s enterprise. We publish that view here. See the Production Index reports published in May and June.
Copy link to headingJuly 2026 summary
The July index reports on AI Gateway data collected in June 2026.
- AI Gateway token volume grew 29% month over month and spend grew 27%. The price per token was flat after rising almost 20% in May.
- Open-weight models ran 29% of gateway tokens, up from 11% in April, on under 4% of spend. DeepSeek reached 22.6% of token volume, in third place and less than two points behind Google, and GLM 5.2 broke into the gateway's top models by volume two weeks after release.
- Anthropic took 61% of gateway spend on 32% of tokens, and captured 72% or more of spend in every high-stakes use case.
- Each modality has a different leader. OpenAI's GPT Image generated 53% of images, ahead of Google's Nano Banana at 39%. Chinese labs took two-thirds of video spend, while xAI's Grok Imagine generated 42% of videos, the most of any model.
- Claude Fable 5 was released June 9 and reached 22% of Opus 4.8's request volume in four days. A US export-control directive suspended access for the rest of the month.
Copy link to headingAI investment kept climbing, and the average token price went flat
In June, token volume and spend grew at nearly the same rate, 29% and 27%, while the price per token remained steady. The month before, spend had grown more than twice as fast as volume, driving the average token cost up by almost 20%.
The market spent more overall, but not more per token. Since April, open-weight models have climbed from a ninth of all token volume to nearly a third, at about a tenth of the average token price on the gateway. That alone should have pulled the price per token down. But as cheap volume rose in June, so too did closed-weight frontier prices, up about 12% per token. They offset each other, so the average price per token was flat.
This is evidence of the routing discipline June's report documented, now visible in the aggregate: high-volume work goes to low-cost models, high-risk work stays on the frontier. Investment in AI is still growing quickly, and companies are getting more strategic about balancing that spend across models.
Copy link to headingOpen weight ran nearly a third of production tokens
Open-weight models ran 29% of gateway tokens in June on just under 4% of spend. That is nearly a third of the tokens for one twenty-fifth of the dollars.
Open-weight token share has almost tripled since April. Most of that volume comes from the continued growth of DeepSeek, now the third-largest source of tokens on the gateway at 22.6%, behind Anthropic and Google. Google's share of token volume slipped to 24% as its April surge unwound, leaving DeepSeek within two points of second place.
Roughly one in eight enterprise customers now runs an open-weight model in production. And on current trajectories, an open-weight lab will soon be the second-largest by volume on AI Gateway.
The newest entrant shows the pattern accelerating. Z.ai released GLM 5.2, an MIT-licensed open-weight model aimed at long-horizon agentic work at approximately a fifth of Opus 4.8 pricing. From its June 16 API availability through month end, its daily token volume grew about 50x. It ranked #11 on AI Gateway by tokens in the final week, and as high as #7 on single days. It took 76% of its family's June tokens in barely two weeks. Gemini 3.1 Pro, the fastest in-family migration we had previously charted, took until its second month to reach a similar share.
Copy link to headingFrontier labs kept nearly all of the spend
The top four frontier US labs took 95% of AI Gateway spend in June. Anthropic alone took 61% of spend on 32% of tokens, down from 65% in May, but in line with April. It also took 72% or more of spend in the most consequential use cases, where mistakes are costly: coding agents, back-office agents, and app generation.
OpenAI’s token share fell from 12.5% to 10.3% while its spend share rose from 13.3% to 16.1%, driving its cost per token up about 50% relative to the market in one month. In contrast to DeepSeek, customers sent OpenAI less volume but costlier work, a divergence that only shows up in production data.
Copy link to headingEach modality has a different leader
Image generation is a two-lab race on AI Gateway. OpenAI's GPT Image family generated 53% of images and took 52% of image spend in June. Google's Nano Banana generated almost 39% of images and took 43% of spend. No other family cleared 5% of either.
Video is the one modality where Chinese labs currently hold the premium tier. ByteDance's Seedance led spend at 49% with only a third of videos generated. xAI's Grok Imagine generated 42% of videos on 19% of spend, the same volume-discount position DeepSeek holds in text. The Chinese labs together, Seedance, Kling, and Alibaba's Wan, took roughly two-thirds of video dollars. The category is early, so we expect these shares to fluctuate.
Anthropic leads text, OpenAI leads images, xAI leads video volume, and ByteDance leads video spend. Running the leading model in every modality requires routing across at least three labs.
Copy link to headingUS export controls suspended Claude Fable 5 only days after launch
Anthropic released Claude Fable 5 on June 9, and it was adopted rapidly. In only four days, Fable usage rose to 22 requests for every 100 sent to Opus 4.8.
On June 12, a US export-control directive took effect covering Fable 5, and Anthropic suspended access to comply. The model stayed offline for the rest of the month. Controls lifted on June 30 and access resumed July 1.
Copy link to headingAlso in June's data
- Back-office agents are the most expensive workload per token on the gateway, running 5% of total tokens on 14% of total spend.
- B2B use cases drove 46% of June tokens but 60% of spend; B2C the reverse, at 43% of tokens and 26% of spend.
- Google's volume concentrates in consumer-shaped workloads. It ran 57% of personal-assistant tokens and 54% of education tokens in June, but less than 2% of coding agent tokens.
Copy link to headingAbout this report
This analysis is based on anonymized, aggregate routing data from the Vercel AI Gateway through June 2026.
A few notes on measurement:
- Spend uses market-rate pricing (published list price) to provide a normalized view across teams that bring their own API keys.
- Volume counts tokens routed through AI Gateway.
- B2C, B2B, and use-case classifications are aggregate. No individual team or workload is identified.
- Image and video figures count generated media (successful requests only), not tokens. Image and video models are not billed per token, so shares reflect images and videos produced.
- Open-weight models are measured two ways. Volume and spend shares (the lab-share charts) classify by provider, counting the four labs that serve their own models on the gateway (DeepSeek, MiniMax, Moonshot, Z.ai). This is conservative. The enterprise-adoption figure classifies by model, counting open-weight models regardless of which provider served them, such as Qwen, Gemma, GPT-OSS, and Kimi.
Copy link to headingPrevious reports
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み