Vercel AI Gateway、DeepSeek が Google をトークン量で上回る
本文の状態
日本語全文を表示中
詳細モードで約12分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Vercel Blog
2026 年 7 月のデータによると、DeepSeek がトークン量で Google を抜いて業界第 2 位となり、Anthropic が支出の大部分を占めるなど、AI エコシステムの競争構造と価格動向に劇的な変化が生じた。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 08:46
AI深層分析
キーポイント
DeepSeek の市場シェア急拡大
DeepSeek は 7 月にトークン量で Google を上回り業界第 2 位となり、特に消費者向けアシスタント機能において Google のシェアを大きく奪った。
Anthropic の高単価支配
Anthropic は全トークン量の 30% を占める一方で、ゲートウェイ支出の 65% を握り、他社平均の 4.4 倍という高い価格で利用されている。
Kimi K3 のエージェント活用
Moonshot が発表した Kimi K3 は長期ホライズンのエージェントワーク向けに設計され、ローンチ直後から前世代比 12 倍のトークン消費を記録した。
オープンウェイトモデルの台頭
7 月の支出におけるオープンウェイトモデルのシェアは 8.6% に倍増し、Moonshot と Z.ai がその成長の大半を担ったが、依然として低単価の傾向にある。
オープンウェイトモデルの台頭と市場シェアの変化
7月にはオープンウェイトモデルのトークンボリュームシェアが3分の1を超え、Google のシェアは24.0%から10.7%へ急落した。Z.ai や Moonshot などの最新モデルが、従来クローズドウェイトラボが担っていたワークロードを事実上引き受けた。
重要な引用
DeepSeek became the second-largest lab by token volume, now running more than twice Google's volume.
Anthropic collected 65% of gateway spending on 30% of token volume, at 4.4 times the average price of every other lab's tokens.
Kimi K3 launched into agent work... its usage was heavy from the start at about twelve times the tokens per request of its predecessor.
"Almost none of the spend it lost went to DeepSeek, whose share barely moved even as its volume grew."
編集コメントを表示
編集コメント
2026 年という近未来のデータにおいて、中国発の AI モデルが市場を支配し始めた事実は、グローバルな競争環境が急速に変化していることを示唆する。特にエージェントワークロードにおける Kimi K3 の台頭は、LLM の用途拡大と価格構造の変化を象徴する重要な転換点である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI Gateway 生産インデックス — 2026 年 8 月版
毎月、AI Gateway は本番環境のアプリケーションと AI ラボの間で数十兆トークンのやり取りを処理しています。このトラフィックデータから、現在の企業における AI の実際の利用状況が見えてきます。私たちはこれを毎月のレポートとして公開しています。5 月、6 月、7 月の生産インデックスレポートもご覧ください。
2026 年 8 月のサマリー
8 月版のインデックスは、2026 年 7 月末までの AI Gateway データに基づいています。
7 月のトークンあたりの平均支払額は前月比で 13.6% 低下しました。これは 5 月に約 20% の上昇があった後、6 月は横ばいだった状況からの転換です。トークン消費量が急激に拡大したため、支出自体は 37% 増加していましたが、結果としてトークンあたりのコストは二桁の減少を記録しました。
DeepSeek はトークンボリュームにおいて世界第 2 位のラボとなり、Google の処理量をすでに 2 倍以上上回っています。
Anthropic は全トークンボリュームの 30% を占めるに留まっていますが、ゲートウェイ全体の支出の 65% を独占しています。これは他社すべてのラボの平均価格と比較して 4.4 倍の高額です。
メディア系リーダーボードでも首位交代が起きました。画像生成では Google の Nano Banana が OpenAI の GPT Image を抜いてトップに立ち、動画生成では ByteDance の Seedance がボリュームと金額の両方で首位となりました。
Kimi K3 がエージェントワークへ参入
Moonshot は 7 月 16 日に Kimi K3 をリリースしました。6 月に公開された Z.ai の GLM 5.2 と同様、このモデルは長期にわたるエージェント作業を目的として設計されています。そのため、利用開始直後から前作の K2.5 に比べて 1 回のリクエストあたりのトークン数が約 12 倍と非常に重厚な負荷がかかる運用となりました。
急速に規模を拡大しました。K3 の日次処理量は、発売週から7月末の最終週にかけて3倍になり、月末には Claude Opus 4.8 に匹敵するほどの重負荷となりました。7月の最後の完全な1日では、ゲートウェイ上のトークン量で8位にランクインしています。
この需要は既存のものが転用されたものではなく、新規のものでした。K3 は2週間で Kimi シリーズ全体のトークンのほぼ3分の2を処理し、最終週には82%を占めました。一方、同シリーズの他のモデルの処理量はわずかに減少するだけにとどまりました。
オープンウェイトモデルがゲートウェイ支出に占める割合は7月に倍増し、8.6%となりました。この成長の90%以上は Moonshot と Z.ai によるものです。Moonshot のシェアは4倍になり、全体の2.3%を占めました。安価なオープンウェイトモデルは数ヶ月にわたり処理量を増やしていましたが、収益には反映されていませんでした。Kimi K3 と GLM 5.2 は、トークンあたりの DeepSeek の料金の11倍以上で、初めて大きな処理量を獲得したモデルです。
DeepSeek が処理量で2位に躍り出ました
今月のレポートでは DeepSeek がトークン量の争いに参入したと伝えましたが、先月にはオープンウェイトの研究所がすぐに2位になると予測していました。そして7月、DeepSeek は Google を抜いてその座を奪いました。
4月にはゲートウェイ上のトークン量の約40%を Google が処理し、DeepSeek は1%未満でした。しかし7月には DeepSeek のシェアは25%に達し、Google の11%の倍以上となりました。DeepSeek V4 Flash(最も安価なモデル)単独で処理したトークン量は、Google 全体の量を上回っています。
この逆転は、消費者向けワークロードに集中しています。Google のパーソナルアシスタント用トークンのシェアは 1 か月で半分以下に減った一方、DeepSeek は 3 倍以上に増え、Google が失ったトークンシェアのほとんどが DeepSeek に流れました。7 月のゲートウェイでは、DeepSeek の V4 Flash が他のどのモデルよりも多くのトークンを処理し、全体の約 5 分の 1 を占め、次点のモデルを 70% も上回りました。
以前は安価な市場をオープンウェイトモデルが独占していました。ゲートウェイでのトークンボリュームにおけるシェアは 4 月から 6 月にかけて 11% から 29% にほぼ倍増しましたが、支出シェアはゲートウェイのドルにつき 4 セント未満にとどまりました。
7 月にはボリューム成長傾向が続き、36% に達しました。一方、支出はトレンドを破り、ドルにつき 9 セント強に倍増し、インデックス史上最高となりました。
主要なフロンティアラボ 4 社のトークン支出シェア合計は 7 ヶ月間 93% を下回ることはありませんでしたが、89% に低下しました。この減少の大部分を Google が占めています。Google が失った支出のほとんどが DeepSeek に回ったわけではなく、DeepSeek のシェアはボリュームが増えたにもかかわらずほぼ変動しませんでした。Z.ai と Moonshot の最新モデルが、オープンウェイトにおける支出増加分のすべてを占めています。
購入者はタスクに必要な機能でモデルを分類し、そのカテゴリ内で価格に応じて選択します。DeepSeek と Opus が競合することは決してありませんでした。GLM 5.2 と Kimi K3 は直近 2 ヶ月間にリリースされたオープンウェイトモデルですが、これまでクローズドウェイトラボが独占していたワークロードの有意なシェアを処理する初のケースです。
トークンあたりの平均価格は 13.6% 下落しました
7 月は、企業が推論利用をさらに拡大し、低価格モデルの利用比率が高まった月でした。利用量は前月比で 59% 増、支出は 37% 増となりましたが、1 トークンあたりの平均支払額は 13.6% 低下しました。
もし 6 月のモデル構成をそのまま維持していた場合、平均価格はむしろ横ばいだったはずです。OpenAI がその好例です。OpenAI のトークン 1 つあたりの平均コストは 6 月水準の 58% にまで下落しましたが、これは追加された利用量の 85% が GPT-5 ファミリーの中で最も安価な「GPT-5-Nano」に割り当てられたためです。
平均価格の低下全体は、企業がどのようにルーティングを選んだかによるものです。両月とも 1000 万トークン以上を利用した数千チームのうち、4 社中 3 社はモデル構成の少なくとも 1 割を変更しました。また、5 社中 3 社は 25% 以上の変更を行いました。
利用チームの中央値における 1 トークンあたりのコストは 2.9% 低下しましたが、スタート地点から±5% の範囲内に留まったのは 6 チームに 1 人だけでした。
コストを 30% 以上削減したチームも全体の 4 分の 1 を占め、これらは通常、月初めに平均水準より遥かに高い料金を支払っていたチームです。一方、少なくとも 20% 増額したチームも同様に 4 分の 1 存在し、これらは当初から平均水準以下の料金で利用していたチームでした。結果として、高価格帯の層は低下し、低価格帯の層は上昇しました。
7 月に利用されたトークンの 81% は、半年前にはゲートウェイに存在していなかったモデル上で実行されました。オープンウェイトモデルが全利用量の約 3 分の 1 を占め、その料金は最先端モデルの約 7 分の 1 でした。Google のシェアはゲートウェイのトークン利用量において 24.0% から 10.7% に低下し、Anthropic は 2 ポイント減の 29.8% となりました。一方、OpenAI のシェアは 12.8% に上昇しました。
5 月以来、ゲートウェイでの支出は 74% 増加し、利用量は倍増しました。特に 7 月のトークン成長率は、6 月のおよそ 2 倍のペースで推移しています。1 ドルで購入できるトークン数が増えているにもかかわらず、推論(インファレンス)への支出は引き続き増加傾向にあります。
Anthropic のプレミアム価格帯が拡大しました
Anthropic は、オープンウェイトモデルの利用量が 3 倍に増えた期間を通じて、測定されたすべての月でゲートウェイ支出の 60% 以上を占め続けています。7 月には、総利用量の 30% を占める中で、全支出の 65.1% を獲得しました。
Anthropic のトークンあたりの平均価格は、他のすべてのラボの平均の 4.4 倍に達し、6 月の 3.4 倍からさらに上昇しました。この価格上昇の一因は、Claude Fable 5 です。同モデルは 3 週間にわたる輸出規制による一時停止を経て 7 月 1 日に復帰するとすぐに、全ゲートウェイ支出の 13.2% を占めるまでに急成長し、Opus 4.8 に次ぐ第 2 位となりました。特にトークン数で最大の利用ケースであるコーディングエージェントにおいては、Anthropic が支出の 80% 以上を独占しています。
Anthropic は市場最安層にモデルを持っていません。最も安価な Haiku 4.5 でさえ、ゲートウェイ全体のトークン平均価格の約 2/3 を要します。一方、OpenAI の GPT-5-Nano はその 1/6、DeepSeek の V4 Flash はさらに 1/16 です。
OpenAI や Google でのコスト削減を図る顧客は、引き続きカタログ内のモデルを利用できます。しかし Anthropic でコストを削減したい場合は、Anthropic を離れる以外に方法がありません。それでも 7 月には、全支出の過半数を Anthropic が占め続けていました。Anthropic のプレミアム価格とは、安価なティアでできることと、顧客が実際に必要としている機能との間のギャップそのものです。
AI ゲートウェイ上でモデルを切り替えるのは一行の変更で済むため、顧客が特定のベンダーにロックインされることはありません。Fable 5 が復活した際、日間の処理量は禁止解除前の水準にほぼ完全に回復しましたが、7 月にこのシステムを利用していたチームの 9 割は、それ以前には利用していませんでした。業務自体は戻ってきたものの、顧客層は変わっていたのです。
市場全体のトークン単価は Anthropic の価格の 4 分の 1 に満たないものが平均でしたが、ゲートウェイでの支出総額の 3 分の 2 を Anthropic が占めています。価格競争は確かに起きていますが、それは資金力が最も乏しい市場セグメントで起こっているに過ぎません。
Google が画像を、ByteDance が動画を席巻
6 月には OpenAI の GPT Image がゲートウェイ上の生成画像の大部分を担っていましたが、7 月には Google の Nano Banana が 45% で首位に立ち、GPT Image は 42% と僅差でした。両者の支出額はほぼ半々です。Nano Banana のシェア拡大は、実質的に Gemini 3.1 Flash Lite Image という一つのモデルによるものです。Google はトークン量では 4 位に転落したのと同じ月に、画像生成において首位を奪いました。
ByteDance の Seedance は、生成された動画数と支出額の両方で動画分野をリードしました。6 月のボリュームリーダーだった xAI の Grok Imagine は 2 位に後退し、中国のラボ企業が動画関連の支出全体の約 7 割を占めています。
7 月のデータからもう一つの事実
トークン量で最大のワークロードであるコーディング分野では、DeepSeek が処理量のほぼ 3 分の 1 を担い、Anthropic は支出額の 5 分の 4 以上を獲得しました。
バックオフィス業務を自動化するエージェントは、トークン単価で見ると最も高コストな仕事です。その支出シェアは、ボリュームシェアの約 2.5 倍に達しています。
本レポートについて
本分析は、2026 年 7 月までの Vercel AI Gateway を経由した匿名化された集計ルーティングデータに基づいています。
測定に関するいくつかの注記:
ボリューム(利用量)は、AI Gateway を経由して処理されたすべてのトークンをカウントします。
支出額は、各リクエストに対してラボが公表しているリスト価格で算出しています。
1 トークンあたりの料金は、「総支出額 ÷ 総利用量」で計算されます。
オープンウェイトのボリュームおよび支出シェアは、大規模に自社モデルを提供する 4 つのオープンウェイト研究機関(DeepSeek、MiniMax、Moonshot、Z.ai)を指します。
画像および動画の数値は、リクエスト数やトークン数ではなく、生成されたメディアの量を表しています。
すべての数値は利用可能な最新データを使用しており、手法が更新されることで過去の月次データも改訂される可能性があります。
原文を表示
AI Gateway Production Index — August 2026
Every month, AI Gateway routes tens of trillions of tokens between production applications and AI labs. That traffic gives us a view of what AI usage actually looks like in today's enterprise, and we publish it here monthly. See the Production Index reports from May, June, and July.
August 2026 Summary
The August index reports on AI Gateway data collected through July 2026.
The average price paid per token fell 13.6% in July, after rising almost 20% in May and holding steady in June. Token consumption grew so quickly that even with 37% growth in spend, cost per token saw a double-digit drop.
DeepSeek became the second-largest lab by token volume, now running more than twice Google's volume.
Anthropic collected 65% of gateway spending on 30% of token volume, at 4.4 times the average price of every other lab's tokens.
Both media leaderboards changed hands. Google's Nano Banana took the lead in image volume from OpenAI's GPT Image, and ByteDance's Seedance led video in both volume and dollars.
Kimi K3 launched into agent work
Moonshot released Kimi K3 on July 16. Like Z.ai's GLM 5.2 released in June, it is built for long-horizon agent work, so its usage was heavy from the start at about twelve times the tokens per request of its predecessor, K2.5.
It scaled quickly. K3's daily volume tripled between launch week and the final week of July, and by month end its requests were as heavy as Claude Opus 4.8's. On the last full day of July it ranked eighth on the gateway by token volume.
The demand was new rather than diverted. K3 processed nearly two-thirds of all Kimi tokens within two weeks and 82% by the final week, while the rest of the family's volume fell only slightly.
Open weight's share of gateway spend more than doubled in July to 8.6%. More than 90% of that growth is Moonshot and Z.ai. Moonshot's share of total gateway spend quadrupled, to 2.3%. Cheap open-weight models have been taking volume for months without taking revenue. Kimi K3 and GLM 5.2 are the first to capture significant volume at more than eleven times DeepSeek's rate per token.
DeepSeek is now second by volume
In June, this report said DeepSeek had entered the fight for token volume. Last month, we said an open-weight lab would soon be second by volume. In July, DeepSeek surpassed Google to take that place.
Google ran nearly 40% of the gateway’s token volume in April, and DeepSeek less than 1%. By July, DeepSeek ran a quarter, more than twice Google’s 11%. DeepSeek V4 Flash, its cheapest model, ran more tokens alone than all of Google.
The reversal is concentrated in consumer-facing work. Google's share of personal-assistant tokens fell by more than half in a month while DeepSeek's more than tripled, and most of the token share Google gave up went to DeepSeek. DeepSeek's V4 Flash ran more tokens than any other model on the gateway in July, nearly a fifth of the total and 70% more than the next model.
Open-weight models previously owned the cheap end of the market. Their share of gateway token volume nearly tripled between April and June, from 11% to 29%, while share of spend stayed under four cents of every gateway dollar.
In July, volume continued its growth trend, increasing to 36%. Spend broke its trend, more than doubling to nearly nine cents of every dollar, the highest in the index's history.
The four largest frontier labs' combined share of token spend, which had not fallen below 93% in seven months, fell to 89%. Google accounted for most of the decline. Almost none of the spend it lost went to DeepSeek, whose share barely moved even as its volume grew. Z.ai and Moonshot's latest models account for the entire open-weight spend increase.
Buyers sort models by what the task requires, then choose inside that tier according to price. DeepSeek and Opus were never competing. GLM 5.2 and Kimi K3, both released in the last two months, are the first open-weight models running a meaningful share of the workloads historically owned by closed-weight labs.
Average price per token fell 13.6%
Companies bought more inference in July and ran more of it on cheap models. Volume grew 59%, spend grew 37%, and the average price paid per token fell 13.6%.
Holding June's mix of models constant, the average price would have held essentially flat instead of declining. OpenAI is one example: the average cost of an OpenAI token fell to 58% of its June level, because 85% of the volume it added went to GPT-5-Nano, the cheapest model in the GPT-5 family.
The entire decline in average price came from what companies chose to route. Among the thousands of teams that ran more than 10 million tokens in both months, three in four changed at least a tenth of their model mix. Three in five changed at least a quarter.
The median team's cost per token fell 2.9%, but only one team in six stayed within five percent of where it started.
A quarter pared cost by more than 30%, typically teams that entered the month paying well above the gateway average per token. Another quarter paid at least 20% more, typically teams that had been paying below the average. Concurrently, the expensive end shifted down and the cheap end migrated up.
81% of July’s tokens ran on models that were not on the gateway six months ago. Open-weight models crossed a third of all volume, at about a seventh of frontier rates. Google fell from 24.0% of gateway token volume to 10.7%. Anthropic slipped two points to 29.8%, and OpenAI’s share rose to 12.8%.
Gateway spend is up 74% since May and volume has doubled, with July's token growth running at nearly double June's pace. Spend on inference continues to increase even as each dollar buys more tokens than it did in the previous month.
Anthropic's premium widened as prices fell
Anthropic has held more than 60% of gateway spend in every month we have measured, through a period in which open weight tripled its share of volume. In July, it collected 65.1% of all spend on 30% of total volume.
The average price per Anthropic token ran 4.4 times the average across every other lab, up from 3.4 in June. Part of that rise was Claude Fable 5, which returned on July 1 after a three-week export-control suspension and quickly grew to 13.2% of all gateway spend, second only to Opus 4.8. In coding agents, the gateway's largest use case by tokens, Anthropic collected more than 80% of spend.
Anthropic has no model at the bottom of the market. Haiku 4.5, its least expensive, runs just over two-thirds of the gateway average price per token. OpenAI's GPT-5-Nano runs a sixth; DeepSeek's V4 Flash a sixteenth.
Customers cutting costs on OpenAI or Google can stay in the catalog. Cutting costs on Anthropic means leaving Anthropic, and in July it still took a majority of all spending. Anthropic’s premium is the gap between what the cheap tier can do and what its customers need done.
Switching models on the AI Gateway is a one-line change, so nothing locks a customer in. When Fable 5 came back, daily volume returned to the pre-ban level almost exactly, but nine in ten of the teams running it in July had not used it before. The work returned; the customers were different.
The rest of the market's tokens averaged less than a quarter of Anthropic's price, and Anthropic still took two of every three dollars spent on the gateway. Price competition is happening, but it's happening in the part of the market with the least money in it.
Google took images, ByteDance took video
In June, OpenAI's GPT Image generated most of the gateway's images. In July, Google's Nano Banana led, at 45% of images to GPT Image's 42%, with spend split almost evenly between them. Nearly all of Nano Banana's gain came from one model, Gemini 3.1 Flash Lite Image. Google took the image lead in the same month it fell to fourth by token volume.
ByteDance's Seedance led video on both counts, in videos generated and dollars spent. xAI's Grok Imagine, June's volume leader, fell to second. Chinese labs collected about seven of every ten dollars spent on video.
Also in July’s data
Within coding, the gateway's largest workload by tokens, DeepSeek ran nearly a third of the volume and Anthropic collected more than four of every five dollars.
Back-office agents remain the most expensive work per token, with a share of spending roughly two and a half times their share of volume.
About this report
This analysis is based on anonymized, aggregate routing data from the Vercel AI Gateway through July 2026.
A few notes on measurement:
Volume counts all tokens routed through AI Gateway.
Spend values every request at the lab's published list price.
Price per token is calculated as total spend divided by total volume.
Open-weight volume and spend shares count the four open-weight labs serving their own models at scale (DeepSeek, MiniMax, Moonshot, and Z.ai).
Image and video figures count media generated, not requests or tokens.
All figures use the most recent data available; prior months may be revised as methodology is updated.
Read more
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み