トークン量争奪戦に DeepSeek が参入、支出支配は Anthropic が継続
本文の状態
日本語全文を表示中
詳細モードで約9分の本文を読めます。
Vercel の AI Gateway データによると、DeepSeek の利用シェアが単月で 1% から 17% に急増し、トークン量の争奪戦に本格参入した。一方、支出面では Anthropic が依然として支配的な地位を維持している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
毎月、AI Gateway は生産アプリケーションと AI ラボの間で数十兆トクンのルーティングを行い、リーダーボードやベンチマークとは別に、実際の AI 利用状況を目視可能にしています。このデータは、AI Gateway 生産インデックスとして毎月公開されています。
2026 年 5 月サマリー
総 AI Gateway トークン数は前月比 +20% 増、総支出額は同 +43% 増でした。顧客が支払ったトークンあたりの平均価格は、4 月に比べてほぼ 20% 高くなりました。
DeepSeek のトークンシェアは、わずか 1 か月の間に 1% 未満から 17% に急上昇しましたが、支出シェアは約 1% のまま推移しました。
Anthropic の支出シェアは 5 月に 61% から 65% に拡大し、AI アプリ生成、バックオフィスエージェント、コーディングエージェントといったすべての高リスクユースケースにおいて、支出の 70〜80% を占め続けました。
コスト意識の高まりにより、低価格モデルと最先端モデルの間でのより賢いルーティングが行われました。顧客はどのモデルがどの作業を担当するかをより意図的に選定するようになり、全体としての利用量は引き続き増加しました。
先月、トークン予算の超過に関するニュースがテック界隈のヘッドラインを席巻しました:Uber は第 1 四半期直後に Claude Code の年間予算を使い果たし、Amazon は非生産的なトークン最大化(tokenmaxxing)を抑えるために KiroRank を停止しました。暴走するコストは確かに深刻な問題ですが、今月のレポートでは、生産ユースケースへの支出が依然として増加していることが示されています。
5 月の AI Gateway データから得られた 2 つの洞察:
低価格モデルが生産環境に参入:新モデルは既存の大手ラボをさらに高価に見せる価格帯で出荷され、生産環境での採用に必要な十分な能力を備えています。
支出は増加していますが、より賢明なモデルミックスによるもの:チームは依然としてトークン予算を増やしていますが、1 ドルあたりからより多くの価値を引き出すために、より賢明なルーティング戦略を実装しています。
低コストモデルが初めて生産量で大きな役割を果たす
2 月から 4 月にかけて、AI Gateway 上の各ラボ間のボリューム分布は緩やかに変化していましたが、5 月に DeepSeek V4 が登場したことでトークンシェアが完全にシフトしました。4 月にはほとんど存在しなかった市場の低コスト層が、全体の支出に大きな影響を与えることなく、5 月には AI Gateway のボリューム別第 3 位のプロバイダーとなりました。
4 月には DeepSeek は AI Gateway のトークンの 1% 未満、支出の 0.2% 未満を占めていました。しかし 5 月にはそのボリュームシェアが急上昇し、トークン全体の 17% を占めるようになり、OpenAI を上回る第 3 位となりました。このボリュームのほとんどは、5 月にリリースされた 2 つのモデル、すなわち deepseek/deepseek-v4-flash と deepseek/deepseek-v4-pro に由来しています。
支出に関するデータはその物語のもう半分を語っています。DeepSeek のトークンシェアが単月で 17% に拡大したにもかかわらず、そのコストシェアは約 1% のまま推移しました。
DeepSeek V4 Flash は、百万トークンあたり入力 0.14 ドル、出力 0.28 ドルという価格で登場し、これは同等の Anthropic モデルと比較して約 20~50 倍低く、Qwen 3.6 Plus や Kimi K2.6 のような他のバリューティアのフラッグシップモデルと比較しても 8~12 倍低い水準です。この大きな節約ギャップにより、チームは V4 Flash を急速に採用しました。
価格だけで DeepSeek の利用量が1ヶ月でこれほど大幅に増加したとは考えにくく、つまり既存の評価基準に対して DeepSeek V4 をテストしたチームは、単に低コストだから試すというだけでなく、出力品質が実装に耐えうるレベルであると判断したことを意味します。
AI Gateway では従来からバリューティアのモデルが利用可能でしたが、これほど大規模なトークンシェアを獲得したのは初めてであり、DeepSeek V4 はその価格帯で生産環境の品質基準をクリアした初のモデルとなります。
フロンティア研究所は引き続き新規支出の過半数を占め続けています
市場の低コスト層がボリューム面で最も急速に成長する一方で、高価格層の方がドル建てではより速く成長しました。
Anthropic のトークンシェアは 26% から 32% に拡大し、支出シェアも 61% から 65% に上昇しました。OpenAI のトークンシェアは約 13% で横ばいでしたが、総量が大幅に大きくなったため支出シェアは 12% から 13% にわずかに増加し、顧客は5月の OpenAI トークンあたりにより多くの費用を支払っています。
DeepSeek が平均値を引き下げているにもかかわらず、5月は平均トークン単価が上昇しました。この増大は、フロンティアモデルを必要とするワークロードの成長速度が、そうでないワークロードよりも速かったためです。AI コーディングエージェントの使用事例において、低コスト層とフロンティア層の明確な分岐が最もよく示されています:
DeepSeek はセグメント全体のトークンボリュームの 49% を占めましたが、コストはわずか 4% でした。
Anthropic はトークンの 28% とコストの 70% を担いました。
低価格モデルはもはや生産ワークフローにおいて重要な一部となっていますが、フロンティアモデルの使用はまだ成長を続けており、これが全体支出の増加を牽引しています。
フロンティアモデルのトークンあたりのコストは上昇しており、顧客はまだ支払いを続けています。Anthropic は支出面で引き続き主導権を握っており、5 月のすべてのゲートウェイ支出の 65% を占め、あらゆる高リスクユースケースにおける支出の 70〜80% を支配しています。
コスト管理がルーティング戦略へ
全体的な支出の増加は、5 月においても AI への需要が続いて成長していることを示しましたが、チームはルーティングを通じて予算に対してより精密な対応を行いました。安価で高ボリュームのワークロードを低価格モデルに割り当て、品質が最も重要となる場面ではフロンティアモデルを活用しています。Google の最新 Flash モデルである Gemini 3.5 Flash の採用が遅れていることは、その明確な例です。
Gemini 3.5 Flash は 5 月に、Gemini 3.0 Flash よりも高い価格帯でリリースされましたが、大規模な移行は起こりませんでした。月末時点では、Flash ファミリー内のトークンにおいて 3.5 がわずか 7% を占めるに過ぎず、一方 3.0 は 90% を占めていました。
2 月と 3 月に Gemini 3.1 Pro が急速に採用されたことと比較すると、3.5 Flash への移行が遅れていることは、3.0 Flash に満足しているチームがまだ高いコストを支払う用意がないことを示しています。
結論:コストパフォーマンスに優れ、能力のある選択肢こそが、より賢明なモデルミックスを意味する
今月のレポートは、全体的な支出とトークンボリュームが増加しているにもかかわらず、市場における価格感応度が高まっていることを示唆しています。つまり、開発者は 1 ドルからより多くの価値を引き出す方法を模索しているのです。
データが明らかにした 2 つの最適化戦略:
- リスクが低くボリュームの高いタスクには、安価ながら能力のある DeepSeek の V4 ファミリー(V4 family)を活用する
- モデルファミリーのアップグレードは、投資対効果(ROI: Return on Investment)が明確になるまで延期する
ルーティング機能により、各チームはラボ間での異なる生産 AI ワークロードの層を巡る競争に応じて、モデル構成と予算を実時間で調整することが可能になります。
付録
B2B クラス別トークン数対コストシェア
B2B アプリケーションは回数が少なく高価な呼び出しを実行する一方、B2C アプリケーションは多数の安価な呼び出しを実行します。1 トークンあたりのコストで見ると、5 月には B2B は B2C よりも約 60% 高くつきました。
トークン数とリクエスト別エージェントツール使用状況
リクエストのほぼ 4 分の 1 がツールの呼び出しで終了しますが、これらのリクエストが全トークンの過半数を占めています。両方の指標は前月比でほぼ横ばいです。
リクエストボリューム別モデル多様性分布
アプリが処理するリクエスト数が多いほど、本番環境で実行されるモデルの数も増えます。最低ボリューム層では単一モデル構成が支配的ですが、100 万回以上のリクエストがある場合、大半のアプリは 11 個以上のモデルにルーティングされています。
ユースケース別コスト対ボリュームシェア
ユースケースのコストシェアは、誤った回答が出た際の費用の高さを示すものであり、消費されるトークンの多さを示すものではありません。パーソナルアシスタントやコーディングエージェントは 1 トークンあたりのコストが安価ですが、バックオフィス業務や採用関連の作業ははるかに高額になります。
過去のレポート
2026 年 4 月の AI Gateway 本番インデックスをお読みください。
データについて
この分析は、2026 年 5 月までの Vercel AI Gateway から収集した匿名化された集計ルーティングデータを基にしています。
測定に関するいくつかの注記:
支出(Spend)は、各チームが独自 API キーを使用するケースを比較可能にするため、市場価格(公開リスト価格)に基づいて計算されています。
ボリュームは、AI Gateway を経由してルーティングされたトークン数をカウントしています。
B2C、B2B、およびユースケース分類は集計値です。個別のチームやワークロードは特定されていません。
続きを読む
原文を表示
Every month, AI Gateway routes tens of trillions of tokens between production applications and AI labs, giving us visibility into what AI usage actually looks like, separate from leaderboards and benchmarks. We publish the data monthly in the AI Gateway production index.
May 2026 summary
Total AI Gateway tokens grew +20% MoM; total spend grew +43% MoM. Customers paid almost 20% more per token on average than in April.
DeepSeek’s share of tokens jumped from under 1% to 17% in a single month, while its share of spend stayed near 1%.
Anthropic’s share of spend grew from 61% to 65% in May, holding 70–80% of spend across every high-stakes use case (AI app generation, back office agents, and coding agents).
Cost-consciousness meant smarter routing between low-cost and frontier models. Customers got more deliberate about which model did which work, while overall usage kept climbing.
Last month, headlines about blown token budgets dominated tech news: Uber burned through its annual Claude Code budget shortly after Q1 and Amazon shut down KiroRank to curb unproductive tokenmaxxing. While runaway cost is a real problem, this month’s report shows that spend on production use cases still increased.
Two insights emerged from AI Gateway data in May:
Low-cost models entered production: New models shipped at price points that made the established labs look even more expensive, and they are capable enough to enter the mix in production.
Spend is increasing, but with smarter model mixes: Teams are still increasing token budgets, but they are implementing smarter routing strategies to get more value out of every dollar.
Low-cost models saw significant production volume for the first time
From February to April, volume distribution across labs on AI Gateway changed slowly, but in May, DeepSeek V4's launch completely shifted token share. The low-cost end of the market that barely existed in April became AI Gateway’s third-largest provider by volume in May, without a significant impact on overall spend.
In April, DeepSeek accounted for less than 1% of AI Gateway tokens and less than 0.2% of spend. In May, its volume share jumped to 17% of tokens, putting it in third place, ahead of OpenAI. Almost all of the volume comes from two models: deepseek/deepseek-v4-flash and deepseek/deepseek-v4-pro, both released in May.
The spend picture tells the other half of the story. Even though DeepSeek’s token share grew to 17% in a single month, its cost share stayed near 1%.
DeepSeek V4 Flash launched at $0.14 input / $0.28 output per million tokens, roughly 20–50× lower than comparable Anthropic models and 8–12× lower than other value-tier flagships like Qwen 3.6 Plus and Kimi K2.6. With a savings gap that big, teams adopted V4 Flash quickly.
Price alone wouldn’t have shifted DeepSeek’s volume that much in a month, meaning teams testing DeepSeek V4 against their existing evals found the output good enough to ship, not just low-cost enough to try.
Value-tier models have always been available on AI Gateway, but have never captured token share at this scale, meaning DeepSeek V4 was the first model at its price point to clear the quality bar for production work.
Frontier labs continued to capture a majority of new spend
Even as the low-cost end of the market grew fastest in volume, the expensive end grew faster in dollars.
Anthropic’s token share grew from 26% to 32%, and its spend share from 61% to 65%. OpenAI’s token share held near 13%, but its spend share ticked up from 12% to 13% on a much larger total, so customers were paying more per OpenAI token in May.
The average token got more expensive in May, even with DeepSeek pulling the average down. That increase happened because the work that demands frontier models grew faster than the work that doesn’t. The AI coding agent use case shows the low-cost/frontier split most clearly:
DeepSeek drove 49% of the segment’s token volume, but only 4% of the cost.
Anthropic drove 28% of tokens and 70% of the cost.
Lower-cost models are now a meaningful part of production workflows, but frontier model use is still growing, driving the increase in overall spend.
The frontier is getting more expensive per token, and customers are still paying. Anthropic continues to lead on spend, taking 65% of all gateway spend in May, and 70–80% of spend across every high-stakes use case.
Cost discipline became a routing strategy
Increased overall spend showed that demand for AI continued to grow in May, but teams applied more precision to their budgets through routing. They sent the cheap, high-volume work to lower-priced models and used frontier models where quality mattered most. Slow adoption of Google's latest Flash model is a clear example.
Gemini 3.5 Flash launched in May at a higher price point than Gemini 3.0 Flash, but migration didn’t happen at scale. By month-end, 3.5 held only 7% of the Flash family’s tokens while 3.0 held 90%.
Compared to the rapid adoption of Gemini 3.1 Pro across February and March, slower migration to 3.5 Flash shows that teams happy with 3.0 Flash aren't willing to pay the higher cost yet.
Conclusion: Cost-effective, capable options mean smarter model mixes
This month's report signals increased pricing sensitivity in the market, even as overall spend and token volume grow. That means developers are looking for ways to get more out of every dollar.
Data revealed two optimization strategies:
Using DeepSeek's cheap, but capable V4 family for lower-risk, high-volume tasks
Choosing to delay model family upgrades until the ROI makes sense
Routing gives teams the ability to adjust their model mix, and budget, in real time as the labs compete for different layers of production AI workloads.
Appendix
Token vs cost share by B2B classification
B2B applications run fewer, more expensive calls, while B2C applications run many cheap ones. On a per-token basis, B2B cost roughly 60% more than B2C in May.
Agent tool use across tokens and requests
Just under a quarter of requests end in a tool call, but those requests carry well over half of all tokens. Both metrics are roughly flat month-over-month.
Model diversity distribution by request volume
The more requests an app serves, the more models it runs in production. Single-model setups dominate the lowest-volume tier, while at 1M+ requests the majority of apps route across 11 or more models.
Cost vs volume share by use case
Use case cost share indicates how expensive a wrong answer is, not how many tokens it burns. Personal assistants and coding agents run cheap per token, while back-office and recruiting work costs far more.
Previous reports
Read the April 2026 AI Gateway production index.
About this data
This analysis is based on anonymized, aggregate routing data from the Vercel AI Gateway through May 2026.
A few notes on measurement:
Spend uses market-rate pricing (published list price) to provide a normalized view across teams that bring their own API keys.
Volume counts tokens routed through AI Gateway.
B2C, B2B, and use-case classifications are aggregate. No individual team or workload is identified.
Read more
同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み