アリババと深思考、中国の AI モデル競争を低コストへ
本文の状態
日本語全文を表示中
詳細モードで約8分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
AI News
アリババが超大規模モデルQwen3.8-Maxを、DeepSeekは低コスト推論に特化したV4-Flashを発表し、中国AI業界における性能と価格の競争激化を招いている。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月7日 22:27
AI深層分析
キーポイント
アリババの新モデル発表と技術的特徴
アリババは最大規模となるQwen3.8-Maxを発表し、2.4兆パラメータのうち950億個のみを活性化させるMixture-of-Expertsアーキテクチャを採用してコスト削減を実現した。
DeepSeekの低価格戦略と推論効率
DeepSeekはV4-Flashで2840億パラメータのうち130億個のみを活性化し、入力トークンあたり0.14ドルという競合他社を圧倒する低価格設定で市場に参入した。
Moonshot AIとの価格競争とベンチマーク
アリババのQwen3.8-MaxはMoonshot AIのKimi K3より安価だが、DeepSeekのV4-Flashはベンチテストあたりのコストを数セントに抑え、OpenAIやAnthropicのモデルを大幅に下回っている。
性能評価とランキングでの位置づけ
Qwen3.8-Maxは中国テキストモデルでArena.AIのトップにランクインしたが、Anthropic製モデルには依然として劣っており、画像分析でも同社製品に次ぐ2位にとどまっている。
1 トンあたりの価格だけでは実際のタスクコストは判断できない
モデルが生成する出力量や必要な対話回数によって、広告された API レートよりも総コストが高くなる可能性がある。
重要な引用
Alibaba has launched Qwen3.8-Max, its largest AI model to date
DeepSeek uses a similar sparse architecture at a smaller scale
V4-Flash costs $0.14 per million input tokens and $0.28 per million output tokens
Qwen3.8-Max moved to the top position among Chinese text models on crowdsourced comparison platform Arena.AI
編集コメントを表示
編集コメント
中国AI業界における価格競争の激化は、開発コストの低下を通じて大規模モデルの利用をより広範な層に開放する可能性を秘めている。技術的なアーキテクチャの最適化が、単なるパラメータ数の増加よりも実用的な価値を生み出す鍵となっていることが明確になった。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
アリババは、これまでで最大の AI モデル「Qwen3.8-Max」を公開しました。一方、ディープシーク(DeepSeek)の最新モデル「V4-Flash」は、競合他社よりも低い推論価格が注目されています。
Qwen3.8-Max は 2.4 兆個のパラメータを持ち、エキスパート混合(Mixture-of-Experts)アーキテクチャを採用しています。この仕組みでは、リクエストごとにモデルの一部のみを活性化します。アリババによると、一度に活性化するパラメータ数は約 950 億個で、フルモデルを起動する場合と比べてコスト削減と応答遅延の短縮が実現されています。
ディープシークも同様のスパースアーキテクチャを採用していますが、規模はより小ぶりです。Artificial Analysis のデータによると、V4-Flash は総パラメータ数が 2840 億個で、推論時に活性化する数は 130 億個です。一方、ムーンショット AI(Moonshot AI)の「Kimi K3」は総パラメータ数 2.8 兆個、活性化数は約 1040 億個となっています。
Qwen3.8-Max はテキスト、画像、動画の処理に対応し、最大 100 万トークンのコンテキストをサポートします。アリババによれば、このモデルはソフトウェアエンジニアリングプロジェクトを 16 日間にわたって完遂した実績もあります。
その規模は、ムーンショット AI が 7 月に発表した Kimi K3 に匹敵するものです。両社は価格面でも競合しており、Qwen3.8-Max の料金は入力トークン 100 万個あたり 2 ドル、出力トークン 100 万個あたり 6 ドルです。これに対し、Kimi K3 はそれぞれ 3 ドル、15 ドルとなっています。
モデルのサイズだけで推論コストが決まるわけではありません。アーキテクチャや活性化するパラメータ数、トークンの消費量、タスク完了に必要な呼び出し回数など、複数の要素が運用コストに影響します。
中国のテキストモデルを対象としたクラウドソーシング比較プラットフォーム「Arena.AI」において、Qwen3.8-Max がリリース直後にトップに躍り出ました。ただし、全体ランキングではいくつかの Anthropic 製モデルにはまだ及びません。また、画像や視覚資料を分析するモデルの分野でも、Anthropic の Claude Fable 5 系の変種に次ぐ第 2 位を獲得しています。
DeepSeek が推論コストを引き下げる
DeepSeek は V4-Flash で異なるアプローチを採用しました。Alibaba や Moonshot AI の最新モデルと同規模を目指すのではなく、広く利用されている多くの AI システムよりも低価格で提供しています。
Artificial Analysis によると、V4-Flash の料金は入力トークン 100 万個あたり 0.14 ドル、出力トークン 100 万個あたり 0.28 ドルです。同調査ファームは、このモデルのコンテキストウィンドウが 100 万トークンで、パラメータ総数は 2,840 億に達すると報告しています。そのうち推論時に実際に動作するのは 130 億パラメータのみです。
Artificial Analysis は、V4-Flash の Max Effort バージョンにおけるキャッシュヒット時の料金も公開しています。これはトークン 100 万個あたり 0.003 ドルで、標準的な入力料率の 98% 引きです。キャッシュされた入力は、以前に処理済みのコンテキストを指し、その後のリクエストでも再利用可能です。
DeepSeek の低価格なトークンレートは、Artificial Analysis が実施したベンチマークテストの結果にも反映されました。Reuters は、同調査ファームが V4-Flash の平均コストを 1 テストあたり 3 セントと推定していることを報じています。これに対し、Kimi K3 は 86 セント、OpenAI の GPT-5.6 Sol は 1.86 ドル、Anthropic の Claude Fable 5 は 3.15 ドルとなっています。
各モデルがベンチマークを完了するために使用する入出力の量を考慮した比較です。トークンあたりの料率が低くても、生成される出力量が多かったり追加の対話が必要だったりすれば、必ずしもタスク全体の費用が安くなるわけではありません。
Artificial Analysis は、DeepSeek V4-Flash の Max Effort 推論版に知能指数で 40 点を付けました。同社のテストでは、出力レートも約 118 トークン/秒を記録しています。
トークン価格だけではコストの全体像は見えません
Moonshot AI の Kimi K3 は、広告された API 価格と実際の長文タスク完了コストが異なる例として挙げられます。Artificial Analysis によると、Kimi K3 の料金は入力トークン 100 万あたり 3 ドル、出力トークン 100 万あたり 15 ドルで、キャッシュされた入力は 100 万あたり 0.30 ドルです。
エージェントによる知識作業を評価する Artificial Analysis の「AA-Briefcase」ベンチマークでは、Kimi K3 はタスクあたりの平均コストが 10.57 ドルとなりました。このモデルは 1 タスクあたり約 12 万トークンの出力を生成し、平均 83 回の対話ターンを要しました。
Artificial Analysis は、このコストは Kimi K3 のトークン価格に加え、出力量とモデルの呼び出し回数が反映された結果だと説明しています。モデルへの繰り返し呼び出しや大量の出力が発生すれば、広告上の API レートが示すよりもタスク完了の総費用が高くなる可能性があります。
Kimi K3 はテスト時点での AA-Briefcase 評価で Claude Fable 5 に次ぐ第 2 位の総合スコアを記録しました。また、Artificial Analysis の広範な知能指数でも 57 点を獲得しています。
DeepSeek との比較から、標準的な API 価格だけでなく「タスクあたりのコスト」を測る指標がなぜ重要なのかが見えてきます。同じタスクを実行する際でも、アーキテクチャや利用パターンが異なるモデルでは、必要な計算資源やトークン数が大きく異なるからです。
オープンウェイトによる展開オプションの拡大
中国の開発者が自社のモデルをどう配布するかという点も、コスト形成に大きな影響を与えています。アリババ、DeepSeek、Moonshot AI は、ホスト型 API アクセスを提供し続ける一方で、オープンウェイト版のリリースも継続的に支援しており、開発者にはモデル展開の方法を選ぶ余地が生まれています。
Artificial Analysis によると、DeepSeek V4-Flash は MIT ライセンスの下でオープンウェイトモデルとしてリストされており、Hugging Face から重み(weights)を入手可能です。また、Kimi K3 も Moonshot AI 独自のライセンスのもとでオープンウェイト版が利用できます。
オープンウェイト版であれば、開発者は特定のベンダーに依存した推論サービスだけでなく、自社のインフラやサードパーティのプロバイダ上でモデルを実行できるようになります。展開コスト自体は使用するハードウェアやインフラに左右されますが、モデルへのアクセスが単一のホスト API に縛られることはありません。
このアプローチは、OpenAI、Anthropic、Google が提供する主要モデルとは対照的です。これらの企業は通常、モデルの重みをクローズドに保っています。
Omdia のシニアアナリストである Lian Jye Su は、多くの業務ワークロードにおいて、モデル選定が「最高性能のシステムへのアクセス」だけによって決まるわけではないと指摘しています。
「多くのビジネスワークフローには、業界最高峰のモデルは必要ない」と蘇氏は語る。「必要なのは、十分高性能で手頃な価格、透明性があり、アクセスしやすいモデルだ。オープンウェイトモデルがその需要に応える。」
(写真:Solen Feyissa 氏提供)
関連記事:アリババはエージェントを中心に AI チップを設計中であり、それが競争の本質を変えている

業界のリーダーから AI やビッグデータについてさらに学びたい方は、アムステルダム、カリフォルニア、ロンドンで開催される「AI & Big Data Expo」にご参加ください。この包括的なイベントは TechEx の一部であり、サイバーセキュリティ&クラウドエキスポなど他の主要なテックイベントと併催されます。詳細はこちら。
本記事「アリババ、DeepSeek が中国の AI モデル競争を低コスト化へ」は、AI News に最初に掲載されました。
原文を表示
Alibaba has launched Qwen3.8-Max, its largest AI model to date, as DeepSeek’s latest V4-Flash model draws attention for inference pricing that is lower than several competing systems.
Qwen3.8-Max has 2.4 trillion parameters and uses a mixture-of-experts architecture, which activates only part of the model for each request. Alibaba said around 95 billion parameters are active at a time, reducing costs and response delays compared with activating the full model.
DeepSeek uses a similar sparse architecture at a smaller scale. Artificial Analysis lists V4-Flash at 284 billion total parameters, with 13 billion active during inference, while Moonshot AI’s Kimi K3 has 2.8 trillion total parameters and about 104 billion active.
Qwen3.8-Max can process text, images, and video and supports up to one million tokens of context. Alibaba also said the model completed a software engineering project over 16 days.
Its size places it close to Kimi K3, which Moonshot AI released in July. The two companies are also competing on price, with Qwen3.8-Max costing $2 per million input tokens and $6 per million output tokens, compared with $3 and $15, respectively, for Kimi K3.
Model size alone does not determine inference cost. Architecture, active parameter count, token consumption, and the number of calls required to complete a task also affect how much a model costs to run.
Qwen3.8-Max moved to the top position among Chinese text models on crowdsourced comparison platform Arena.AI following its release, although it remained behind several Anthropic models in the overall rankings. It also ranked second on Arena.AI’s leaderboard for models that analyse images and other visual material, behind an Anthropic Claude Fable 5 variant.
DeepSeek pushes down inference pricing
DeepSeek has taken a different approach with V4-Flash. Rather than matching the overall scale of Alibaba’s and Moonshot AI’s latest models, it has priced the model below several widely used AI systems.
V4-Flash costs $0.14 per million input tokens and $0.28 per million output tokens, according to Artificial Analysis. The research firm lists the model with a one-million-token context window and 284 billion total parameters, of which 13 billion are active during inference.
Artificial Analysis lists cache-hit pricing of $0.003 per million tokens for the Max Effort version of V4-Flash, 98% below its standard input rate. Cached input covers previously processed context that can be reused across subsequent requests.
DeepSeek’s lower token rates also carried through to Artificial Analysis’ benchmark testing. Reuters reported that the research firm estimated V4-Flash’s average cost at three cents per test, compared with 86 cents for Kimi K3, $1.86 for OpenAI’s GPT-5.6 Sol, and $3.15 for Anthropic’s Claude Fable 5.
The comparison accounts for the amount of input and output each model uses to complete the benchmark. A lower per-token rate does not necessarily result in a lower task cost if a model generates more output or requires additional interactions.
Artificial Analysis gave the Max Effort reasoning version of DeepSeek V4-Flash a score of 40 on its Intelligence Index. The research firm also recorded an output rate of about 118 tokens per second during testing.
Token prices tell only part of the cost story
Moonshot AI’s Kimi K3 provides another example of how advertised API prices can differ from the cost of completing longer workloads. Artificial Analysis lists the model at $3 per million input tokens and $15 per million output tokens, with cached input priced at $0.30 per million tokens.
On Artificial Analysis’ AA-Briefcase benchmark for agentic knowledge work, Kimi K3 averaged $10.57 per task. It generated around 120,000 output tokens and used an average of 83 turns per task.
Artificial Analysis said the cost reflected Kimi K3’s token pricing, output volume, and number of model interactions. Repeated model calls and larger outputs can therefore raise the total cost of completing a workload beyond what the headline API rate suggests.
Kimi K3 recorded the second-highest overall score on the AA-Briefcase evaluation at the time of testing, behind Claude Fable 5. It also scored 57 on Artificial Analysis’ broader Intelligence Index.
The comparison with DeepSeek shows why cost-per-task measurements add useful context to standard API pricing. Models with different architectures and usage patterns can consume substantially different amounts of compute and tokens while working through the same type of task.
Open weights add another deployment option
Cost is also being shaped by how Chinese developers distribute their models. Alibaba, DeepSeek, and Moonshot AI have continued to support open-weight releases alongside hosted API access, giving developers more options for how the models are deployed.
Artificial Analysis lists DeepSeek V4-Flash as an open-weight model under an MIT licence, with weights available through Hugging Face. Kimi K3 is also available as an open-weight model under Moonshot AI’s own licence.
Open weights allow developers to run models on their own infrastructure or through third-party providers instead of relying solely on a developer-hosted inference service. Deployment costs still depend on the hardware and infrastructure used, but access to the model is not tied to a single hosted API.
The approach differs from the main models offered by OpenAI, Anthropic, and Google, which generally keep their model weights closed.
Lian Jye Su, chief analyst at Omdia, said model selection for many business workloads does not depend solely on having access to the highest-performing system.
“Many business workflows do not need the industry’s very best model,” Su said. “They need models that are good enough, affordable, transparent and accessible, and open-weight models help meet that demand.”
(Photo by Solen Feyissa)
See also: Alibaba is designing AI chips around agents, and that changes what the race is actually about

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
The post Alibaba, DeepSeek push China’s AI model race towards lower costs appeared first on AI News.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み