中国製オープン大規模モデル3種を比較
Moonshot AI、DeepSeek、Zhipu AI の 3 つの中国発オープンモデルがトリリオンパラメータ規模で競合し、ベンチマークとコスト構造における新たな基準を提示した。
キーポイント
トリリオン級 MoE モデルの台頭
Kimi K3(2.8T)、DeepSeek V4 Pro(1.6T)、GLM-5.2(744B)の 3 つがオープンウェイトリーダーボードを席巻し、すべて百万トークンのコンテキストウィンドウと長期的コーディング・エージェントワークロードに対応している。
ベンチマークにおける性能格差
Artificial Analysis の中性指標では Kimi K3 が 57 点で他社をリードし、DeepSeek V4 Pro は 44 点、GLM-5.2 は 51 点であり、コーディングベンチでも K3 が GLM-5.2 を大幅に上回っている。
アーキテクチャとコストの多様性
各社は異なる MoE 構成(K3 は 896 中 16 活性化、V4 Pro は 49B アクティブ)を採用し、Vision/Video 対応の有無や推論モードの違いにより、用途とコスト効率の最適解が分かれている。
性能とコストのトレードオフ
Kimi K3 は測定された能力で最も強力だが、DeepSeek V4 Pro はコードタスクに特化し、圧倒的な低コスト(1ドルで約1.15Mトークン)を提供する。
ライセンスと利用形態の違い
DeepSeek と GLM-5.2 は MIT ライセンスで即時商用・自己ホスト可能だが、Kimi K3 の重みは 2026 年 7 月まで非公開で API 利用のみに限定される。
自己ホストの実用性
GLM-5.2 は VRAM 要件が大きいものの実装可能だが、K3 と DeepSeek V4 Pro はそれぞれ 64 基以上のアクセラレータや巨大なメモリを必要とし、多くのチームにはローカル展開が困難である。
Kimi K3 の性能と制限
人工知能分析インデックスで最高位を獲得しているが、API のみで重みのダウンロードは7月27日まで不可であり、出力価格も競合の5倍から17倍高い。
重要な引用
Three Chinese labs now hold the top of the open-weight leaderboard.
Kimi K3 is a 2.8-trillion-parameter Stable LatentMoE model activating 16 of 896 experts per token.
On that index, Kimi K3 scores about 57... GLM-5.2 held the top open-weight spot until K3 shipped.
"DeepSeek V4 Pro is the cost leader by a wide margin. At list output rates, one dollar buys roughly 1.15M output tokens from V4 Pro..."
"Kimi K3 is heaviest: Moonshot recommends 64 or more accelerators, putting local serving out of reach for most teams."
Kimi K3 leads, but at 5x to 17x the output price and no downloadable weights until July 27.
影響分析・編集コメントを表示
影響分析
この比較は、中国 AI ラボがオープンソース分野で急速に技術的優位性を確立し、米国大手モデルと互角以上に戦えるレベルに至ったことを示しています。特にトリリオンパラメータ規模でありながら MoE 技術により推論コストを管理可能な点は、企業における大規模 LLM の実装戦略や、オープンソースエコシステムの競争構造に大きな影響を与えるでしょう。
編集コメント
2026 年という未来のリリース日付が記載されている点は、この記事が将来シナリオまたは予測ベースの分析であることを示唆していますが、技術的な比較軸(MoE、コンテキストウィンドウ、コスト)は現在の業界トレンドを鋭く反映しています。中国勢によるオープンモデルの高性能化は、開発者がローカル環境やカスタムインフラで高機能 AI を構築する選択肢を劇的に広げるでしょう。
中国のテック企業 3 社が、オープンウェイトモデルのリーダーボードを席巻しています。Moonshot AI の「Kimi K3」、DeepSeek の「V4 Pro」、そして Zhipu AI の「GLM-5.2」はすべて、百万トークン単位のコンテキストウィンドウを持つスパースな Mixture-of-Experts(MoE)モデルです。これらはすべて、長期にわたるコーディングタスクやエージェントワークロードの処理を目的としています。
本記事では、AI チームが実際に意思決定を行う際の 3 つの軸——測定された能力、ライセンス条項、そして運用コスト——に基づいて、これら 3 モデルを比較します。
「トリリオンパラメータ」という表現は、Kimi K3(2.8T)と DeepSeek V4 Pro(1.6T)に当てはまります。一方、GLM-5.2 は総パラメータ数が 744B で、3 つの中では最も小型です。しかし、Kimi K3 が登場する以前からオープンウェイト分野をリードしてきた実績があるため、このリストへの選出は妥当です。
(function(){
var frame=document.getElementById("mtpc-embed-frame");
window.addEventListener("message",function(e){
if(e&&e.data&&e.data.mtpcHeight&&frame){
frame.style.height=e.data.mtpcHeight+"px";
}
});
})();
3 つの主要候補
Kimi K3 は、2.8 兆パラメータを持つ「Stable LatentMoE」モデルです。トークンごとに 896 個ある専門家のうち、16 個が活性化されます。Moonshot AI は正確なアクティブパラメータ数については公開していません。Kimi K3 はネイティブの視覚機能と 100 万トークンのコンテキストウィンドウ、そして常時稼働する推論機能を備えています。Moonshot AI はこれを「オープン化された最初の 3T クラスモデル」と称しています。
(当社のローンチ記事はこちら)
DeepSeek V4 Pro は、1.6 兆パラメータの MoE モデルで、アクティブなパラメータ数は 490 億です。384 のルーティング済みエクスパートと 1 つの共有エクスパートを採用しています。コンテキストウィンドウは 100 万トークン、最大出力長は 38.4 万トークンをサポート。より軽量な V4 Flash バリアント(総パラメータ 2,840 億、アクティブ 130 億)も用意されており、コスト効率を重視するワークロードに対応します。モデルの重みは Hugging Face で公開されています。
GLM-5.2 は約 7,440 億パラメータの MoE モデルで、アクティブ数は約 400 億です。コンテキストウィンドウも 100 万トークンに対応します。Zhipu AI が提供するこのモデルには、「High」と「Max」の推論モードが用意されており、API 経由での利用も可能です。
Kimi K3、DeepSeek V4 Pro、GLM-5.2 の主要スペックを比較すると以下のようになります。
| モデル | Kimi K3 | DeepSeek V4 Pro | GLM-5.2 |
|---|---|---|---|
| 総パラメータ数 | 2.8T | 1.6T | 744B(Artificial Analysis によると 753B) |
| アクティブパラメータ数 | 非公表(16/896 エクスパート) | 490 億 | 約 400 億 |
| コンテキストウィンドウ | 1M | 1M(最大出力 38.4K) | 1M(最大出力 13.1K) |
| モダリティ | テキスト+ビジョン+ビデオ | テキストのみ | テキストのみ |
| リリース日 | 2026 年 7 月 16 日 | 2026 年 4 月 24 日 | 2026 年 6 月 13 日 |
ベンチマーク結果について
ベンダーが報告するスコアは評価環境(ハルネス)が異なるため、各ラボ間で数値を単純比較することは困難です。中立な指標として「Artificial Analysis Intelligence Index」が挙げられます。このインデックスでは、3 つのモデルを同一の評価セットで測定しています。
同インデックスでのスコアは、Kimi K3 が約 57、DeepSeek V4 Pro(Max 推論モード)が 44、GLM-5.2 が 51 です。総合ランキングでは K3 は 3 位に位置し、Claude Fable 5 や GPT-5.6 Sol に次ぎ、Opus 4.8 や GPT-5.5 と同等の性能を誇ります。なお、Kimi K3 のリリースまで、オープンウェイトモデルとしての最高位は GLM-5.2 が維持していました。
コーディング関連のベンチマークでも同様の傾向が見られますが、いくつか注意が必要です。Moonshot 社自身が公開した表では、K3 と GLM-5.2 を同一の評価環境で比較しています。その結果、K3 はすべての共通ベンチマークにおいて GLM-5.2 を大幅に上回っています。
ベンチマーク結果(Moonshot ハーネス)
| モデル | DeepSWE | Program Bench | Terminal Bench 2.1 | FrontierSWE | SWE Marathon | Automation Bench | GPQA-Diamond |
|---|---|---|---|---|---|---|---|
| Kimi K3 | 67.5 | 77.8 | 88.3 | 81.2 | 42.0 | 30.8 | 93.5 |
| GLM-5.2 | 46.2 | 63.7 | 82.7 | 67.3 | 13.0 | 12.9 | 91.2 |
Moonshot の公式テーブルには DeepSeek が含まれていないため、同社の数値は別テストからのものです。DeepSeek-V4-Pro-Max は SWE-bench Verified で 80.6% を記録し、リリース時点でオープンウェイトモデルとして最高スコアをマークしました(Gemini 3.1 Pro と並ぶ)。また MRCR 1M では 83.5 を達成し、強力な長文脈処理能力を実証しています。一方、GLM-5.2 は SWE-bench Pro で 62.1 を記録し、GPT-5.5(58.6)をわずかに上回っています。
総合的な測定能力においては K3 が三モデル中最も優れています。DeepSeek V4 Pro は特定のコーディングタスクにおいて競争力がありますが、GLM-5.2 は K3 に劣るものの、依然として有能なオープンウェイトオプションです。
ライセンス
三モデルとも「オープンウェイト」として公開されていますが、現状の実際の利用状況には違いがあります。
DeepSeek V4 Pro は MIT ライセンスで、リリース初日から Hugging Face で重み(ウェイト)を入手可能です。GLM-5.2 も同様に MIT ライセンスであり、zai-org 組織の下で全重みが Hugging Face に公開されています。両モデルとも現在、商用利用の制限なく微調整やセルフホスティングが可能です。
例外は Kimi K3 です。Moonshot は 2026 年 7 月 27 日までに重みの公開を約束しており、その際は Modified MIT ライセンスが適用される見込みです。それまでは API と Kimi アプリ経由での利用のみとなります。Moonshot が最近発表した Modified MIT の条件には、1 つの帰属表示条項が含まれています。これは月間アクティブユーザー数が 1 億人を超えた場合に発動するものです。
提供コスト
API の価格設定は、これらのモデルを明確に区別しています。
| モデル | 入力 ($/MTok) | 出力 ($/MTok) | キャッシュ済み入力 |
|---|
Kimi K3、DeepSeek V4 Pro、GLM-5.2 の 3 つのオープンなトリリオン規模 MoE モデルを、ベンチマーク結果、ライセンス、推論コストの観点から比較します(続き 4/5)。
まずコスト面では、DeepSeek V4 Pro が圧倒的なリーダーです。リスト価格での出力レートを見ると、1 ドルで DeepSeek V4 Pro は約 115 万トークン、GLM-5.2 は約 22.7 万トークン、Kimi K3 は約 6.7 万トークンを購入できます。
Artificial Analysis は、ベンダーの主張を排するために、キャッシュ・入力・出力の比率を 7:2:1 に統一したブレンド価格で各モデルを評価しています。この基準では、Kimi K3 が 100 万トークンあたり 2.31 ドル、GLM-5.2 が 0.90 ドル、DeepSeek V4 Pro はわずか 0.18 ドルとなります。タスクあたりのコストで見ても、同様のデータによると Kimi K3 は 0.94 ドル、GLM-5.2 は 0.32 ドル、DeepSeek V4 Pro は 0.04 ドルです。
速度にも差があります。Artificial Analysis の測定では、GLM-5.2 は秒間約 168 トークンの処理速度を記録し、DeepSeek V4 Pro と Kimi K3(ともに秒間約 62 トークン)を大きく引き離しています。Moonshot によると、コーディングタスクでのキャッシュヒット率は 90% を超えており、これにより Kimi K3 の実質的な入力コストは 100 万トークンあたり 0.30 ドルまで低下します。
自己ホスト(オンプレミス)の場合は制約が異なります。744B パラメータの GLM-5.2 は、BF16 精度で VRAM を 1TB 以上必要とし、FP8 精度でも H200 グラフィックボードを約 8 枚分必要とします。DeepSeek V4 Pro(1.6T)はさらに多くのリソースを要します。最も重いのは Kimi K3 で、Moonshot は 64 基以上のアクセラレータの使用を推奨しており、ほとんどのチームがローカル環境での運用を断念せざるを得ません。なお、Kimi K3 はより幅広いハードウェアに対応するため、MXFP4 形式の重みと MXFP8 形式の活性化値を採用しています。
各モデルに最適な用途は?
コード品質を維持しつつ、1 トークンあたりのコストを最小化したいなら、DeepSeek V4 Pro が明確な選択肢です。重み(weights)のダウンロードが可能でライセンスもクリア、さらに出力価格が他 2 社を下回っています。
測定された能力の高さでは Kimi K3 が首位ですが、その価格は他社の 5 倍から 17 倍に達し、かつ 7 月 27 日までは重みのダウンロード不可です。GLM-5.2 はその中間に位置します。K3 よりも安価で、両社よりも高速、今日からセルフホスト可能、そしてサイズに見合わない高い性能を備えています。
検証の深さやライセンスの明確さを重視して選定するなら、今すぐ DeepSeek と GLM を検討すべきです。ベンチマークスコアの最大化を目指す場合は、K3 の重み公開を待つか、API 利用料の高いプランを選ぶことになります。
Key Takeaways
- Kimi K3 は Artificial Analysis Intelligence Index で首位(約 57 点、総合 3 位)ですが、7 月 27 日までは API のみの提供です。
- DeepSeek V4 Pro はコストリーダー。1 タスクあたり約 0.04 ドル、1 ドルで約 115 万トークンの出力が可能です(定価ベース)。
- GLM-5.2 (744B) は最小サイズでありながら最速(約 168 トークン/秒)、かつ MIT ライセンスの下で今日からセルフホスト可能です。
- 3 社とも 100 万トークンのコンテキスト長に対応しますが、現在オープンな重みが利用可能なのは DeepSeek と GLM のみです。
本記事「Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost」は、MarkTechPost に掲載されたものです。
原文を表示
Three Chinese labs now hold the top of the open-weight leaderboard. Moonshot AI’s Kimi K3, DeepSeek V4 Pro, and Zhipu AI’s GLM-5.2 are all sparse Mixture-of-Experts (MoE) models with million-token context windows. Each targets long-horizon coding and agent workloads. This article compares them on three axes an AI team actually decides on: measured capability, license terms, and serving cost.
‘Trillion-parameter’ fits Kimi K3 (2.8T) and DeepSeek V4 Pro (1.6T). GLM-5.2 is 744B total, so it is the smallest of the three by total parameters. It earns its place because it led the open-weight field before K3 shipped.
(function(){
var frame=document.getElementById("mtpc-embed-frame");
window.addEventListener("message",function(e){
if(e&&e.data&&e.data.mtpcHeight&&frame){
frame.style.height=e.data.mtpcHeight+"px";
}
});
})();
The three contenders
Kimi K3 is a 2.8-trillion-parameter Stable LatentMoE model activating 16 of 896 experts per token. Moonshot has not published the exact active-parameter count. K3 adds native vision, a 1M-token context window, and always-on reasoning. Moonshot calls it the first open 3T-class model. Our launch coverage is here.
DeepSeek V4 Pro is a 1.6-trillion-parameter MoE with 49B active parameters, using 384 routed experts plus one shared expert. It carries a 1M-token context window with 384K max output. A smaller V4 Flash variant (284B total, 13B active) covers cheaper workloads. Weights are on Hugging Face.
GLM-5.2 is a 744-billion-parameter MoE with roughly 40B active parameters and a 1M-token context window. Zhipu ships it with High and Max reasoning modes. It comes with API access
SpecKimi K3DeepSeek V4 ProGLM-5.2
Total parameters2.8T1.6T744B (753B per Artificial Analysis)
Active parametersNot disclosed (16/896 experts)49B~40B
Context window1M1M (384K max output)1M (131K max output)
ModalityText + vision + videoTextText
ReleasedJuly 16, 2026April 24, 2026June 13, 2026
Benchmarks
Vendor-reported scores use different harnesses, so per-benchmark numbers rarely line up cleanly across labs. The neutral comparator is the Artificial Analysis Intelligence Index, which scores all three on the same suite.
On that index, Kimi K3 scores about 57, DeepSeek V4 Pro (Max reasoning) scores 44, and GLM-5.2 scores 51. K3 ranks #3 overall, behind only Claude Fable 5 and GPT-5.6 Sol, and comparable to Opus 4.8 and GPT-5.5. GLM-5.2 held the top open-weight spot until K3 shipped.
Coding benchmarks tell a similar story with caveats. Moonshot’s own table runs K3 and GLM-5.2 through matched harnesses. There, K3 leads GLM-5.2 on every shared benchmark by wide margins.
Benchmark (Moonshot harness)Kimi K3GLM-5.2
DeepSWE67.546.2
Program Bench77.863.7
Terminal Bench 2.188.382.7
FrontierSWE81.267.3
SWE Marathon42.013.0
Automation Bench30.812.9
GPQA-Diamond93.591.2
DeepSeek does not appear in Moonshot’s table, so its numbers come from separate testing. DeepSeek-V4-Pro-Max scores 80.6% on SWE-bench Verified, the highest open-weight result at its release and tied with Gemini 3.1 Pro. It also posts 83.5 on MRCR 1M, confirming serious long-context ability. GLM-5.2 scored 62.1 on SWE-bench Pro, edging GPT-5.5 at 58.6.
So, K3 is the strongest of the three on measured capability. DeepSeek V4 Pro is competitive on isolated coding tasks. GLM-5.2 trails K3 but remains a capable open-weight option.
License
All three ship as open-weight models, but the practical status differs today.
DeepSeek V4 Pro is MIT-licensed, with weights on Hugging Face from day one. GLM-5.2 is also MIT-licensed, with full weights on Hugging Face under the zai-org organization. Both allow unrestricted commercial use, fine-tuning, and self-hosting now.
Kimi K3 is the exception. Moonshot has committed to publishing weights by July 27, 2026, expected under a Modified MIT license. Until then, K3 is usable only through the API and Kimi apps. Moonshot’s recent Modified MIT terms add one attribution clause. It triggers only above 100M monthly active users.
Serving cost
API list pricing separates these models sharply.
ModelInput ($/MTok)Output ($/MTok)Cached input
Kimi K33.0015.000.30
DeepSeek V4 Pro0.4350.87~0.0036
GLM-5.21.404.400.26
DeepSeek V4 Pro is the cost leader by a wide margin. At list output rates, one dollar buys roughly 1.15M output tokens from V4 Pro, about 227K from GLM-5.2, and about 67K from K3.
Artificial Analysis prices every model on one blended 7:2:1 cache/input/output basis, which removes vendor framing. On that basis it lists K3 at $2.31 per 1M tokens, GLM-5.2 at $0.90, and DeepSeek V4 Pro at $0.18. On cost per task, the same source reports K3 at $0.94, GLM-5.2 at $0.32, and DeepSeek V4 Pro at $0.04.
Speed also differs. Artificial Analysis measures GLM-5.2 at about 168 tokens/sec, well ahead of DeepSeek V4 Pro and Kimi K3 at about 62 each. Moonshot reports above 90% cache hits in coding workloads, which drops K3’s effective input cost to $0.30 per million.
Self-hosting is a different constraint. GLM-5.2 at 744B needs over 1TB of VRAM in BF16, or roughly 8x H200 at FP8. DeepSeek V4 Pro at 1.6T needs more still. Kimi K3 is heaviest: Moonshot recommends 64 or more accelerators, putting local serving out of reach for most teams. K3 uses MXFP4 weights with MXFP8 activations for broader hardware support.
(function(){
var frame=document.getElementById("mtph-embed-frame");
window.addEventListener("message",function(e){
if(e&&e.data&&e.data.mtphHeight&&frame){frame.style.height=e.data.mtphHeight+"px";}
});
})();
Which model for which job
For lowest cost per token at strong coding quality, DeepSeek V4 Pro is the clear pick. Its weights are downloadable, its license is clean, and its output price undercuts both rivals.
For the highest measured capability, Kimi K3 leads, but at 5x to 17x the output price and no downloadable weights until July 27. GLM-5.2 sits between them: cheaper than K3, faster than both rivals, self-hostable today, and more capable than its size suggests.
If you are planning to choose based on verification depth and license clarity favor DeepSeek and GLM now. Buyers chasing peak benchmark scores wait for K3 weights or pay the API premium.
Key Takeaways
Kimi K3 leads the Artificial Analysis Intelligence Index (~57, #3 overall) but stays API-only until July 27.
DeepSeek V4 Pro is the cost leader: ~$0.04 per task and ~1.15M output tokens per dollar at list rates.
GLM-5.2 (744B) is the smallest yet fastest (~168 t/s) and self-hostable today under MIT.
All three ship 1M-token context; only DeepSeek and GLM have open weights available now.
The post Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost appeared first on MarkTechPost.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み