Z.ai、コーディング限定 GLM-5.3 を発表し API・Hugging Face 公開へ
本文の状態
日本語全文を表示中
詳細モードで約13分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Interconnects
Z.ai がコード生成に特化した GLM-5.3 を発表し、少数のパラメータ数で主要な米国製モデルや競合他社のモデルを上回る性能を示したことで、中国のAI研究が最先端と互角であることが示された。
AI深層分析を開く2026年8月15日 06:45
AI深層分析
キーポイント
GLM-5.3 の性能と公開スケジュール
Z.ai はコード生成プランに限定して GLM-5.3 を発表し、API や Hugging Face での公開を予定している。このモデルは Kimi K3 や Claude Fable 5 などの主要なベンチマークで上位のスコアを記録した。
効率的なパラメータ数と学習手法
GLM-5.3 は約 750B パラメータという比較的少数で、Kimi K3 の約 3 分の1の規模ながら最先端のコード生成ベンチマークに匹敵する。同社によると、この成果は事前学習ではなく、ポストトレーニングの拡張が主な要因である。
中国 AI 業界の競争力と歴史的背景
Z.ai は GLM モデルの開発を長年継続しており、その経験が現在の高性能モデルの実現に寄与している。この発表は、中国企業がいかにして米国の最先端モデルと互角の成果を出し続けているかという議論に新たな事実を提供する。
中国ラボのベンチマーク最適化戦略
米国の大手企業が公開までに数ヶ月のテスト期間を費やす間、中国のラボは同じ時間を活用してベンチマークで性能を向上させている。この迅速なリリースサイクルにより、実世界のパフォーマンスと論文上のスコアの乖離が生じている可能性がある。
強化学習中心の訓練アプローチ
Z.ai は単純な知識蒸留ではなく、多様な環境やタスクを用いた大規模な強化学習(RL)主導の訓練 regime を採用している。この手法は環境やインフラをスケーラブルに運用する能力が不可欠であり、単なるモデルの縮小では達成できない。
重要な引用
Scaling post-training is all we did for GLM-5.3.
This model looks exceptional, with a somewhat astounding increase in scores.
The determining factors are much more big picture than technical
Chinese labs use all the time that American labs do pre-release testing to keep hillclimbing on benchmarks
編集コメントを表示
編集コメント
Z.ai の GLM-5.3 は、パラメータ数と性能のバランスにおいて新たな基準を提示している。ポストトレーニング技術の進化が、大規模な事前学習に依存しないモデル開発の可能性を示唆しており、業界全体のパラダイムシフトを促す可能性がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
お断り:私は現在、出張中のため音声解説は作成できません。追記:メール送信後に中国のデータ産業に関する項目 5 を追加しました。
本日、Z.ai が GLM-5.3 モデルを発表しました。現時点ではコーディングプランでのみ利用可能ですが、API への展開は間もなく、2 週間後には Hugging Face でオープンウェイト版が公開される予定です。このモデルは非常に優秀で、スコアに驚くべき向上が見られます。多くのベンチマークにおいて Moonshot AI の Kimi K3 を上回り、一部では Claude Fable 5 や GPT-5.6-Sol も凌駕しています。

より詳細な比較は以下の通りです。

これにより、このモデルはエージェント型コーディングベンチの最前線に位置することになりました。パラメータ数は約 750B と非常に少なく、Kimi K3 のわずか 3 分の 1 です。Z.ai のブログ記事は非常に明快で、冒頭には力強い一文が掲げられています。
GLM-5.3 で我々が取り組んだのは、ポストトレーニングのスケーリングのみです。
GLM-5.3 は GLM-5.2 と同じベースモデルですが、ポストトレーニングの期間が大幅に延長されています。あえて大まかに言い換えれば、Z.ai は Kimi のような「事前学習の傑作」を作る企業とは異なり、ポストトレーニングにおいて強みを持っているようです。
今回のリリースを機に、「中国がいかにしてこれほどまでに最先端を追えているのか」「なぜこれほど小さなモデルが、米国トップクラスの公開モデルと互角に渡り合えるのか」といった議論が活発化しています。また、これらの成果は本当に信頼できるものなのかという疑問も投げかけられています。
最もシンプルな答えは、Z.ai が自らの分野で極めて卓越した能力を持っているということです。彼らがこの種のモデル開発に取り組んできたのは、業界のほぼ他社よりも長い期間にわたります。以下に GLM モデルシリーズの主な歴史をまとめます。
Zhipu AI 設立 – 2019 年
- GLM (General Language Model) – 2021 年 3 月 – THUDM(清華大学データマイニング/知識工学グループ)がリリース。重み公開済み。
- GLM-130B – 2022 年 8 月 – スケーリング版。GLM-130B から GLM-4 までの技術レポートおよび重みが公開されています。
- ChatGLM – 2023 年 3 月 14 日 – 初のチャット対応バージョン。重み公開済み。
- ChatGLM2 – 2023 年 6 月 25 日 – 重み公開済み。
- ChatGLM3 – 2023 年 10 月 27 日 – 重み公開済み。
- GLM-4 – 2024 年 1 月 16 日 – 「GLM」への名称統一。同年 6 月にオープンウェイトの GLM-4-9B が追加リリースされました。重み公開済み。
- GLM-5 – 2026 年 2 月 11 日 – 最新世代。重み公開済み。
今年6月22日に公開されたGLM-5.2は大きな話題を呼びました。リリースから数週間経った今でも、私が知るAI研究者たちはこのモデルを使い続けています。その理由の一つが速度です。一部のチームでは、パブリックな提供よりも高速な処理を実現するために、社内のクラスタにこのモデルをデプロイしています。もう一つの理由はシンプルさです。ロールバック機能がないなどの特徴は、フロンティア(最先端)AIシステムを開発する際に役立っています。GLM-5.2は、周囲の過剰な期待に応えるだけの実力を持っていました。
私も最初は同じように懐疑的でした。「どうして彼らはこれほど連続して成果を出せるのか?モデルが思われているほど優秀ではないはずだ」と考えました。米国の企業が圧倒的なリソースの優位性を持ちながら、能力面で差をつけられないという事実に違和感を覚えます。よく言われるのは「知識蒸留(distillation)」ですが、私はこれを主要な要因とは考えていません。以前にも詳しく書いた通りです。
最近ある論文で、最先端モデルから推論プロセスを抽出する単純な手法が紹介されました。これは中国のラボがスケールして活用できるような技術です。なぜ米国のラボはこうした挙動に対して素早く対策を講じないのでしょうか?むしろ政府に政策支援を求める動きがあります。私にはその理由が納得できません。
Z.aiのブログ記事は非常に直接的で、強化学習(RL)中心のトレーニング体制と合致しています。彼らは「より多くの環境、多様なタスク、そしてそれらに対する学習に投入した計算資源を増やした」と述べています。強化学習の環境、それをスケールして実行するためのインフラ、あるいは効果的に組み合わせるためのアルゴリズムを、単に「蒸留」するだけで実現できるわけではありません。
Interconnects AI は読者支援型の出版物です。購読をご検討ください。
では、中国のラボは蒸留なしでどうやってそれを成し遂げているのでしょうか?ベンチマーク最適化(benchmaxxing)を行っているのでしょうか?「ベンチマーク最適化」とは、モデルをテストセットに特化させることを指し、実世界でのパフォーマンスが紙上のスコアと大きく乖離する状態を意味します。この現象の決定要因は、技術的な詳細よりも大局的な要素にあります(もちろん技術的詳細も重要ですが、ラボ間でそれを明確に区別するのは難しいのです)。
Z.ai のモデルリリースまでの期間は数ヶ月ではなく、数日です。OpenAI や Anthropic が数ヶ月を要するのと対照的です。おそらく OpenAI と Anthropic は、Z.ai や Moonshot AI よりもはるかに優れた内部モデルを持っています。それでもこれらの米企業は、一般公開までに数ヶ月を要するため、中国のラボが最前線の採用判断において圧倒的に有利に働くことになります。つまり、中国のラボは、米国のラボがリリース前のテストに費やす時間をすべて活用してベンチマークで上位を維持し続けているのです(SpaceXAI もおそらくこの点では中国のラボに近い立場にあります)。進歩のスピードがこのように速い中で、これが中国のラボが最前線を維持し続ける最大の要因である可能性が高いです。これまでのところ、米国のラボにとってもこれは経済的に許容できる状況でした。なぜなら、彼らのモデルに対する需要は依然として莫大だったからです。
LLM を開発する各ラボでモデルの自己改善ループが加速する中、これらのフィードバックループにユーザーデータが必要となる場合、この迅速なリリースサイクルは中国のラボにとって圧倒的な優位性をもたらす可能性があります。その結果、次世代の大幅に優れたモデルが登場するまでの間、自社の製品が市場に長く残存し続けることで、競合他社への需要を削ぐことにもつながりかねません。
業界関係者の多くが懸念しているのはまさにこの競争の力学です。主要な能力範囲内で多くのラボが最先端モデルの開発を進めている現状では、こうした状況が近々収束するとは考えにくいです。
確かに Z.ai は、OpenAI や Anthropic に比べてパブリックベンチマークの結果により敏感に反応している可能性があります。Artificial Analysis Intelligence Index などの主要な集約指標で高いスコアを獲得することは、株価に直接的な影響を及ぼします。資金調達を継続し、チームの士気を維持するためには、こうした成果を示すことが不可欠です。アメリカの巨人たちと互角に渡り合う「不屈のアンダードッグ」という物語は、投資家や社員の心を掴むのに非常に効果的だからです。
ベンチマークでの微妙なスコア操作(ベンチマックス)が、必ずしも切迫した状況や同様の圧力から生じるわけではありません。これは驚くほど多くのラボで業界標準となっています。多くの企業が採用しているデータ獲得戦略の一つは、自らが立ち遅れているベンチマークのデータを積極的に購入することです。
Z.ai は、GLM-5.3 がベンチマークで破綻するほど無茶なテストを意図的に行っているわけではありません(少なくとも意図的なものではなく、検知されるようになっています)。現在、すべてのラボが強化学習のスケールアップにおける荒削りな課題に直面しています。Anthropic の Opus 5 や Sonnet 5 モデルは驚異的なベンチマークスコアを記録しているにもかかわらず、評判は賛否両論です。業界全体がこの状況にあるため、モデルの重みによっては使いやすさに差が生じるものの、リリースブログに記載されたベンチマークスコア自体は本物です。
GLM-5.3 は、Claude Fable や GPT Sol に比べて用途が限定されたモデルである可能性が高いです。GPT-5.2 がリリースされた際も、エージェントによるコーディング分野を除けば評価は分かれていました。同時に、OpenAI と Anthropic は莫大な規模のビジネスを支え、無数のユースケースで自社のモデルを活用しています。これは採用曲線の初期段階にある企業ならではのメリットであり、最も価値の高いユースケースに特化できる点です。ポストトレーニングにおいて、あえて範囲を絞ることで、最終的なモデルを組み立てる工程が格段に楽になります。
これは少し大げさに表現しましたが、Z.ai は堅牢なオンプレミス展開事業の成果として、年間収益(ARR)10 億ドルを達成したと報じられています。
同様に、フラッグシップの GLM モデルにはこれまで視覚機能がありませんでした。テキスト専用であることは Z.ai がより高い競争力を持つスコアを獲得する上で有利に働きますが、その分、競合環境は過酷です。一方、Inkling-Small のようにオムニモーダル(多様なモダリティに対応)な設計を目的としたモデルも存在します。
(追記)中国における強化学習(RL)データの産業が急成長しています。私たちがフォローしている多くの情報源や噂筋でも、この業界の活況が報じられています。その背景には、米国のデータ企業が中国のモデル開発ラボへ販売を行うという動きが大きく影響しています。
具体的には、米国の最先端ラボで利用されているのと同じ強化学習環境を中国のラボも多数購入し、それを用いて学習させたモデルをより早くリリースするといったケースが想定されます。この市場の規模や影響力についてはまだ不確実な点が多いものの、確実に重要性を増していることは間違いありません。
Z.ai は極めて能力の高い大規模言語モデル(LLM)開発組織です。OpenAI や Anthropic と比較しても、はるかに計算資源を効率的に活用できている可能性が高いと言えます。この点は繰り返し強調する価値があります。彼らはまさにプロフェッショナルであり、その業務において卓越したスキルを持っています。
同社は清華大学と非常に密接な関係にあり、中国の優秀なコンピュータサイエンティストの多くが在籍しています。こうした豊富で意欲的な人材プールは、西側の対抗組織と同様に、彼らの成功にとって不可欠な要素となっています。
総合的に見ると、GLM シリーズモデルにおいて実行している戦略は非常に優れたものと言えます。リリースおめでとうございます。重み(weights)の公開を心待ちにしており、より詳細なテストを行いたいと考えています(私は通常、Fireworks や Baseten といった米国のオープンウェイト推論サービスを利用しています)。
コメントを残す
これは、経済全体にわたって強力なサイバー能力が普及していく必然的な過程における、もう一つのステップです。Z.ai もこれを認識しており、以下のように述べています:
GLM-5.3 は、サイバーセキュリティタスクにおいて現時点で最も能力が高いモデルです。脆弱性の発見やエクスプロイト分析、複雑な多段階のセキュリティタスクにおいて大幅な改善が実現されています。これらの機能により、防御側は弱点を早期に特定し、リスクを検証して対策を加速させることが可能になります。
一方で、この技術には明確な二重利用(デュアルユース)のリスクも伴います。そのため、当社は段階的なリリースアプローチを採用しています。まず、選抜されたセキュリティパートナーが制御された環境で GLM-5.3 を評価します。その後、より広範なアクセス権限と API の提供を開始します。必要な安全性評価とリリース準備が完了次第、GLM-5.3 の完全なモデル重み(weights)を公開する予定です。
さらに、彼らはプラットフォーム上での推論実行を、リクエスト分類器や思考の連鎖(Chain of Thought)モニタリングを通じて監視していることを認めています。ただし、肝心なのは詳細部分であり、どの AI ラボがどこまでの実行レベルを持っているかは依然として不明です。この種の能力拡散は、最も低い共通基準によって決定されてしまいます。
結局のところ、真のオープンウェイト(重み公開)モデルが登場すれば、こうした安全対策の意味は薄れてしまいます。GLM-5.3 でなくても、別のモデルが現れるでしょう。これらの機能を持つモデルのサイズは時間とともに縮小し、改変やデプロイが容易になっています( safeguards を欠いた状態での利用も想定されます)。Z.ai は脆弱性発見の促進や積極的な管理など、正しい取り組みの一部を実行していますが、単独の企業がこれらをすべて処理できる状況にはほど遠いです。
この移行に備え、すべてのソフトウェア分野で直ちに準備を進めるためには、政府や産業界の連携による産業規模のガイドラインが必要です。
原文を表示
Housekeeping: I’m traveling so cannot make a voiceover for this post. EDIT — I added a bullet point 5 on the Chinese data industry after sending the email out.
Today, Z.ai announced their GLM-5.3 model, currently only available in the coding plan, coming soon to their API and in two weeks’ time to Hugging Face (open weights). This model looks exceptional, with a somewhat astounding increase in scores. On many benchmarks the model has surpassed Moonshot AI’s Kimi K3 and on some it’s surpassed Claude Fable 5 or GPT-5.6-Sol.

Here’s a more complete comparison:

This puts the model more or less at the frontier of agentic coding benchmarks, with only ~750B parameters – a third of Kimi K3! The Z.ai blog post is rather straightforward, and starts with a bold sentence:
Scaling post-training is all we did for GLM-5.3.
GLM-5.3 is the same base model as GLM-5.2 with substantially extended post-training. To risk a broad oversimplification, Z.ai seems to have a strength in post-training when compared to Kimi, which is more of a pretraining masterpiece. Following this release there have been a lot of discussions wondering how China can keep up so well? How can such a small model be matching the leading public American models? Are these results real?
Subscribe now
The simplest explanation is that Z.ai is very good at what they do – it’s worth recalling that they’ve been working on this line of models longer than almost anyone in the industry. Here’s a brief history of the GLM models.
Zhipu AI Founded – 2019
GLM (General Language Model) — March 2021 — released by THUDM, Tsinghua University’s Data Mining / Knowledge Engineering group. Weights
GLM-130B — August 2022 — Scaled version. Technical report for GLM-130B through GLM-4 — Weights
ChatGLM — March 14, 2023 — first chat version. Weights
ChatGLM2 — June 25, 2023 — Weights
ChatGLM3 — October 27, 2023 — Weights
GLM-4 — January 16, 2024 — rebranded as just GLM; open-weight GLM-4-9B followed in June. Weights
GLM-5 — February 11, 2026 — latest major generation. Weights
GLM 5.2, released on June 22 of this year, was a big deal – weeks after the release, I regularly heard from AI researchers I know who still used the model due to its speed (some deploy the model on internal clusters for faster speeds than public offerings) and simplicity (as a model with no rollbacks, etc., when working on frontier AI systems). GLM-5.2 altogether stood up to the hype.
I’ve been going through some of the same denial myself, thinking “how do they keep doing this? Surely the models aren’t as good as they look.” There’s something a bit off-putting with how the American companies have such a commanding resource lead, but can’t seem to pull away in capabilities. The common answer is distillation, which I’ve written at length about, but I deem not to be the major factor. On that note, there was a recent paper that showed simple methods for extracting the reasoning traces from frontier models – this is the sort of thing that Chinese labs could definitely use at scale. I’m confused why the labs in the U.S. haven’t patched this behavior faster; instead they’re running to the government asking for policy help. It doesn’t add up for me.
Z.ai’s blog is direct and matches with an RL-dominated training regime. They say they used “more environments, more diverse tasks, and more compute spent training on them.” One does not simply “distill” RL environments, infrastructure to run them at scale, or algorithms to mix them together effectively.
Interconnects AI is a reader-supported publication. Consider becoming a subscriber.
So, how do the Chinese labs do it if not distillation? Are they benchmaxxing? An accepted definition of benchmaxxing is focusing the model on the test sets, such that the real-world performance meaningfully differs from the on-paper scores. The determining factors are much more big picture than technical (yes, the technical details definitely matter, but are harder to differentiate from lab to lab):
The time to release for Z.ai is likely days, not months as with OpenAI or Anthropic. It is very, very likely that OpenAI and Anthropic have far better internal models than Z.ai and Moonshot AI. Still, these American companies tend to take months to release their models to the public, which massively flatters the Chinese labs in adoption decisions at the frontier. To put it simply – the Chinese labs use all the time that American labs do pre-release testing to keep hillclimbing on benchmarks (SpaceXAI is likely far closer to the Chinese labs here). With the pace of progress being so fast, this is likely the largest determining factor of why Chinese labs stay at the frontier. This, so far, has been economically acceptable for the American labs, as they’ve still had massive demand for their models.
As model self-improvement loops ramp up within the labs building LLMs, if any of these feedback loops require user data, this faster release cycle could massively favor the Chinese labs, giving their offerings longer lifespans before the next vastly superior model comes out, undercutting demand for their models.
These are very clearly the race dynamics that many in the industry worry about. With so many labs building frontier models in the envelope of leading capabilities, it is hard to see this abating in the near future.
Yes, Z.ai probably cares slightly more about public benchmarks than OpenAI or Anthropic. These benchmarks, e.g. scoring highly on the Artificial Analysis Intelligence Index, or similar aggregators, have a very direct impact on their stock price. They in many ways need to do this to keep raising capital and maintain team morale, as being the scrappy underdog matching American giants is a wonderful story.
Subtle benchmaxxing does not need to come out of desperation or any similar pressures. It’s the industry standard across a remarkable number of labs. Many companies’ data acquisition strategy is to buy data on the benchmarks they’re behind on.
Z.ai is not benchmaxxing to the point where GLM-5.3 is fried (at least not intentionally, and they’ll check for it). Every lab is dealing with the rough edges of scaling RL right now. Anthropic’s Opus 5 and Sonnet 5 models have very mixed reputations, despite the incredible benchmark scores. Everyone in the industry is in the same boat, so some model weights end up being easier to use than others, but the benchmark scores in their release blogs are the real deal.
GLM-5.3 is likely a narrower model than Claude Fable or GPT Sol. When GPT-5.2 was released, it had mixed reviews outside of agentic coding. At the same time, OpenAI and Anthropic support very large businesses with countless use-cases for their models. This is a benefit of being a company earlier in their adoption curve – you can target the most valuable use-cases. Within post-training, caring about a bit less will make assembling the final model far easier.
I’m overstating this a bit, as Z.ai reportedly reached $1B of ARR on the back of a strong on-premises deployment business.
Similarly, the flagship GLM models have not had visual capabilities. Being text-only definitely helps Z.ai get more competitive scores, but it is a more competitive space. On the other side of things are models like Inkling-Small, which is designed to be omnimodal.
(ADDED) The RL data industry is taking off in China. Many sources and rumor-mills we’re following have been mentioning how the data industry is taking off in China — very much driven by American data companies selling to Chinese model labs. This could look like Chinese labs buying many of the same RL environments that are used by American frontier labs, and releasing the downstream RL’d model sooner. We still have large error bars on the scale and impact of this market, but it is certainly becoming important.
Z.ai is an extremely skilled LLM organization – one that is likely far more compute efficient than OpenAI / Anthropic. This needs repeating. These folks are very good at what they do. The company has very close ties to Tsinghua University, which is home to many of the best Chinese computer scientists. This abundant, eager talent pool is as central to their success as it is for any Western counterpart.
Altogether, it seems like a perfectly good strategy they’re executing with the GLM line of models. Congrats on the release! I’m excited for the weights to be out so I can do more extended testing (I tend to use American open-weight inference services like Fireworks or Baseten).
Leave a comment
This is another step towards the inevitable proliferation of very strong cyber capabilities across the economy. Z.ai has acknowledged this, saying:
GLM-5.3 is our most capable model to date for cybersecurity tasks. It delivers substantial improvements in vulnerability discovery, exploit analysis, and complex multistep security tasks. These capabilities can help defenders identify weaknesses earlier, validate risks, and accelerate remediation.
They also create clear dual-use risks. We are therefore taking a staged approach to release. Selected security partners will first evaluate GLM-5.3 in controlled settings. Broader access and API availability will follow. Once the necessary safety evaluations and release preparations are complete, we will publish GLM-5.3’s complete model weights.
They go on to acknowledge how they’re monitoring inference on their platforms via a request classifier and chain of thought monitoring (on top of model alignment). The devil is in the details here, and it is unclear the level of execution every AI lab will have here. The capability diffusion is determined by the lowest common denominator.
At the end of the day, this type of safety barely matters when true open-weights are coming. If not GLM-5.3, then another model. The size of the models with these capabilities is reducing over time, becoming easier to modify and deploy (potentially without safeguards). Z.ai does some of the right things, including pushing for more vulnerability discovery and proactive management, but any single company is far from being able to handle this on their own.
We need industrial-scale guidance led by the government or industry coalitions to immediately prepare for this transition across all software.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み