Z.ai が GLM-5.3 を公開、ポストトレーニングによる改善を主張
本文の状態
日本語全文を表示中
詳細モードで約11分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The New Stack AI
中国の Z.ai が GLM-5.3 をリリースし、ポストトレーニングによるコード能力や長期タスク処理の向上を主張したが、開発者からはベンチマークの実世界妥当性について懐疑的な見方が示されている。
AI深層分析を開く2026年8月19日 20:30
AI深層分析
キーポイント
GLM-5.3 の技術的特徴と改善点
Z.ai は GLM-5.2 と同じコードベースから派生し、ポストトレーニングの強化により複雑なコーディングや長期ホライズンのタスク処理能力を大幅に向上させたと発表している。
独自ベンチマークと実世界ワークフロー
同社は公共テストセットからの汚染リスクを避けるため「Z.ai Code Bench」を導入し、エンジニアの実際の業務環境(計算クラスター、ドキュメントなど)に準じたタスクで 50% の改善を示した。
開発者による懐疑的な評価
NonBioS.ai の共同創設者 Nishant Soni は、GLM-5.3 が実世界の長期タスクで優れているという主張に対して客観的なベンチマークが存在しないため、懐疑的な見方を示している。
公開された外部ベンチマーク結果
Z.ai は TerminalBench 3.0 や HLE(Humanity's Last Exam)など複数の独立した学術・業界ベンチマークでの結果も公開し、モデルの性能を裏付けようとしている。
GLM-5.3の真の能力源は産業規模の蒸留
Soni氏は、同モデルがAnthropic製モデルからの産業規模の蒸留によるものだと推測している。内部テストではKimiやGLMの出力がClaudeと驚くほど類似しており、GeminiやGrokとは異なる多様性を示さなかった。
重要な引用
Over the past month we kept scaling on this [GLM-5.2] stack: more environments, more diverse tasks, and more compute spent training on them
To my knowledge, no benchmarks objectively demonstrate the claimed superior long-horizon capability of GLM-5.3 in real world tasks
"To my knowledge, no benchmarks objectively demonstrate the claimed superior long-horizon capability of GLM-5.3 in real world tasks," Soni says.
"I suspect this is largely an effort to deflect from the model's true source of frontier capability – which shows a pattern consistent with industrial-scale distillation of Anthropic models," he adds.
編集コメントを表示
編集コメント
Z.ai は実世界での複雑なタスク処理能力を強調したが、その主張に対する業界からの懐疑的な反応も同時に示されている。開発者はベンチマークの数値だけでなく、実際の業務環境におけるモデルの振る舞いを慎重に検証する必要があるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

中国のフロンティアモデル企業 Z.ai は金曜日、GLM-5.3 をリリースしました。このモデルは前作 GLM-5.2 と同じコードベースから派生しており、すべての性能向上はトレーニング後の調整によって実現されています。最新バージョンでは、複雑なコーディングや長期にわたるタスクの処理能力が大幅に強化されたと主張されています。
ポストトレーニングによる最適化プロセスは、楽譜のように単一の不変のスクリプトに従う明確な工程とは通常考えられていません(コードスクリプトの意味ではありません)。しかし、この文脈では特に重要であり、推論アライメント、教師あり微調整(SFT)、そして人間のフィードバックからの強化学習(RLHF)などが含まれる可能性があります。
Z.ai は匿名のブログ記事でこう述べています。「過去1ヶ月間、私たちは [GLM-5.2] のスタック上でスケーリングを続けました。より多くの環境、多様なタスク、そしてそれらに対するトレーニングに投入する計算リソースを増やしました」。
はるかに幅広い生産ワークフローで訓練されたモデル
同社はさらに詳細を明らかにし、拡大され複雑化したモデルのトレーニング環境が、「はるかに幅広い生産ワークフロー」をカバーするようになったと説明しています。ここで扱われる「多様なタスクカテゴリ」は、エンジニアリングや研究の実務が実際に行われているプロセスに基づいて設計されています。一部のタスクには、熟練したエンジニアであれば数日かかる作業が含まれていました。
例えば、機械学習インフラのタスクにおいて、モデルはエンジニアと同じ作業環境を与えられ、計算クラスターやストレージシステム、内部ドキュメント、コードベース、実験結果へのアクセス権限を付与されます。モデルはトレーニングスタック全体のボトルネックを特定し、最適化を実装し、実験を実行しながら、正確性を保ったまま測定可能なエンドツーエンドの速度向上を実現しなければなりません。
社内ベンチマークに価値があるとするなら、組織が新たに導入した「Z.ai Code Bench」では、GLM-5.3 のコーディング能力は GLM-5.2 よりも 50% 向上していると評価されています。同社は厳格な倫理観を誓う立場から、このプライベートベンチマークが「公開テストセットからの汚染リスクを低減し、現実世界のユーザー体験をより忠実に反映する」と考えています。
その他の主要なパブリックベンチマークの結果も、同社によって公然と紹介されています。これには TerminalBench 3.0、DeepSWE、Agents' Last Exam、AutomationBench、そして独立した学術ベンチマークである HLE w/ Tools: Humanity's Last Exam (HLE)、さらに OpenAI の GDPVal-AA v2 が含まれています。
これらの動きを AI デベロッパーたちはどう捉えているのでしょうか?
こうした指標やベンチマークの枠を超えて、オープンウェイトモデルの世界で何が起きているかについて、現場の実際の開発者たちはどのように考えているのでしょうか。
長期的自律型ソフトウェアエンジニアリング企業 NonBioS.ai の共同創業者であるニシャン・ソニー氏は、The New Stack に対し、Z.ai が現実世界の長期タスクにおけるスケーリングで何を行ったかについて考える際、「自身の実際の経験に基づけば、その成果は少し疑わしい」と述べています。
「私の知る限り、GLM-5.3 の claimed superior long-horizon capability(主張される優れた長期能力)を現実のタスクで客観的に証明するベンチマークはありません」とソニー氏は指摘します。「Z.ai が説明する特定のオーケストレーション手法では、最先端モデルにそのような能力をもたらすために必要な、差別化された包括的なデータセットを提供するのは難しいでしょう」。
「これはおそらく、このモデルの真の最先端能力の源泉から目を逸らさせるための努力であり、その源泉は Anthropic モデルを産業規模で蒸留した結果と一致するパターンを示している」とソニー氏は付け加えます。
NonBioS 社内のテストにおいて、ソニー氏によると、同チームは Kimi/GLM の出力と Claude の出力の間に「驚くほど類似性がある」ことを確認しました。一方、Gemini や Grok といった他の最先端モデルは、Claude の出力と比較してより多様な結果を示しています。
「これはおそらく、このモデルの真の最先端能力の源泉から目を逸らさせるための努力であり、その源泉は Anthropic モデルを産業規模で蒸留した結果と一致するパターンを示している」とソニー氏は付け加えます。
現実の結果を得るためには、モデルとの妥協点を見つける必要がある。
AI ベンチマーク専門企業 Megaton の創業者である Sherif Higazy 氏は、The New Stack に対し、あらゆる最先端モデル開発企業(あるいはユーザーやチーム)が、タスク(サイバーセキュリティを含むその他諸々)に対してトークン数とコストを測定する内部評価システムを構築することには一定の合理性があると考えていると語った。
「ベンチマークのためにモデルを訓練することと、約束された能力を実際に発揮させることは、一般的には良いことだ」と Higazy 氏は言う。「しかし、エージェントから一貫して有用な成果を引き出すためには、モデル側も歩み寄る必要がある。つまり、チームはエージェントの動作特性に合わせて作業環境を適応させなければならないのだ(例えば、反復的なテスト実行を強制したり、入力・出力を機械可読かつ検証可能にしたりする)。そうすることで最良の結果が得られるようになる」。
「Z.ai は、この負担をモデル訓練の段階へ前倒しすることで、より高いタスクパフォーマンスを実現できると主張している。エージェントを多様で現実的な環境の中で動作するように訓練できれば、監督の手間を減らしつつ、より長く複雑なタスクでも優れたパフォーマンスを発揮できるようになるだろう。おそらく、まだ誰もがこのアプローチに完全に賛同しているわけではないが、議論に加える価値は十分にある」と彼は付け加えた。
AI ツールのデータ精度を強化する企業 Sphinx の共同創業者兼 CEO、ローハン・コディヤラムは The New Stack に対し、モデルの進化スピードがあまりにも速いため、現時点でトップに立つ特定のモデルを中心としたエンタープライズ基盤を構築することはないと語った。
「OpenAI、Anthropic、Google、そして中国製のオープンウェイトモデルに至るまで、各社は互いに常に飛躍的な進歩を遂げている」とコディヤラムは指摘する。「ビジネスの文脈は、基盤となるモデルから独立して存在すべきだと考えている。そうすれば、企業は新しいシステムに自社の仕組みを再教育することなく、モデルを自由に切り替えることができるからだ。モデル非依存性はもはや単なる技術的な好みを越え、6 ヶ月後にどのモデルが最良になるか誰も予測できない AI 市場におけるリスクヘッジの手段へと進化している。」
彼はさらに、複数のモデルがほぼ同等の能力に達した段階では、選択基準はリーダーボード上の順位よりも実務上の負荷に左右されると強調する。コスト、レイテンシ、プライバシー、デプロイメントモデル、ツール連携、そして特定のタスクにおけるパフォーマンスといった要素が、ベンチマークの数値をわずかに上回る重要性を持つ可能性があるのだ。
「企業は自らのモデル選択を不可逆的なものにしてはならない。能力、経済性、インフラストラクチャが進化する中で、相互運用性を備えた設計を目指すべきだ。」
しかし、コディヤラムはさらに付け加える。「個々の企業の判断を超えたグローバルなモデル選定も進行中である。計算資源の可用性が、実際に大規模で動作可能なモデルを決定づけるからだ。現在、OpenAI や Anthropic といった大手プロバイダー向けに割当てられるインフラストラクチャの方が、Z.ai の GLM などの新興選択肢よりもはるかに多いのが実情だ。」
仮にすべての企業が明日にもモデルを切り替えたいと考えても、市場全体がそれに追随できるとは限りません。そのためコディヤラム氏は、開発チームは自社のモデル選択を不可逆的なものにしてはならず、機能や経済性、インフラの進化に合わせてポータビリティ(移植可能性)を意識した設計を行うべきだと指摘しています。
「端的に言えば、中国の研究機関は、アメリカの研究所がリリース前のテスト期間中に費やす時間をすべて活用し、ベンチマークで常に上位を維持しようとしています」
Z.ai はベンチマークの最適化を行っているのでしょうか?
GLM 5.3 について Hacker News で議論された投稿数は限定的ですが、ML 研究者であるネイサン・ランバート氏の洞察は特に鋭いものです。彼は自身のサイト「Interconnects」で投稿し、中国の研究機関がどのようにしてシリコンバレーの最先端を追い続けているのかという疑問を投げかけています。
まず、東洋の最先端プレイヤーたちが「ベンチマーク最適化(benchmaxxing)」を行っているかどうか、つまりテストセットにモデルを最適化させて高いスコアを達成するものの、実際の運用環境では再現できない結果になっているのではないかという点について疑問を呈しました。Z.ai は同社が最新のモデルを複雑な実世界の開発ワークフローに基づいて構築したと主張しているにもかかわらずです。
ランバート氏は、もしそのような行為が行われているとしても、それは極めて微妙なレベルに留まっていると考えています。
「OpenAI や Anthropic が、Z.ai や Moonshot AI よりもはるかに優れた内部モデルを持っている可能性は極めて高い。それでも、これらの米企業はモデルを一般公開するまでに数ヶ月を要するため、最先端における採用判断において中国のラボが有利に働くことになる。端的に言えば、中国のラボは米国のラボがリリース前のテストに費やす時間をすべて活用し、ベンチマークでのスコアを継続的に引き上げているのだ」とラムバートは記している。
自己回帰的に設計されたクロージング・ギャップ埋め込み手法
控えめなタイトルを持つ「GLM」は、現在のモデルの基盤となっている「General Language Model(汎用言語モデル)」トレーニングアルゴリズムに由来する名称である。このモデルは、クロージングテストを用いてデータの一部を削除または隠蔽し、モデルの語彙力、理解力、推論能力を育成するための訓練手法である「自己回帰的空白埋め込み(autoregressive blank infilling)技術」に依存している。
完全オープンソースとして公開される Z.ai は、安全性評価とハードニングが完了した後に GLM-5.2 の重みデータをリリースすると確認しており、その時期はリリースから 2 週間後(2026 年 8 月の最終日頃)となる予定だ。
※本記事は The New Stack に掲載された「An industrial-scale distillation of models, or subtle benchmaxxing: What developers really think of GLM-5.3」の翻訳です。
原文を表示

Chinese frontier model outfit Z.ai released GLM-5.3 on Friday, a model hewn from the same codebase as its predecessor GLM-5.2, with every gain engineered as a result of post-training. The latest release is claimed to be much better at complex coding and long-horizon tasks.
Rarely considered to be a defined process that follows a single immutable script (as in song sheet, not as in code script), post-training model optimization processes are especially significant here and will have included reasoning alignment, supervised fine-tuning, and Reinforcement Learning from Human Feedback (RLHF).
“Over the past month we kept scaling on this [GLM-5.2] stack: more environments, more diverse tasks, and more compute spent training on them,” stated Z.ai, in an anonymously authored blog.
A model trained on a much broader range of production workflows
The company clarified further, saying that the widened and more complex model training environments now cover “a much broader range of production workflows”, with “diverse task categories” designed around how engineering and research work is actually carried out in practice. Some tasks constituted what would represent several days of work for an experienced engineer.
“In an ML infrastructure task, for example, the model may be given the same working environment as an engineer, with access to compute clusters, storage systems, internal documentation, codebases, and experiment results. It must diagnose bottlenecks across the training stack, implement optimizations, run experiments, and deliver a measurable end-to-end speedup while preserving correctness,” stated Z.ai.
If in-house benchmarks are of any worth, the organization’s newly introduced Z.ai Code Bench measures GLM-5.3’s coding abilities at a 50% improvement over GLM-5.2. Pledging to be scrupulously virtuous, the company thinks that private benchmark “reduces the risk of contamination from public test sets” to provide a more faithful measure of real-world user experience.
Other public benchmark positions are also openly showcased by the company, here spanning TerminalBench 3.0, DeepSWE, Agents’ Last Exam, AutomationBench, HLE w/ Tools: Humanity’s Last Exam (HLE), an independent academic benchmark, as well as OpenAI’s GDPVal-AA v2.
What do AI developers make of these moves?
Outside these yardsticks and benchmarks, what do real world developer practitioners think of what’s happening in the open-weight model universe?
Co-founder of long-horizon autonomous software engineering company NonBioS.ai, Nishant Soni, tells The New Stack that when he considers what Z.ai has done in relation to scaling for real-world long-horizon tasks, he “would take it with a pinch of salt” based upon his own real-world experiences.
“To my knowledge, no benchmarks objectively demonstrate the claimed superior long-horizon capability of GLM-5.3 in real world tasks,” Soni says. “The specific orchestration that Z.ai describes seems unlikely to provide the differentiated and comprehensive datasets required to engender such capability in frontier models.”
“I suspect this is largely an effort to deflect from the model’s true source of frontier capability – which shows a pattern consistent with industrial-scale distillation of Anthropic models,” he adds.
In internal testing at NonBioS, Soni explains that his team has seen “striking similarity” between Kimi/GLM outputs and Claude’s outputs. In contrast, other frontier models – notably Gemini and Grok – show greater diversity compared to Claude’s outputs.
“I suspect this is largely an effort to deflect from the model’s true source of frontier capability – which shows a pattern consistent with industrial-scale distillation of Anthropic models,” he adds.
Meet the model halfway to real results
Founder at AI benchmarking specialist Megaton, Sherif Higazy, tells The New Stack that he can see sense in any frontier model company (or user, or team) laying down an internal evaluation system that measures tokens and spend against tasks (cybersecurity or otherwise).
“When a model is trained for benchmarks, versus delivering promised capability, that’s generally a good thing,” Higazy says. “But ultimately, getting consistently useful work from an agent requires meeting the model halfway so that teams adapt their working environments around how agents work (for example, forcing repeated test runs or making inputs & outputs machine-readable and verifiable) to get the best results.”
Getting consistently useful work from an agent requires meeting the model halfway so that teams adapt their working environments around how agents work (for example, forcing repeated test runs or making inputs & outputs machine-readable and verifiable) to get the best results.”
“Z.ai is claiming that developers can shift this burden upstream into model training to get better task performance. If you can train the agent to work within a more diverse and realistic environment, it will perform better on longer and more complex tasks without as much supervision. Perhaps not everybody has drunk the Kool-Aid in terms of this approach yet, but it’s worth bringing it into the mix,” he adds.
Co-founder and CEO at AI tool data accuracy enforcement company Sphinx, Rohan Kodialam tells The New Stack that the pace of improvement in models is exactly why he wouldn’t build enterprise infrastructure around whichever model happens to lead today.
“OpenAI, Anthropic, Google and increasingly Chinese open-weight models are leapfrogging each other constantly,” Kodialam says. “We think business context should live independently of the underlying model, so companies can move between models without having to reteach the new system how their organization works. Model agnosticism is becoming less of a technical preference and more of a hedge against an AI market where nobody knows who will have the best model six months from now.”
He reminds us that once several models reach roughly comparable capability, choosing between them becomes much more about the workload than the leaderboard. Cost, latency, privacy, deployment model, tool use, and performance on your specific tasks can matter more than a few points on a benchmark.
“Enterprises should avoid making their own model choices irreversible and build for portability as capabilities, economics, and infrastructure evolve.”
“But there’s also a global model choice happening beyond any individual enterprise: availability of compute shapes which models can actually operate at massive scale, and today far more infrastructure is dedicated to providers like OpenAI and Anthropic than newer alternatives like Z.ai’sn GLM,” adds Kodialam.
Even if every company wanted to switch tomorrow, the market as a whole couldn’t necessarily move with them, which is why Kodialam says developer teams should avoid making their own model choices irreversible and build for portability as capabilities, economics, and infrastructure evolve.
“To put it simply – the Chinese labs use all the time that American labs do pre-release testing to keep hillclimbing on benchmarks.”
Is Z.ai benchmaxxing the bencmarks?
Of the comparatively limited posts on Hacker News discussing GLM 5.3, ML researcher Nathan Lambert is perhaps the most insightful. He cross-references a post on his own Interconnects site which questions just exactly how the Chinese labs keep up with Silicon Valley’s frontier glitterati.
First questioning whether or not the Far East frontier players are “benchmaxxing” i.e. pointing models at test sets to achieve good benchmark scores that fail to replicate in real world deployment scenarios (despite Z.ai telling us that it has built the latest model based on complex real world developer workflows), Lambert thinks that if there is any going on, it’s only to a subtle level.
“It is very, very likely that OpenAI and Anthropic have far better internal models than Z.ai and Moonshot AI. Still, these American companies tend to take months to release their models to the public, which massively flatters the Chinese labs in adoption decisions at the frontier. To put it simply – the Chinese labs use all the time that American labs do pre-release testing to keep hillclimbing on benchmarks,” wrote Lambert.
Autoregressively engineered to cloze the gap
The modestly titled GLM derives its name from the General Language Model training algorithm upon which the current model is now built. It relies on autoregressive blank infilling techniques, a model training technique that deletes or occludes sections of model data using cloze tests to develop model vocabulary, comprehension and reasoning.
Fully open source, Z.ai confirmed it will release the GLM-5.2 weights two weeks after launch (approximately the last day of Aug 2026), once safety evaluation and hardening are complete.
The post An industrial-scale distillation of models, or subtle benchmaxxing: What developers really think of GLM-5.3 appeared first on The New Stack.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み