Z.ai、ベースモデル再学習なしで GLM-5.3 を公開
本文の状態
日本語全文を表示中
詳細モードで約6分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
Z.ai はベースモデルの再学習なしに事後学習を拡張することで、複雑なコーディングおよび長期タスクにおいて大幅な性能向上を実現した GLM-5.3 をリリースし、セキュリティ分野でも競合他社を上回る結果を示した。
AI深層分析を開く2026年8月14日 17:20
AI深層分析
キーポイント
ベースモデルの再使用と事後学習の強化
GLM-5.3 は GLM-5.2 と同じ 743B ベースモデルを使用しており、すべての性能向上はタスク環境やトレーニング期間のスケーリングによる事後学習の結果である。
コーディング能力の劇的な向上
Terminal-Bench 3.0 で 4.6 から 28.3 へ、DeepSWE v1.1 では 46.2 から 66.9 へと大幅なスコア上昇を記録し、長期ホライズンのタスク処理能力が顕著に高まった。
セキュリティ分野での予期せぬ成果
Z.ai は単一バグ推論の改善を期待していたが、トレーニングのスケーリングに伴い完全なエクスプロイトチェーンにおける計画形成が可能となり、CyberGym で 84.5% を達成した。
競合モデルとの比較と公開スケジュール
セキュリティベンチマークでは Mythos 5 や GPT-5.6 Sol を上回る結果となったが、重みは現時点で非公開であり、安全評価完了後約 2 週間での公開を予定している。
ベースモデルの再学習なしでの性能向上
GLM-5.3 は GLM-5.2 のベースモデルを流用しており、すべての性能向上はポストトレーニングのスケーリングによるものである。
重要な引用
Every reported gain comes from scaled post-training: more task environments, more environment types, longer training.
The model began forming coherent plans across complete exploitation chains.
GLM-5.3 is live through the Z.ai API, the GLM Coding Plan, and ZCode.
GLM-5.3 reuses the GLM-5.2 base model; all gains come from post-training scaling.
編集コメントを表示
編集コメント
ベースモデルを再学習せずに事後学習の規模と質を高めるだけで、複雑な推論タスクにおいて競合他社を凌駕する結果を出した点は技術的に極めて興味深い。ただしセキュリティ分野での躍進が意図せぬ成果であったことは、大規模トレーニングにおける予測不能な能力発現の可能性を示唆しており、今後の研究動向に注目が必要である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Z.ai が GLM-5.3 をリリースしました。GLM-5.3 は、GLM-5.2 と同じ 743B のベースモデル上で動作します。報告されている性能向上はすべて、スケールされたポストトレーニングによるものです。具体的には、タスク環境の増加、環境タイプの多様化、そしてより長い学習期間が寄与しています。
その成果は主に二つの領域で顕著です。コーディング能力、特に最長ホライズンのベンチマークでは劇的な向上が見られ、Terminal-Bench 3.0 のスコアは 4.6 から 28.3 に跳ね上がりました。また、サイバーセキュリティ分野でも Z.ai が想定していた以上に大きな進歩があり、CyberGym では 84.5% を達成しています。
なお、モデルの重み(ウェイト)はまだ公開されていません。
導入は可能か?
部分的には可能です。GLM-5.3 はすでに Z.ai の API や GLM Coding Plan、そして ZCode を通じて利用できます。ただし、重みの公開はまだ先です。Z.ai によると、安全性評価とハードニング(堅牢化)が完了するLaunch からおよそ2週間後に公開される見込みです。
今すぐ導入を検討できる企業は?
スタートアップや中規模のエンジニアリング組織であれば、Coding Plan や API を通じて今日から利用可能です。一方、データ所在地規制やベンダー審査を厳格に求める大企業は、重みの公開を待つべきでしょう。セキュリティベンダーや MSSP(マネージド・セキュリティ・サービス・プロバイダ)にとっては、最も重要なシグナルであり、同時に政策上の影響も大きいと言えます。
注目すべき業界:
- 開発者向けツール
- クラウドインフラ
- アプリケーションセキュリティ
- ファイナンスや EC のエンジニアリング
- カーネル、ブラウザエンジン、ネットワークスタックを提供するベンダー
主な活用事例:
- リポジトリ規模のリファクタリング
- 長期間にわたる CLI エージェントの運用
- CI(継続的インテグレーション)の失敗原因調査
- ホワイトボックスでの脆弱性発見
- クラッシュの原因特定
- セキュアなコードレビュー
Terminal-Bench 3.0 では、GLM-5.2 のスコアが 4.6 から 28.3 に向上しました。DeepSWE v1.1 は 46.2 から 66.9 へ、Agents' Last Exam (CLI) は 23.8 から 28.5 へとそれぞれ大幅な伸びを示しています。また、44 の職業分野を網羅する GDPval-AA v2 では、GLM-5.3 が 1,769 というスコアを記録しました。
Z.ai 独自の評価基準である Z.ai Code Bench においても、同社によると GLM-5.2 よりも 50% の改善が見られました。このベンチマークでは、タスクあたり約 50,000 トークンの出力で 31.4% を達成しています。一方、Claude Opus 4.8 は 120,000 トークンで 29.5%、Claude Fable 5 は最大限の努力を払って 39.5% と依然として首位に立っています。Z.ai は、この非公開ベンチマークがデータ汚染(コンタミネーション)のリスクを低減できると主張しています。
一方、一般的な評価スイートでは、GLM-5.3 は GPT-5.6 Sol や Fable 5 に比べて、いくつかの難易度の高いコーディング評価においてやや劣っています。なお、すべての数値はベンダーが報告したものであり、使用されたハーン(harness)、コンテキスト長、サンプリング設定などは発表文書に明記されています。
セキュリティ分野での結果
Z.ai はこの成果について「意図せぬもの」と位置付けています。同社は単一バグの推論能力向上を期待して脆弱性発見データを追加しましたが、実際にはトレーニング規模の拡大に伴い、その能力が複合的に強化されました。その結果、モデルは完全な攻撃チェーン全体にわたって整合性の高い計画を立てられるようになったのです。
ホワイトボックスソースコードから脆弱性の発見と検証を試す CyberGym では、スコアが 77.2% から 84.5% に向上しました。これにより、Mythos 5 の 83.8% や GPT-5.6 Sol の 83.6% をわずかに上回っています。
根本原因の推論と実際に動作するエクスプロイト(攻撃コード)の実装が求められる ExploitBench では、スコアが 24.4% から 54.4% に跳ね上がりました。Mythos 5 は 78.0% を記録しています。
ExploitGym におけるタスク完了数では、GLM-5.3 は 2 時間で 105 件、6 時間で 130 件のタスクを完了しました。これに対し、GLM-5.2 はそれぞれ 29 件と 39 件でした。Mythos 5 は 181 件と 247 件を達成しています。
ベンチマークの性質を詳しく見ると、その評価が「実装(exploitation)」の連鎖に深く入り込むほど、GLM-5.2 に対する性能向上幅は大きくなり、クローズドな最先端モデルとの差も広がっていくという傾向が一貫して見られます。
主なポイント
- GLM-5.3 は GLM-5.2 のベースモデルをそのまま流用しており、すべての性能向上はポストトレーニングの規模拡大によるものです。
- Terminal-Bench 3.0 ではスコアが 4.6 から 28.3 に、DeepSWE v1.1 では 46.2 から 66.9 にそれぞれ大幅に改善されました。
- CyberGym では 84.5% を達成し、Mythos 5(83.8%)や GPT-5.6 Sol(83.6%)を抜いて首位となりました。
- ExploitBench のスコアは 54.4% と倍増しましたが、依然として Mythos 5 の 78.0% には及びません。
重み付きモデルの公開は、安全性評価と堅牢化(ハードニング)を経てから約 2 週間後を予定しています。
詳細は Z.ai の GLM-5.3 技術ブログや、Zai_org の公式発表、セキュリティ開示台帳、そして GitHub の zai-org/GLM-5 リポジトリでご確認ください。Twitter や 15 万人以上の ML 関係者が集まる SubReddit、ニュースレターへの登録もぜひご検討ください。Telegram ユーザーの方にも、今すぐ Telegram チャンネルにご参加いただけるようになりました。
GitHub リポジトリや Hugging Face ページ、製品リリース、ウェビナーなどのプロモーションをご希望の場合は、お気軽にお問い合わせください。
原文を表示
Z.ai just released GLM-5.3. GLM-5.3 runs on the same 743B base model as GLM-5.2. Every reported gain comes from scaled post-training: more task environments, more environment types, longer training. The results land in two places. Coding jumps most on the longest-horizon benchmarks, with Terminal-Bench 3.0 moving from 4.6 to 28.3. Cybersecurity moved further than Z.ai says it expected, with CyberGym reaching 84.5%. Weights are not public yet.
Is It Deployable?
Partially, GLM-5.3 is live through the Z.ai API, the GLM Coding Plan, and ZCode. Weights are not out. Z.ai says it will publish them roughly two weeks after launch, once safety evaluation and hardening finish.
Which companies can move now: Startups and mid-market engineering orgs can adopt it today via the Coding Plan or API. Enterprises with data-residency or vendor-review rules should wait for weights. Security vendors and MSSPs get the most signal, and the most policy exposure.
Industries: Developer tooling, cloud infrastructure, application security, fintech and e-commerce engineering, and vendors shipping kernels, browser engines, or network stacks.
Applications: Repository-scale refactors, long-horizon CLI agents, CI failure triage, white-box vulnerability discovery, crash triage, and secure code review.
Coding Results
Terminal-Bench 3.0 moves from 4.6 to 28.3 against GLM-5.2. DeepSWE v1.1 moves from 46.2 to 66.9. Agents’ Last Exam (CLI) moves from 23.8 to 28.5. On GDPval-AA v2, which spans 44 occupations, GLM-5.3 scores 1,769.
On Z.ai Code Bench, an internal evaluation, the company reports a 50% improvement over GLM-5.2. It reports 31.4% at roughly 50,000 output tokens per task. Claude Opus 4.8 scores 29.5% at 120,000 tokens. Claude Fable 5 still leads at 39.5% at maximum effort. Z.ai argues a private benchmark reduces contamination risk.
On public suites, GLM-5.3 trails GPT-5.6 Sol and Fable 5 on several harder coding evaluations. All figures are vendor-reported, with harness, context length, and sampling settings documented in the announcement.
The Cybersecurity Result
Z.ai flags this one as unplanned. It added vulnerability-discovery data expecting better single-bug reasoning. Instead, capability kept compounding as training scaled. The model began forming coherent plans across complete exploitation chains.
CyberGym, which tests discovery and validation from white-box source, moves from 77.2% to 84.5%. That edges past Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. ExploitBench, which requires root-cause reasoning and a working exploit, moves from 24.4% to 54.4%. Mythos 5 sits at 78.0%. On ExploitGym, GLM-5.3 completes 105 tasks in two hours and 130 in six. GLM-5.2 completes 29 and 39. Mythos 5 completes 181 and 247.
The pattern is consistent. The deeper into the exploitation chain a benchmark sits, the larger the gain over GLM-5.2. The gap to closed frontier models also widens.
Interactive Explainer
Key Takeaways
GLM-5.3 reuses the GLM-5.2 base model; all gains come from post-training scaling.
Terminal-Bench 3.0 moves from 4.6 to 28.3; DeepSWE v1.1 from 46.2 to 66.9.
CyberGym hits 84.5%, ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%).
ExploitBench more than doubles to 54.4%, but trails Mythos 5 at 78.0%.
Weights ship in about two weeks, after safety evaluation and hardening.
Check out the Z.ai GLM-5.3 technical blog, Zai_org announcement, Z.ai Security Disclosure Ledger and zai-org/GLM-5 on GitHub. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks appeared first on MarkTechPost.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み