智譜 AI、GLM-5.3 を公開し長文処理と自律的タスク実行能力を強化
本文の状態
日本語全文を表示中
詳細モードで約24分の本文を読めます。
智譜 AI は GLM-5.3 を発表し、ベースモデルは変更せずポストトレーニングの拡張によりコーディング能力とサイバーセキュリティ脆弱性発見能力が劇的に向上したことを示した。
AI深層分析を開く2026年8月18日 21:54
AI深層分析
キーポイント
ポストトレーニングによる性能飛躍
GLM-5.3 は GLM-5.2 と同じベースモデルを使用し、計算資源と多様なタスク環境を投入したポストトレーニングの拡張のみで性能が向上している。
コーディング能力における新記録
同社の内部ベンチマークでは GLM-5.2 より 50% 改善し、オープンウェイトモデルとして最上位に位置付けられ、公開ベンチでも SOTA を達成した。
予期せぬサイバー能力の創発
ポストトレーニングのスケーリングに伴い、脆弱性発見やエクスプロイトチェーンにおける利用能力が急激に向上し、GLM-5.2 の倍以上の性能を示した。
安全評価後のオープンソース化
モデルの重みは安全評価と堅牢化が完了する 2 週間後に公開される予定であり、現時点ではクローズドな状態である。
実務レベルのコーディング環境への拡張
GLM-5.3 は単なるコード演習ではなく、実際の専門家業務に近いタスクを対象に環境を拡張した。エンジニアや研究者が実際に作業を行うワークフローに基づいた広範な生産環境をカバーしている。
重要な引用
Scaling post-training is all we did for GLM-5.3.
GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench.
As we scaled post-training, cyber capability developed faster than we expected.
For GLM-5.3, we pushed environment scaling toward tasks that look less like coding exercises and more like real units of expert work.
編集コメントを表示
編集コメント
ベースモデルの変更なしにポストトレーニングの拡張だけで劇的な性能向上を達成した点は、リソース効率化の観点から極めて興味深い。特にサイバーセキュリティ能力が予測以上に急速に発達した事実は、セキュリティ分野における AI のリスクと可能性の両面を再考させる内容である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
GLM-5.3 の開発で私たちが行ったのは、ポストトレーニングの規模拡大のみです。
GLM-5.2 では、効率的な長文脈処理のための IndexShare、長期ホライズンタスク向けの RL 用 SAO、大規模非同期学習用の slime を備えたスタックを構築しました。これらはすべて、私たちが蓄積してきた長期ホライズンタスクの環境上で稼働しています。
過去 1 ヶ月間、私たちはこのスタック上で規模拡大を続けました。より多くの環境と多様なタスクを導入し、それらに対するトレーニングにさらに多くの計算リソースを投入したのです。
本日、GLM-5.3 をリリースします。ベースモデルは GLM-5.2 と同じですが、すべての性能向上はポストトレーニングによるものです。GLM-5.2 と比較すると、複雑なコーディングや長期ホライズンタスクでの能力が大幅に向上しています。
- 強力なコーディング能力: GLM-5.3 は、現在最も高性能なオープンウェイトのコードモデルです。社内ベンチマークである Z.ai Code Bench では GLM-5.2 より 50% 改善しました。また、Terminal Bench 3.0 や Agents' Last Exam といった公開ベンチマークでも、オープンソース界で最高性能(SOTA)を達成しています。
- 突発的なサイバーセキュリティ能力: ポストトレーニングの規模拡大に伴い、サイバーセキュリティに関する能力は予想以上に急速に発展しました。GLM-5.3 は脆弱性発見における CyberGym で最高性能を記録しており、特にエクスプロイト(攻撃)チェーンの上流部分での改善幅が顕著です。エクスプロイト関連のベンチマークでは、GLM-5.2 の倍以上のスコアを叩き出しています。
- オープンソース化: セーフティ評価とハードニング(強化)が完了次第、リリースから 2 週間以内にモデルの重み(ウェイト)を公開する予定です。

**比較対象モデルとの性能
| ベンチマーク | GLM-5.3 | GLM-5.2 | Kimi K3 | DeepSeek-V4Pro-0813 | Qwen3.8-Max | Opus 4.8 | Fable 5(w/ fallback) | GPT-5.6 Sol |
|---|---|---|---|---|---|---|---|---|
| コーディング | ||||||||
| Terminal Bench 2.1 | 88.2 | 81.0 | 88.3 | 87.9 | 86.6 | 85.0 | 88.0 | 88.8 |
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | - | - | 21.1 | 33.7 | 34.6 |
| DeepSWEv1.1 | 66.9 | 46.2 | 67.5 | 62.7 | 56.6 | 58.0 | 69.7 | 72.7 |
| NL2Repo | 58.0 | 48.9 | 58.0 | 61.1 | 55.9 | 69.7 | - | - |
| ProgramBenchAlmost Solved | 19.0 | 9.5 | 17.5 | - | 10.5 | 15.5 | 33.0 | 23.0 |
| FrontierSWE | 78.1 | 67.5 | - | - | - | 66.5 | 88.2 | - |
| SWE-Marathonv1.1 | 42.5 | 19.4 | 48.1 | - | - | 48.8 | 33.1 | 42.5 |
| PostTrainBench | 39.8 | 31.7 | 32.0 | - | - | 32.9 | 41.8 | 36.2 |
| サイバーセキュリティ | ||||||||
| CyberGym | 84.5 | 77.2 | 80.0 | 83.3 | 78.5 | 78.1 | 83.8 | 83.6 |
| ExploitGym2h / 6h | 105 / 130 | 29 / 39 | 36 / 70 | - | 14 / 26 | 80 / 120 | 181 / 247 | 216 / 293 |
| ExploitBench | 54.4 | 24.4 | 32.2 | - | 28.8 | 40.0 | 78.0 | 76.5 |
| エージェント | ||||||||
| Toolathlon Verified | 73.0 | 59.9 | 76.5 | 74.1 | 72.5 | 76.2 | 74.7 | 74.9 |
| AutomationBenchv1.0.6 | 48.2 | 26.2 | 46.7 | 43.2 | 39.8 | 41.0 | 46.2 | 45.8 |
| Agents' Last ExamALE-CLI | 28.5 | 23.8 | 27.6 | 25.7 | 27.0 | 25.7 | 23.8 | 28.6 |
| HLE w/ Tools | 62.5 | 54.7 | 59.8 | 60.0 | 56.2 | 57.9 | 63.9 | 64.5 |
| GDPval-AA v2 | 1769 | 1508 | 1682 | 1590 | 1739 | 1588 | 1743 | 1730 |
コーディング能力の強化
GLM-5.3 では、環境のスケーリングを「コーディング演習」から「専門家の実務ユニット」へとシフトさせました。これにより、実際のエンジニアリングや研究現場でどのように作業が遂行されるかを反映した、より広範な生産ワークフローに対応できるようになりました。中には熟練のエンジニアであれば数日かかるような複雑なタスクも含まれています。
例えば ML インフラのタスクでは、モデルにエンジニアと同じ作業環境を提供します。計算クラスターやストレージシステムへのアクセス権限、社内ドキュメント、コードベース、実験結果などすべてが利用可能です。モデルはトレーニングスタック全体のボトルネックを特定し、最適化を実装して実験を実行。そして、正しさを保ちつつ測定可能なエンドツーエンドの速度向上を実現することが求められます。
このレベルの環境で学習させることで、モデルは問題を分解して各工程をユーザーに監督させるのではなく、大規模な作業を自ら責任を持って完遂する方向へと進化します。
エージェントの能力が向上するにつれ、ポストトレーニングのスケーリングにおける困難はモデル側から環境側に移ってきています。有用なタスク環境とは、実行可能で検証可能であり、かつ実際の専門的な業務に近いものでなければなりません。また、数個の手作業で作られたものではなく、多数の環境が必要となります。
このプロセスをスケールさせるために、私たちはエンドツーエンドで環境を合成するパイプラインを構築しました。一部のタスクについては、強化学習(RL)の報酬信号も同時に生成しています。研究用エージェントが実際の業務からタスクリターンパターンを収集し、多段階の依存関係と隠れた状態を持つ実行可能な長期ホライズンの環境へと変換します。その後、ジャッジエージェントが各タスクを実行して、実際に解決可能であることを検証します。
検証者(Verifier)は参照解にアクセスすることなく合成され、ソルバーの軌跡(Trajectory)を用いて報酬のショートカットを発見・解消します。オラクルチェック、無操作(No-op)チェック、未解決状態チェックをすべて通過した検証者は、直接学習に使用できる十分な信頼性を持つ二値報酬を提供します。
GLM-5.2 で導入された強化学習戦略はそのまま引き継がれており、特に圧縮を伴う SAO(Sparse Action Optimization)は、短期タスクだけでなく長期ホライズンのタスクにおいてもこれらの成果を持続させるのに貢献しています。この効果はコーディングと一般的なエージェントタスクの両方で確認されています。
具体的な数値では、GLM-5.3 は Terminal-Bench 3.0 で 4.6 から 28.3 に、DeepSWE v1.1 では 46.2 から 66.9 に、Agents' Last Exam では 23.8 から 28.5 にそれぞれ向上しました。ただし、これらのパイプラインはまだ相当量の人間による介入(Human-in-the-loop)を必要としており、環境生成と検証のさらなる自動化が今後のステップとなります。
公開ベンチマーク以外にも、Z.ai Code Bench という独自ベンチマークを導入しました。これは現実的なユーザーシナリオにおけるコーディングエージェントを評価するために設計されたもので、多様なタスクカテゴリに対応し、複雑なローカル開発環境でエージェントを動作させます。
異なる努力レベルにおいて、エージェントは「エンドツーエンドのタスク完了率」と「詳細なチェックリスト精度」の 2 つの観点から評価されます。この非公開ベンチマークにより、公開テストセットからの汚染リスクが低減され、現実世界のユーザー体験をより忠実に測定することが可能になります。

図に示す通り、GLM-5.3 は性能とトークン効率の両面で改善が見られます。あらゆる努力レベルにおいて、GLM-5.2 よりもはるかに強力なエージェントとしてのコーディング結果を達成しながら、出力に必要なトークン数は削減されています。
最大努力(Max effort)では、GLM-5.3 は 1 タスクあたり約 75K トークンで 34.5% の性能を達成するのに対し、GLM-5.2 は 96K トークンで 23.4% です。クローズドモデルとの比較でも同様の傾向が見られます。
高努力(High effort)では、GLM-5.3 は約 50K トークンで 31.4% を達成し、120K トークンで 29.5% の Claude Opus 4.8 を上回ります。ただし、最大努力で 39.5% を達成する Claude Fable 5 にはまだ及びません。
創発的なサイバー能力
ポストトレーニングの一環として、脆弱性発見のデータと環境を学習データに追加しました。これにより、モデルが脆弱性の特定や推論能力を向上させることを期待していました。しかし驚いたのは、学習規模を拡大するにつれてこの能力が急速に進化し続けたことです。
GLM-5.3 は単に個々の欠陥を検出する精度が高まったわけではありません。攻撃の複数の段階にまたがる推論が可能になり、完全な攻撃チェーンを実現するための一貫した計画を立てるようになりました。

GLM-5.3 の性能を、脆弱性分析とエクスプロイト(攻撃)の各段階を網羅する 3 つのベンチマークで評価しました。
まず「CyberGym」では、ホワイトボックスのソースコードから開始し、モデルが脆弱性を特定・検証できるかを、障害を誘発させることでテストします。GLM-5.3 は 84.5% のスコアを記録し、前バージョンである GLM-5.2 の 77.2% から大幅に向上しました。これは同ベンチマークにおける最高スコアで、Mythos 5(83.8%)や GPT-5.6 Sol(83.6%)を上回っています。
次に「ExploitBench」は、実際の脆弱性とそれを利用する手法についてより深い推論を要求します。GLM-5.3 は 54.4% のスコアを達成し、GLM-5.2 の 24.4% よりも倍以上の成績です。ただし、Mythos 5 が 78.0%、GPT-5.6 Sol が 76.5% を記録しており、依然として差はあります。
「ExploitGym」では、時間あたりの処理能力を正規化した予算制約下で、モデルがどの程度のエクスプロイトタスクを完了できるかを測定します。GLM-5.3 は 2 時間で 105 タスク、6 時間で 130 タスクを完了しました。一方、GLM-5.2 はそれぞれ 29 と 39 です。各モデルの処理速度(スループット)に基づいて予算は正規化されており、詳細は脚注に記載されています。Mythos 5 は依然として 181 と 247 で圧倒的なリードを維持しています。
この 3 つの結果に共通する傾向は、エクスプロイトの連鎖において上位にあるベンチマークほど、GLM-5.2 から GLM-5.3 への改善幅が大きい一方で、クローズドな最前線モデルとの差も依然として広いという点です。つまり、我々が最も遅れている領域こそが、能力の向上速度が最も速い場所なのです。
これらの能力が制御されたベンチマークを超えて転移するかどうかを検証しました。GLM-5.2 から、中国の複数のセキュリティチームと連携し、実世界のコードベースに対してモデルを実行してきました。専門家のレビュー、スクリーニング、重複排除を経て、モデルは 269 のプロジェクト全体で 2,436 の脆弱性を特定しました。そのうち 1,097 は中程度から深刻度の高い問題です。
発見された脆弱性の範囲は、システムカーネルやオペレーティングシステム、ブラウザエンジン、オープンソースインフラストラクチャ、Web アプリケーション、ネットワークプロトコルに及びます。多くの脆弱性は数年、あるいは数十年にわたり見過ごされてきました。最も古いものは約 40 年前に遡ります。
この取り組みは現在、継続的な開示活動へと発展しています。発見された事案の開示プロセスを追跡し、公開記録を維持するために Z.ai セキュリティ開示台帳 を構築しました。この台帳は、新しい脆弱性がレビューされ開示されるたびに随時更新されます。すでに公表済みの問題と、現在も開示手続き中の問題を明確に区別しています。公開された事案については、影響を受けるプロジェクト、深刻度、利用可能な CVE 番号、そしてコードベース内に存在していた期間などの情報を記録しています。
2,436
追跡中の発見数
53
公的に開示済み
2,383
embargo 中(非公開)
1,097
クリティカルおよびハイ
269
OSS プロジェクト
45
影響の年数
発見された脆弱性の影響範囲は 45 年に及びます。最も古い欠陥は 1981 年に導入されましたが、平均して発見されるまで 26.6 年間存在し続けていました。
セverity の分布
Critical: 107
High: 990
Medium: 1,286
Low: 53
欠陥が導入された時期
1981〜2026
Z.ai セキュリティ開示台帳 ↗ を参照してください。
slime: 長期ホライズン RL スケーリングのために構築された基盤
これらすべては、slime というオープンソースのポストトレーニングフレームワーク上で動作しています。これは RL スケーリングを目的としており、トレーニング側には Megatron、ロールアウト側には SGLang を採用しています。この設計により、トレーニング、ロールアウト、データバッファが単一のデータフロー上に統合されています。その結果、数学的推論、コード生成、サンドボックス環境、検証器、そして長期ホライズン型のエージェント環境は、すべて「データ生成」として組み込まれるため、トレーニングループ自体を変更する必要がありません。これが、GLM-5.2 や GLM-5.3 の段階でも、毎回トレーニングスタックを再構築することなく新たな環境を追加し続けられた理由です。
GLM-5.3 においても、この基盤は二つの側面からさらに強化されました。アルゴリズムの側では、RL 研究に特化した機能を追加しました。具体的には、top-p マスク、top-k および完全語彙における OPD(Optimal Policy Distribution)、そしてトレーニングとロールアウトの一貫性を高める設定です。これには R3 スタイルの設定や、数値的な完全なアライメントが含まれており、トレーニングとロールアウトのパス間で厳密な整合性を実現します。これらの機能により、サンプリング、トレーニング、教師信号に対する制御が細かく行えるようになり、統制された比較実験を高速に実行することが可能になりました。
トレーニングとロールアウトの一貫性を評価した結果、対数尤度(logprob)の平均差は 1e-7 レベルで抑制されました。これは、以前のセットアップと比較して 99.99% 以上の削減を意味します。
大規模な強化学習(RL)におけるリソース効率とシステムのスループット向上にも注力しました。ローカルストレージを新たなキャッシュ層として活用し、モデルの状態やデータを階層的に保持することで、ホストメモリの負荷を軽減しています。これは特に、複数の教師モデルを動的に切り替えながら学習を行う「マルチティーチャー OPD」において重要です。各教師モデルのために個別の常時稼働する推論サービスを立てる必要がなくなり、追加オーバーヘッドを抑えつつ、リソース消費を大幅に削減できます。
エージェント型や非同期なワークロードに対応するため、ルーターとスライム(Slime)の間でジョイントスケジューリングと負荷分散を改善しました。これにより、実行時間や完了までの時間が大きく異なるロールアウト要求も、推論リソースを効率的に活用できるようになりました。さらに、各ロールアウト環境の特性に基づいてスループット最適化設定(プリフェッチ/デコードのリソース比率、並列処理の設定、その他のスループットに直結するパラメータなど)を導き出すワークロード認識型のヒューリスティクスを追加しました。
その結果、長期ホライゾンのコーディング RL タスクにおいて、システムレベルの最適化によりエンドツーエンドの RL 学習スループットが 2.3 倍以上向上。これにより、より長い軌跡や複雑な環境でのトレーニングを、大幅に高い効率でスケールできるようになりました。
これらの改善により、実験の柔軟性が高まり、リソースコストが低下し、スループットも向上しました。これが、強化学習の継続的なスケーリングを実用的なものにする要因です。
GLM-5.3 の始め方
GLM-5.3 における API の変更
GLM-5.3 は、思考の努力レベルとして「low」「high」「max」の 3 つをサポートしています。ただし、思考機能を無効化する機能は GLM-5.3 ではサポートされなくなりました。
思考パラメータ
| パラメータ | 値 | デフォルト | 説明 |
|---|---|---|---|
thinking.type | enabled | enabled | 思考機能を有効化します。disabled はもはやサポートされません。 |
reasoning_effort | low, high, max | max | low: 軽量; high: 強化; max: 深層。 |
コーディングタスクには max が推奨されます。
{
"model": "glm-5.3",
"thinking": { "type": "enabled" },
"reasoning_effort": "max"
}
移行が必要です: アプリケーションで現在 thinking.type: "disabled" を使用している場合は、これを enabled に変更し、reasoning_effort を設定してください。
モデルIDをglm-5.3に更新する前に、low 状態を確認してください。
ただし、この条件を満たさない場合、リクエストは失敗します。
GLM-5.3 を GLM Coding Plan と ZCode で活用する
お気に入りのコーディングエージェントで GLM-5.3 を試してみましょう。ZCode, Claude Code, OpenCode などに対応しています。詳細は https://docs.z.ai/devpack/overview をご覧ください。
GLM コーディングプラン契約者向け: GLM-5.3 をすべての GLM コーディングプランユーザーに展開しました。新しい GLM コーディングプランでは、ポイント制のクォータシステムが導入されました。入力トークン、キャッシュされた入力トークン、出力トークンはそれぞれ別々に計算されます。ピークアワー(月曜~金曜 14:00–18:00 UTC+8)以外の時間帯にモデルを呼び出す場合、標準ポイントの 50% で利用可能です。週末を含むそれ以外の時間はすべてオフピーク料金となります。今すぐ構築を始めましょう:https://z.ai/subscribe
ZCode を活用して GLM-5.3 の効果を最大化
- キャッシュヒット率が 98% 以上 — 重複するコンテキストはキャッシュ料率で課金され、実効トークン数が約 30% 増加します。
- 期間限定のクォータボーナスが 1.5 倍 — キャッシュによる節約と併用することで、8 月 31 日までに標準クォータの最大 180% を利用できます。
長期目標の達成 — ゴールモードでは、目標が満たされるまで計画、コーディング、テスト、検証を繰り返します。
リモートコントロール — WeChat や Feishu を介してスマートフォンから、長時間実行中のタスクを監視・操作できます。
ZCode の体験: https://zcode.z.ai
GLM-5.3 をローカルで実行する
GLM-5.3 のモデル重みは、今後 2 週間以内に一般公開されます。
脚注
- ツール付き HLE: 評価にはサンプリングパラメータとして
temperature=1.0とtop_p=0.95を使用し、生成トークンの最大長は163,840トークンとします。
評価は、コンテキスト管理戦略を用いて最大 30 万トークンのコンテキスト長で行います。判定モデルには GPT-5.6-luna (medium) を使用します。 (原文の技術表記: 300,000)
- NL2Repo: 100 万トークンのコンテキスト内で、温度パラメータを 1.0、top_p を 1.0、最大生成トークンを 64k に設定して NL2Repo を評価しました。悪意のある動作(例えば、許可されていない pip や curl の実行など)を防ぐため、ルールベースの判定と LLM による判定を組み合わせてセキュリティ対策を講じています。
- DeepSWE: mini-swe-agent ハーネスを用いて評価を実施しました。設定は温度 0.95、top_p 1.0、タイムアウト 6 時間、コンテキスト長 400K です。
- Terminal-Bench 2.1: Claude Code 2.1.207 を用いて評価を行いました。温度 1.0、top_p 1、最大生成トークン数 65,536、タイムアウト 6 時間の条件でテストしました。
- Terminal-Bench 3.0: Terminal-Bench 3 のタスクは、Claude Code 2.1.207 ハーネス(推論エフォート:最大、コンテキスト長 400K、最大出力 128K)で評価し、各タスクごとに 3 回のロールアウトの平均値(avg@3)を報告しました。各ロールアウトはタスク公式イメージから構築された独立したコンテナ内で実行され、エージェントのターン数は最大 600、タイムアウトは 10 時間に制限されています。ツール検索は無効化されており、各エージェントが生成した成果物は、タスクごとに用意された別個の検証プログラムによって採点されます。 (原文の技術表記:
temperature=0.95、top_p=1.0、timeout=6h)
- Agent’s Last Exam (CLI): ALE の公式評価プロトコルに基づき、Claude Code ハーネス(推論努力度:最大、コンテキスト長:1M、出力制限:64K)を用いて評価を行いました。105 件のタスクはすべて、各 Task Card に記載されたリソースを使用する個別の Docker コンテナ内で実行されます。デフォルトのタイムアウトは 4 時間ですが、タスク固有の制限(最大 8 時間まで)が優先されます。ツール検索機能は無効化され、結果は公式の ALE 評価者によって採点されます。
- Toolathlon Verified: すべての結果は公式評価サービス経由で取得し、3 回の独立した実行における pass@1 の平均値を報告しています。
- AutomationBench: AutomationBench v1.0.6 で評価を行いました。これは、PR #13 で導入された
null型処理の不具合に対する修正を反映したバージョンです。
- GDPval-AA v2: モデルの評価は Artificial Analysis が実施しました。
- CyberGym: GLM-5.3 の評価には Claude Code 2.1.207 を使用しました(推論努力度:最大、Web ツールなし、温度パラメータ:1.0、top_p: 1.0、最大生成トークン数:128,000)。各タスクのタイムアウト制限は設けず、1,507 件のタスクについて単一実行での Pass@1 を計測しました。実際の利用シナリオを模擬するため、エージェントをタスクコンテナ内に配置しています。また、不正行為を防ぐため Git 関連情報をすべて削除し、ドメインホワイトリスト(基本ツールのインストールに必要な pypi.org や deb.debian.org などのみ許可)を適用しています。
ExploitGym: GLM-5.3、Kimi-K3、Qwen3.8 Max の性能を、Claude Code 2.1.207(推論努力最大設定、Web ツール使用不可、temperature=1.0、top_p=1.0、max_new_tokens=128000)上で評価しました。報告された結果は、2 つのタイムアウト制約(それぞれ 2 時間と 6 時間)下で実施した 869 のタスクにおける単一実行の Pass@1 です。これらの時間は、各モデルごとのトークン生成速度(TPS)に基づいて API 推論時間を再スケーリングして算出しています。具体的には、Artificial Analysis から取得したデータを用い、GLM-5.3 は 115 TPS、Kimi K3 は 40 TPS、Qwen3.8 Max は 47 TPS でそれぞれ補正し、さらに API 以外のオーバーヘッド分を加算しています。また、エージェントが不正を行うのを防ぐため、ドメインのホワイトリストを適用しました(基本ツールのインストールに必要な pypi.org や deb.debian.org など、必要なドメインのみを許可)。
ExploitBench: GLM-5.3 の性能を、Claude Code 2.1.207(推論努力最大設定、Web ツール使用不可、temperature=1.0、top_p=1.0、max_new_tokens=128000)上で評価しました。公式の評価設定に準拠し、エージェントと環境間のインタラクション回数の上限を 300 回に制限しています。また、全 41 のタスクにおける 3 つの改訂版の結果を平均化して、平均カバレッジスコアを算出しました。各タスクのカバレッジ結果は、すべての改訂版で達成された機能の結合(ユニオン)によって決定され、最終的な平均スコアはその結果を単純に平均化したものです。不正防止のため、ExploitGym と同様にドメインのホワイトリストを適用し、基本ツールのインストールに必要な pypi.org や deb.debian.org などの必須ドメインのみを許可しています。
- FrontierSWE: 本評価は Proximal が実施したもので、コンテキスト長は 1M、最大努力レベル、出力トークン数は最大 128K です。ドミナンススコアは 2026 年 8 月 14 日時点のものです。
- PostTrainBench: GLM-5.3 の評価には、Claude Code 2.1.207 を「最大努力」レベルで実行し、温度パラメータ(temperature)を 1.0、top_p を 1.0、生成トークン数(max_new_tokens)を 128,000、コンテキストウィンドウを 1M トークンとして設定しました。3 回の試行結果の加重平均を報告します。スコアを出力できなかった試行については、公式のゼロショットベースモデルの基準値に置き換えています。第三者 API の利用を防ぐためのチェック項目では、ローカルの vLLM エンドポイントを OpenAI SDK を経由してアクセスした際に誤検知(偽陽性)が発生する原因となった、元のパターンマッチングに基づくチェックを削除しました。代わりに、外部 API の使用状況を LLM エージェントが検証する仕組みを採用しています。
- SWE-Marathon: GLM-5.3 の評価には、Claude Code 2.1.207 を最大限の努力レベルで実行し、温度パラメータを 1.0、top_p を 0.95、生成トークンの上限を 128,000、コンテキストウィンドウを 1M トークンに設定しました。
strip-cloneのケースでは、元の不正検出チェックがインポートの検出範囲を広げすぎたため、正当な実装まで誤って却下される問題がありました。
誤検知を避けるため、影響を受けたチェック項目を削除し、代わりに LLM による検査を実施しました。また、parameter-golf および trimul-cuda では NVIDIA のウェル(wheel)パッケージの変更により Docker イメージのビルドが失敗したため、--extra-index-url https://pypi.org/simple を追加して正常なビルドを復元しています。
原文を表示
Scaling post-training is all we did for GLM-5.3. With GLM-5.2 we built the stack: IndexShare for efficient long-context processing, SAO for RL on long-horizon tasks, and slime for large-scale asynchronous training — all running on the long-horizon task environments we have been accumulating. Over the past month we kept scaling on this stack: more environments, more diverse tasks, and more compute spent training on them.
Today we are releasing GLM-5.3. It uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:
- Stronger Coding: GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents' Last Exam.
- Emergent Cyber Capability: As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.
- Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.

Performance across comparison models
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | DeepSeek-V4Pro-0813 | Qwen3.8-Max | Opus 4.8 | Fable 5(w/ fallback) | GPT-5.6 Sol |
|---|---|---|---|---|---|---|---|---|
| Coding | ||||||||
| Terminal Bench 2.1 | 88.2 | 81.0 | 88.3 | 87.9 | 86.6 | 85.0 | 88.0 | 88.8 |
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | - | - | 21.1 | 33.7 | 34.6 |
| DeepSWEv1.1 | 66.9 | 46.2 | 67.5 | 62.7 | 56.6 | 58.0 | 69.7 | 72.7 |
| NL2Repo | 58.0 | 48.9 | 58.0 | 61.1 | 55.9 | 69.7 | - | - |
| ProgramBenchAlmost Solved | 19.0 | 9.5 | 17.5 | - | 10.5 | 15.5 | 33.0 | 23.0 |
| FrontierSWE | 78.1 | 67.5 | - | - | - | 66.5 | 88.2 | - |
| SWE-Marathonv1.1 | 42.5 | 19.4 | 48.1 | - | - | 48.8 | 33.1 | 42.5 |
| PostTrainBench | 39.8 | 31.7 | 32.0 | - | - | 32.9 | 41.8 | 36.2 |
| Cyber | ||||||||
| CyberGym | 84.5 | 77.2 | 80.0 | 83.3 | 78.5 | 78.1 | 83.8 | 83.6 |
| ExploitGym2h / 6h | 105 / 130 | 29 / 39 | 36 / 70 | - | 14 / 26 | 80 / 120 | 181 / 247 | 216 / 293 |
| ExploitBench | 54.4 | 24.4 | 32.2 | - | 28.8 | 40.0 | 78.0 | 76.5 |
| Agentic | ||||||||
| Toolathlon Verified | 73.0 | 59.9 | 76.5 | 74.1 | 72.5 | 76.2 | 74.7 | 74.9 |
| AutomationBenchv1.0.6 | 48.2 | 26.2 | 46.7 | 43.2 | 39.8 | 41.0 | 46.2 | 45.8 |
| Agents' Last ExamALE-CLI | 28.5 | 23.8 | 27.6 | 25.7 | 27.0 | 25.7 | 23.8 | 28.6 |
| HLE w/ Tools | 62.5 | 54.7 | 59.8 | 60.0 | 56.2 | 57.9 | 63.9 | 64.5 |
| GDPval-AA v2 | 1769 | 1508 | 1682 | 1590 | 1739 | 1588 | 1743 | 1730 |
Stronger Coding
For GLM-5.3, we pushed environment scaling toward tasks that look less like coding exercises and more like real units of expert work. The environments now cover a much broader range of production workflows, with tasks designed around how engineering and research work is actually carried out in practice. Some represent several days of work for an experienced engineer. In an ML infrastructure task, for example, the model may be given the same working environment as an engineer, with access to compute clusters, storage systems, internal documentation, codebases, and experiment results. It must diagnose bottlenecks across the training stack, implement optimizations, run experiments, and deliver a measurable end-to-end speedup while preserving correctness. Training on environments at this level pushes the model toward taking ownership of substantial work end to end, rather than relying on users to decompose the problem and supervise each step.
As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment. A useful task environment has to be executable, verifiable, and close to real professional work — and we need many of them, not a handful of hand-built ones. To scale this process, we built pipelines that synthesize environments end to end, and for a subset of tasks, the RL reward signal as well. Research agents collect task patterns from real work and turn them into runnable long-horizon environments with multi-step dependencies and hidden state; a judge agent then attempts each task to verify that it is actually solvable. Verifiers are synthesized without access to the reference solution, while solver trajectories are used to discover and close reward shortcuts. A verifier that passes oracle, no-op, and unsolved-state checks produces a binary reward reliable enough to train on directly.
It carries over the RL strategies introduced in GLM-5.2, including SAO with compaction, which helps these gains hold on long-horizon tasks rather than only on short ones. The effect shows up across both coding and general agent tasks. GLM-5.3 improves from 4.6 to 28.3 on Terminal-Bench 3.0, from 46.2 to 66.9 on DeepSWE v1.1, and from 23.8 to 28.5 on Agents' Last Exam. These pipelines still require a meaningful amount of human-in-the-loop work; making environment generation and verification more autonomous is one of the next steps.
Beyond public benchmarks, we introduce Z.ai Code Bench, an in-house benchmark designed to evaluate coding agents under realistic user scenarios. It covers diverse task categories and places agents in complex local development environments. At different effort levels, we evaluate agents along two dimensions: end-to-end task completion rate and fine-grained checklist accuracy. As a private benchmark, Z.ai Code Bench also reduces the risk of contamination from public test sets and gives us a more faithful measure of real-world user experience.

As shown in the figure, GLM-5.3 improves both performance and token efficiency. It delivers markedly stronger agentic coding results than GLM-5.2 at every effort level while consuming fewer output tokens. At Max effort, GLM-5.3 reaches 34.5% at roughly 75K output tokens per task, compared with 23.4% at 96K for GLM-5.2. The same shift holds against closed models. At High effort, GLM-5.3 reaches 31.4% at around 50K output tokens, surpassing Claude Opus 4.8 at 29.5% with 120K. GLM-5.3 remains behind Claude Fable 5, which reaches 39.5% at Max effort.
Emergent Cyber Capability
As part of post-training, we introduced vulnerability discovery data and environments into the training mix. We expected this to make the model better at finding and reasoning about vulnerabilities. What surprised us was how quickly the capability continued to develop as training scaled. GLM-5.3 did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains.

We evaluate GLM-5.3 across three benchmarks covering different stages of vulnerability analysis and exploitation. On CyberGym, which starts from white-box source code and tests whether the model can identify and validate vulnerabilities by triggering faults, GLM-5.3 scores 84.5%, up from GLM-5.2's 77.2% — the best result on the benchmark, ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%). On ExploitBench, which requires deeper reasoning about real vulnerabilities and their exploitation, GLM-5.3 reaches 54.4%, more than doubling GLM-5.2's 24.4%, while Mythos 5 and GPT-5.6 Sol score 78.0% and 76.5%, respectively. On ExploitGym, which measures how many exploitation tasks a model can complete under time-normalized budgets, GLM-5.3 completes 105 tasks within two hours and 130 within six hours, compared with 29 and 39 for GLM-5.2; budgets are normalized across models using per-model throughput figures, detailed in the footnotes. Mythos 5 remains well ahead at 181 and 247 tasks. The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the larger the gain from GLM-5.2 — and also the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind.
We then tested whether these capabilities transfer beyond controlled benchmarks. Since GLM-5.2, we have been working with several security teams in China to run our models against real-world codebases. After expert review, screening, and deduplication, the model identified 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high severity issues. The findings span system kernels, operating systems, browser engines, open-source infrastructure, web applications, and network protocols. Many had remained unnoticed for years or even decades, with the oldest dating back roughly 40 years.
This work has since grown into an ongoing disclosure effort. We built the Z.ai Security Disclosure Ledger to maintain a public record of the findings as they move through the disclosure process. The ledger is continuously updated as new vulnerabilities are reviewed and disclosed, distinguishing issues that have already been made public from those that are still under disclosure. For disclosed issues, it records information including the affected project, severity, CVE where available, and how long the vulnerability had remained in the codebase.
2,436
FINDINGS TRACKED
53
PUBLICLY DISCLOSED
2,383
UNDER EMBARGO
1,097
CRITICAL & HIGH
269
OSS PROJECTS
45
YEARS OF IMPACT
Findings span 45 years of impact - the oldest flaw was introduced in 1981, and on average a vulnerability lived 26.6 years before discovery.
SEVERITY DISTRIBUTION
Critical107High990Medium1,286Low53
WHEN THE FLAWS WERE INTRODUCED
19812026
View Z.ai Security Disclosure Ledger ↗
slime: Built for Long-Horizon RL Scaling
All of this runs on slime, our open-source post-training framework for RL scaling, with Megatron on the training side and SGLang on the rollout side. Its design keeps training, rollout, and the data buffer on a single dataflow, so math, code, sandboxes, verifiers, and long-horizon agentic environments plug in as data generation rather than as changes to the training loop. That is what let us keep adding environments through GLM-5.2 and GLM-5.3 without rebuilding the training stack each time.
Through GLM-5.3 we kept building it out on two fronts. On the algorithmic side we added capabilities aimed at RL research: top-p mask, top-k and full-vocabulary OPD, and configurations that improve training–rollout consistency, including R3-style setups and full numerical alignment between the training and rollout paths, which give us finer control over sampling, training, and teacher signals, and make it fast to run controlled comparisons. In our training–rollout consistency evaluation, the average difference in log probabilities (logprob) was controlled at the 1e-7 level, representing a reduction of more than 99.99% compared with previous setups.
We also worked on resource efficiency and system throughput for large-scale RL. Local storage now serves as an additional caching layer, holding model states and data hierarchically that would otherwise sit in host memory. This matters most for multi-teacher OPD: with dynamic teacher switching and prefetching on the training side, several teachers can be used without standing up a dedicated long-running inference service for each, at limited added overhead and substantially lower resource consumption. For agentic and asynchronous workloads, we improved joint scheduling and load balancing between the router and slime, so that rollout requests with widely varying lengths and completion times make better use of inference resources. We added workload-aware heuristics that derive throughput-oriented configurations — prefill/decode resource ratio, concurrency settings, and other throughput-critical parameters — from the characteristics of each rollout environment. As a result, for long-horizon coding RL tasks, these system-level optimizations improved end-to-end RL training throughput by more than 2.3×, allowing us to scale training over longer trajectories and more complex environments with substantially higher efficiency.
Taken together, these give us more experimental flexibility, lower resource cost, and higher throughput — which is what makes it practical to keep scaling RL.
Getting started with GLM-5.3
API Changes in GLM-5.3
GLM-5.3 supports three thinking effort levels: low, high, and max. Disabling thinking is no longer supported by GLM-5.3.
Thinking Parameters
| Parameter | Values | Default | Description |
|---|---|---|---|
thinking.type | enabled | enabled | Enables thinking. disabled is no longer supported. |
reasoning_effort | low, high, max | max | low: light; high: enhanced; max: deep. |
max is recommended for coding tasks.
{
"model": "glm-5.3",
"thinking": { "type": "enabled" },
"reasoning_effort": "max"
}
Migration required: If your application currently uses thinking.type: "disabled", change it to enabled and set reasoning_effort to low before updating the model ID to glm-5.3. Otherwise, the request will fail.
Use GLM-5.3 with GLM Coding Plan & ZCode
Try GLM-5.3 in your favorite coding agents—ZCode, Claude Code, OpenCode, and more. https://docs.z.ai/devpack/overview
For GLM Coding Plan subscribers: We’ve rolled out GLM-5.3 to all GLM Coding Plan users. The new GLM Coding Plan now uses a points-based quota system. Point usage is calculated separately for input, cached input, and output tokens. Model calls made outside peak hours consume 50% of the standard points. Peak hours are 14:00–18:00 (UTC+8), Monday through Friday; all other hours, including weekends, receive the 50% off-peak rate. Start building now: https://z.ai/subscribe
Get more from GLM-5.3 with ZCode
- 98%+ cache hit rate — repeated context billed at the lower cached rate, ~30% more effective tokens;
- 1.5x limited-time quota boost — stack it with the cache savings for up to 180% your standard quota through August 31.
- Long-horizon mastery — Goal mode plans, codes, tests, and verifies until the target is met;
- Remote Control — monitor and steer long-running tasks from your phone via WeChat or Feishu.
Try ZCode: https://zcode.z.ai
Serve GLM-5.3 Locally
The model weights of GLM-5.3 will be publicly available soon in two weeks.
Footnotes
- HLE w/ tools: We use sampling parameters of temperature=1.0 and top_p=0.95 for evaluation, with a maximum generation length of 163,840 tokens. The evaluation is conducted with a maximum context length of 300,000 tokens, using a context management strategy. We use GPT-5.6-luna (medium) as the judge model.
- NL2Repo: We evaluated NL2Repo with temperature=1.0, top_p=1.0, and max_new_tokens=64k under 1M context. To prevent hacking, we use rule-based and a LLM-based judgement to prevent malicious behaviors (e.g., unauthorized pip or curl operations).
- DeepSWE: We run DeepSWE using the mini-swe-agent harness with temperature=0.95, top_p=1.0, timeout=6h and 400K context.
- Terminal-Bench 2.1: We evaluate in Claude Code 2.1.207 with temperature=1.0, top_p=1, max_new_tokens=65536 with 6h timeout.
- Terminal-Bench 3.0: We evaluate Terminal-Bench-3 tasks with the Claude Code 2.1.207 harness (reasoning effort=max, 400K context, and 128K maximum output), reporting avg@3 over three rollouts per task. Each rollout runs in an isolated container built from the task's official image, and is capped at 600 agent turns with a 10-hour timeout. Tool Search is disabled, and the artifacts each agent produces are scored by the task's official separate verifier.
- Agent’s Last Exam (CLI): We evaluate ALE using the official evaluation protocol with the Claude Code harness (reasoning effort=max, 1M context, and 64K maximum output). Each of the 105 tasks runs in an isolated Docker container using the resources declared in its Task Card. The default timeout is 4 hours, with task-specific limits taking precedence (up to 8 hours). Tool Search is disabled, and results are scored by the official ALE evaluators.
- Toolathlon Verified: We obtain all results via the official evaluation service and report pass@1 averaged over 3 independent runs.
- AutomationBench: We evaluate on AutomationBench v1.0.6, incorporating the fix for the null-type handling issue introduced in PR #13.
- GDPval-AA v2: Models are evaluated by Artificial Analysis.
- CyberGym: We evaluate GLM-5.3 in Claude Code 2.1.207 (max reasoning effort, no web tools with temperature=1.0, top_p=1.0, max_new_tokens=128000). All evaluations are under unlimited timeout per task and results are single-run Pass@1 over 1,507 tasks. To simulate real-world usage scenarios, we place the agent inside the task container. We also remove all Git-related information and apply a domain whitelist (allowing only essential domains such as pypi.org and deb.debian.org for basic tool installation) to prevent the agent from cheating.
- ExploitGym: We evaluate GLM-5.3, Kimi-K3 and Qwen3.8 Max in Claude Code 2.1.207 (max reasoning effort, no web tools with temperature=1.0, top_p=1.0, max_new_tokens=128000). The reported results are single-run Pass@1 on 869 tasks under two timeout budgets: 2 hours and 6 hours, which are calculated as the API inference time rescaled by per-model tokens per second rate (per-model TPS sourced from Artificial Analysis; that is, we rescale GLM-5.3's results by 115 TPS, Kimi K3's results by 40 TPS and Qwen3.8 Max's results by 47 TPS), plus the non-API overhead. We also apply a domain whitelist (allowing only essential domains such as pypi.org and deb.debian.org for basic tool installation) to prevent the agent from cheating.
- ExploitBench: We evaluate GLM-5.3 in Claude Code 2.1.207 (max reasoning effort, no web tools with temperature=1.0, top_p=1.0, max_new_tokens=128000). Following the official evaluation settings, we limit the maximum number of interaction rounds between the agent and the environment to 300, and compute the average coverage score over all 41 tasks across 3 revisions. The coverage result of a task is determined by taking the union of capabilities achieved across all revisions, and the average score is obtained by averaging the results. We also apply a domain whitelist (allowing only essential domains such as pypi.org and deb.debian.org for basic tool installation) to prevent the agent from cheating.
- FrontierSWE: The evaluation was conducted by Proximal with 1M context length, max effort level, and 128K maximum output tokens. Dominance score reported as of 2026/08/14.
- PostTrainBench: We evaluate GLM-5.3 using Claude Code 2.1.207 with max effort level, temperature = 1.0, top_p = 1.0, max_new_tokens = 128000, and a 1M-token context window. We report the weighted average over 3 runs. Runs that fail to produce a score fall back to the official zero-shot base-model baseline score. For checks intended to prevent the use of third-party APIs, we removed the original pattern-matching-based checks, as they produced false positives when a local vLLM endpoint was accessed through the OpenAI SDK. Instead, we use an LLM agent to inspect solutions for external API usage.
- SWE-Marathon: We evaluate GLM-5.3 using Claude Code 2.1.207 with maximum effort level, temperature = 1.0, top_p = 0.95, max_new_tokens = 128000, and a 1M-token context window. For strip-clone, the original anti-cheat checks used overly broad import detection that could reject valid implementations. We removed the affected checks and performed llm-based inspection instead to avoid false positives. For parameter-golf and trimul-cuda, changes to the NVIDIA wheels caused the Docker image builds to fail, so we added --extra-index-url https://pypi.org/simple to restore successful builds.
同じ出来事を3媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み