中国製 AI モデル、コーディングタスクで低消費電力を実現
本文の状態
日本語全文を表示中
詳細モードで約8分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
IEEE Spectrum AI
北京のZ.aiが公開重みモデル「GLM 5.2」をリリースし、米国大手企業に匹敵するベンチマークスコアと圧倒的な低価格を実現したことで、中国企業のAI競争力向上と米国の優位性への懸念が浮上している。
AI深層分析を開く2026年8月4日 13:49
AI深層分析
キーポイント
GLM 5.2の技術的特徴とコスト構造
Z.aiは7530億パラメータを持つモデルを400億パラメータで動作させる最適化技術を導入し、API利用料を100万トークンあたり4.40ドルに設定した。
米国企業との性能比較とベンチマーク
GLM 5.2はFrontierSWEやPostTrainBenchなどのコーディングベンチでAnthropicのOpus 4.8とほぼ同点の結果を出し、セキュリティ分野でもAnthropicのMythosと比較される高性能を示した。
中国AI企業の台頭と市場動向
Stanford AI Indexによると、2025年の中国企業による注目すべきAIモデル数は米国の約半分まで増加しており、2023年や2020年と比較して顕著な成長を示している。
開発者のコスト意識と市場の影響
多くの企業がトークン予算を厳格に管理していない現状があり、安価で高性能なモデルの登場が米国のフロンティアAIラボに対する最大の参入障壁となり得ると指摘されている。
データ懸念への対応策として独自ホストが可能
GLM 5.2はオープンウェイトであるため、機密データの流出を恐れる組織が自社のハードウェア上でモデルをホストできる。これは通常API経由でしか利用できない最先端モデルとは対照的な特徴だ。
重要な引用
GLM 5.2 is an open-weights model, meaning any organization with sufficient hardware can download and host the model for free.
That price-be-damned habit... may now be the widest moat protecting the U.S. frontier AI labs.
The model nearly ties Opus 4.8's score on some agentic coding benchmarks.
"This one, I noticed that I could be using it for hours, and it would still have a coherent train of thought."
編集コメントを表示
編集コメント
Z.aiが公開したGLM 5.2は、性能とコストの両面で米国大手企業に挑戦する強力な存在であり、AI開発の民主化を加速させる可能性が高い。業界全体として、米国の優位性が揺らぐ中、中国企業の技術力向上に対する認識を改めて深める必要がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

Together AI の AI エンジニアである Zain Hasan 氏は、コストを意識しながら AI コーディングアシスタントの活用を習得しました。彼は難易度の高い課題には、推論能力や機能面で最先端に近い「フロンティアモデル」を割り当てます。Anthropic の Fable がその一例です。一方、Hasan 氏が外部委託するタスクが比較的単純な場合は、性能は低くてもコストも安い言語モデルを利用します。
現在、彼にとっての安価な選択肢は GLM 5.2 です。北京に拠点を置く Z.ai ラボが 6 月 16 日にリリースしたこのモデルはオープンウェイトであり、十分なハードウェアを備えた組織であれば誰でも無料でダウンロードしてホストできます。
Z.ai にアクセス料を支払う場合でも、コスト削減が可能です。同社の API は出力トークン 100 万個あたり 4.40 ドルです。これは Anthropic の Opus 4.8 モデルへのアクセス料の比較価格の 5 分の 1 未満であり、Anthropic の Fable コーディングモデルの 10 分の 1 です。出力トークンとは、プロンプトに対するモデルの応答として生成されるテキストの基本単位を指します。
しかし Hasan 氏によると、世界中の多くのソフトウェアエンジニアは、特定のコーディングプロジェクトにおける AI の実質的なコストについて、まだ十分に意識していません。
「現在、多くの企業がまだこの技術を理解しようとしている最中であり、トークン予算に余裕がないのが実情です」とハサン氏は語る。そして、誰かがコストを負担している場合、多くのソフトウェアエンジニアにとって合理的な選択は、コスト計算を完全に省略することだ。「最も強力なモデルを選ぶのが一番簡単だからです」。
この価格を無視する習慣は、現在のソフトウェア企業における緩いトークン予算によって強化されており、これが今や米国の最先端 AI ラボを守る最大の堀となっている可能性がある。
Z.ai が米国競合他社とのベンチマーク格差を縮小
Z.ai の GLM 5.2 は、7,530 億パラメータを持つ大規模言語モデル(LLM)だが、同時にアクティブになるのは 400 億パラメータのみだ。この最適化により、モデルの応答速度が向上している。同社は MIT オープンソースライセンスの下でこのモデルを公開しており、誰でも配布、複製、改変、利用することが可能だ。
GLM 5.2 の発表は、米国の AI 企業が競争優位性を失う可能性への懸念をさらに高めた。このモデルは、FrontierSWE や PostTrainBench といったエージェント型コーディングベンチマークにおいて、Opus 4.8 とほぼ互角のスコアを記録している。サイバーセキュリティ研究者たちもまた、GLM 5.2 がサイバーセキュリティベンチマークで高い評価を得ていることを発見しており、この能力は Anthropic の Mythos と比較されるきっかけとなった。
Z.ai の登場は、より広範なトレンドの一環だ。スタンフォード大学の AI インデックス(AI 動向を年次調査する 300 ページ以上の報告書)によると、中国企業の 2025 年に生み出した「注目すべき」AI モデルの数は、米国の競合他社の約半分だった。これは 2023 年の約 3 分の 1、2020 年の 5 分の 1 から大幅に増加した数字である。
GLM 5.2 の驚異的なベンチマークスコア、特にオープンウェイトモデルおよび中国製モデル全体における新記録達成により、一部の米国人観測者の間で懸念の声が上がりました。さらに、このモデルが中国発であるという事実は、機密データを中国関連のインフラ経由で処理することに警戒感を抱く米国やその他の国の企業にとって、利用を複雑にする要因となっています。
しかし、同モデルのオープンウェイト化により、回避策が存在します。データの送受信先を懸念する組織であれば、自社のハードウェア上でモデルをホストすることが可能です。これは、API の背後に閉じ込められ、セルフホスティングの選択肢がないことがほとんどである最先端レベルのモデルとは対照的です。
Z.ai は GLM 5.2 のリリースに合わせて自社データを裏付け、同日に研究レポートを発表しました。
このレポートでは、発表はされたもののまだ一般公開されていない Anthropic の Mythos や Fable には言及していません。代わりに焦点を当てているのは、Anthropic の Opus 4.8 と OpenAI の GPT-5.5 です。GLM 5.2 はベンチマークにおいて Opus 4.8 に匹敵する性能を示すことがほとんどありますが、同レポートが勝利と主張しているのは、難易度の低い推論タスクの 2 つだけであり、コーディング分野では勝敗がついていません。
GLM 5.2 の性能を評価するベンチマークの多くでは、エージェント型コーディングタスクにおいて Opus 4.8(場合によっては OpenAI の GPT-5.5)に劣るとされています。例えば、難易度の高い長期実行型のエージェント型コーディングベンチ「SWE-Marathon」では、GLM 5.2 が完了できたのは全タスクのわずか 13% でした。このベンチにおいて Opus 4.8 のスコアは GLM 5.2 の倍に達しています。また、Opus 4.8 は「NL2Repo」「DeepSWE」「Tool-Decathlon」といったコーディングベンチでも 10% 以上の差をつけて勝利を収めています。
コーダーたちは GLM 5.2 をどう使っているか?
GLM 5.2 を自らのワークフローと比較したソフトウェアエンジニアからは、多様な評価が寄せられています。
「私が [GLM 5.2] で気づいたのは、長時間にわたるタスクをこなせる点です」と語るのは、同社で GLM 5.2 を北米のインフラ上で運用している Hasan です。彼は以前のオープンウェイトモデルでは、往復のやり取りが 5〜15 回程度で思考の糸が切れてしまうことが多かったと指摘します。「一方、このモデルなら数時間使っても一貫した思考の流れを維持できることに気づきました」
デンバーに拠点を置く MetaRouter のシニアソフトウェアエンジニアである David Nix は、本業でも個人的なサイドプロジェクトでも LLM を活用しています(Nix は AI エンジニア向けの求人情報ボードも運営しています)。Nix によれば、GLM 5.2 は Anthropic の Opus や OpenAI の GPT-5.5 といった最先端モデルに「非常に近い」性能を持ち、彼のローテーションの中で恒久的なポジションを確保するに十分なレベルだということです。
「フロントエンド開発においては非常に優れています。例えば、Opus や Fable を毎回利用する必要がないからです」と Nix は語ります。Nix によれば、GLM 5.2 は LLM に依頼する業務の 10〜20% を処理しており、フロントエンドデザインなど特定のタスクではもはや最初の選択肢となっています。
同様の評価を寄せる声もあります。Hasan は GLM 5.2 のウェブデザインにおける「優れたセンス」を高く評価しています。ポーランド・クラクフの Screen Studio でソフトウェアエンジニアを務める Kacper Michalik 氏も、GLM 5.2 を活用して Web サイト用のフォームを作成する際、良好な結果を得ています。
一方で、インド・ベンガルールに拠点を置く Indhic AI のソフトウェアエンジニア、Sai Kiran Myadaram 氏は Z.ai についてはやや否定的な見解を示しています。GLM 5.2 がリリースされた週に登録したものの、トークンの割り当てがすぐに枯渇してしまったといいます。「Z.ai が提供する週次クォータは、私にとっては 2〜3 日以内に使い果たされてしまいます」と Myadaram 氏は話します。Z.ai に直接アクセスした Michalik 氏もモデルの品質に大きな問題はないものの、レート制限に occasionally 遭遇しました。ただし彼のケースでは無料プランを利用していたため、深刻な影響はありませんでした。
レート制限に加え、Myadaram 氏は Z.ai が簡単なフロントエンド修正を依頼された際に、ハルシネーション(幻覚)や過剰な計画立案の問題を経験しています。「コードベースが壊れてしまいます」と Myadaram 氏。その後、彼は再び OpenAI の Codex に戻っています。
原文を表示

Zain Hasan, an AI engineer at Together AI, has taught himself to use AI coding assistants while still keeping an eye on cost. He directs difficult problems to a frontier model, meaning one near the current state of the art in reasoning and capability, such as Anthropic’s Fable. But if the task that Hasan is outsourcing is more straightforward, he directs it to a less capable—and less expensive—language model.
Right now, the cheaper model, for him, tends to be GLM 5.2. Released on 16 June by the Beijing-based lab Z.ai, GLM 5.2 is an open-weights model, meaning any organization with sufficient hardware can download and host the model for free.
Those that pay Z.ai for GLM access still can save money, because the company’s API costs US $4.40 per million output tokens. That’s less than a fifth of the comparable price for access to Anthropic’s Opus 4.8 model, and a tenth the price of Anthropic’s Fable coding model. An output token is the basic unit of text a model generates in response to a prompt.
Yet many software engineers around the world, Hasan said, aren’t yet fully mindful of the net AI price tag for a given coding project.
“A lot of companies right now—they’re still trying to figure this technology out, and so there isn’t really a token budget,” said Hasan. And when someone else is paying, the rational move for many software engineers is to skip tabulating costs entirely. “The easiest thing is to pick the most powerful model.”
That price-be-damned habit, reinforced by loose token budgets in software companies today, may now be the widest moat protecting the U.S. frontier AI labs.
Z.ai Narrows Benchmark Gap With U.S. Rivals
Z.ai’s GLM 5.2 is an AI large language model (LLM) with 753 billion parameters, though it has only 40 billion parameters active at once—an optimization that improves the speed at which a model can respond. Z.ai released the model under an MIT open-source license, which means anyone can distribute, copy, modify, and use it.
GLM 5.2’s release added to fears that U.S. AI companies could lose their competitive edge. The model nearly ties Opus 4.8’s score on some agentic coding benchmarks, such as FrontierSWE and PostTrainBench. Cybersecurity researchers have also found that GLM 5.2 scores well in cybersecurity benchmarks, a capability that spurred comparisons to Anthropic’s Mythos.
Z.ai arrives amid a broader trend. According to Stanford’s AI Index (an annual, 300-plus-page survey of AI trends) Chinese companies produced just over half as many “notable” AI models in 2025 as their U.S. counterparts. That’s up from roughly a third in 2023, and a fifth in 2020.
GLM 5.2 caused hand-wringing among some U.S. observers due to its outstanding benchmark scores, which set new records for both open-weights models and Chinese-developed models generally. The model’s Chinese origin also complicates its use for companies in the U.S. and elsewhere that are wary of routing sensitive data through Chinese-linked infrastructure.
However, the model’s open weights provide an out. Any organization worried about where its data is being sent can instead host the model on its own hardware. This stands in contrast to most frontier-level models, which are gated behind an API with no self-hosting option.
Z.ai backed up GLM 5.2’s release with the company’s own numbers, publishing a research report the same day as GLM 5.2’s launch.
The report doesn’t mention Anthropic’s Mythos or Fable, which were announced but not yet publicly available at the time of its release.The report instead focuses on Anthropic’s Opus 4.8 and OpenAI’s GPT-5.5. And while GLM 5.2 often performs almost as well as Opus 4.8 in benchmarks, the report claims a win only in two less-difficult reasoning benchmarks—and none in coding.
Many of the report’s benchmarks place GLM 5.2 behind Opus 4.8 (and, at times, OpenAI’s GPT-5.5) in agentic coding. For example, GLM 5.2 completed just 13 percent of tasks in SWE-Marathon, a difficult long-duration agentic coding benchmark. Claude Opus 4.8 doubled GLM 5.2’s score in this benchmark. Opus 4.8 also notched wins of 10 percent or more in the coding benchmarks NL2Repo, DeepSWE, and Tool-Decathlon.
How Do Coders Use GLM 5.2?
Software engineers who’ve pitted GLM 5.2 against their own workflows report a wide range of results.
“The main thing that I realized with [GLM 5.2], was that it can do more long-horizon tasks,” said Hasan, whose company hosts GLM 5.2 on North American infrastructure. Earlier open-weights models, he said, often lost the thread after around 5 to 15 back-and-forth exchanges. “This one, I noticed that I could be using it for hours, and it would still have a coherent train of thought.”
David Nix, a principal software engineer at the Denver-based MetaRouter, puts LLMs to work at both his day job and for personal side projects. (Nix also operates a jobs board of AI engineers.) Nix said GLM 5.2 comes “really close” to frontier models like Anthropic’s Opus and OpenAI’s GPT-5.5—close enough to earn a permanent spot in his rotation.
“It’s pretty great at front-end development, for example, where I don’t need to always go to Opus or Fable for those things,” said Nix. He estimates that GLM 5.2 handles 10 to 20 percent of the work he sends to an LLM on a given day, and it’s now his first stop for some specific tasks, such as front-end design.
Others reported the same strength. Hasan said GLM 5.2 has “really good taste” in web design. Kacper Michalik, a software engineer at Kraków, Poland–based Screen Studio, received good results while using GLM 5.2 to create forms for use on a website.
On the other hand, Sai Kiran Myadaram, a software engineer at Bengaluru, India–based Indhic AI, reports less positive results with Z.ai. He signed up for Z.ai’s subscription plan the week GLM 5.2 launched and found the model burned through its token allotment quickly. “The weekly quota that Z.ai provides has been exhausted for me in less than two to three days,” he said. Michalik, who also accessed Z.ai directly, had no significant issues with the model’s quality but occasionally bumped into rate limits, though in his case he stuck to the free plan.
In addition to rate limits, Myadaram experienced problems with model hallucinations and overplanning when asked to tackle minor front-end fixes. “It’s messing up my code base,” said Myadaram. He’s since drifted back to OpenAI’s Codex.
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み