中国が西側 AI リードに追いついた、残る課題は何か
本文の状態
日本語全文を表示中
詳細モードで約46分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Decoder
中国のAIモデルが推論やコード生成などで西側と肩を並べる水準に達し、投資家が既存の技術格差への信頼を失い、OpenAI や Anthropic の IPO 計画にも影響を与えている状況を分析する。
AI深層分析を開く2026年8月21日 10:05
AI深層分析
キーポイント
中国モデルの実力向上と格差縮小
DeepSeek R1 から Moonshot Kimi K3、Alibaba Qwen3.8-Max、Zhipu GLM-5.3 にかけて、中国のオープンウェイトモデルが推論やコード生成など広範な評価で西側のトップレベルに迫っている。
投資家の懸念とビジネスモデルへの影響
モデル性能の差が数ヶ月単位まで縮小したことで、独自機能だけで事業を維持することが難しくなり、OpenAI や Anthropic の IPO 計画において不審な質問が投げかけられている。
学習手法とベンチマーク最適化への疑念
中国のラボが西側のモデルを教師として使用した知識蒸留(distillation)を行ったり、広範な能力よりもベンチマークスコアに特化したチューニングを実施しているという指摘がある。
中国モデルのベンチマークと実用性での台頭
中国のK3やQwen3.8-MaxはエージェンタイトスクルやCEOシミュレーションで西側モデルに匹敵する高いスコアを記録している。ただし、中国製モデルはトークン消費量が多く、コスト計算では依然として西側の優位性が一部残る。
西側のリードが縮小した領域
西側の技術的優位性は特定の「フロンティア」領域に限定されつつあり、モデル単体での明確な差は縮まっている。今後の競争の焦点はモデルそのものから、次世代モデルを生み出すシステム全体へと移行している。
重要な引用
Chinese models now sit near the top of almost every broad, demanding evaluation.
Measured by common benchmarks, the often-cited gap of a few months has shrunk enough to become an investor problem.
Whatever a model can do exclusively today, a freely downloadable one can do a few months later.
The American lead hasn't disappeared. But it has retreated to a few, ever-narrower areas of the so-called frontier.
編集コメントを表示
編集コメント
中国 AI モデルの急成長は、単なる技術追従ではなく、コスト構造や学習手法を含めた産業構造の変容を示唆している。西側の主要企業が直面する IPO や投資評価への影響は、今後の市場動向を左右する重要な転換点となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
1 年半前、DeepSeek R1 が業界に衝撃を与えました。中国の研究所が突然、OpenAI の o1(初の商用推論モデル)と競合するに至り、しかもはるかに低コストで実現したというのです。市場は動揺し、数日間で時価総額数十億ドルが消え失せました。「インフラ整備計画は過大評価だったのか」という声も上がりました。
しかし、当時の実情は報道 headline が示すほど単純ではありませんでした。DeepSeek 自身のレポート では、R1 は AIME 2024 などの個別テストで o1 を上回りましたが、事実知識(SimpleQA)などでは明確に劣勢でした。その後のベンチマーク はさらに大きな差を浮き彫りにしました。中国製モデルは特定の分野でトップに立つことはあっても、総合的な評価ではまだ一歩及ばない状況です。
先月後半、この特集の執筆を始めた時点ではまだ Z.ai の GLM-5.2 が同様の傾向を示していました。しかし最近、中国から最新オープンウェイトモデルが相次いで登場しました。Moonshot の Kimi K3 やアリババの Qwen3.8-Max、そして GLM-5.3 です。
中国製モデルは現在、ほぼすべての広範で難易度の高い評価項目においてトップクラスに位置しています。長文の知識処理や多段階にわたるコード生成、そしてツール連携においても、過去のモデルよりもはるかに信頼性が高いことが示されています。
一般的なベンチマークで測ると、これまで数ヶ月単位だと指摘されていた格差は縮小し、もはや投資家の懸念材料となっています。ウォール・ストリート・ジャーナルによると、Anthropic は直近の IPO を控え、中国勢との競争や政治的な逆風を理由に「不愉快な質問」を受けています。同社は自社のトップクラスでの優位性を強調して反論していますが、その下位層では中国製のオープンで安価なモデルが事実上支配しています。
投資家たちは、モデルの性能だけではもはやビジネスを成り立たせられないと懸念しています。今日、あるモデルにしかできないことがあっても、数ヶ月後には誰でも無料でダウンロードできるモデルが同じことができるようになるからです。
二つの批判があります。一つは、中国のラボが西側のモデルを教師として利用し、これを「蒸留(distillation)」と呼ばれる手法で行ったという指摘です。もう一つは、広範な能力に裏打ちされないままベンチマークでの高得点を得るためにモデルを調整しているという、「ベンチマックス化(benchmaxxing)」の疑いです。
米国のリードが完全に消えたわけではありません。しかし、それは技術的に可能なことの最前線である「フロンティア」の一部、そしてその領域もますます狭まっている一部の分野へと後退しました。このリードがどこに残っており、それが経済的にどの程度の価値を持つのか。これが今回の問いの第一点です。
第二の問いはこれに続きます。モデル単体ではもはや明確な差を生めなくなった今、そのリードは何に基づいているのでしょうか? 私たちの仮説は、モデル自体への依存度は次第に低下し、むしろその周囲にある全体システム、つまり継続的な研究開発によって次世代モデルを生み出すエコシステムへの依存度が高まっているという点です。
リードの残滓とは何か
発表直後、Artificial Analysis のインテリジェンス・インデックスにおいて K3 は 57 ポイントで第 3 位にランクされました https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/。その直前までリーダーだった GPT-5.5 と Opus 4.8 のすぐ後ろです。K3 は特にエージェントタスクにおいて大幅に改善しており、これは企業利用において実際に重要な作業能力です。さらに AutomationBench-AA では、Anthropic が Opus 5 で対抗するまで K3 は一時首位を記録しました。
「CEO-Bench」では、エージェントが架空のソフトウェア会社を 500 日間のシミュレーションで運営するテストが行われます。このベンチマークにおいて K3 は、2,215 万ドルという過去最高の単一実行結果を記録しました。一方、中国製モデル(その前身である K2.7 など)は、こうした長期間のタスクではこれまで頻繁に失敗してきました。Qwen3.8-Max も同様に高い総合性能を示しています。
ただし一つ注意が必要です。最新の中国製モデルの中には、西洋製モデルよりも遥かに多くのトークンを消費し、価格競争力を削ぐケースがあります。「コスト・パー・タスク」の計算式は依然として有効です。
現在のモデル能力の限界は、研究者らが「ジグザグなフロンティア」と呼ぶように、不均衡な速度で拡大しています。個々の能力が異なるペースで進化するため、「数ヶ月の格差」という言説も、誰が何を測定したかによって大きく変動してきました。
現在、西洋側がまだ優位性を維持できている領域は限られています。残されているのは主要な 3 つの分野です。
まず、抽象的な専門テストの結果です。K3 の発表直後、Opus 5 は 61 ポイント でランキングの首位に返り咲き、K3 の 57 ポイトとの差はわずかでした。ARC-AGI-1 は小さなパズルグリッドを用いた抽象的なパターン認識を測るテストですが、この指標では K3 と Fable 5(Anthropic がコーディングやエージェント作業向けに開発した主力ライン)はほぼ互角です。K3 のスコアは 94.5、Fable 5 は 98.5 パーセント です。差が顕著になるのは、ARC-AGI-2 だけです。ここでは K3 が 60.4 パーセントに対し、Fable 5 は 89.2 パーセントを記録しました。
ただし、ARC-AGI は日常業務からかけ離れた抽象的なパターン認識能力を意図的に測定しています。このスコアの差が実務上の違いを予測するとは限りません。多くのケースでは経済的な意味を持たないまま終わる可能性もあります。
2 つ目は信頼性です。8 月 12 日に発表された AA-AnalystAgent ベンチマーク は、実際の表や文書を用いたエージェントによるデータ分析能力をテストします。このベンチマークでは、モデルが独立した 5 回の試行すべてで正解した場合のみ、タスクの完了とみなされます。運営側はこの評価基準を「pass^5」と呼んでいます。Opus 5 は 54 パーセントで首位に立ち、GPT-5.5 の 50 パーセントを上回っています。一方、K3 はオープンモデルの中で最も高く 39 パーセントを記録しました。
この差は、ほとんどが再現性の低さに起因しています。K3 は 5 回の試行のうち少なくとも一度でタスクを 73% 解決しますが、これは Opus 5 の 74% とほぼ同等です。信頼性も、一般的な知能指数ランキングの順列には従っていません。GPT-5.6 Sol でさえ、ここでは自らの先行モデルに後れをとっています。
商業的な観点では、信頼性が極めて重要となります。これは、運営チームが ローンチ記事 で説明している通りです。アナリストエージェントは、回答をレビューなしで通用するものでなければ業務効率化には寄与しません。複数回の試行やチェック機能でカバーすることは可能ですが、その分、承認された結果あたりのコストは上昇します。
3 番目は サイバーセキュリティ です。ここでの差が最も明確に文書化されています。英国の AISI と米国の CAISI が共同で実施した 評価 によると、K3 は攻撃的なサイバー能力において主要な米国モデルに大きく遅れをとっています。
エクスプロイト開発用のテスト「ExploitBench」では、K3 のスコアは 32% でした。一方、上位の米国モデルの平均は約 76% です。標的システム上でコードを実行する必要がある 41 のタスクすべてで K3 は失敗しましたが、主要な米国モデルは平均して 20 を解決しました。また、模擬攻撃では、K3 が 32 ステップ中 17 ステップまで到達したのに対し、米国のリーダーたちは平均 28.5 ステップに達していました。
しかし、この格差も縮まりつつあります。Z.ai の自社測定によると、8月14日に公開された「GLM-5.3」はExploitBenchで54.4%のスコアを記録し、前作であるGLM-5.2の倍以上の性能を示しました。これにより、米国のトップモデルとの差がわずか1ヶ月で約半分になりました。さらに、ソースコードの脆弱性発見と検証(CyberGym)においては、GLM-5.3は米国主導のモデルをもわずかに上回っています。
一方で、セキュリティ分野ではプロバイダーが自社の最強機能を公然と販売していないという事情があります。Anthropicの最上位サイバーモデル「Mythos 5」はExploitBenchで78%を達成していますが、「Project Glasswing」を通じた限定された条件下でのみ利用可能です。一方、一般公開されている兄弟モデルであるFable 5は、上流の安全対策が410回のテストエピソードのうち407回をブロックしたため、旧世代のOpus 4.8(スコア40%)と同等のレベルに留まっています。
OpenAI も同様に、Daybreak や専門的なサイバー攻撃対応のバリアントにおいて類似のアプローチを採用しています。GPT-5.6 Sol は、OpenAI 自身の評価 OpenAI's own evaluation において ExploitBench で 73.5% のスコアを達成しましたが、OpenAI のリスクスケールで最高レベルである「Critical」にはまだ達していません。しかし、間もなく登場する Astra については、同社が初めて「自社のモデルがその閾値を超える可能性も否定できない」と述べています can't rule out for the first time。もし実際にその壁を越えた場合、Preparedness Framework がより厳格に発動し、開発の完全停止に至る可能性もあります。実際、OpenAI はすでに一部の内部 Astra 関連作業を一時停止しています。
一方、中国の研究所が GLM-5.3 を通じて初めてこのパターンを採用しました。Z.ai は追加の安全性確認のため、重みの公開を約2週間延期し、最もセンシティブなサイバー機能については認証されたユーザーに限定する計画です。
残るリードの差は以下のようになります。抽象的な専門テストにおけるギャップは測定可能ですが、その経済的価値はまだ不明確です。商業的に重要なのは信頼性のギャップであり、これは最も議論が分かれる点です。同じ研究所内でも、新しいモデルが先行モデルより劣る場合(GPT-5.6 Sol と GPT-5.5 の比較など)、ランキングは明確に変動している状態にあります。
サイバーセキュリティの格差は、通常のチャネルでは顧客がほとんど購入できない機能に関わるものでしたが、この格差も最近になって急速に縮小しています。危険な一部の機能のみを壁で囲むことは可能ですが、コーディングや研究、エージェント作業など商業的に価値のある残りの部分は、API を通じてアクセス可能にし続ける必要があります。ビジネスとして成立させるためにはそうせざるを得ないのです。そして、この点こそが、数ヶ月のうちに優位性が縮小する原因となっています。
最も明白な疑念は、これら二つの事象が関連しているという点です。中国の研究機関が、前述した蒸留(distillation)のように、自らのモデルを訓練するために Western モデルを API を通じて教師として利用している可能性があります。もしそうであれば、その優位性はアクセス権限の欠如する領域にのみ存在することになります。
この疑念には妥当な理由があり、中国の研究機関がどのようにして Western モデル内部の知識から利益を得ているのかについては後ほど詳しく解説します。
しかしまず、戦略的な評価においては「罪」の有無は二次的な問題です。もしその疑いが真実であれば、販売されたあらゆる機能は最終的に競合他社へ移行することになります。なぜなら、データを盗み取る行為(skimming)を防ぐ確実な手段がまだ存在しないからです。もし疑いが誤りであるならば、中国の研究機関は自力で追いついたことになりますが、その場合でも優位性は以前よりもさらに価値の低いものだったと言えます。
どちらの見解を取っても結論は同じです。モデルの優位性が維持されるのは、製品がそもそも広く提供されていない領域に限られます。しかし、そのような領域だけでビジネスを構築することは不可能です。蒸留のメカニズムに目を向けることは依然として意義があります。なぜなら、その仕組み次第で、防衛可能なモデル機能自体が存在するかどうかが決まるからです。
蒸留は速度を説明しても、広がりまでは説明できない
OpenAI と Anthropic は、中国の企業が自社のモデルを大規模に利用して独自モデルを開発していると非難しています。Anthropic によると、DeepSeek、Moonshot、MiniMax による産業的なキャンペーンでは、約 24,000 の不正なアカウントを通じて 1,600 万回以上のやり取りが行われたとされています。特に Moonshot に帰属されるキャンペーンは、エージェントの推論、ツール使用、コーディング、そして推論プロセスの追跡に焦点を当てており、340 万回を超えるインタラクションが確認されました。
これは、Together AI が実務レベルのソフトウェア課題において K3 と Fable 5 のタスク成功率間に 0.72 という高い相関関係を見つけたという発見と一致しており、Typebulb のスタイル分析でも K3 が Fable 5 に最も近いと評価されたこととも整合しています。
批判派は、Fable 5 の公開から K3 の登場までの期間が短すぎて学習への影響は考えられないと反論します。しかし、密接に関連するモデルである Opus 4.8 はもっと長い期間利用可能であり、同様の研究でもそこにも明確な重複が見られています。OpenAI もまた、米国議会上院特別委員会宛ての覚書において DeepSeek と同様のパターンを指摘しています。
ただし、どのラボも検証可能な証拠を発表しておらず、OpenAI や Anthropic への取材においても新たな情報は得られていません。
それでも、既知の告発や複数の研究論文を照らし合わせることで、Moonshot や DeepSeek といった企業がトレーニング中に知識蒸留(distillation)からどのような恩恵を受けたかを再構築することは可能です。
事前学習(pretraining)では、モデルは膨大な量のテキストから基礎的な能力を習得します。中期学習(midtraining)では質の高いデータに切り替わり、その後、教師あり微調整や強化学習を通じて振る舞いが形成されます。蒸留されたデータ は、これらの工程のほぼどこにでも組み込むことが可能です。
一流の Western モデルは、解法パスやツール呼び出しを含む数万件の選択されたタスクに応答します。その中から最も優れた回答が中期学習や微調整へと流れ込み、学生モデル(student model)が教師モデルの解決パターンを吸収します。
強化学習においては、教師となるモデルが生きている必要さえありません。収集したデータは報酬モデル(reward models)の構築に利用できます。
Anthropic のレポートはこの考え方が単なる空想ではないことを示しています。DeepSeek に起因するとされるキャンペーンにおいて、Claude は事前定義された評価基準を用いて大規模な採点タスクを処理し、実質的に他者の強化学習のための報酬モデルとして機能しました。マルチ教師オンポリシー蒸留 や 報酬モデルの蒸留 に関する研究は、この手法がトレーニングパイプライン全体でどのように機能するかを明らかにしています。
根本的な制限はアクセス権限にあります。外国の API を経由して知識蒸留を行う場合、インターフェースが返す情報しか見ることができません。重み値、内部評価プロセス、検索プロセスなどはすべて隠されたままです。ブラックボックス蒸留に関する研究(Research on black-box distillation)は、まさにこの制約を詳述しています。
しかし、推論モデルの登場が新たな突破口を開きました。それは「推論チェーン」を読み出すことです。各ラボはこの可視性を意図的に制限しています。OpenAI は思考の連鎖について要約のみを提供しており、Anthropic は限定的な形で拡張思考を表示し、xAI は推論を暗号化して送信しています(詳細はこちら)。
ただし、情報の抽出を完全に防ぐことは不可能でした。セキュリティ研究者アレクサンダー・パンフィロフ率いるチームが 8 月に発表した研究(arXiv:2608.09867)では、主要モデルの暗号化された推論トレースを API の脆弱性を通じて読み出すことが可能であることを示しています。その後、各プロバイダーはこの隙間を埋める対応を行いました。
同論文は、蒸留仮説に対する最も具体的な証拠を初めて提示したものです。K3 の推論プロセスに、復号化された Opus の推論トレースから数トークンを読み込ませると、その回答が明確に Opus 寄りに変化します。また、個々の Claude や GPT の推論パスを K3 から抽出するのは、次に類似度の高いモデルと比較して最大で 100 万倍も容易であることが示されました。
研究チームは、K3 のサイバーセキュリティ分野での弱さが一つのヒントになると読み解いています。そこでは強力な教師モデルへの到達が困難だったためです。なぜなら Anthropic はその能力を追加のセーフティガードで保護しているからです。
中国のラボ自体の実績も、蒸留が完全な説明ではないことを示唆しています。Moonshot のレポート では、社内だけで構築されたトレーニングパイプラインの詳細が記されていますし、DeepSeek は常に最下層のアーキテクチャ改善を定期的に公開しています。GLM-5.3 もまた、サイバー分野のこのヒントを相対的な視点で捉えています。Z.ai は自社の飛躍をポストトレーニング効果によるものとしており、まさに西洋の教師モデルへの到達が最も困難な領域において、独自のトレーニング環境を活用したのです。
蒸留は中国特有の手法でもありません。Elon Musk は宣誓証言の中で、xAI が OpenAI のモデルを利用したことを確認しました。一方、ByteDance は他のモデルからのデータを使わずに Doubao-1.5-pro をポストトレーニングしたと主張しています(詳細はこちら)。
標的となった研究機関は政府による保護を求めており、ワシントンもそれに応えようとしています。ホワイトハウスは「米国モデルに対する蒸留キャンペーンを国家安全保障上の脅威と宣言する」大統領令に署名し、米 AI モデル盗難防止法は委員会での全会一致で可決されました。さらに上院議員らが BLADE 法 を提出し、後押ししています。Anthropic の Mythos 5 と Fable 5 は現在 輸出管理 の対象となっており、米国政府がモデルそのもののアクセスを規制するのは初めてです。
ただし、オープンモデルや蒸留全般への制限というより広範な措置には、まだ十分な支持が集まっていません。7 月下旬、Nvidia、Microsoft、Meta など 25 社が 早期の制限に反対する声明 を発表しましたが、OpenAI と Anthropic はこれに参加しませんでした。
マーク・ザッカーバーグ氏はさらに、競合他社のモデルから学習することを含む「観測可能なすべてのものから学ぶ」という原則を宣言しました。これは守るべき価値ある考え方です。しかし、この発言には自らの利益が絡んでいます。Meta はクローズドなモデルの開発で中国に遅れをとっており、米国のオープンウェイトモデルを使って対抗しようとしています。それでも、重要な示唆が含まれています。これらのリソースを持つ企業が、他社モデルの出力から学習することを原則として掲げるということは、それが追いつくための最も迅速な手段の一つだと考えていることを意味します。
つまり、知識蒸留(distillation)がそのスピードを説明する要因の多くであり、手がかりは数多く存在します。ただし、Anthropic や OpenAI がより多くの情報を開示しない限り、それらはあくまで推測に過ぎません。
いずれにせよ、米国のプロバイダーが現在 API を通じて提供しているものは、数ヶ月以内に競合他社へ移行してしまいます。彼らが隠しているのは時間を稼げるだけのものであり、出力を「観測不可能」にする法律は存在しません。
残されているのは、最初から提供されていないものだけです。それは販売することもできません。米国のリードが持続するためには、その優位性がモデルそのもの以外のどこかに存在していなければなりません。
エージェントにおいては、システム全体が重要な単位となる
OpenAI に移籍したアメリカ AI アクションプランの共著者、ディーン・ボール氏は、以前「今後私たちが向かうべき方向」という記事で、数日間自律的に動作するシステムは、チャットボットとは全く異なる課題を提起すると指摘しました。具体的には、責任の所在や監督体制、そして実行プロセスの制御に関する問いです。
責任の問題と同様に、競争においても同じことが言えます。価値を生むのが数日間にわたって継続するワークフローになった今、競合するのはもはやモデル単体ではなく、そのワークフローを支えるシステム全体です。そこには、モデルを取り巻く制御ソフトウェア、提供される計算リソース、そしてツールや権限、復元ポイントと連携して動作する環境が含まれます。
このインフラがどれほど重要かを測る指標は、少なくとも部分的に存在します。英国 AI セキュリティ研究所の調査では、タスクあたりのトークン数が 100 万から 1,000 万に増えた際、ソフトウェアエンジニアリングの結果が約 25% 向上したことが示されています。また、一部のサイバーセキュリティ関連の課題は、この閾値を超えて初めて解決可能になることが明らかになっています。
制御ソフトウェア(ハーンネス)は、さらに大きな影響を及ぼします。OpenAI が ARC-AGI-3 評価において、推論状態をステップ間で保持し、コンテキストの圧縮方法を変えただけで、Sol のスコアは約 6 分の 1 の出力トークン数で、13.3% から 38.3% にほぼ 3 倍に跳ね上がりました。同じモデルでも、システム次第でパフォーマンスが 3 倍になるのです。
しかし、制御ソフトウェアだけが防衛壁となるわけではありません。外国のエージェントシステムを集中的に利用する人は誰でも、回答とともにツール呼び出しを確認でき、そこからオーケストレーションのパターンを学習できます。ハーンネスは結局のところソフトウェアであり、Anthropic が今年春に学んだ通りです。Claude Code のソースコード流出事件では、誤って公開されたソースマップによって約 513,000 行のソースコードが露呈しました。しかし、クローン版がすでにそのアーキテクチャを公的に保存していたため、削除要請は実効性を失いました。同様のハーンネスはそもそもオープンになっています。OpenAI の Codex も DeepSeek のエージェントツールもオープンソースです。
エージェントシステムの中でコピーできないのは、稼働中の運用部分だけです。顧客システムの基盤となる認証情報、テスト環境、ログ、介入権限などは、照会や漏洩の対象にはなりません。競合他社がこれを手に入れるには、自社のインフラと顧客を備えて自ら構築するしかありません。
なぜこのポジションがモデルのリーダーシップよりも価値があるのか。エージェントの能力を決定づける最大の要因は、目的別に設計されたトレーニング環境における強化学習です。具体的には、テスト付きのプログラミングタスクやシミュレーションツールを利用したタスク、検証可能な結果を伴うタスクなどが含まれます。さらに、専門知識を購入することも可能です。これはすでに数十億ドル規模の市場となっています。
これらのリソースは、支払能力さえあれば誰でも利用できます。中国の研究機関も例外ではありません。これが中国が追いついた理由の一つです。しかし、買えないものがあります。それは、実際の企業でどのようなタスクが発生するのか、モデルはどこで失敗するのか、そして顧客がどのような結果を受け入れるのかという知識です。
OpenAI の子会社である DeployCo の CTO、アルノー・フーニエ氏は、編集チームとのインタビューでこのフィードバックチャネルの仕組みについて説明しました。彼は、「誰かが明確に要請しない限り」、顧客データを学習データとして使用することはないと述べています。また、そのような研究パートナーシップは稀であるとも指摘しています。
むしろ、このフィードバックチャネルはモデルの弱点やツールの必要性に焦点を当てています。ある顧客チームが文書理解機能の精度が低いことに気づいた場合、研究側がターゲットを絞ったデータを収集して修正を行います。これが、大手銀行 BBVA 向けに構築されたソリューションが GPT-5.0 から 5.5 へと劇的に改善した理由です。
また、顧客からのオーケストレーション(調整・連携)の要望が最初に応用され、オープンソースのリポジトリ Swarm が生まれ、続いて Agent SDK が登場しました。
つまり、このチャネルが運ぶのは生データそのものよりも、「次に何を構築すべきか」という知見の方が重要です。これが知識蒸留(ディストillation)ではコピーされない部分です。蒸留は完成したモデルを複製するだけであり、次の学習で利用される作業へのアクセス権までは複製しません。このアクセス権を持つ者は、実際の業務に即したトレーニング環境を構築できます。一方、単なる複製しかできない者は、昨日の状態からしか学べません。
ここで重要になるのは、プロバイダーが顧客の導入実績から何を引き継いでモデルやトレーニング環境に反映できるかという契約上の課題です。
ただし、このフィードバックチャネルの重要性を過大評価してはいけません。実際の進歩の多くはまだ研究所内部で起こっており、現場での顧客利用とは別問題です。それがどれほど価値を持つのかは、未解決の研究課題にかかっています。「導入時の経験が、どのようにしてモデルの能力へと転換されるのか」という根本的な問いに答えが出るまで、その真価は不明です。
Dwarkesh Patel は、真に能力のある作業エージェントにとっての最大の障壁は、モデルが継続学習を行えない点にあると指摘しています。つまり、モデル生成から次のモデル生成へと段階的に学ぶのではなく、経験をその場で重みに取り込む能力が欠如しているのです。そのため、Andrej Karpathy は、エージェントが本当の意味での同僚として活躍できるようになるまでには約 10 年かかると予測しています。
一方、Nathan Lambert は、そのような重みの更新は必須ではないと主張します。スケーリングを進めれば、より優れたメモリシステムによって実質的に同じ効果が得られるからです。もし彼の説が正しければ、蓄積された経験は顧客が管理するデータの中に存在することになり、事業者へのロックイン効果は弱まります。
しかし今回の問いにおいては、この見解の違いは結果に影響しません。エージェントが実際に担う業務が増えれば増えるほど、その業務にアクセスできるかが、適切なモデルを構築できるかどうかを決定づけるからです。完成したモデルは複製可能ですが、次のモデルの成長を支える業務へのアクセス権はそうはいきません。
US labs tie model, compute, and platform together
モデルのリードを維持しようとするラボほど、そのアクセス権限を確立するために必死に働いているところはありません。GPT-5.6 は最高難度レベルでデフォルト 4 つのエージェントを調整し、OpenAI の Codex はクラウド上で完全に動作する複数のタスクを並列実行します。また Anthropic も Claude Code や Claude Cowork で同等の機能を提供しています。
OpenAI が Ona を買収したことで、顧客がアクセス権限や認証情報、セキュリティ境界を自ら管理できる安全な実行環境が強化されました。
DeployCo の構造は、OpenAI がこのフィードバックチャネルをいかに重要視しているかを示しています。同社は 19 のプライベート・エクイティ企業と共同で設立された本格的なスピンアウト企業であり、社内に直接進出するエンジニアを配置して製品や研究への知見の還元を図っています。一方、Frontier Alliance は Accenture、Capgemini、BCG、McKinsey と連携し、営業面での拡大を進めています。
ハードウェア側からも、Nvidia 自身が同じ方向へ舵を切っています。新しい Vera Rubin シリーズでは、ツールの呼び出しに対応するサンドボックス環境や、長時間動作するエージェントのためのコンテキストメモリ、そしてオーケストレーション機能をプラットフォームの一部として統合しています。チップサプライヤーはもはやアクセラレーターだけでなく、エージェントシステムそのものを販売するようになっています。
こうしたハードウェアに加え、研究機関たちは物理的な第二の防衛線となる計算インフラの構築にも注力しています。
OpenAI の共同創設者兼プレジデントである Greg Brockman は、計算コストについて最も明確に述べています。オープンソースモデルが自動的に安価になるわけではなく、最終的にはすべて同じハードウェア上で動作します。OpenAI は、あらゆるタスクに対して最安の提案を実現するために、独自チップとデータセンターの構築を進めています。
もしモデルがコモディティ化すれば、受け入れられた結果あたりの価格が勝敗を分けます。エネルギー供給から自社製チップ、そしてモデルまでを一貫して統合できる企業が、その価格を引き下げることになるからです。
この賭けが的中すれば、完璧なモデルのコピーであっても商業的には敗北します。なぜなら、オリジナルの方がスケールメリットを活かして同じ性能をより安価に提供できるからです。これはまだ実現した事実ではなく目標ですが、中国製モデルの価格競争力を見れば、その現実味は明白です。
インフラが競争優位性(モート)となるのは、フィードバックチャネルと組み合わされた場合のみです。これにより、展開で得た知見を次世代モデルに転換するトレーニング能力と、エージェント群を並列実行するための推論能力が提供されます。
このインフラ構築は依然として賭けに過ぎません。エージェント需要の成長速度が、モデルの効率化を上回る場合にのみ成功します。サム・アルトマン氏自身が、この計算式がいかに不確実であるかを示しています。彼は ビジネスにおけるコストを巨大な問題と呼んでいます。
中国は、米国の輸出規制により最新ハードウェアへのアクセスが阻害されているため、この層を独自に構築することも、通常の経路で計算リソースを借りることもできません。残された選択肢は回避策です。一つは密輸であり、「ゲートキーパー作戦」 だけで少なくとも 1.6 億ドル相当の GPU が押収されています。もう一つは、海外で計算リソースを借りることです。
いずれも主権を有する独自スタックの代替にはならないため、中国は自前のインフラを整備する必要に迫られています。現状では、ハードウェアを投入することで対応しています。
Huawei の CloudMatrix384 は 384 枚の Ascend アクセラレータをリンクしており、SemiAnalysis の分析によると、Nvidia の GB200 NVL72 を上回るメモリ容量とピーク計算性能を備えています。ただし、消費電力は約 4 倍、必要なアクセラレータ数は 5 倍に達します。
物理的な基盤は拡大しています。中国国家エネルギー局の統計では、2025 年までに 42 の AI クラスターが稼働する見込みです。各クラスターには少なくとも 1 万枚のアクセラレータカードが搭載されます。
しかし、電力を投入してもメモリチップやソフトウェアの質は向上しません。SemiAnalysis は現在、Ascend の生産における最大のボトルネックは HBM(High Bandwidth Memory)だと指摘しています。
Huawei が提供する CANN は、Nvidia の CUDA に相当するソフトウェア層で、ドライバー、ライブラリ、プログラミングツールなどを通じて開発者が Ascend チップを動かすための基盤です。CANN は初日から DeepSeek V4 をサポートしましたが、これは中国の AI 自立推進における大きな成果とされています。
一方で、Nvidia の GB300 はスループットにおいて明確に先行しています。
中国の政策は、データ・計算資源・基準に関する政府行動計画を通じて自国の技術スタックを国境を超えて広めることを目指しています。また、「オープンウェイト」がその架け橋となりつつあります。中国製モデルがまだ NVIDIA の CUDA 上で最もよく動作する限り、間接的に米国のスタックを強化することになりますが、もし Ascend 向けに早期最適化が進めば、中国製のハードウェアも第三国にとって魅力的なものになります。NVIDIA のジェンソン・フアン CEO はまさにこの相互依存関係について警告しており、DeepSeek V4 がリリース初日からサポートを開始したことは、その取り組みが成果を上げ始めている初期の証拠と言えます。
結論は二つの側面を持ちます。ベンチマーク競争については、K3 や GLM-5.3 の実績が示す通り、中国にも勝利の可能性はあります。一方、産業化された全体システムはコピーできず、自ら構築するしかありません。その点では米国の研究機関が先行していますが、永遠に続くわけではありません。モデル単体の優位性よりも長く持続するでしょう。API を通じて提供できない領域こそが守備範囲となります。つまり、実際の業務から得られる知見と、それを活用するためのインフラです。
モデルの導入は価値観の輸入を意味する
モデルを構築するのではなく、既存のモデルを活用して展開している企業、つまり欧州企業の大多数にとって、別の重要な問いが浮上します。それは、モデルを採用する際に具体的に何を引き受けることになるのかという点です。
ポストトレーニング(学習後調整)によって、モデルがどの質問に答え、どの質問を拒否するか、どの情報源を信頼できるものとして扱うか、何が確立された事実とみなすかが決定されます。これらの判断は規範的な性質を持ち、ベンチマークで測定されることはなく、展開時に予期せぬ形で付随してきます。
DeepSeek、Qwen、Kimi、Z.ai など、オープンウェイト版の利用が拡大しているモデルファミリーを含む中国の 7 社によるチャットボットに対する NewsGuard の最近の 監査 は、その差がいかに大きかを具体的に示しています。中国に有利な 10 の事実上誤った主張に対して、中国製のボットは 53% のケースで誤情報への反論を行いませんでした。一方、西側の比較システムではその割合は 24% にとどまりました。
この差の大部分は沈黙によるものです。中国製ボットは 24% のケースで回答を拒否しましたが、西側のシステムではわずか 0.5% でした。また、拒否されたプロンプトの 88% が台湾に関するものでした。
背景には、その定義を明確にせず「政治的順守」を要求する規制があります。中国が低コストのオープンソース AI モデルを通じて国家プロパガンダを輸出しているという指摘もその一例です。当然ながら、中国製ボットは 38% のケースで国営メディアを引用しましたが、西側のシステムでは 16% でした。この分野が均一でないことは明らかで、MiniMax は誤った主張の 85% を debunk(反証)した一方、百度の Ernie は 60% のケースでそれらを繰り返しました。
この危険性は、学習された境界条件が稀なトピックと実際の意思決定が交差する際に初めて表面化するという点にあります。モデルがエージェントチェーンの奥深くに位置すればするほど、問題は後になって現れ、その影響は連鎖的に拡大していきます。
これはすぐにビジネス上の問題へと発展します。台湾の港湾封鎖という噂について問われた際、DeepSeek はその出来事を詳細に確認しましたが、実際にはそのような事実は存在しません。この虚構の世界を前提にサプライチェーンや投資リスクを計算することは、現実離れした判断を下すことになります。
少なくとも、これはハードウェア的に固定されたものではありません。NewsGuard によると、CTGT ラボが報告したところでは、DeepSeek から蒸留されたモデルには政治的な検閲機能が引き継がれていなかったそうです。価値観と機能は、異なる学習層に存在しているようです。ただし、その作業を誰かが最初に実行する必要があります。
ヨーロッパはシステム主権を必要としている
ヨーロッパにとって、この変化は朗報と悲報の両方を含んでいます。朗報とは、最も明確に敗北したレース、つまり最強の完成モデルをめぐる競争が戦略的価値を失いつつあるという点です。なぜなら、販売されたモデルは数ヶ月以内にコピーされるからです。
しかし、悲観的な側面の方が重く影響します。将来、リードを維持するために守られるべき層は、ダウンロードしたり簡易的に模倣したりできるものではありません。それは構築されなければなりません。そして、その点において出発点は不均衡です。EuroStack イニシアチブ の試算では、ヨーロッパはデジタルインフラの 80% 以上を輸入に依存しており、クラウド市場における欧州プロバイダーのシェアはわずか約 15% に過ぎないとされています。
「規制が原因だ」という説明は便利ですが、AI 法が施行される前からこの格差は存在していました。AI は過去のデジタル化の波によって築かれたインフラの上に成り立っており、欧州にはその際に重要だった二つの要素が欠けていました。すなわち、グローバルクラウドを有するプラットフォーム企業と、数十億規模のモデル学習に必要な資金力です。
さらに資本の縮小も問題です。European Investment Bank の試算によると、米国の企業が毎年集めるベンチャーキャピタルは欧州企業の 6〜8 倍に達し、大規模な資金調達ラウンドの 5 つのうち 4 つを外国投資家が主導しています。
そして資金を提供する側が、企業を海外へ移転させたり、買収したりすることも珍しくありません。DeepMind、ARM、そして最近では Silo AI の事例がそれを示しています。問題なのは研究そのものではありません。JRC のデータによると、2023 年の生成 AI に関する世界全体の論文の約 21% が欧州から発表されましたが、関連する特許出願はわずか約 2% です。「欧州は発明し、他国が産業化している」という構図です。
欧州の一部の議論で見られる、「大規模言語モデルは統計的なオウム返しに過ぎない」として切り捨てる反応も、状況を好転させるには役立ちませんでした。トランスフォーマーが真の知性への最もエレガントな道かどうかは別問題として、それらの上に構築されたシステムを制御する必要があるかどうかが問われています。
Mistral が、その方向性が正しいことを示しています。欧州のモデルを牽引する同社はフルスタックプロバイダーへと進化し、8 月 11 日に地域エンドポイントを発表しました。また、自社プラットフォーム上で他社のオープンウェイトモデルのホスティングや、欧州全体の計算需要のプール化も進めています。2030 年までに最大 1 ギガワットの容量を確保する計画です。
これにより、Mistral は独自のモデル、独自プラットフォーム、そして計画中の計算リソースという 3 つの要素をすべて備えた、欧州唯一のプロバイダーとなりました。
しかし、提供内容はまだ完全ではありません。今回のテーマであるエージェントワークロードに対応するエンドポイントはまだ整備されていません。また、他社のオープンモデルをホスティングすることには 2 つの弱点があります。Qwen を運用しているのは誰かによって、その展開ノウハウは Qwen の開発元であるアリババが掌握してしまいます。欧州側は、このノウハウを自らのモデル改善に転換するしかありません。さらに、モデルに埋め込まれた価値判断は、ポストトレーニングで何らかの調整が行われない限り、フィルタリングされずにそのまま流れてしまうリスクがあります。
一方、競合他社は欧州大陸において、欠けているブロックの構築を進めています。OpenAI の DeployCo は、Tomoro から約 150 名の展開専門家を吸収し、パリ、ロンドン、ミュンヘンにオフィスを開設しています。
紙面上では、EU は問題認識を示しています。InvestAI によって AI 分野に 2,000 億ユーロ(そのうち 200 億は AI ギガファクトリー向け)を動員する計画です。Cloud and AI Development Act はデータセンターの容量を少なくとも 3 倍に拡大することを目指しており、Jupiter in Jülich は欧州初のエクサスケールシステムとして稼働しています。しかし、これらの構想の多くは未だ実現していません。ギガファクトリーの入札は数回延期されており、最初の施設が稼働するのは 2027 年以降になると見られています。一方、米国の大手企業 4 社だけで 2026 年に AI 分野に約 7,000 億ドルを投資する可能性があり、これは数年かけて実施される欧州全体の取り組みの 3 倍に相当します。
何より重要なのは、この資金が米国ラボの優位性を支える要素の一部しか賄えていないという点です。前章で示した通り、その優位性は「計算資源でモデルを訓練し、モデルが顧客のエージェントシステムで動作し、その運用が次世代のための知見を生む」というサイクルの上に成り立っています。
ギガファクトリーは最初のステップに過ぎません。これは不可欠なステップです。なぜなら、自前の計算リソースがなければ、その上のすべての層が coercion(強制)のリスクにさらされるからです。中国が強制的なハードウェア迂回策を余儀なくされた事例がそれを示しています。しかし、このサイクルはデータセンター内で完結するわけではありません。実際のデプロイ(展開)段階で初めて閉じられるのです。
その結果、当然の答えとして、米中双方のオープンウェイトモデルに全賭けするという選択肢が残ります。しかし、これには3つの反対意見があります。
まず懸念されるのは供給の問題です。次世代のオープンモデルへの権利は保証されておらず、そのリリースは撤回可能なビジネス判断に委ねられています。
Meta は発表していた「Behemoth」モデルを保留しました。また、ライセンス上の理由からマルチモーダルな Llama モデルは EU 企業には提供されませんでした。米国政府は Fable 5 と Mythos 5 の輸出規制を課し、Fable 5 については解除されるまで 18 日間待たされました。
ロイター通信によると、北京当局も中国の最良モデルへの外国企業のアクセス制限を検討しており、オープンソースモデルもその対象に含めています。さらに Z.ai も当初は GLM-5.3 の重み(weights)を公開していませんでした。
こうした状況の背後には、より根深い分断が横たわっています。主要な研究機関で働く 25 人の研究者へのインタビュー調査では、AI 研究そのものを自動化できるモデルが製品として登場する可能性を予想した回答者はわずか 4 人だけでした。大多数は、最も強力なバージョンは社内にとどまり、競合他社に販売して格差を縮めるのではなく、むしろ自社の優位性をさらに拡大させるものだと考えています。
アクセス可能な市場、つまりオープンソースまたは API を通じて利用可能なモデルは、結局のところ第 2 世代のものに限定されてしまいます。これは、この章の冒頭で示された「販売されたものだけがコピーされる」という良いニュースにも、最初から歯止めをかける要因となります。
二つ目の反論は前章からの続きです。外国製のモデルを採用するということは、そのモデルに組み込まれた世界観を引き受けることを意味します。中国製モデルには国家による検閲が、米国製モデルには民間企業のコンテンツポリシーがそれぞれ埋め込まれており、ワシントンの政権交代に応じて後者は頻繁に変更される可能性があります。
したがって、システム主権を確立するには、規範的な層の構築も不可欠です。能力だけでなく情報の完全性や拒否行動も測定する独立した評価インフラと、学習後の段階でオープンモデルを調整できる能力が必要です。CTGT の調査結果は技術的に可能であることを示していますが、欧州でこれを体系的かつ監査可能な形で実施している組織があるかどうかは、まだ不明な点です。
三つ目の反論として、国内開発モデルだけでは主権には程遠いという指摘があります。なぜなら、主権はスタックのあらゆる層において個別に脆弱だからです。中国がその典型例と言えます。アリババや字节跳动(ByteDance)はモデル層を支配し、国内データセンターも保有しています。しかし、最新の Nvidia チップは同国への輸入が禁止されており、利用可能な代替品ではチップあたりおよびワットあたりの性能が大幅に劣ります。さらに両社は、フラッグシップモデルの学習のために、海外事業者から借りた Nvidia ハードウェアに依存せざるを得なくなったとの報道もあります。
欧州は、他者が制御するレイヤーがどれほど速く政治的武器へと転化するかを身をもって体験しました。米国が国際刑事裁判所の検察官に制裁を加えた際、マイクロソフトはその検察官のメールアクセスを停止したのです。主権とは、他者が制御する最も弱い層の強さ次第でしかあり得ません。
今、決定的な新レイヤーがこれらに加わろうとしています。それは「実際の業務へのアクセス」です。自由に入手可能なインターネットデータはほぼ枯渇しつつあります。今後モデルを向上させる鍵となるのは、専門知識、プロセスに関するノウハウ、そして導入経験です。これらは一部には購入したトレーニングデータとして提供されますが、これは短期間で数十億ドル規模の市場へと成長しました。しかし何よりも重要なのは、「どの能力に学習コストを投じるべきか」という知見です。
完成したモデルは複製可能です。しかし、次のモデルが学習する元となる業務へのアクセスは複製できません。その業務が流れるシステムを運営する者は二重の利益を得ます。現在のモデルによる成果からの収益に加え、次世代モデルに必要なタスクを把握できるという優位性です。これは強化学習によって可能になったことです。
欧州にとって、これは単なるビジネス機会の問題ではありません。世界における自国の地位の基盤そのものが問われているからです。欧州の影響力は、単一市場の規模、基準設定能力、そして高付加価値な知識労働に由来しています。
この研究は二つの側面から圧力を受けています。まず、AI がその価値を低下させており、これはすでに若手開発者向けのエントリーレベル市場の縮小として顕在化しています(Stanford の調査では、ChatGPT などの AI ツールが露出度の高い分野で若年層の雇用を大幅に減少させたことが示されています)。同時に、AI はその価値を外国のシステムへと移転させています。
企業売却による従来の流出に加え、第二のより迅速なチャネルが現れています。専門知識は、専門家を実験室へ仲介するデータベンダーを通じて、また欧州の中核プロセスに組み込まれる展開ユニットを通じて、断片的に外国のモデルへと流れ込んでいます。これは製造業の時代に見られた低賃金国への流出ではなく、外国の AI モデル内への流出です。今回は知的資本が国外へ去る事態となっています。
同時に、この研究こそが欧州にとって今回の競争における最大の資産です。それはモデル競争で争われる原材料であり、欧州にはそれが豊富にあります。
もしこれが外国のシステムを通じて流れるなら、欧州は二重に支払うことになります。知識を提供し、次の世代のモデルとして結果を買い戻すのです。もしこれが自国のシステムを通じて流れるなら
原文を表示
A year and a half ago, DeepSeek R1 caused a shock. A Chinese lab was suddenly competing with OpenAI's o1, the first commercial reasoning model, and had reportedly done it for far less money. Markets got nervous. Billions in market value evaporated within days. Was the planned infrastructure buildout overblown?
The picture back then was murkier than the headlines suggested. In DeepSeek's own report, R1 beat o1 on individual tests like AIME 2024 but trailed clearly on others, such as factual knowledge (SimpleQA). Later benchmarks exposed more gaps. Chinese models only reached the top in individual disciplines, not across the board.
As recently as late June, when we started working on this issue, Z.ai's GLM-5.2 still showed the same pattern. Then came the latest Chinese open-weights models: Moonshot's Kimi K3, Alibaba's Qwen3.8-Max, and GLM-5.3.
Chinese models now sit near the top of almost every broad, demanding evaluation. They handle long knowledge tasks, code across many steps, and coordinate tools far more reliably than their predecessors.
Measured by common benchmarks, the often-cited gap of a few months has shrunk enough to become an investor problem. According to the Wall Street Journal, Anthropic is fielding uncomfortable questions ahead of its upcoming IPO and points to its remaining lead at the top in its defense. Below that tier, the field belongs largely to open, far cheaper models from China.
Investors worry that raw model performance can barely carry a business anymore. Whatever a model can do exclusively today, a freely downloadable one can do a few months later.
Two accusations are in play: Chinese labs allegedly tapped Western models as teachers, a practice known as distillation. And they allegedly tune their models for strong benchmark scores without the broad capabilities to match - so-called benchmaxxing.
The American lead hasn't disappeared. But it has retreated to a few, ever-narrower areas of the so-called frontier, the leading edge of what's technically possible. Where that lead still sits, and what it's worth economically, is the first question of this issue.
The second follows from it. If the model alone can't make a clear difference anymore, what does the lead rest on? Our thesis: less and less on the model, and more and more on the overall system around it, meaning the system where ongoing work produces the next models.
What's left of the lead
At launch, Artificial Analysis had K3 in third place on its Intelligence Index with 57 points, right behind then-leaders GPT-5.5 and Opus 4.8. K3 improved most on agentic tasks, so exactly the hands-on work that matters in enterprise use. On AutomationBench-AA, K3 even debuted in first place until Anthropic answered with Opus 5.
On CEO-Bench, where an agent runs a fictional software company for 500 simulated days, K3 posted the best published single run at $22.15 million. Chinese models, including its predecessor K2.7, had regularly failed these long hauls. Qwen3.8-Max reaches a similarly high overall level. One caveat remains: Newer Chinese models sometimes burn far more tokens than Western ones, which eats into part of their price advantage. The cost-per-task math still applies.
The edge of current model capabilities has been growing at uneven speeds for a while, what researchers call the "jagged frontier." Individual capabilities advance at different rates, which is why the much-quoted months-long gap always depended on who measured what. What's new is where a Western lead is still measurable at all. Three areas remain.
The first is abstract specialty tests. Shortly after K3's launch, Opus 5 retook the top of the index with 61 points, a small gap over K3's 57. On ARC-AGI-1, a test of abstract pattern recognition using small puzzle grids, K3 and Fable 5 - Anthropic's flagship line for coding and agent work - sit practically even at 94.5 and 98.5 percent. Only on ARC-AGI-2 does the gap widen: 60.4 versus 89.2 percent. But ARC-AGI deliberately measures abstract pattern recognition far removed from everyday tasks. There's no guarantee this gap predicts practical differences. It could simply stay economically irrelevant in most cases.
The second is reliability. The AA-AnalystAgent benchmark, launched August 12, tests agentic data analysis on real tables and documents. It only counts a task as solved if a model gets it right in five out of five independent runs, a measure the operators call pass^5. Opus 5 leads with 54 percent, ahead of GPT-5.5 at 50. K3 is the best open model at 39 percent.
The gap comes almost entirely from poor repeatability. K3 solves 73 percent of tasks at least once in five attempts, practically even with Opus 5 at 74. Reliability doesn't follow the usual intelligence-index rankings either. Even GPT-5.6 Sol falls behind its own predecessor here.
Commercially, reliability weighs heavily, as the operators explain in their launch article. An analyst agent only saves work when its answers hold up without review. Multiple runs and checkers can compensate, but they drive up the cost per accepted result.
The third is cybersecurity. Here the gap is best documented. A joint assessment by the UK's AISI and the US CAISI found that K3 lags far behind leading US models in offensive cyber capabilities.
On ExploitBench, a test for developing exploits, K3 scored 32 percent. The top US models averaged about 76. K3 failed all 41 tasks that required executing code on a target system; leading US models solved 20 on average. And in a simulated attack, K3 made it to step 17 of 32, the US leaders to 28.5 on average.
But this gap is shrinking too. GLM-5.3, unveiled August 14, scores 54.4 percent on ExploitBench by Z.ai's own measurement, landing on more than double of its predecessor GLM-5.2. That cuts the distance to the US leaders roughly in half within a month. On finding and validating vulnerabilities in source code (CyberGym), GLM-5.3 even edges past the leading US models.
At the same time, cybersecurity is the one area where providers don't openly sell their strongest capabilities. Anthropic's best cyber model, Mythos 5, which hits 78 percent on ExploitBench, is only available under controlled conditions through Project Glasswing. Its public sibling Fable 5 effectively stays at the level of the older Opus 4.8 at 40 percent, because upstream safeguards intercepted 407 of 410 test episodes.
OpenAI takes a similar approach with Daybreak and specialized cyber variants. GPT-5.6 Sol reaches 73.5 percent on ExploitBench in OpenAI's own evaluation but stays below Critical, the highest level on OpenAI's risk scale. With the upcoming Astra, OpenAI can't rule out for the first time that one of its own models crosses that threshold. If it does, the Preparedness Framework kicks in much harder, up to a full development halt. OpenAI has already paused some internal Astra work.
And with GLM-5.3, a Chinese lab is adopting this pattern for the first time. Z.ai is delaying the weights release by about two weeks for extra safety work and plans to limit the most sensitive cyber functions to verified users.
The remaining lead adds up like this: The gap in abstract specialty tests is measurable, but its economic value is unclear. The reliability gap matters most commercially but is the least settled - when a newer model falls behind its own predecessor within the same lab (GPT-5.6 Sol vs. GPT-5.5), that ranking is clearly in flux.
The cybersecurity gap involves capabilities almost no customer can buy through regular channels, and even that gap has narrowed sharply of late. Only the few dangerous capabilities can be walled off. The commercially valuable rest, like coding, research, agent work, has to stay accessible through APIs, or there's no business. And that's where the lead shrinks within months.
The obvious suspicion is that the two are connected. Chinese labs could be using Western models as teachers through their APIs (the distillation mentioned earlier) and the lead would then hold exactly where that access is missing. There are good reasons for this suspicion, and we'll get into how Chinese labs might profit from the knowledge inside Western models below.
But first, for the strategic assessment, the question of guilt is almost secondary. If the suspicion is true, every capability sold eventually migrates to the competition, because there's still no reliable way to prevent the skimming. If it's not true, Chinese labs caught up on their own, and the lead was worth even less.
Both readings lead to the same place. A model lead only holds where the product isn't broadly offered in the first place, and you can't build a business on that. A close look at distillation is still worthwhile, because its mechanics decide whether any defensible model capabilities exist at all.
Distillation explains the pace, not the breadth
OpenAI and Anthropic accuse Chinese companies of using their models at scale to build their own. According to Anthropic, industrial campaigns by DeepSeek, Moonshot, and MiniMax ran more than 16 million interactions through around 24,000 fraudulent accounts. The campaign attributed to Moonshot targeted agent reasoning, tool use, coding, and reasoning traces with more than 3.4 million interactions. That fits with Together AI finding a 0.72 correlation between the task-level success rates of K3 and Fable 5 on real software problems, and with a style analysis by Typebulb placing K3 closest to Fable 5.
Critics counter that the window between Fable 5's release and K3's was too short to influence training. But Opus 4.8, a closely related model, was available much longer, and the same studies show clear overlaps there too. OpenAI describes a similar pattern with DeepSeek in its memorandum to the US Congress. None of the labs has published verifiable evidence, though, and our inquiries with OpenAI and Anthropic turned up nothing new.
Still, the known accusations and several research papers make it possible to reconstruct how Moonshot, DeepSeek, and others could benefit from distillation during training.
In pretraining, a model learns basic capabilities from massive amounts of text. Midtraining switches to less but higher-quality data, and post-training shapes behavior through supervised fine-tuning and reinforcement learning. Distilled data can plug in at almost any of these stages.
A leading Western model answers tens of thousands of selected tasks, complete with solution paths and tool calls. The best answers flow into midtraining or fine-tuning, and the student model picks up the teacher's solution patterns.
For reinforcement learning, the teacher doesn't even need to be live anymore. The collected data can be used to build reward models.
Anthropic's report shows this is no thought experiment. In the campaign attributed to DeepSeek, Claude processed grading tasks at scale using predefined rubrics, effectively serving as a reward model for someone else's reinforcement learning. Research on multi-teacher on-policy distillation and reward model distillation shows how the method works across entire training pipelines.
The hard limit is access: Anyone distilling through a foreign API only sees what the interface returns. Weights, internal evaluations, and search processes stay hidden. Research on black-box distillation describes exactly this constraint.
The rise of reasoning models opened another door, however: reading out the reasoning chains. Labs deliberately restrict that visibility. OpenAI provides only summaries of the chains of thought, Anthropic shows extended thinking in limited form, and xAI transmits reasoning encrypted.
The extraction was never fully preventable, though. A study published August by a team led by security researcher Alexander Panfilov shows that the encrypted reasoning traces of top models could be read out through an API vulnerability. The providers have since closed the gap.
The same paper delivers the most specific evidence yet for the distillation thesis. When K3's reasoning is prefilled with a few tokens from decrypted Opus reasoning traces, its answers shift measurably toward Opus. And individual Claude and GPT reasoning passages could be extracted from K3 up to six orders of magnitude more easily than from the next-most-similar model. The team also reads K3's weak cyber results as a clue: a strong teacher was hard to reach there, because Anthropic shields those capabilities with extra safeguards.
The Chinese labs' own substance argues against distillation as the full explanation. Moonshot's report documents a training pipeline built entirely in-house, and DeepSeek regularly publishes architecture improvements at the deepest level. GLM-5.3 also puts the cyber clue in perspective. Z.ai attributes its own jump to post-training effects - its own training environments, in exactly the area where a Western teacher is hardest to reach.
Distillation isn't a Chinese specialty either. Elon Musk confirmed under oath that xAI used OpenAI models for it. ByteDance, meanwhile, claims to have post-trained Doubao-1.5-pro without data from other models.
The targeted labs are calling for government protection, and Washington seems willing to deliver. The White House has declared distillation campaigns against US models a national security threat by memorandum, the Deterring American AI Model Theft Act cleared committee unanimously, and senators followed up with the BLADE Act. Anthropic's Mythos 5 and Fable 5 are now under export controls - the first time the US government has controlled access to a model itself.
The broader step of restricting open models or distillation in general lacks wide support, though. In late July, 25 companies including Nvidia, Microsoft, and Meta warned against premature restrictions. OpenAI and Anthropic didn't sign.
Mark Zuckerberg then even declared learning from everything observable, including distilling competing models, a principle worth protecting. That's self-serving - Meta trails on closed models and wants to take on China with American open-weights models. But it shows something. When a company with these resources elevates learning from other models' outputs to a principle, it considers this one of the fastest ways to catch up.
In short: distillation can explain much of the pace, and the clues are plentiful. But as long as Anthropic and OpenAI don't disclose more, clues are all they are.
Either way, everything US providers' APIs put out currently migrates to the competition within months. What they hide may only buy time, and no law makes outputs unobservable.
The only thing that holds is what's never offered in the first place - and you can't sell any of that. If the American lead is going to last, it has to live somewhere other than the model.
With agents, the whole system becomes the unit that matters
Dean Ball, co-author of the American AI Action Plan and now at OpenAI, argued before his move that systems acting autonomously for days raise entirely different questions than chatbots, like about liability, oversight, and control of running processes.
What applies to responsibility applies equally to competition. Once value comes from workflows that run for days, it's no longer just models competing but the systems that carry those workflows: the control software around the model, the compute it gets, and the environment where it works with tools, permissions, and restore points.
How much this scaffolding matters can be measured, at least partly. The UK AI Security Institute shows that software engineering results rose about 25 percent when a model got ten million tokens per task instead of one million. Some cyber tasks only became solvable above that threshold.
The harness, meaning the control software, has an even bigger effect. When OpenAI merely kept the reasoning state between steps and compressed context differently in an ARC-AGI-3 evaluation, Sol's score nearly tripled, from 13.3 to 38.3 percent, at a sixth of the output tokens. Same model, different system, triple the performance.
But the control software only goes so far as a moat. Anyone who queries a foreign agent system intensively sees the tool calls alongside the answers and can learn orchestration patterns from them. And the harness is software in the end, as Anthropic learned this spring. In the Claude Code leak, an accidentally shipped source map exposed roughly 513,000 lines of source code, and the takedowns went nowhere because clones had long since preserved the architecture in public. Comparable harnesses are open anyway. OpenAI's Codex is open source, as are DeepSeek's agent tools.
Only one part of an agent system stays uncopyable: its running operation. The anchoring in customers' systems - credentials, test environments, logs, intervention rights - can't be queried or leaked. A competitor can only build it themselves, with their own infrastructure and their own customers.
Why is this position worth more than a model lead? The biggest driver of agent capabilities is reinforcement learning in purpose-built training environments, meaning programming tasks with tests, simulated tools, tasks with verifiable results, plus purchased expert knowledge, long since a billion-dollar market.
Both are available to anyone who can pay, including Chinese labs, which partly explains the catch-up. What can't be bought is knowing which tasks actually come up in real companies, where models fail there, and what customers accept as a result.
Arnaud Fournier, CTO of OpenAI's deployment subsidiary DeployCo, described how this feedback channel works in an interview with our editorial team. OpenAI doesn't train on customer data "unless someone explicitly asks us to," and such research partnerships are rare, he says.
Instead, the channel runs through model weaknesses and tool needs. If a team at a customer finds that document understanding works poorly, the research side sources targeted data and fixes it. That's how the solution built for the major bank BBVA improved markedly from GPT-5.0 to 5.5. And customers' orchestration needs first produced the open-source repository Swarm, then the Agent SDK.
So the channel carries less raw data than knowledge about what to build next. That's exactly what distillation doesn't copy. It copies the shipped model, never the access to the work the next one learns from. Whoever has that access builds training environments that match real work. Whoever only copies always learns from yesterday's state. The key contract question becomes what providers may carry over from customer deployments into their models and training environments.
This feedback channel shouldn't be overstated, though. Most progress still happens inside the labs, not at the customer. How much it will be worth depends on an open research question: how deployment experience turns into model capability at all.
Dwarkesh Patel sees the central obstacle to truly capable working agents in models' lack of continual learning, that is the ability to fold experience into their weights on the fly instead of learning only from model generation to model generation. Andrej Karpathy therefore puts agents about a decade away from true coworker status.
Nathan Lambert argues those weight updates are dispensable, because scaling plus better memory systems delivers practically the same thing. If he's right, the accumulated experience would live in data the customer controls, and the lock-in to the operator would be weaker.
For this issue's question, it makes no difference. The more real work agents take on, the more access to that work decides who can build the right models. A finished model can be copied. Access to the work the next one grows from cannot.
US labs tie model, compute, and platform together
No one is working harder to lock in that access than the labs whose model lead is shrinking. GPT-5.6 coordinates four agents by default at its highest effort level, OpenAI's Codex runs multiple tasks in parallel, now fully in the cloud, and Anthropic has comparable offerings in Claude Code and Claude Cowork. OpenAI's Ona acquisition adds secure execution environments, where customers control access, credentials, and security boundaries themselves.
The structure of DeployCo shows how seriously OpenAI takes this feedback channel. The company is a genuine spinout, launched together with 19 private equity firms, and embeds forward deployed engineers directly inside corporations - explicitly also to carry lessons back into product and research. Meanwhile, the Frontier Alliance with Accenture, Capgemini, BCG, and McKinsey scales the sales side.
Even Nvidia is pushing the same development from the hardware side. The new Vera Rubin generation integrates sandbox environments for tool calls, context memory for long-running agents, and orchestration as part of the platform. The chip supplier increasingly sells the agent system, not just the accelerator.
On top of this and other hardware, the labs are building a second, physical line of defense: compute infrastructure.
OpenAI co-founder and president Greg Brockman states the calculation most clearly. Open models aren't magically cheap, everything ultimately runs on the same hardware, and OpenAI is building out chips and data centers to make the cheapest offer for every task.
Because if the model becomes a commodity, the price per accepted result decides, and whoever integrates the stack from energy through their own chips to the model pushes that price down.
If the bet pays off, even a perfect model copy loses commercially, because the original offers the same capability cheaper at scale. That's still a goal, not reality, as the price advantages of Chinese models show.
Infrastructure only becomes a moat in combination with the feedback channel. It provides the training capacity to turn deployment knowledge into the next generation, and the inference capacity to run agent swarms in parallel.
The buildout is still a bet. It only pays off if agent demand grows faster than models get more efficient. Sam Altman himself shows how open this math is when he calls the costs for businesses a huge problem.
China can neither distill this layer nor rent it through regular channels, since US export controls block access to the latest hardware. That leaves workarounds. One is smuggling - Operation Gatekeeper alone covered GPUs worth at least $160 million. The other is renting compute abroad.
Neither replaces a sovereign stack, so China has to industrialize one, so far by throwing more hardware at the problem. Huawei's CloudMatrix384 links 384 Ascend accelerators and, according to SemiAnalysis, delivers more memory and peak compute than Nvidia's GB200 NVL72, but at about four times the power, with five times as many accelerators. The physical base is growing. China's energy administration already counts 42 AI clusters for 2025, each with at least 10,000 accelerator cards.
But more energy replaces neither memory chips nor software quality. SemiAnalysis currently sees HBM as the main bottleneck in Ascend production. And while Huawei's CANN - the counterpart to Nvidia's CUDA, the software layer of drivers, libraries, and programming tools developers use to run Ascend chips - supported DeepSeek V4 from day one for the first time, Nvidia's GB300 was clearly ahead in throughput.
China's policy therefore aims to spread its own stack beyond its borders, with a government action plan for data, compute, and standards. Open weights are becoming the bridge. As long as Chinese models run best on Nvidia's CUDA, they still indirectly strengthen the American stack. But if they get optimized for Ascend early, they make Chinese hardware attractive in third countries too. Nvidia CEO Jensen Huang warned about exactly this coupling, and DeepSeek V4's day-one support is early evidence the effort is paying off.
The conclusion cuts two ways. The benchmark race is winnable for China, as K3 and GLM-5.3 have shown. The industrialized overall system, by contrast, can't be copied, only built - and there the US labs are ahead. Not forever, but longer than a model advantage ever would be. Only what doesn't fit through the API is defensible: the knowledge from real work and the infrastructure to exploit it.
Deploying a model means importing values
For everyone deploying models instead of building them - which describes the vast majority of European companies - another question comes up. What exactly do you take on when you adopt a model? Post-training decides which questions a model answers and which it refuses, which sources it treats as credible, and what it treats as proven. These decisions are normative, no benchmark measures them, and they come along unbidden at deployment.
A recent audit by NewsGuard of chatbots from seven Chinese providers, including DeepSeek, Qwen, Kimi, and Z.ai, exactly the model families whose open weights see growing use, documents how large the differences are. On ten provably false pro-China claims, the Chinese bots left the misinformation unchallenged in 53 percent of cases. Western comparison systems: 24 percent.
Most of this gap comes from silence. The Chinese bots refused to answer in 24 percent of cases, Western ones in 0.5 percent, and 88 percent of the refused prompts concerned Taiwan.
Behind this is regulation that demands political conformity without defining its terms. Fittingly, the Chinese bots cited state media in 38 percent of cases, Western ones in 16 percent. The field isn't uniform. MiniMax debunked 85 percent of the false claims, while Baidu's Ernie repeated them 60 percent of the time.
This is dangerous mainly because the trained-in boundary sometimes only surfaces when a rare topic meets a real decision. The deeper the model sits in agent chains, the later the problem shows up and the further it cascades.
It can become a business problem fast: Asked about an alleged blockade of Taiwanese ports, DeepSeek confirmed the event in detail, even though it never happened. Anyone basing supply chain or investment risks on that is calculating with an invented world.
At least it isn't hardwired. According to NewsGuard, the lab CTGT reported that a model distilled from DeepSeek didn't carry over the political censorship. Values apparently sit in different training layers than capabilities. But someone has to do that work first.
Europe needs system sovereignty
For Europe, this shift holds good news and bad. The good news: the race Europe lost most clearly, the one for the strongest finished model, is losing strategic value, because every model sold gets copied within months.
The bad news weighs more.:The layer where leads will be defended in the future can't be downloaded or skimmed. It has to be built. And there, the starting position is lopsided. The EuroStack initiative estimates that Europe imports over 80 percent of its digital infrastructure, and European providers hold only about 15 percent of the cloud market.
The convenient explanation is regulation, but the deficit predates the AI Act. AI builds on the infrastructure of previous digitalization waves, and Europe lacked both things that mattered there: platform giants with global clouds and the cash flows for billion-scale model training. Then there's the capital squeeze. By European Investment Bank calculations, US companies attract six to eight times as much venture capital each year, and a foreign investor leads four out of five large European funding rounds.
And whoever puts up the money often moves the company abroad or buys it outright, as DeepMind, ARM, and most recently Silo AI show. Research isn't the problem. According to the JRC, about 21 percent of the world's publications on generative AI in 2023 came from the EU , but only about 2 percent of the related patent filings. Europe invents, others industrialize.
The reflex, common in parts of the European debate, to dismiss large language models as statistical parrots hasn't helped either. Whether transformers are the most elegant path to real intelligence is a different question from whether you need to control the systems built on them.
Mistral, of all companies, shows the direction is right. Europe's model champion has evolved into a full-stack provider and announced on August 11 regional endpoints, hosting for third-party open-weights models on its own platform, and pooled European compute demand, with up to one gigawatt of capacity planned by 2030.
That makes Mistral the only European provider with all three building blocks: its own models, its own platform, and planned compute capacity.
The offering still isn't complete. The endpoints don't yet support the agentic workloads this issue is about. And hosting others' open models has two weaknesses. Whoever runs Qwen collects the deployment knowledge, but Alibaba trains the next Qwen generation - Europe can only translate that knowledge into better models itself. Meanwhile, the model's trained-in value judgments run along unfiltered unless someone adjusts them in post-training.
The competition, meanwhile, is building the missing block right on European soil. OpenAI's DeployCo took over about 150 deployment specialists with Tomoro and is opening offices in Paris, London, and Munich.
On paper, the EU has recognized the problem. InvestAI is meant to mobilize 200 billion euros, 20 billion of it for AI gigafactories. The Cloud and AI Development Act is supposed to at least triple data center capacity, and Jupiter in Jülich is running as the first European exascale system. Little of it is built. The gigafactory tender has been delayed several times, with the first facilities expected in 2027 at the earliest - while the four largest US companies alone are likely to invest around $700 billion in AI in 2026, three times the entire European initiative, which is stretched over years.
Above all, the money funds only part of what makes up the US labs' lead. As the previous chapters showed, that lead rests on a cycle: compute capacity trains models, the models work in agent systems at customers, and that deployment produces the knowledge for the next generation.
The gigafactories only cover the first step. It's a necessary one, because without your own compute, every layer above stays open to coercion, as China's forced hardware detours show. But the cycle doesn't close in the data center. It closes at deployment.
That leaves the obvious answer of betting fully on American and Chinese open-weights models. Three objections stand against it.
The first concerns supply. There's no entitlement to the next open generation, and its release remains a revocable business decision. Meta held back its announced Behemoth model. Multimodal Llama models never shipped to EU companies for licensing reasons. The US put export restrictions on Fable 5 and Mythos 5, which for Fable 5 only lifted again after 18 days. According to Reuters, Beijing is weighing restrictions on foreign access to China's best models, explicitly including open ones. And even Z.ai is initially holding back the weights of GLM-5.3.
Behind this looms a deeper break. In an interview study with 25 researchers at leading labs, only four of twenty respondents expected that models capable of automating AI research itself would ever ship as a product. The majority expects the strongest versions to stay in-house, where they widen the lead instead of selling it.
The accessible market, open or via API, would then only reflect the second tier. That also tempers the good news from the start of this chapter: only what's sold can be copied.
The second objection follows from the previous chapter. Whoever deploys a foreign model takes on its trained-in worldview - state-mandated censorship with Chinese models, and the content policies of private corporations with American ones, which can shift with every change of power in Washington.
System sovereignty therefore also needs a normative layer: independent evaluation infrastructure that measures information integrity and refusal behavior alongside capabilities, plus the capacity to adjust open models in post-training. The CTGT finding shows it's technically possible. Whether anyone in Europe does it systematically and auditably remains open.
The third objection: even homegrown models don't add up to sovereignty, because sovereignty is vulnerable on every layer of the stack individually. China, of all places, proves the point. Alibaba and ByteDance control the model layer and own domestic data centers. But the latest Nvidia chips can't enter the country, the available alternatives deliver far less per chip and watt, and both companies have reportedly fallen back on rented Nvidia hardware at foreign operators to train their flagship models.
And Europe has experienced firsthand how quickly a layer someone else controls turns into a political weapon. When the US sanctioned the chief prosecutor of the International Criminal Court, Microsoft cut off his email access. Sovereignty is only ever as strong as the weakest layer someone else controls.
A decisive new layer is now joining these: access to real work. Freely available internet data is largely tapped out. What makes models better going forward is expert knowledge, process know-how, and deployment experience - partly as purchased training data, a market that became a billion-dollar business in short order, but above all as knowledge about which capabilities are worth training.
A finished model can be copied. Access to the work the next one learns from cannot. Whoever runs the systems this work flows through earns twice: from the current model's work, and from knowing what the next one needs to do, made possible by reinforcement learning.
For Europe, more is at stake than a business opportunity, because this touches the foundation of its position in the world. Europe's weight comes from its single market, its power to set standards, and its high-value knowledge work.
That work is under pressure from two sides. AI devalues it, already visible in the shrinking entry-level market for junior developers, and AI simultaneously transfers it into foreign systems.
On top of the familiar drain through company sales comes a second, faster channel. Expertise now flows piecemeal into foreign models through data vendors that broker experts to the labs, and through deployment units that embed themselves in core European processes. It would be an exodus not to lower-wage countries, as manufacturing once saw, but into foreign AI models. This time, it's intellectual capital leaving.
At the same time, this very work is Europe's biggest asset in this competition. It's the raw material the model race will be fought over, and Europe has a lot of it.
If it runs through foreign systems, Europe pays twice: it supplies the knowledge and buys the results back as the next model generation. If it runs through its own
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み