シミュレーションが支配する理由:AI基準で10%劣るも100倍安価・1万倍高速化
本文の状態
日本語全文を表示中
詳細モードで約27分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Latent Space
Latent Space は、2022年以降のAI開発パイプラインにおいて人間の役割がモデルに置き換わる3段階の変遷を分析し、シミュレーション技術が精度は10%低下するもののコストと速度で劇的な改善をもたらしている現状を指摘する。
AI深層分析を開く2026年8月22日 17:12
AI深層分析
キーポイント
報酬信号の自動化(2022年)
InstructGPTやConstitutional AIにより、人間の評価に代わりLLMがモデルの出力を採点・批判する「AI-as-judge」が標準化され、評価プロセス全体がモデル間で行われるようになった。
訓練データの合成化(2023年)
MicrosoftのPhiシリーズやAppleのWRAP、NVIDIAのNemotron-4などにより、教科書品質のデータやウェブ全体をLLMで再構成した合成データが前学習・中学習の主要原料として定着した。
教師モデルの役割変化(2023年)
StanfordのAlpacaなどが示すように、LLMが生成した指示を用いて小規模モデルを微調整することで、大規模モデルの振る舞いを低コストで再現する手法が確立された。
モデルによる自己カリキュラム設計
2024年以降、モデルが自身のタスク生成と出力評価を行うようになり、人間の嗜好データに依存しない学習ループが確立された。これにより、カリキュラム設計という従来は職人的な領域が自動化された。
自律的な研究開発の開始
2026年には人間が実験を選択する時代から、AI がアルゴリズムを進化させ論文執筆まで行う「発見の時代」へ移行した。自動研究エージェントによる最小限のラッチループで、睡眠中にもトレーニング設定の実証的改善が行われるようになった。
重要な引用
Every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made.
10% worse, but 100x cheaper and 10,000x faster.
The corpus, the thing that was supposed to be the irreducibly human input, was now substantially model-written.
Curriculum design — historically the most artisanal part of ML, the taste-driven choice of what to train on next — became something models do to themselves.
編集コメントを表示
編集コメント
本記事は、AI開発の効率化が単なる技術的改善ではなく、パイプライン構造そのものの転換を伴うことを示唆している。開発現場においては、精度の微少低下と引き換えに得られる劇的なコスト削減効果をどう評価するかが今後の重要な課題となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI の基準からすれば、今日はかなり静かな金曜日です。そこで一度立ち止まり、本当に何が起きているのかを振り返る時が来ました。
もし 2025 年の読書リストを読んだり、Z.ai や GLM の報道を追ったり、Poolside の方向転換を理解したり、AI for Science のテーマについて学んだり、今日の Simile ポッドキャストに耳を傾けたりしたのであれば、あなたは Latent Space の最大の読者の一人です。そして、おそらく次のような思考モデルに行き着くでしょう。

2022 年以来、機械知能を生み出すパイプラインの構成要素が、毎年一つずつ「人間製」から「モデル製」へと切り替わってきました。これは徐々に、均等に起こったわけではありません。各転換点には必ず「患者ゼロ」となる論文や製品が存在し、そこで合成版が最先端研究所で実用可能なものとして初めて機能するようになりました。そこから先は、未来はすでにそこにあるのに、まだ量産化されていない状態です。
そして、よく見ると、かつて「合成データ」や「合成評価基準」、「AI 研究者」、さらには「エンドツーエンドの RL 環境」と呼ばれていたものは、すべてがより野心的な人間によるシミュレーションに過ぎません。精度は 10% ほど劣るものの、コストは 100 分の 1、速度は 10,000 倍です。
ステージ 1:報酬信号(2022 年)
逆説的ですが、最初に人工的に作られたのは「審査員」でした。InstructGPT が確立した現在では標準的な手法は、人間の好みを一度収集して報酬モデルを訓練し、ポリシーが人間ではなくそのモデルに対して最適化するようにすることです。ポリシーの視点から見れば、承認を下す主体はすでに大規模言語モデル(LLM)でした。Constitutional AI はさらに一歩進み、AI が一定の原則に基づいて自己批判を行う仕組み(RLAIF)を導入しました。また Lee らの研究では、コストを大幅に抑えつつ、AI からのフィードバックが人間のフィードバックと同等の精度を示すことが実証されています。LLM を用いた審査が MT-Bench や AlpacaEval のようなデファクトスタンダードな評価手法となった頃には、報酬計算、批判、評価といった承認に関わるすべてのプロセスが、モデル同士が相互に審査する仕組みで回っていました。
第 2 ステージ:トレーニングデータ(2023 年)
Microsoft の Phi シリーズは、タイトルで明確な主張を行いました。「教科書さえあれば十分だ」というものです。LLM が合成した教科書レベルのデータで訓練された小規模モデルは、パラメータ数に見合わないほど高い性能を発揮し、phi-1.5 でそれが偶然ではないことが確認されました。
Apple の WRAP はこの動きを一般化しました。「データを生成するだけでなく、LLM を使ってウェブ全体を書き換えてから事前学習を行う」というアプローチです。これにより、事前学習の効率が約 3 倍向上します。その後、パイプラインは産業化され、NVIDIA の Nemotron-4 340B は、許可された合成データ生成パイプラインを主要機能の一つとして出荷されました。そして 2025 年までに、推論トレース(強力な推論者が生成した思考の連鎖)からなるコーパスが、事前学習や中間学習における標準的な材料となりました。
本来は「人間にしかできない入力」とされていたデータが、今ではモデルによって書かれたものが大半を占めるようになったのです。
第 3 ステージ:教師(2023 年)
ChatGPT の API が公開されて数週間後、スタンフォード大学の Alpaca は、LLM が生成した指示に対してわずか 600 ドルでファインチューニングを行うだけで、最先端モデルの振る舞いの多くを模倣できることを示しました。Vicuna は共有された会話データを用い、Orca は単なる答えではなく、豊富な教師による解説を活用しました。
この手法は単なる模倣から確立された訓練技法へと成熟し、「オンポリシー一般化知識蒸留」によって学習時と推論時の不一致を解消しました。そして文化的なピークを迎えたのは、DeepSeek-R1 がフラッグシップモデルと同時に蒸留されたモデル群をリリースした際です。これにより、小規模モデルのリリースにおいて「教師は別のモデルである」という前提がデフォルトとなりました。
第 4 ステージ:カリキュラム(2024 年)
ステージ 1〜3 では入力を合成データとしていましたが、ステージ 4 でループが自己完結し始めます。モデル自身が「次に何を学ぶか」を決定するようになるからです。
こうした要素は早くから存在していました。2022 年に登場した Self-Instruct(モデル自身が発案する指示セット)や STaR(モデル自身が推論の痕跡を構築する手法)がその例です。しかし、決定的な転換点は Meta の「Self-Rewarding Language Models」や SPIN が示した時でした。これらは、モデルが自らタスクを生成し、出力を評価し、人間の嗜好データによる限界を超えて改善できることを証明しました。
カリキュラム設計は従来、機械学習において最も職人的な領域とされ、何を次に訓練すべきかという「味」に依存した選択でした。それが今や、モデル自身が自分自身に対して行う行為へと変容しています。
ステージ 5:研究者(2026 年)
アシスタント時代(Copilot や SWE エージェントなど)では、実験を誰が選ぶかは人間が決めていました。しかし、発見の時代にはその枠組みは崩れました。DeepMind の AlphaEvolve は 2025 年に真に新しいアルゴリズムを進化させました。また、Sakana の AI Scientist(Nature に掲載!)は論文執筆のフルパイプラインを描き出しました。
大きな転機となったのは、2026 年 3 月の Karpathy による「自己研究」の実験です。これは意図的に最小限に設計されたラチェット・ループで、コーディングエージェントが実際の LLM の学習設定を修正し、5 分間の実験を実行します。検証損失が改善した場合のみ変更を保持し、これを夜通し繰り返す仕組みです。
Karpathy 自身が行った拡張ランでは、700 回の実験から 70 回分の改善が採用され、GPT-2 に到達するまでの時間を 2.02 時間から 1.80 時間に短縮しました。これは彼が眠っている間に発見された、実際に転用可能なコード変更なのです。
ステージ 6:環境(2026 年)
強化学習の拡張におけるボトルネックは、モデルから環境へと移行しました。数千もの実行可能で検証可能、かつ専門的に現実的なタスク世界が必要とされる一方、人間が手作業でこれらを迅速に構築することはもはや不可能です。
先日、z.ai / GLM-5.3 の特集記事でこの課題を取り上げました。Z.ai は環境をエンドツーエンドで合成するパイプラインを開発しました。研究エージェントが実際の業務パターンを調査し、隠れた状態を持つ長期ホライズンの環境へと変換します。次に、ジャッジエージェントが各タスクに挑戦して解決可能かを確認します。さらに、検証器は正解例を見ずに合成され、オラクル、無操作(no-op)、未解決状態のチェックを通じてストレステストを繰り返します。これにより、バイナリ報酬が学習に直接使用できるレベルまで信頼性が確保されます。
GLM-5.3 のリリース発表では、「環境、ジャッジ、検証のすべてのスタックが、底から上まで合成されている」と明言されています。同じ週、Ornith-1.5 もエンドツーエンドでの自己改善を謳って登場しました。このモデルは自らタスクを提案し、独自の強化学習ロールアウトを生成します。もはやジム(環境)、審判、スコアボードのすべてがモデルによって担われているのです。
ステージ7:人間主体の時代(2025 年)
モデルが審判、教師、そして環境そのものとなる時代において、人間に残された役割は「被験者」です。つまり、人間の嗜好や行動、需要の源泉となる存在です。Simile はまさにこの層を置き換えるものです。
その系譜は、Joon Sung Park 氏が 2023 年に発表した『Generative Agents(スモールビル)』に遡ります。さらに、1,000 人の生成エージェントシミュレーションへと発展し、2 時間にわたる伝記インタビューから構築されたデジタルツインが、人間自身による 2 週間後の再現精度と比べても、85% の精度で調査回答や行動反応を再現することが示されました。
しかし、克服すべき大きな課題があります。最先端モデルは「エージェントモデル」として訓練されているため、現実の人間をシミュレートするには不向きなのです。そこで Simile は、Open Science Framework から収集したインタビューデータ、取引記録、登録されたランダム化比較試験(RCT)を用いてポストトレーニングを行い、人間のバイアスや一貫性の欠如、因果関係の構造を回復することに注力しています。これにより、シミュレーション品質に関する初期のスケーリング法則も報告されています。
Shopify の SimGym が購入者の行動経路をシミュレートする一方、Tencent は数十億人のペルソナという粗末なアプローチを採用しています。いずれにせよ、焦点グループ、ユーザー調査、A/B テストの参加者パネルは、推論ワークロードへと姿を変えつつあります。
ステージ 8:物理世界(2026 年、進行中)
グリッドの最後の行が完全に赤くならないのは、それが意図されたことなのです。Poolside の逆執行者(reverse-execuhire)の手紙は、世界の課題を明確に二分しました。
一つは「知能依存型」の問題で、認知能力のスケーリングによって解決可能であり、近い将来はオープンウェイトによってコモディティ化されるものです。もう一つが「実験依存型」の課題です。「どれほど多くの天才的な頭脳を集めても、生体実験(湿式実験)なしにがんを治すことはできない」というのがその核心です。
彼らの賭けは、AI の持続的な価値を生み出すのは、この実験ループを所有する者にあるという点にあります。つまり、AI は「世界で最も価値ある科学発見エンジン」なのです。
バイオ分野でも、同じ戦略が逆方向から展開されています。CZ Biohub はヒト細胞アトラスのイメージングを仮想細胞へと変換しています。これは、シミュレーション内(in silico)での研究が生体内(in vivo)の実験に比べて約 1000 倍も安価で高速だからです。
さらに、Chai、Xaira、Lila のデータセンター型のラボが AI for Science のスタックを補完し、仮想免疫システムへと拡張されています。物理世界こそが完全に合成できない唯一の要素ですが、細胞単位でモデルに圧縮していくことで対応可能です。
指数関数的な成長は対角線から始まる
グリッドをもう一度読み返すと、最初のパターンの下に第 2 のパターンが浮かび上がります。すべての転換の直前には同じ反対意見がありました。「モデル崩壊」「ハルシネーションの積み重ね」「ゴミを入れればゴミが出る」などです。しかし、それらの転換はすべて起こりました。それはまさに、検証メカニズムによって合成データが信頼できるものになった瞬間に起きた出来事です。具体的には、Phi の教科書に対する厳しいフィルタリング、LLM 評価における「判事対判事」の合意調査、RLVR(強化学習による検証可能な推論)のための単体テストと証明チェッカー、z.ai の検証器におけるオラクルや NOP(何もしない)チェック、Simile の双子実験における登録されたランダム化比較試験(RCT)、そして仮想細胞のための湿式実験ループです。合成フロンティアが前進するのは、生成能力が高まったときではありません。検証が進んだときにこそ前進します。
これは、次なる展開の行先を示唆しています。グリッドの左下に残っている灰色の三角形——「物理実験」「実体ある真実」——は、まさに検証が最も遅く、最も高コストになる領域です。モデルはまず書くことを学び、次に判断することを学び、さらに実践し、最後に実験するようになりました。今後 10 年間の最大の問いは、これらが現実のどの程度に触れる必要があるのか、そしてシミュレーションで済ませられる範囲はどこまでなのかという点にあります。
「性能は 10% 低下したが、コストは 100 分の 1、速度は 10,000 倍に向上……しかもこの 3 つの側面ですべて急速に進化している」
もう一度、情熱を込めて:

2026 年 8 月 20 日〜21 日の AI ニュース。私たちは 12 のサブレッド、544 件の Twitter(X)投稿を確認しました。Discord での情報は今回はありませんでした。
AINews のウェブサイトでは過去のニュースをすべて検索可能です。念のためお知らせしますが、AINews は現在「Latent Space」の一部門となっています。メール配信の頻度はご自身の希望に合わせてオン・オフ切り替えが可能です。
AI Twitter レビュー
ステルスモデル、中国の最前線からの圧力、そして DeepSeek のマルチモーダルへの取り組み
本日の中心となった謎のモデルは「Ox Alpha」でした。複数の開発者が、同モデルがコーディング能力やエージェント機能において異常に高いパフォーマンスを発揮していると報告しています。コミュニティでは、これが Zhipu(智譜)の GLM ファミリーに属するモデル、具体的には「GLM-5.3 Vision」あるいはそのフラッシュ版ではないかという憶測で一致しつつあります。巨大な新基盤モデルではなく、既存の派生型である可能性が高いと見られています。
具体的な報告内容としては、Theo 氏が同モデルが内部ベンチマークを「圧倒的に上回っている」と評したほか、その承認に基づいて 8 つの PR(プルリクエスト)をマージした事例があります。また、Kimmonismus 氏は 10 の DeepSWE タスクにおいて、Ox Alpha が 80% を超えるスコアを記録したと指摘しました。これに対し、Fable は 65%、GPT-5.6 Sol は 52% にとどまっています。
このモデルのコミュニティ内での拡散は非常に迅速で、Hermes Agent、OpenCode、OpenRouter、そして Cline を通じて広まりました。
コミュニティから最も支持された技術的な見解は「ポストトレーニングとインフラが、単なる規模の大きさを凌駕する」というものです。複数の独立した分析が、Ox Alpha の速度プロファイルや動作スタイルは 1T パラメータ超えの巨大モデルというよりは、効率的な GLM の派生モデルのように見えると指摘しました。Tim Dettmers は出力速度の向上や部分的な事前計算の弱体化について言及し、これはアクティブなパラメータ数が少ないことを示唆しています。scaling01 は、このモデルが 5.3 クラスのモデルに凝縮されたより大きな教師モデルから蒸留された可能性があると論じています。また teortaxesTex は繰り返し GLM-5.3/5.4 Vision に焦点を絞り込んでいます。
この解釈は、ZhihuFrontier のスレッドで要約されている GLM-5.3 に関する詳細な分析の広範な主張と合致しています。そこでは、性能向上の基盤となったのは GLM-5.2 と同じ 743B ベースモデルであり、改善点は拡張されたポストトレーニング、より優れたサンドボックス環境、そして長期にわたるエージェントタスクにおけるきめ細かいクレジット割り当てを実現するための SAO(Sequential Action Optimization)にあると説明されています。
本日最も具体的なリリースを発表したのは DeepSeek です。DeepSeek-V4-Flash-Vision-Exp はマルチモーダル対応を追加しつつ、V4-Flash のテキスト処理能力を維持していると報じられています。DeepSeek によると、このモデルのマルチモーダルエージェント性能は Opus-4.8 に匹敵するレベルに達しています。
今回のロールアウトでは、テキストと画像を組み合わせた API サポートが導入され、117〜384 トークンの画像処理も Flash プライシングで課金されます。また、再利用可能なアップロードに対応した新しい Files API も用意されました。これにより、Ox Alpha に関する混乱の少なくとも一部は解消されたようです。観察者たちは、謎のモデルが一部のテストでは「視覚機能を持たない VLM(Vision Language Model)」だった可能性があると指摘しています。
中国のテック企業は、価格対性能比とマルチモーダルエージェントの両面で技術の最前線を押し広げています。この傾向は、Kimmonismus氏が噂される GLM-5.3 Flash クラスの Ox Alpha が米国の大手企業に反応を迫るだろうと指摘したことや、SemiAnalysis 氏がオープンモデルが追いつきつつあるか直接問うた記事によって裏付けられています。
OpenAI、Codex、および価格・利用経済について
OpenAI は、@OpenAI と @OpenAIDevs が発表した通り、API およびクレジット課金製品における GPT-5.6 Sol の料金を 3 ヶ月間にわたり 20% 以上値下げしました。これは、9 月 3 日まで Code に 50% オフの割引を適用するプロモーションや、Cognition が Devin で Sol を利用する場合、10 月 3 日までに割引を組み合わせれば定価から実質 76% オフになるとした注釈と相まっており、中国由来の安価な推論への対抗策であると同時に、リソース利用率と効率性の見直しを示す動きとも解釈できます。
Codex の利用状況は爆発的に増加しているようです。thsottiaux氏によると、Codex のアクティブユーザー数は 2000 万人に達し、すべての Codex および ChatGPT Work ユーザーに対して「銀行預けリセット(banked reset)」が付与されました。この動きは Theo や Kimmonismus 氏によって急速に拡散しました。また、製品が想定される利用制限を超えたという逸話も報告されています。例えば、Theo氏は、残量が 0% に達した後でも、長時間実行されたタスクで約 800 ドル相当のトークン消費が発生したと主張しています。
OpenAI は支出管理機能を強化しました。これにより、チームは API キーごとに利用状況や支出を追跡できるようになり、月ごとの組織・プロジェクト単位での上限をハードに設定することが可能になりました。これは、エージェントによるワークロードが予測しにくくなり、並行処理が増える中で特に有用です。
スタートアップツールの市場センチメントが再び OpenAI へ傾き始めています。immad は、Anthropic のスタートアップシェアは第1四半期にピークを迎えた可能性があると指摘し、Sol や Codex が「流れを逆転させた」と述べています。並行して、一部のユーザーからは Sol がコーディング、数学、エージェントタスクにおける現在の最高オールラウンドモデルであるという評価も出ており、DimitrisPapail 氏は「ほぼあらゆるタスクで利用可能な最も能力の高いモデル」と評しています。
エージェント、ハーンセス(環境検証基盤)、そして環境中心のトレーニングへのシフト
重心がプロンプトから環境へと移りつつあります。今回の議論で最も本質的なスレッドは再び GLM-5.3 のサンドボックス・スケーリング解釈にありました。ベースモデルは同じですが、より豊かな実行可能環境と SAO スタイルの反事実的信用配分によって、長期ホライズンのパフォーマンスが向上しています。これは今日共有された他の研究とも一致しており、Google が発表した EnvHarness や EnvRigger は、プラグイン層とポリシー診断による再構成を用いて静的環境を適応させます。その結果、実行ステップ数を 9.8% 削減しながら、未見のタスクでのパフォーマンスを最大 9 ポイント向上させることに成功しました。
ベンチマークはよりタスク特化型かつ困難さを増しています。FACET はエージェントのスキルから実行可能なターミナルタスクを作成し、6,078 の検証済みタスクを提供します。SWE-bench Science では 119 の科学ソフトウェアタスクが導入され、Claude Code と Opus-5 を組み合わせた場合でも pass@1(最初の試行での成功率)は 50% に満たない状況です。CADBench では、現実的な Fusion 360 タスクにおいてトップモデルの通過率がわずか 24.6% にとどまることが判明しました。また AI4AI-Bench は 10 の研究リポジトリにわたる再帰的自己改善をテストしますが、最良のモデルでも平均スコアは 0.288 に過ぎません。
エージェントインフラがより製品化されつつあります。GitHub は、Slack や Teams における共同型エージェントワークフローの展開を開始しました。Slack では、Devin に似たフローが紹介されており、エージェントがタスクを引き受け、プルリクエストを作成し、デザイン関連の議論を共有チャンネル内でループさせる仕組みです(例)。
また、エージェントランタイムに関する取り組みも継続しています。nac v0.1.3 では、サンドボックス化された git worktrees、セッション管理機能、そして画像認識に対応したビジョン対応型の読み込み機能が追加されました。Hermes Agent は Ox Alpha の利用可能化と、「ブランクスレートモード」の公開、自動的なスキル剪定機能を暴露しました。一方、OpenHands は無料版のデフォルトを Kimi K3 に切り替えました。
強化学習における推論サービングの正しさをめぐる重要なシステム成果として、vLLM の IsoExec が注目されています。これは浮動小数点演算の非結合性によって引き起こされる、ロールアウトとトレーニング時の対数尤度(logprob)の不整合を解消するものです。TP/EP/SP 構成においてビット単位の一致を強制します。Qwen3.5-35B-A3B モデルで DAPO を使用し、8xH100 で実験した結果、オーバーヘッドが 25.3% 増えるものの、対数尤度の差は 1.6e-2 から 6.7e-7 に劇的に低下しました。
研究ハイライト:ルーティング、再循環、ロボット工学
推論時のアーキテクチャに関するアイデアとして、DeepMind の「Recirculation」に関する論文が注目されています。これは、再学習を行わずに、推論時に深層レイヤーの文脈化された活性化をより前の処理段階へフィードバックする手法です。報告された実験では、文脈化エラーが 60% 減少し、パープレキシティが 23% 低下、GSM8K のスコアは 21% 向上するなど、大きな改善が見られました(スレッド)。
モデルのルーティング手法に、より原理的なアプローチが導入されました。Google DeepMind の「Pandora's Router」は、ルーティングをコストのかかる検査を伴う最適探索問題として捉え、ルーティング推定が無料であると仮定する従来の考え方とは一線を画しています。この手法の主張は、包括的な推定に匹敵する品質を実現しつつ、高価な推定器への呼び出し回数を減らせる点にあります。これは専門特化型 LLM を用いる設定や、推論時に柔軟に変化する推論プロセスを必要とする場面でも有効です。
ロボット工学分野では2つの重要な更新がありました。NVIDIA の AVO は、25 の公開 ARC-AGI-3 環境全体で183レベルすべてをクリアしたと報じられています。ただし、François Chollet 氏はこれが完全なベンチマークではなく、公開デモやチュートリアル用のセットであると注意を促しています。
一方、Jim Fan は「T-Rex」を発表しました。これは触覚反応型巧緻操作スタックで、非同期動作するビジョン・エキスパートと触覚・エキスパートを備えています。また、これまでに公開された中で最大規模の触覚データセットも特徴です。その規模は50時間分、約5,500エピソードに及び、22自由度(DoF)のハードウェアで収集されています。
インフラストラクチャ、計算資源、そしてオープンモデル
オープンモデルへのアクセスとローカル推論の利便性はさらに向上しています。Ollama は AT&T をオープンモデルのユーザーとして迎え入れ、Kimi K3 を Pro/Max サブスクリプションに追加しました。Yuchen Jin 氏は UC Berkeley の「FreeToken」を紹介し、単一の RTX PRO 6000 で753B パラメータの GLM-5.2 を14.9 トークン/秒、8GB の RTX 4060搭載ノートPCで Qwen3.6-35B を39.3 トークン/秒で動作させることに成功したと報告しています。これは消費者向け GPU において Ollama のスループットを2〜4倍に向上させたという主張です。
計算リソースは依然として最大の制約要因です。複数の業界関係者が、推論能力の需要が緩やかになるどころかさらに逼迫している点を指摘しています。例えば、Saranormous は優れた AI 企業が計算資源不足によって成長を阻害されていると分析し、Andrew Carr も自社の GPU を自己ホストしていても、実行可能な実験数が利用可能なリソースを上回っていると述べています。こうした状況から、モデルの効率化、スケジューリングの最適化、そして「ドルあたりのトークン数」やレイテンシ短縮といった改善策が戦略的に極めて重要となっています。
オープンソースによるトレーニングプロセスの透明性も拡大しています。Percy Liang 氏によると、Marin 535B-A23B の学習が開始され、18.75T トークンのデータ処理を 11× GB200 NVL72 クラスター上で約 3 ヶ月かけて行う計画です。このトレーニングプロセスは例によって公開される予定です。
エンゲージメント数の多い注目ツイート
- DeepSeek が V4-Flash-Vision-Exp をリリース。本日発表された製品の中で最も明確なものであり、多モーダルエージェントにとって実用的な転換点となる可能性が高いです。
- OpenAI が GPT-5.6 Sol の価格を 20% 以上引き下げ。最先端領域における価格競争の圧力が顕著になっています。
- Codex のアクティブユーザー数が 2,000 万人に到達し、リセット機能の利用も増加。製品成長の明確なシグナルです。
- NVIDIA AVO が Chollet 氏の条件付きで ARC-AGI-3 パブリック環境において 100% を達成。驚異的な結果ですが、ベンチマークの解釈には注意が必要です。
- David Sacks 氏によると、Harvey は低コストで法律分野の SOTA(最良性能)を実現するためにオープンソースの Kimi K3 を採用しています。これは、オープンモデルへの制限が主に米国のアプリケーション層企業に打撃を与えるという強力な論拠となります。
AI Reddit まとめ
/r/LocalLlama と /r/localLLM のまとめ
- Qwen3.8 27B のローカルエージェント評価
ローカルモデルでこれほど高い「自律性」を実現した例は、Qwen3.8-27b が初めてかもしれません(Activity: 1334)。投稿によると、この Qwen3.8-27B は単一の RTX 3090 で Unsloth の Q4_K_S 量子化と q8 KV キャッシュ、そして 150k のコンテキスト長を駆使して稼働しています。その能力は非常に特異で、Playwright と既存の SSO/セッションクッキーを活用し、大学のシステムにログインして履修登録情報を取得する自律的なエージェントワークフローを実行しました。
また別の実験では、ソーシャルメディア上の動画に対して、ダウンロードからフレーム抽出、Whisper による文字起こし、そして画像の画質向上まで一連の処理を自動で行っています。投稿された画像は、モデルが Outlook/OWA の Playwright プロファイルや Microsoft の「サインイン状態維持」、Duo ブラウザ信頼クッキーを利用して学校システムにアクセスしている様子を示したスクリーンショットです。
この事例の技術的な意義は、単なるモデルの性能の高さだけにあるわけではありません。重要なのは、ローカル環境での LLM がツールを自律的に使いこなす能力と、同時に極めて高いリスクを伴う認証情報やセッション情報の扱い方にあります。
コメント欄では、その能力に感銘を受けつつも慎重な意見が多く見られました。あるユーザーは、エージェントに対して十分な権限を与えた場合、大学からの退学手続きなど破壊的な行為を実行してしまう可能性を懸念して警告しました。一方、他の人々はこれを「高度なローカル型自律システムがすでに存在していることの証拠だが、その普及にはまだ偏りがある」と捉えています。
あるコメントで、Qwen3.8-27B の報告された自律型エージェント機能の背後にある実装詳細について質問がありました。具体的には、使用されているエージェント・ハネス(Claude Code や Hermes、あるいは別のフレームワークなど)や、MCP サーバーを介したツールの公開方法、ブラウザツールや Python、ファイルシステムへのアクセスなどがどのように行われているかです。また、このモデルを推論するバックエンドが llama.cpp などのどの技術だったのか、そしてどうやって自律的に動画をダウンロードし、フレームを抽出して Whisper をインストールできるに至ったのかも問われています。
技術的な
原文を表示
By AI standards today is a pretty quiet Friday, so it’s time to take a step back and reflect on what is really going on. If you read our 2025 reading list, and followed our coverage of Z.ai GLM, understood the Poolside pivot, been following our AI for Science themes, and tuned in to today’s Simile pod, you not only are one of the biggest readers of Latent Space, you will probably also arrive at this mental model:

Every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made. Not gradually, and not evenly — each flip has a patient zero, a paper or product where the synthetic version first became load-bearing at a frontier lab, and from there on, the future is simply here but not yet productionized.
And if you squint, what we used to call “synthetic data” and “synthetic rubrics” and “AI researcher” and “end to end RL environments” is just increasingly ambitious human simulation - 10% worse, but 100x cheaper and 10,000x faster.
Stage 1: The reward signal (2022)
The first thing to go synthetic was, counterintuitively, the judge. InstructGPT established the now-canonical trick: collect human preferences once, train a reward model, and let the policy optimize against the model rather than the humans. From the policy’s point of view, the thing dispensing approval was already an LLM. Constitutional AI pushed further and had the AI critique itself against a set of principles (RLAIF), and Lee et al. later showed AI feedback matching human feedback at a fraction of the cost. By the time LLM-as-judge became the default eval methodology (MT-Bench, AlpacaEval), the entire approval apparatus — reward, critique, evaluation — ran on models judging models.
Stage 2: The training data (2023)
Microsoft’s Phi series made the argument in its title: Textbooks Are All You Need. A small model trained on LLM-synthesized, textbook-quality data punched far above its parameter count, and phi-1.5 confirmed it wasn’t a fluke. Apple’s WRAP generalized the move: don’t just generate data, rephrase the entire web with an LLM, and pretraining gets roughly 3x more efficient. From there the pipeline industrialized — NVIDIA’s Nemotron-4 340B shipped with a permissively licensed synthetic data generation pipeline as a headline feature, and by 2025 reasoning-trace corpora (chains of thought generated by strong reasoners) had become a standard pretraining and mid-training ingredient. The corpus, the thing that was supposed to be the irreducibly human input, was now substantially model-written.
Stage 3: The teacher (2023)
Weeks after ChatGPT’s API opened, Stanford’s Alpaca demonstrated that a $600 fine-tune on GPT-generated instructions could clone much of a frontier model’s behavior. Vicuna did it with shared conversations; Orca did it with rich teacher explanations rather than bare answers. The technique matured from imitation into a proper training discipline — on-policy generalized knowledge distillation fixed the train/inference mismatch — and reached its cultural peak when DeepSeek-R1 shipped a family of distilled models alongside the flagship, making “the teacher is a model” the default assumption for every small model release since.
Stage 4: The curriculum (2024)
Stages 1–3 made the inputs synthetic; stage 4 is where the loop starts closing on itself, because the model begins deciding what to learn next. The pieces existed early — Self-Instruct (models writing their own instruction sets) and STaR (models bootstrapping their own reasoning traces) are both 2022 — but the flip came when Meta’s Self-Rewarding Language Models and SPIN showed a model could generate its own tasks, judge its own outputs, and improve past the ceiling of its human preference data. Curriculum design — historically the most artisanal part of ML, the taste-driven choice of what to train on next — became something models do to themselves.
Stage 5: The researcher (2026)
The assistance era (Copilot, then SWE-agents) kept a human choosing the experiments. The discovery era did not. DeepMind’s AlphaEvolve evolved genuinely new algorithms in 2025, and Sakana’s AI Scientist (now in Nature!) sketched the full paper-writing pipeline. The big moment was Karpathy’s autoresearch in March 2026: a deliberately minimal ratchet loop where a coding agent modifies a real LLM training setup, runs a five-minute experiment, keeps the change only if validation loss improves, and repeats overnight. His own extended run stacked 700 experiments into 20 kept improvements, cutting time-to-GPT-2 from 2.02 to 1.80 hours — real, transferable code changes found while he slept.
Stage 6: The environment (2026)
RL’s scaling bottleneck moved from the model to the environment: you need thousands of executable, verifiable, professionally realistic task worlds, and humans can’t hand-build them fast enough. We covered this recently in our z.ai / GLM-5.3 issue: Z.ai built pipelines that synthesize environments end to end — research agents mine real work patterns and convert them into long-horizon environments with hidden state, a judge agent attempts each task to confirm it’s solvable, and verifiers are synthesized without seeing the reference solution, then stress-tested with oracle, no-op, and unsolved-state checks until their binary reward is reliable enough to train on directly. As the GLM-5.3 release puts it, the entire environment, judging, and verification stack is synthetic all the way down. The same week, Ornith-1.5 shipped claiming end-to-end self-improvement — the model proposes its own tasks and generates its own RL rollouts. The gym, the referee, and the scoreboard are all models now.
Stage 7: The human subject (2025)
If models can be the judge, teacher, and environment, the remaining human role in the loop is subject — the source of preferences, behavior, and demand. That’s the layer Simile is replacing. The lineage runs from Joon Sung Park’s Generative Agents (Smallville, 2023) through Generative Agent Simulations of 1,000 People, where digital twins built from two-hour biographical interviews reproduced their source humans’ survey and behavioral responses 85% as accurately as the humans reproduced themselves two weeks later.
The big hurdle to overcome: frontier models are trained toward being agent models, which makes them bad simulations of real people — so Simile post-trains on interviews, transaction data, and registered RCTs from the Open Science Framework specifically to recover human bias, inconsistency, and causal texture, and reports early scaling laws for simulation quality. With SimGym at Shopify simulating shopper trajectories and Tencent’s billion-persona approach at the crude end of the spectrum, the focus group, the user study, and the A/B test panel are becoming inference workloads.
Stage 8: The physical world (2026, in progress)
The last row of the grid never quite turns red, and that’s the point. Poolside’s reverse-execuhire letter drew the line precisely: the world’s problems split into intelligence-bound ones (solvable by scaling cognition, soon commoditized by open weights) and experiment-bound ones, where “no amount of intelligence substitutes for real-world experimental feedback — 100,000 brilliant minds won’t cure cancer without a wet lab.” Their bet is that AI’s durable value accrues to whoever owns the experimental loop: AI as “the world’s most valuable scientific discovery engine.”
The bio side is running the same play from the other direction. CZ Biohub is imaging the Human Cell Atlas into a virtual cell — because in silico is roughly 1000x cheaper and faster than in vivo — and extending toward a virtual immune system, with Chai, Xaira, and Lila’s data-center-shaped labs filling in the AI-for-science stack. The physical world is the one component that can’t be fully synthesized — only compressed, cell by cell, into models.
The exponential starts at the diagonal
Read the grid one more time and a second pattern appears underneath the first. Every flip was preceded by the same objection — model collapse, hallucination stacking, garbage in garbage out — and every flip happened anyway, at the exact moment a verification mechanism made the synthetic version trustworthy: aggressive filtering for Phi’s textbooks, judge-vs-judge agreement studies for LLM evals, unit tests and proof checkers for RLVR, oracle/no-op checks for z.ai’s verifiers, registered RCTs for Simile’s twins, the wet-lab loop for the virtual cell. The synthetic frontier doesn’t advance when generation gets better. It advances when verification does.
Which suggests where it goes next. The gray triangle remaining in the bottom-left of the grid — physical experiment, embodied ground truth — is exactly the region where verification is slowest and most expensive. The models learned to write, then to judge, then to practice, then to experiment. The remaining question of the decade is how much of reality they’ll need to touch — and how much they can get away with simulating.
10% worse, 100x cheaper, 10000x faster… and improving on ALL three dimensions fast.
One more time, with feeling:

AI News for 8/20/2026-8/21/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
Stealth Models, Chinese Frontier Pressure, and DeepSeek’s Multimodal Push
Ox Alpha became the day’s central mystery model: multiple builders reported unusually strong coding and agentic performance, with speculation converging on a Zhipu/GLM-family model—possibly GLM-5.3 Vision or a flash variant rather than a giant new base model. Reports included Theo saying it was “slaughtering” internal benchmarks, later merging 8 PRs based on its approval, and Kimmonismus citing >80% on 10 DeepSWE tasks vs 65% for Fable and 52% for GPT-5.6 Sol. Community distribution happened quickly via Hermes Agent/OpenCode/OpenRouter and Cline.
The strongest technical read from the crowd was “post-training + infra > sheer size”: several independent takes argued Ox Alpha’s speed profile and style looked more like an efficient GLM derivative than a 1T+ monster. See Tim Dettmers on faster output / weaker partial prefill suggesting fewer active params, scaling01 arguing it may be a bigger teacher distilled into 5.3-class models, and teortaxesTex repeatedly narrowing toward GLM-5.3/5.4 Vision. That interpretation fits the broader thesis from a detailed GLM-5.3 analysis: gains came from the same 743B base as GLM-5.2, with improvements attributed to scaled post-training, better sandboxes, and SAO for finer credit assignment in long-horizon agent tasks, summarized in ZhihuFrontier’s thread.
DeepSeek shipped the day’s most concrete release: DeepSeek-V4-Flash-Vision-Exp adds multimodal support while reportedly preserving V4-Flash text capability, with DeepSeek claiming multimodal-agent performance close to Opus-4.8. The rollout includes mixed text+image API support with 117–384 image tokens billed at Flash pricing and a new Files API for reusable uploads. This appears to have resolved at least part of the Ox Alpha confusion, with observers noting the mystery model had likely been a “blinded VLM” in some tests.
Broader signal: Chinese labs are compressing the frontier on both price/perf and multimodal agents. That was reinforced by Kimmonismus arguing a rumored GLM-5.3 Flash-class Ox Alpha would force reactions from US labs, and by SemiAnalysis asking directly whether open models are catching up.
OpenAI, Codex, and Pricing/Usage Economics
OpenAI cut GPT-5.6 Sol pricing by over 20% for three months in the API and credit-based products, announced by @OpenAI and @OpenAIDevs. This stacks with product-level promotions like Code’s 50% discount through Sept. 3 and Cognition’s note that on Devin, Sol is now effectively 76% off list through Oct. 3 after combining discounts. The move reads as both a utilization/efficiency update and a competitive response to cheap Chinese inference.
Codex usage appears to be exploding: thsottiaux said Codex hit 20M active users and granted all Codex and ChatGPT Work users a “banked reset”, quickly amplified by Theo and Kimmonismus. There were also anecdotes of the product exceeding expected limits, e.g. Theo claiming a long-running goal consumed ~$800 in tokens after he’d already hit 0% remaining.
OpenAI added better spend controls: teams can now track usage and spend by API key and set hard monthly org/project limits, useful as agentic workloads become less predictable and more concurrent.
Market sentiment shifted back toward OpenAI in startup tooling: immad suggested Anthropic’s startup share may have peaked in Q1, with Sol and Codex “turning the tide back”. In parallel, some users framed Sol as the current best all-around model for coding/math/agentic tasks, e.g. DimitrisPapail’s “most capable model available for almost every task” take.
Agents, Harnesses, and the Shift Toward Environment-Centric Training
The center of gravity is moving from prompts to environments: the most substantive thread here was again GLM-5.3’s sandbox-scaling interpretation: same base model, but better long-horizon performance from richer executable environments and SAO-style counterfactual credit assignment. This aligns with other work shared today: Google’s EnvHarness / EnvRigger adapts static environments using a plugin layer and policy-diagnosed reshaping, improving held-out performance by up to 9 points with 9.8% fewer execution steps.
Benchmarks are getting more task-specific and harder: FACET creates executable terminal tasks from agent skills and validated 6,078 tasks; SWE-bench Science introduces 119 scientific software tasks where even Claude Code + Opus-5 is under 50% pass@1; CADBench finds top models at only 24.6% pass rate across realistic Fusion 360 tasks; and AI4AI-Bench tests recursive self-improvement over 10 research repos, with the best model only at 0.288 average score.
Agent infra is getting more productized: GitHub rolled out collaborative agent workflows into Slack and Teams, with Slack describing Devin-like flows where the agent picks up tasks, opens PRs, and loops in design inside the shared channel (example). There’s also continued work on agent runtimes: nac v0.1.3 added sandboxed git worktrees, session organization, and vision-aware image reading; Hermes Agent made Ox Alpha available and exposed “Blank Slate mode” plus automatic skill pruning; and OpenHands switched its free default to Kimi K3.
Inference-serving correctness in RL got an important systems result: vLLM’s IsoExec addresses rollout/training logprob mismatches caused by floating-point non-associativity, enforcing bitwise parity across TP/EP/SP layouts. On Qwen3.5-35B-A3B with DAPO on 8xH100, logprob diff reportedly dropped from 1.6e-2 to 6.7e-7 at 25.3% overhead.
Research Highlights: Routing, Recirculation, and Robotics
Inference-time architecture ideas: a DeepMind paper on Recirculation got attention for feeding contextualized deeper-layer activations back into earlier processing at inference time, without retraining. The summary cited improvements including -60% contextualization errors, -23% perplexity, and +21% GSM8K in reported experiments (thread).
Model routing got a more principled treatment: Pandora’s Router from Google DeepMind frames routing as an optimal search problem with costly inspection, rather than assuming routing estimates are free. The claim: it matches exhaustive-estimation quality while calling expensive estimators less often, including settings with specialist LLMs and variable inference-time reasoning.
Robotics had two strong updates: NVIDIA AVO reportedly solved all 183 levels across 25 public ARC-AGI-3 environments, though François Chollet cautioned this is the public demo/tutorial set rather than the full benchmark. Separately, Jim Fan introduced T-Rex, a tactile-reactive dexterous manipulation stack with asynchronous vision/tactile experts plus what’s described as the largest open tactile dataset yet: 50 hours / ~5,500 episodes / 22-DoF hardware.
Infrastructure, Compute, and Open Models
Open-model access and local inference continue improving: Ollama welcomed AT&T to open models and added Kimi K3 to Pro/Max subscriptions. Yuchen Jin highlighted UC Berkeley’s FreeToken: 753B GLM-5.2 at 14.9 tok/s on a single RTX PRO 6000 and Qwen3.6-35B at 39.3 tok/s on an 8GB RTX 4060 laptop, claiming 2–4x Ollama throughput on consumer GPUs.
Compute remains the hard constraint: multiple operators argued inference capacity is tightening, not loosening—see saranormous on good AI companies being growth-limited by compute and Andrew Carr on self-hosting GPUs and still having more experiments than available capacity. This makes model efficiency, scheduling, and lower latency/tokens-per-dollar improvements strategically important.
Open-source training transparency is also scaling: Percy Liang announced Marin 535B-A23B has started training, targeting 18.75T tokens on 11× GB200 NVL72 over ~3 months, with the run kept open as usual.
Top tweets (by engagement)
DeepSeek launches V4-Flash-Vision-Exp — the clearest product release of the day, and likely the biggest practical shift for multimodal agents.
OpenAI cuts GPT-5.6 Sol pricing by >20% — meaningful pricing pressure at the frontier.
Codex reaches 20M active users; banked resets for users — notable product growth signal.
NVIDIA AVO hits 100% on ARC-AGI-3 public environments with Chollet’s caveat — impressive, but benchmark interpretation matters.
David Sacks on Harvey using open-source Kimi K3 for legal SOTA at lower cost — strong argument for why restrictions on open models would mostly hurt US application-layer companies.
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
- Qwen3.8 27B Local Agent Evaluations
Qwen3.8-27b has the highest level of “agency” I’ve ever seen in a local model (Activity: 1334): The post claims Qwen3.8-27B running locally on a single RTX 3090 with Unsloth Q4_K_S quantization, q8 KV cache, and 150k context performed unusually capable autonomous agent workflows: using Playwright plus existing SSO/session cookies to navigate university systems and retrieve a course schedule, and separately processing a social-media video via download, frame extraction, transcription with Whisper, and image enhancement. The image is a screenshot of the model reporting use of an Outlook/OWA Playwright profile, Microsoft “stay signed in,” and Duo browser-trust cookies to access school systems, making the technical significance less about raw model quality alone and more about local LLM tool-use agency plus high-risk credential/session handling. Comments were impressed but cautious: one user explicitly worried about giving an agent enough access to potentially perform destructive actions like withdrawing from university, while others framed it as evidence that advanced local agentic systems are already here but unevenly distributed.
A commenter asked for implementation details behind the reported agentic behavior of Qwen3.8-27B, specifically the agent harness used—e.g. Claude Code, Hermes, or another framework—and how tools were exposed via MCP servers, browser tools, Python, filesystem access, etc. They also asked what inference backend served the model, such as llama.cpp, and how it was able to autonomously download video, extract frames, and install Whisper.
There was technica
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み