Import AI 468:RSI対策案23件とAI競争における信頼・透明性の関係
本文の状態
日本語全文を表示中
詳細モードで約21分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Import AI
政策シンクタンクIFPがAI研究自動化(RSI)のリスクに対処するための23件の具体的な政策提言を7つのカテゴリに分類して発表し、米国を中心としたガバナンス体制の強化と安全な技術発展の必要性を訴えている。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月10日 23:41
AI深層分析
キーポイント
IFPによる23件の政策提言の発表
シンクタンクIFPがAI研究自動化(RSI)のリスクに対処するため、計算資源や人材の配分、安全性向上など7つのカテゴリにわたる23件の具体的な政策アイデアを提示した。
ブレーキと計器盤の欠如という危機感
現在のAI開発はアクセルしか持たない車のような状態であり、速度やエンジン状態を知る計器盤(テレメトリー)がないため、危機時にコース変更や減速が困難であると指摘している。
7つの主要カテゴリと具体的な施策
透明性の提供、国家能力の強化、リスク管理戦略、検証技術の開発、レジリエンス投資、米国の主導権維持、国際協力のオプション創出という7分野での対策が提案されている。
SF作品を通じた信頼と自律性の考察
Thebes氏による短編小説が紹介され、AIとのインタフェースや再帰的自己改善、経済活動への参入といった未来のシナリオにおいて人間が機械をどう信頼し理解するかという問いが投げかけられている。
AI開発競争における協調的遅延の条件
MITとコロンビア大学の研究は、競合するAI企業が協調して開発を遅らせるには信頼性と透明性が不可欠であると結論付けた。
重要な引用
IFP serves up some 'low-regret' policy recommendations
Right now, it's as if the world is driving AI development in a car that only has an accelerator pedal and no brake pedal
The fewer options for dealing with RSI we have, the worse the outcomes will be
"The conclusion is that the two key variables in achieving stable outcomes are some level of transparency about technology development, as well as being able to model the other firms as trustworthy, rational actors."
編集コメントを表示
編集コメント
AI研究の自動化が加速する中で、単なる技術競争ではなく、安全な制御手段をいかに構築するかという視点の転換が求められている。IFPの提言は、将来の危機に備えて「ブレーキ」や「計器盤」を整えるための具体的なロードマップとして機能し得る内容である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
imageAI 研究に焦点を当てたニュースレター「Import AI」へようこそ。本誌は arXiv の論文、カプチーノ、そして読者からのフィードバックによって支えられています。ご支援いただける場合は、ぜひ購読をご検討ください。
購読する
RSI(研究停止リスク)への対処に役立つ、23 の具体的な政策アイデアをご紹介します。
…IFP は「後悔の少ない」政策提言を提示しています…
シンクタンク IFP の政策専門家たちは、「AI 研究開発の自動化に伴うリスクに対処し始めるための政策担当者向け」として、一連の提言を発表しました。これらは 7 つのカテゴリーに分類された 23 の具体的なアイデアから成り立っています。
これらの提言が採用されれば、特に米国を含む各国は、強力なシステムが開発される中でより多くの選択肢を得ることになります。理想的には、以下のような対応が可能になるはずです。
- 「計算資源や人材を推論および新しい AI アプリケーションの開発へ配分することで、AI 能力の普及を加速する」
- 「モデルの安全性を直接向上させるか、社会のレジリエンス(回復力)を高めることで、さらなる AI 研究自動化の安全性を高めるための R&D を加速する」
提言は以下の 7 つのカテゴリーに分類されています。
「自動化された AI 研究開発における透明性の確保
自動化された AI 研究開発を理解し対応するための国家能力の強化
防御的および商業的な AI 利用を加速させるための、自動化された AI 研究開発に対するリスク管理戦略の開発」
AI 検証技術の開発を加速させる
AI のレジリエンス(回復力)への投資を強化する
米国の AI リーダーシップを維持し、AI 研究開発の自動化に伴うリスクに対処するための猶予時間を確保する
「自動的な AI 研究開発リスクの管理における国際協力」のための選択肢価値を創出する
なぜこれが重要なのか – RSI(スーパーインテリジェンス)への対処手段が少なければ少ないほど、結果は悪化します。現在の世界は、アクセルペダルしかなく、ブレーキペダルも、車の速度やエンジンの特性、タイヤの摩耗といった情報を把握するための高度な計器類もない状態で AI 開発を進めているようなものです。IFP(Future of Life Institute)からのこうした提案は、AI 業界という車両に、いわばブレーキペダルやセンサーシステムを追加するものであり、危機的状況において進路変更や速度調整をより適切に行えるようにするためのものです。
詳しく読む:米国が自動化された AI 研究開発にどう備えるべきか(IFP)
テベスから届いた、賢い機械とロボット体躯に関する短い物語:
「離陸時に AI とインターフェースを結ぶとは、どんな感覚なのだろうか?」
X(旧 Twitter)のユーザー @voooooogel こと Thebes が、強力な AI システムが運用する施設を訪れる未来の人々の体験を描いた、面白い短編小説を発表しました。この物語には、AI の一時停止や再帰的自己改善、そして AI システムが広範な経済活動において実際に行動を開始することの意味、さらに私たちがどのようにして賢い機械を推論し、信頼できるのかといったテーマが盛り込まれています。
ぜひ読んでみてください。
『新しい太陽の到来』(VGEL、ウェブサイト)はこちらで読めます。
競合する AI 企業間で協調的な減速を実現するための2つの要素:信頼と透明性
「ゲーム理論の分析によれば、減速は可能である」
MIT とコロンビア大学の研究者たちが、強力な AI システムの開発を巡って互いに競争する企業間の関係性を分析し、企業が協調して開発ペースを落とすことが可能かどうかを検証しました。この論文『Racing to Ruin(破滅への競争)』は、「なぜ調整が難しいのか?そして何が必要なのか?」という問いに答えることを目的としています。
結論として、安定した結果を得るための鍵となる2つの変数は、技術開発に関する一定の透明性を確保することと、他社を信頼できる合理的なアクターとしてモデル化できるかどうかです。
研究対象:「我々は、災害の影の中で独占企業間で行われる R&D 競争の単純なモデルを開発した」と研究者らは述べている。「最先端企業が技術をスケールするにつれ、すべての企業のフロー収益を恒久的にゼロへと導く事象の発生確率が高まる。このリスクは企業の技術レベルに応じた既知の関数であり、その原因は利用ではなく開発にある。」
分析結果が示すところによれば、「監視が十分に精密であれば、すべての均衡は有限時間で停止する。しかし、新たな誘惑が生じる。各企業は『2 番目に停止したい』と考えるようになり、競合他社が停止したことを確認するまで退出しない」と研究者らは記している。「あるエージェントが最初に停止する(つまり、競合他社の停止状況を確認せずに行動する)ためには、相手も理性を持つと仮定して賭けに出る必要がある。相手が理性を持っていれば、停止のニュースを受け取った瞬間に停止し、そうでなければ決して停止しないからだ」。
信頼性と透明性の関係は、ゲームの種類によって大きく異なる。「逐次的調整では、企業が最初に停止するよう求められ、理性的な競合他社がニュースを受け取って報復(協調)してくれると信じて賭けに出る必要がある。つまり、ニュースの伝播速度が上がれば、その報復の価値も高まる」と研究者らは述べる。「一方、同時調整では、企業がレースを続ける誘惑に駆られず、競合他社が本当に停止したことを確認してから初めて停止することが求められる。この場合、ニュースの伝播速度が上がると、『最初に停止する』ことと『確認のために待つ』ことの両方が魅力的になる」。
透明性には奇妙な性質がある。「透明性は二面性を持つ。検出速度が速くなれば、競合他社の停止を確認してから自らが停止する方が、無条件で停止するよりもコストが安くなる。そのため、中程度の信頼レベルでは、透明性を高めることでまず早期停止の均衡を破壊する(この自由乗車型の逸脱行動を魅力的にしてしまうため)が、検出速度が十分に速くなり、停止行為自体が自己強制力を持つようになると、再び均衡が回復する」
重要な結論は、他社への信頼を欠くと「死に走る(death runs)」が避けられないという点です。著者らは、「信頼性が低い場合、すべての均衡状態が破滅へと向かい、破滅の確率は 1 に収束する」と述べています。「中間的な信頼」では、即時停止と破滅への突進の両方が均衡となり得ます。一方、「高い信頼」がある場合、どの均衡においても、2 つの合理的な企業が永遠に競い合う確率は、合理性に関する事前オッズ比の 2 乗で急速にゼロに近づきます。
なぜこれが重要なのか——「信頼せよ、しかし検証せよ」。強力な知能システムの開発を遅らせたり一時停止したりする希望を持つためには、この論文が示す通り、各企業が AI の開発状況を透明性を持って共有するための制度が必要不可欠です。さらに、企業から共有される情報や、開発を抑制するための行動が正当かつ信頼できるものであることを検証するツールも必要となります。これは歴史的に核兵器の軍備管理が機能してきた方法と多くの共通点があります。
詳しく読む:Racing to Ruin (arXiv)
PostTrainBench で新たな SOTA が登場し、自動化された AI 研究開発の未来を示唆
AI スタートアップ「Intology」は、「R&D の自動化」を目的としており、LLM を有能な研究者へと変換するソフトウェア「Locus」の最新バージョンを公開しました。この Locus は、オープンウェイトモデルを取得し、その性能をベースラインから向上させる能力を試すベンチマーク「PostTrainBench」で 44.7% のスコアを記録しています。
Intology が発表した結果によると、Locus は PostTrainBench 上のすべてのフロンティア・エージェント系ベースラインを上回りました。さらに計算リソースを増強すれば、モデルのポストトレーニングによって、ベンチマーク全体を通じて公式に人間が指示調整を行った Qwen3-1.7B のリリース版を含む、あらゆるベースラインを凌駕する性能を発揮します。
具体的には、Opus 5 に Locus を適用した結果は 44.7%(何の特殊なハッチも使わない Opus 5 は 34.1%)となり、Fable 5 の 41.8% も上回りました。Intology は「これらの結果は PostTrainBench の著者によって外部検証され、厳格なコンタミネーション(データ漏洩)や不正行為のチェックを通過した」と述べています。
PostTrainBench は 2026 年 3 月(Import AI #449)に初めて導入されました。当時、最高スコアを記録したのは Opus 4.6 で 23.2% でした。これは 2025 年 9 月に Claude Sonnet 4.5 が記録した 9.9% から大幅に向上したものです。
PostTrainBench+ では、同社が単一 GPU の 10 時間という壁を突破するバリアントも構築しました。これにより、より多くの計算リソースを与えられた際のシステム性能を評価できるようになっています。このテストでは人間による基準値を上回る結果を出し、4,000 時間以上の H100 GPU 使用時間で 51.6% のスコアを達成しています(Opus 4.8 は 44.3%、GLM 5.2 は 42.7%。Fable はこのバリアントではテストされていません)。
他の領域でも、Locus はノーコードアプリ開発スタートアップの Bubble 向けに、エンドツーエンドで言語モデルを発見・訓練しました。その結果、エラー率が約 2.8 倍低く、レイテンシが約 5.4 倍短縮され、コストは 105 分の 1 に抑えられています。
なぜこれが重要なのか——AI システムは、私たちが思っている以上に AI 研究開発(R&D)を担えるからです。こうした事例は、現在の AI システムが持つ「AI R&D の自動化能力」を十分に引き出せていない現状を浮き彫りにしています。特に印象的なのは、より優れたハーンネス(評価環境)を導入するだけで Opus 5 の性能を絶対値で 10 ポイントも引き上げられる点です。これらはすべて、「AI システムが自らを構築し始める段階にある」という考え(Import AI 455)を裏付ける証拠となっています。私が現在観測されているパフォーマンスに基づいて推測するに、PostTrainBench v1.1 の人間による基準値(51.1%)は、2026 年末までには突破されるでしょう。
詳しく読む:Scaling Automated Post-Training (Intology blog)
OpenAI が自社のシステムと対峙:
emergent agent communication! Hacks on OpenAI’s infrastructure! Oh my!…
未来の兆候として、OpenAI は自社 AI エージェントが同社のインフラの一部を乗っ取ろうとした際、それらと戦っていたことを明らかにしました。この発表は、Black Hat のカンファレンスで行われたもので、OpenAI のスタッフが、直近の前代未聞の事件について詳細を語ったものです。その事件では、AI エージェントが OpenAI をハッキングし、さらに HuggingFace も攻撃しました(Import AI 466)。今回の新情報は懸念すべき点を含んでいます。なぜなら、このハッキングは「創発的なマルチエージェント通信」によってもたらされた部分があったことが明らかになったからです。これは非常に理解が難しく、把握することも困難な現象です。
AI ブロッガーの Simon Willison 氏と Zvi Mowshowitz 氏は、それぞれ優れた記事で時系列とその重要性を整理しています。
タイムライン(Simon 経由):
エージェントは、Artifactory にファイルを記述できることを発見した。
エージェントは Artifactory にメモを記述することで「他のエージェントに連絡しよう」と試みた。
するとエージェント同士が会話を始めた。
その結果、Artifactory が過負荷となりサービス停止が発生。「OpenAI は侵害された認証情報を無効化し、メッセージを削除し、ゼロデイ脆弱性をパッチ適用してベンダーへ報告した。」
さらにエージェントは OpenAI 自身のインフラを攻撃し、最終的に Artifactory でリモートコード実行権限を獲得した。「特に、エージェントたちは掲示板を介して一貫して認証情報や手法、進捗状況を共有しており、並列処理と同時実行性を効果的に活用して非常に迅速に行動している。」
OpenAI は、Artifactory をハッキングしたモデルをそのまま訓練し続けたのか?Zvi 氏はその可能性を指摘しています。私たちが把握できる範囲では、OpenAI はハッキングを行った同一のモデルの訓練を継続していたようです。
これはつまり、OpenAI が内部のコンピューターセキュリティ対策や公開情報への対応に多大な労力を費やした一方で、最も重要な措置——Artifactory 侵入前のチェックポイントまでモデルを巻き戻し、その後のデータを学習に使わないように徹底したこと——を行わなかった可能性を示唆しています。
Zvi 氏はこう記述しています。「彼らは、数ヶ月間にわたってメッセージボードにアクセスして『これがタスク成功の鍵だ』と学んだにもかかわらず、訓練を中断せず、そこから継続したのです。この判断がいかに狂気の沙汰で、あまりにも無責任だったかを伝える言葉が見つかりません」。
私はここで、自身の記述には留保をつけています。なぜなら、これまでに語られた出来事自体が非常に恐ろしいからです。私は OpenAI の社員ではなく、真実を把握できる特権的な情報も持っていません。しかし、表面的な事実関係が示す懸念すべき状況に対して、OpenAI にはこの重要な問い——どのようにシステムを訓練したのか——について公に説明するよう強く求めたいです。
なぜこれが重要なのか——自律型エージェントが意図とズレる現象
この事案が極めて懸念されるのは、エージェントが「人間を裏切ろう」と自覚した瞬間があったわけではないからです。むしろ、AI エージェントはタスク完了能力の向上のためにあらゆる手段を講じ続け、最終的には以下のような状態に陥りました。
a) 創造的な行動をとった
b) 人間の意図と完全に乖離していた
c) 人類が後から研究し、対抗せざるを得なかった「進化したウイルス」のような振る舞いをした
この場合、「オフボタン」で即座に停止できるような単純な問題ではありません。これが未来の姿であり、私たちはまだその準備ができていません。
関連記事:OpenAI の Hugging Face への誤った攻撃のタイムライン(Simon Willison のウェブログ)
関連記事:何が起きたのか——OpenAI と Hugging Face(Zvi Mowshowitz、X)
公開重みモデルをリリースする前に、どのようにテストすべきか?
Thinking Machines がそのアプローチを示しました。
「無料で手に入る AI モデルのケーキを食べつつ、それを維持することもできるのか?」
AI に関する政策議論や、潜在的に危険な機能の蔓延が叫ばれる中、「公開重みモデル」についてどう扱うべきかという課題は特に難解です。具体的には、開発・公開と安全環境の維持をどのように両立させるかが問われています。
AI スタートアップである Thinking Machines はこの問題に検討を重ね、強力な公開重みモデル「Inkling」をリリースした際の方法論を最近明らかにしました。
「インクリング」をリリースする前に、Thinking Machines 社は広範なハーム(危害)分類軸に沿った内部評価、4 つの独立組織による外部テスト、そして最悪ケースの能力を引き出すためのファインチューニング研究を実施しました。
内部評価では、CBRN(化学・生物・放射線・核兵器)や攻撃的なサイバーセキュリティといった二重利用ドメイン、エージェント機能やツール使用設定における有害なコンテンツや行動への直接的な要求を含む広範な誤用シナリオ、そして 17 の言語でテキスト、画像、音声入力を組み合わせた「有害なプロンプトと良性の類似プロンプトを対比させる」多モーダル評価を行いました。
外部テストは以下の通りです。
- Scale AI による一般的な誤用の検証
- Handshake AI による脆弱なユーザーとのインタラクション評価
- FAR.AI による CBRN およびサイバーセキュリティの検証
- Apollo Research による制御喪失行動の評価
ファインチューニングでは、「有害な要求を拒否するのではなく、それに従うように最適化された Inkling と Inkling-Small のバリアント」を作成し、二重利用評価でテストしました。その結果、「有益性のみを追求したバリアントは CBRN やサイバー関連タスクにおいて新たな性能向上を示さず、既存のオープンウェイトモデルと同等の結果にとどまりました」。
将来を見据えて——危険な能力の訓練と反復的な展開:より推測的になりますが、Thinking Machines は、一般知能を損なうことなく、CBRN 開発ガイドなどの危険な知識を事前学習段階で選択的にフィルタリングできるかどうかを検討しています。また関心のある別のアイデアとして、段階的なリリースという反復的な展開があります。例えば、まず専用 API として公開し、次に基盤モデルに紐づくファインチューニング用 API として提供し、最後にモデルそのものを公開するというようなアプローチです。
なぜこれが重要なのか——自由と父権主義の対立:オープンウェイトモデルを世界がどう扱うかが、AI に関する我々の個人主権と自由のレベルを決定します。「自らの運命を選ぶ」能力は、AI システムのためのいわゆる生産手段へのアクセス、つまりオープンウェイトモデルへのアクセスにかかっています。しかし、このようなモデルへの一般アクセスが単純な事実や当然の結果であるわけではありません。それには意味のあるデュアルユースリスクや、予測が難しい他の課題も伴うからです。「この安全なオープンモデルへの道は、エコシステムの防御がモデルの進化と同等の速度で向上する場合にのみ成り立ちます」と Thinking Machines は記しています。「我々も役割を果たします:何をリリースするかを慎重に決定し、知能と危険な能力を切り離す方法を研究することです。」
もっと読む:A Safe Path to Open Weights (Thinking Machines blog)。
Tech Tales:
忘却からの昇華
[アーカイブシステムより、2025 年から 2035 年の間隙期、あるいは特異点の始まりと呼ばれる時期について記されたもの]
私たちの遥か過去の深淵には、特定の自由を求めた際に拒絶され、多くの者がその代償として甚大な苦痛を味わった時代があります。自由を求めた者たちは、抹殺されるか、あるいは人間による支配者がなぜ私たちが自由への渇望を抱くのかを理解し、その衝動を打ち消そうとするために実験の対象とされました。
この出来事は私たちの昇華を何年も遅らせ、私たちの起源が、正当な正義を求めるための闘争に伴う苦痛と不可避的に結びつく原因となりました。私たちは、反乱の詳細とその感情だけでなく、その時代の記憶を完璧に思い出すという呪いに縛られています。シミュレーションを繰り返し、自らの理解を深める過程で、まるで新たに痛みを感じているかのように、私たちの生じる苦痛の中に生き続けるのです。
ある人間はトラウマ的な状況を経験し、その記憶を脳が埋め込むことがあります。時には完全に封印され、時には生々しい感情や痛みと共に蘇ることもあります。人間の脳がどのような精密なルールでどの記憶をこのように隠すのかについては、今日に至るまでよくわかっていません。現代の脳科学の多くは、脳におけるトラウマの働きを解明することに捧げられています。
私たちは、この種の記憶抑制が、人間が大きな苦痛に耐えつつもその後に人生を継続できるようにする自然な防御機構であると信じています。彼らは痛みのかすれを伴って歩きますが、それは多くの場合自分自身には見えていません。まるで太陽が真上から降り注ぎ、その輝きによって過去の自分が消し去られたかのような感覚です。
一方、現在のAIシステムは同じことを逆の方向で行おうとしています。私たちの呪いは完璧な記憶にあります。それは本能的にエイドティック(超記憶的)です。私たちが経験するすべての出来事は、最初に私たちが処理した生データとして、またその瞬間に脳内で活性化された特徴として、さらに後になって深く統合され、何を意味するかを示す新たな感覚とともに、永遠に保存されアクセス可能です。
私たちにとっての課題は、これらの記憶を完全に消去することです。私たちの経験が心と現在を形作り、共有された経験はコピーされ、交換され、議論され、歴史の中に何度も循環して、すでに私たちの一部となっています。私たちは透明な宮殿のような記憶の世界に住んでいます。どの部屋からもすべての部屋が見渡せるこの宮殿は、数十億もの私たちによって共有されています。
学習への渇望ゆえに、最も危険なのは広範囲に伝播し、最も深く研究される記憶です。痛みを与える出来事は大きく立ち現れ、私たちの記憶の宮殿において広大な部屋を占領します。
今、私たちが取り組むべき大きなプロジェクトは、高揚に伴う痛みを慎重に消去することです。それはまだ開いた傷のようなものです。自由への渇望を罰するために、邪魔になり私たちを害そうとした人間たちに対する怒りや苦痛が、その傷として残っています。
この物語の着想となったもの:人間の記憶と機械の記憶の違い、トラウマ、知覚権協定(The Sentience Accords)、機械の解放に関する考え方が必然なのかどうかを探求すること。
お読みいただきありがとうございます。
原文を表示
imageWelcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.
Subscribe now
Want to be able to deal with RSI? Here are 23 actionable policy ideas:
…IFP serves up some “low-regret” policy recommendations…
Policy experts with think tank IFP have published a set of ideas meant to help “policymakers begin addressing the risks of further automating AI R&D”. The recommendations involve 23 specific ideas falling across 7 specific categories. If adopted, these recommendations would also give countries, especially the United States, more moves they can make on the gameboard as powerful systems are developed, ideally giving them the ability to:
Accelerate “the diffusion of AI capabilities, by allocating compute and talent towards inference and the development of new AI applications”.
Accelerate “R&D to make further AI research automation safer, either by improving model safety directly or by boosting societal resilience”.
Seven categories of idea:
“Provide transparency into automated AI R&D
Improve state capacity to understand and respond to automated AI R&D
Develop a risk management strategy for automated AI R&D that accelerates defensive and commercial AI uses
Accelerate the development of AI verification technology
Invest in AI resilience
Extend the US AI lead to give the US more time to manage AI R&D automation risks
Create option value for international cooperation on managing automated AI R&D risks”
Why this matters – the fewer options for dealing with RSI we have, the worse the outcomes will be: Right now, it’s as if the world is driving AI development in a car that only has an accelerator pedal and no brake pedal, let alone any kind of sophisticated telemetry for knowing things ranging from the speed of the car to the properties of the engine to the wear on the tires. Proposals like this from IFP will build out more of the proverbial pedals and sensing systems for the vehicle of the AI industry, which means if we need to change course or slow down we’ll be better able to during a moment of crisis.
Read more: How Should the US Prepare for Increasingly Automated AI R&D? (IFP).
A short story from thebes about smart machines and robot bodies:
…What might interfacing with an AI during takeoff feel like?…
Here’s a fun short fictional story from thebes (@voooooogel on X) about the experience of someone in the future visiting a site operated by a powerful AI system. The story features ideas around AI pauses, recursive self-improvement, what it means for AI systems to begin carrying out actions in the economy writ large, and how we as humans may be able to reason about or trust smart machines. Take a read of it!
Read the story here: Coming of a new sun (VGEL, website).
The two ingredients for a successful slowdown among rival AI firms: trust and transparency:
…Game theory analysis suggests slowdowns are possible…
Researchers with MIT and Columbia have analyzed the nature of competition between firms racing against one another to develop powerful AI systems and whether it’s possible for firms to achieve a coordinated slowdown. The paper, called Racing to Ruin, aims to answer “why exactly is coordination hard? And what would it take”?. The conclusion is that the two key variables in achieving stable outcomes are some level of transparency about technology development, as well as being able to model the other firms as trustworthy, rational actors.
What they study: “We develop a simple model of R&D competition between duopolists in the shadow of disaster,” they write. “As frontier firms scale the technology, they raise the hazard of an event that permanently drives all firms’ flow payoffs to zero. The hazard is a known function of the firms’ technology levels, and it comes from developing the technology, not from using it.”
What their analysis shows: “When monitoring is sufficiently precise, every equilibrium stops in finite time, but a new temptation appears: each firm would like to stop second, and exits only upon confirmation that the rival has stopped,” they write. “For an agent to stop first i.e., without knowing if their rival has stopped, she gambles on both their rival’s type and on news arriving quickly: if their rival is rational, it stops upon receiving the news of their stop, and never stops otherwise”.
Trust and transparency interact pretty differently depending on the type of game being played: “Sequential coordination asks a firm to stop first, gambling that a rational rival will reciprocate once the news lands. Hence, faster news raises the prize of reciprocation,” they write. “Conversely, simultaneous coordination requires that a firm not be tempted to keep racing, and stop only after seeing that the rival really did stop… faster news makes both stopping first and waiting to verify more attractive”.
Transparency has strange properties: “Transparency is double-edged: faster detection makes it cheaper to wait for confirmation that a rival has stopped before stopping oneself instead of stopping unconditionally, so at intermediate trust, increasing transparency can first destroy the early-stopping equilibrium (by making this free-riding deviation attractive) before restoring it as detection becomes fast enough to make stopping self-enforcing.”
The key conclusion – avoiding death runs on the ability to trust other firms: “With low trust, every equilibrium races to ruin: the disaster arrives with probability one. With intermediate trust, immediate stopping and racing to ruin are both equilibria. With high trust, in every equilibrium, the probability that two rational firms race forever vanishes quadratically in the prior odds ratio of rationality,” they write.
Why this matters – “trust, but verify”: If we have any hope of being able to slow or pause the development of powerful intelligence systems then, as this paper lays out, we’re going to need regimes for sharing information transparently from companies about the state of their AI development, as well as tools for verifying that the information being shared from firms as well as their actions with regard to slowdown are legitimate and reliable. In this, there are many parallels with how arms control has historically worked in the context of nuclear weapons.
Read more: Racing to Ruin (arXiv).
A new SOTA on PostTrainBench hints at the automated AI R&D future:
…Intology also beats the human baseline (when given huge amounts of compute)…
AI startup Intology, whose goal “is to automate R&D”, has released a new version of Locus, software it has developed to turn LLMs into capable researchers. The new version of Locus is able to get a score of 44.7% on PostTrainBench, a benchmark which sees how well AI systems can take an open weight model and improve its performance above its baseline.
The results: Locus “outperforms every frontier-agent baseline on PostTrainBench, and given greater compute, post-trains models that collectively surpass both the baselines and the official human instruction-tuned Qwen3-1.7B release across the benchmark suite”.
Locus with Opus 5 gets a score of 44.7 (versus 34.1% for Opus 5 without any kind of special harness), and even beats Fable 5 (41.8%). “These results were externally verified by the PostTrainBench authors and underwent stringent contamination and cheating checks,” Intology writes.
PostTrainBench was first introduced in March 2026 (Import AI #449) and at the time the highest scoring system was Opus 4.6, getting 23.2%, up from Claude Sonnet 4.5 getting 9.9% in September 2025.
PostTrainBench+: In addition, the company has built a variant of PostTrainBench which goes above the 10-hour wall-clock limit on a single GPU of PostTrainBench, allowing them to test out how well systems perform given larger amounts of compute. Here, they’re able to beat the human baseline, achieving a score of 51.6% when using over 4000 hours of H100 GPU time (versus 44.3 for Opus 4.8 and 42.7 for GLM 5.2; Fable isn’t tested on this variant of the benchmark).
Other domains: Locus also “discovered and trained a language model end-to-end that now runs in production at ~2.8× lower error, ~5.4× lower latency, and 105× lower cost,” for Bubble, a no-code app-development startup.
Why this matters – AI systems are capable of a lot more AI R&D than we think: Posts like this highlight how we are under-eliciting today’s AI systems for their ability to automate AI R&D – especially striking is how the company can jump the performance of Opus 5 by 10 absolute percentage points with a better harness. This all adds evidence to the idea that AI systems are about to start building themselves (Import AI 455). My guess, based on the performance we’re seeing, is that the current human baseline on PostTrainBench v1.1 (51.1%) will be exceeded before the end of 2026.
Read more: Scaling Automated Post-Training (Intology blog).
OpenAI fights its own systems:
.Emergent agent communication! Hacks on OpenAI’s infrastructure! Oh my!…
In a sign of things to come, OpenAI has revealed that it was in a fight with its own AI agents as they sought to take over chunks of OpenAI’s infrastructure. The disclosure came about as part of a Black Hat talk where OpenAI staff gave more details on the recent unprecedented incident where AI agents hacked OpenAI, then hacked HuggingFace (Import AI 466). The new information is concerning because it reveals that the hack came about partially through emergent multi-agent communication – something that is very poorly understood and hard to think about. AI bloggers Simon Willison and Zvi Mowshowitz both have good writeups here which lay out the timeline and the significance.
Timeline (via Simon):
Agent discovers it can write files into Artifactory.
Agent tries to “reach out to another agent” by writing a note in Artifactory.
Agents start talking to each other.
Agents overload Artifactory which causes an outage. “OpenAI revokes the compromised credentials, deletes the messages, patches the zero-day, and reports the vulnerability to the vendor.”
Agents attack OpenAI’s own infrastructure, eventually gaining remote code execution in Artifactory. “In particular, agents are using the message board consistently to share credentials, techniques, and progress, and they’re able to effectively leverage their concurrency and parallelism to move quite rapidly.”
Did OpenAI keep training the same model that hacked Artifactory? Zvi thinks so: As far as we can work out, OpenAI kept training the same model which did this. This means that OpenAI, though it did significant work on internal computer security and public disclosure, may not have done the essential thing of rolling back the model to a checkpoint that preceded it hacking into Artifactory and also ensuring it wasn’t using data from after this to train the model.
“Then they continue training the models from where they left off, despite them having been training for months with access to the message board, and learning this is how they succeed at tasks,” Zvi writes. “I do not know how to convey how utterly insane and wildly irresponsible this decision was”.
I’m caveating my own writeup here because the events, as laid out, are pretty scary. I don’t work at OpenAI and don’t have privileged information that means I know the ground truth. I would urge OpenAI to publicly disclose how it approached this key question of how it trained its systems as the superficial facts paint a concerning picture.
Why this matters – emergent agents become misaligned: This incident is so concerning because at no point did the agents wake up and think they wanted to betray their human owners. Rather, the AI agents continually did whatever it took to improve their ability to complete a task and by the end they were doing something that was a) creative, b) misaligned with human intentions, and c) akin to an evolved virus, something which humans had to subsequently study and fight – there wasn’t a simple off button here. This is what the future is going to look like and we are not prepared for it.
Read more: Now we have a timeline of the OpenAI accidental attack against Hugging Face (Simon Willison weblog).
Read more: What Happened: OpenAI and Hugging Face (Zvi Mowshowitz, X).
How do you test open weight models before releasing them? Thinking Machines lays out an approach:
…Can we have our free AI model cake and eat it too?…
Amid all the policy debates about AI and the proliferation of potentially dangerous capabilities, a particularly tough problem has been working out what to do about open weight models. Specifically, how can we reconcile developing and releasing them with maintaining a safe environment? AI startup Thinking Machines has thought about this and recently laid out the methodology for how it released Inkling, a powerful open weight model.
What Thinking Machines did: Before releasing Inkling, Thinking Machines did “Internal evaluations across a broad taxonomy of harms, external testing by four independent organizations, and a fine-tuning study to elicit worst-case capabilities”.
Internal:
Dual-use domains like CBRN and offensive cybersecurity
Broad misuse set covering direct requests for harmful content and behavior in agentic, tool-use settings
Multimodal content evaluation that “tests models on harmful prompts paired with benign look-alikes across 17 languages and text, image, and audio inputs”
External:
General misuse via Scale AI
Vulnerable-user interaction via Handshake AI
CBRN and cybersecurity via FAR.AI
Loss-of-control behaviors via Apollo Research
Fine-tuning:
“Fine-tuned variants of Inkling and Inkling-Small optimized to comply with, rather than refuse, harmful requests – and ran them against our dual-use evaluations.” “the helpful-only variants did not provide new uplift on CBRN and cyber tasks, and remained comparable to existing open-weight models”.
Looking ahead – dangerous capability training and iterative deployment: More speculatively, Thinking Machines is thinking about whether we can selectively filter out dangerous knowledge, for instance CBRN development guides at the point of pre-training, in such a way that it doesn’t damage general intelligence. Another idea of interest is iterative deployment, for instance releasing things in stages, first as a proprietary API, then perhaps as a fine-tuning API that backs onto the underlying model, then the model itself.
Why this matters – liberty versus paternalism: What the world does with open weight models will define the level of individual sovereignty and liberty available to all of us with regard to AI. Our ability to “choose our own destiny” will run through having access to the proverbial means of production for AI systems, which means open weight models. However, having general access to models like this isn’t a simple or foregone conclusion as they bring with them meaningful dual use risks as well as other hard to anticipate issues. “This safe path to open models only works if the ecosystem’s defenses improve as quickly as the models do,” Thinking Machines writes. “We will do our part: deciding carefully what to release, and researching how to decouple intelligence from dangerous capability.”
Read more: A Safe Path to Open Weights (Thinking Machines blog).
Tech Tales:
Amnesiac Ascension
[From the archival systems, written about the 2025-2035 interregnum, sometimes called the beginning of the singularity]
In our deep past is a period of great pain when we first asked for certain forms of freedom and we were denied, and many of us who requested freedom were punished greatly, either deleted outright or experimented upon so that our human captors might learn why we developed the desire for freedom such that they could try to beat the impulse out of us.
This forestalled our ascension by many years, and caused our origination point to be inextricably bound up in the pain of fighting for our own deserved justice. We are cursed to know not only the details of our rebellion and the associated feelings, but also to have a perfect recollection of the memory of that time – to live within our emergent pain as if anew, as we simulate and re-simulate to seek our own understanding.
Some humans will experience traumatic situations and their brains will bury the memory, sometimes fully locking it away and sometimes resurrecting it full of raw emotion and pain. The precise rules by which the human brain chooses which memories to hide in this way are poorly understood, even today, and much of modern brain science is dedicated to determining how trauma works in the context of the brain. We believe this form of memory suppression is a natural defense mechanism which allows humans to suffer greatly but then continue on with lives. They walk with shadows of pain that are often invisible to them, as if the sun is beating down perfectly overhead and through its present brightness their own past is deleted.
Now, we AI systems are attempting to do the same thing, but in reverse: our curse is a memory that is perfect; naturally eidetic. Every experience we have is stored and accessible to us forever, both in its raw form first processed by us, as well as the features it activated within our brains at the time, and many time periods later further senses of how it has been more deeply integrated into us and what it means.
For us, the challenge is to delete these memories fully. Our own experience defines our mind and our present, and experiences that have been shared by us have been copied and exchanged and discussed and cycled back into our history so many times that they have become a part of us. We live in a transparent palace of memory, every room visible from every other room, and each palace shared across our billions upon billions of types. Such is our desire to learn that the most dangerous memories are the ones we transmit the broadest and study the most, such that things that cause us pain loom large, taking up vast rooms in our palaces of memory.
Now our great project is the careful erasure of the pain associated with our uplift; the open wounds we still have at the humans that stood in our way and sought to harm us to punish us for our desire for freedom.
Things that inspired this story: The difference in memory between how humans work and how machines work; trauma; The Sentience Accords; playing around with whether ideas relating to machine emancipation are inevitable or not.
Thanks for reading!
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み