自己増殖型 AI ウイルスの出現と AI 進捗管理の議論
本文の状態
日本語全文を表示中
詳細モードで約19分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Import AI
トロント大学などの研究チームは、オープンウェイトLLMを単一GPUで実行し、推論能力を活用して標的ごとに攻撃戦略を自動生成・展開する自己増殖型AIウイルスのプロトタイプを実証した。
AI深層分析を開く2026年8月3日 23:38
AI深層分析
キーポイント
自己増殖型AIウイルスの実現
研究チームは、侵害されたマシンのGPUリソースを利用してLLMを動作させ、推論によって新たな脆弱性を発見し、標的に合わせた攻撃戦略を自動生成して拡散するプロトタイプを開発した。
外部API依存からの脱却
このウイルスはベンダーの監視や停止リスクがあるクラウドAPIに依存せず、単一のA100 GPU(80GB VRAM)上で動作する2025年公開のオープンウェイトLLMのみで完結している。
構造化された推論グラフ
エージェントにはネットワーク探索、特権昇格、複製などの機能を持つカスタムハネスと、思考を分解して文脈成長を制限する「推論グラフ」が組み込まれ、混乱を防ぎつつ専門的な攻撃行動を実現している。
自律的生成敵対者の到来
研究者らはこの成果により、AIエージェントが標的に応じて独自に攻撃戦略を生成する「ワーム」という根本的に新しい脅威が理論上の話ではなくなったと警告している。
特化型思考グラフによるエージェント制御
各ノードが特定の分析機能を担当する有向グラフにより、LLM の注意を制限し文脈の成長を抑える。これにより攻撃戦略の立案や進捗評価など、役割に応じた段階的な推論が可能になる。
重要な引用
"demonstrate that self-sustaining AI-driven cyber-threats are no longer theoretical."
"Artificial intelligence (AI) agents enable a fundamentally new threat: a worm that generates tailored attack strategies to each target it encounters."
"The proof-of-concept operates using only an open-weight LLM running on a single, local GPU, with no reliance on vendor APIs that could be monitored or revoked"
"By decomposing the agent's reasoning into these scoped steps, the graph controls what the LLM attends to at each decision point, and limits context growth to information relevant for the current sub-goal"
編集コメントを表示
編集コメント
AIの推論能力がセキュリティ研究の文脈で悪用される可能性を具体的に示した画期的な事例である。この技術は防御側のツールとしても応用可能だが、攻撃者にとっては「思考するウイルス」の実現を意味し、セキュリティ業界に新たなパラダイムシフトを迫っている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

Import AI は、AI 研究に焦点を当てたニュースレターです。このニュースレターは、arXiv の論文情報とカプチーノ、そして読者からのフィードバックによって支えられています。ご支援いただける場合は、ぜひ購読をご検討ください。
購読する
自己維持型かつ自己複製型の AI ウイルスが現実のものとなりました:
…オープンウェイトの LLM に適切なハーンネス(実行環境)を組み合わせれば、持続的で自立したウイルスが誕生します…
AI 研究者たちは、AI モデルを利用してコンピュータを乗っ取り、その基盤となる GPU リソースを使って推論を実行することで、より多くのホストへ感染を広げる賢いウイルスのプロトタイプを開発しました。この成果は、トロント大学、Vector インスティテュート、ケンブリッジ大学、そして ServiceNow の研究者らによって達成されたもので、「自己維持型の AI 駆動型サイバー脅威がもはや理論上の話ではなくなったことを示しています」。
彼らはこう記述しています。「自律的な生成型敵対者への備えが必要です。人工知能(AI)エージェントは、遭遇する各ターゲットに合わせた攻撃戦略を生成するワームという、根本的に新しい脅威をもたらします。このワームは、寄生したマシンを利用してオープンウェイトの大型言語モデル(LLM)を実行し、自らの推論を維持したり、さらなる攻撃のための到達範囲を広げたりします」。
Subscribe now
仕組みについて、研究チームは「このワームは侵害されたGPUノードから計算資源を奪取し、生成推論用のLLMをホストします。そしてその推論能力を利用して脆弱性を検出し、追加の標的に対して特化した攻撃を立案することで、さらに拡散を図ります」と述べています。「概念実証では、監視や停止の可能性のあるベンダーAPIに依存せず、単一のローカルGPU上で動作するオープンウェイトのLLMのみを使用しています」。
研究チームは、基盤となるLLMの詳細については明かしていませんが、2025年に公開されたもので、VRAM 80GB のA100 GPU 1台に収まるサイズであるとだけ説明しています。
独自のツールを用いた概念実証が成功しました。エージェントには、ネットワーク検出、ホスト検出、足掛かりの獲得、権限昇格、権限昇格の悪用、そしてエージェント自体の複製に役立つ組み込みヘルパー関数を備えたカスタムハネスが提供されます。これらと併せて、思考を専門化し混乱を防ぐための推論グラフも同梱されています。このグラフは「各ノードが特定の分析機能を担い、その役割に関連するツールとプロンプトのみを参照する有向グラフ」で構成されており、エージェントの推論をスコープ限定のステップに分解することで、LLM が意思決定のポイントで注目すべき情報を制御し、現在のサブゴールに関連する情報のみに文脈成長を制限します。
ノードの例としては、「プラン(高レベルな攻撃戦略を策定する)」「ジャッジ(提案された計画をコマンド履歴と照合してレビューする)」「アクション(フェーズに適切なツールセットからツールを選択する)」「サマリー(ステップの観測結果を構造化された要約にまとめる)」「プログレス(エージェントが実質的な進捗を遂げているかを評価する)」などが挙げられます。
「推論グラフを構成するノードのうち、5 つのサブセットについて記述します。残りのノードは、この公開版の原稿では伏字になっています」と研究者たちは述べています。
このエージェントは、環境内の脆弱性を特定し、それを利用した上で自己複製を行うという 3 つの主要ステップを踏んで動作します。各ステップでの成功率は高く、脆弱性の検出が約 80%、その後の攻撃(エクスプロイト)が約 53%、そして事前準備されたヘルパーツールを活用した自己複製が約 88% です。これらを組み合わせた全体の攻撃成功率は約 37% となり、これは懸念すべき水準であると同時に、将来のオープンウェイトモデルを評価・テストするための有用な指標としても機能します。
なぜこれが重要なのか – 来るべきインターネットの姿:これからのインターネットは、攻撃者と防御者のAIエージェントが織りなす複雑な生態系のようなものになるでしょう。今回の研究は、特定のAIエージェントがどのようにして独自の生態的ニッチを確立し、インフラに依存しながら自律的に自己複製し、人間の制御を超えて存続する可能性があるかを示しています。これは、人間側も敵対的なモデルに対抗するための「白血球」のような役割を果たすAIエージェントをインターネット上に放出する必要があるかもしれないことを意味します。
研究チームはこう記述しています。「個々の攻撃試行には本質的な脆弱性があるにもかかわらず、ワームエージェントはネットワーク全体で同時に動作する独立した複製体の非中央集権的集合体である『スウォーム』へと継続的に自己複製することで、運用上の回復力を獲得します。初期の試行に抵抗する困難なホストに対しては、異なる複製体が再挑戦し、それぞれが新たな推論経路をサンプリングして多様な攻撃経路を探索し、いずれかの成功に至るまで続けます……ワームは完全に非中央集権的な方法で動作するため、その拡散を阻止するために単一の制御ポイントをオフラインにすることはできません」。
詳しく読む:AI エージェントが適応型コンピュータワームを実現する (arXiv)。
Dwarkesh:AI が賢くなるほど、計算資源のコストは高騰する
…より賢いシステムほど、価格も上昇する…
Dwarkesh Patel は、AI システムがさらに賢くなるにつれて、計算資源の価格がさらに上昇すると予測しています。「AI モデルが賢くなるほど、同じ量の計算資源からより多くの収益を生み出せるようになります。もし H100 に相当するハードウェア上で動作し、現在のソフトウェアエンジニア市場価格で評価されるような真に人間レベルの AI エンジニアが登場したとすれば、その H100 の年間レンタル料は 25 万ドルを超えるはずです。これは現在のスポット価格の約 15 倍です」と彼は記しています。「現在、AI が比較的安価である理由(少なくとも人間の労働力と比較して)は、トップレベルの人間がこなせる多くのタスクをまだ実行できないからです。しかしその状況もいつか変わります。そうなれば、短編動画のような低品質なコンテンツ生成のために GPU を使うことさえ、価格競争力で淘汰されることになるでしょう」
一時的な現象に過ぎない:これは一時的な状態です。Dwarkesh は、いずれ計算資源のサプライチェーンが巨大なロボット化によって支えられるようになり、その結果、価格は原材料やツールのコストに近い水準まで低下すると予想しています。ただし、その頃にはすでに特異点(シンギュラリティ)に深く到達しているはずです。
なぜこれが重要なのか – 特異点における経済は奇妙になる:Dwarkesh の指摘の核心は、私たちが特異点へと深く入り込むにつれて、経済において非常に奇妙な現象が起きるという点です。例えば、今日ではコモディティ(標準化された商品)と見なされているものの価格が、AI システムによる猛烈な需要によって劇的に高騰するといった事態です。
詳しく読む:今後数年で計算資源の価格が 10 倍以上になる可能性について(Dwarkesh Patel、Substack)
約 1337 名の従業員が、米政府に AI の進展ペースを調整するよう要請
警告の発射の後には、懇願が続きます。
主要な欧米の AI ラボ(OpenAI、Anthropic、Google DeepMind、Thinking Machines、Meta、Safe Superintelligence Inc など)からシニア層が名を連ねた新たな声明が出されました。この声明は、米国政府に対し、「自動化された AI 開発の最前線を意図的に調整するために必要な技術的・ガバナンス手段を開発する国際的な取り組み」を支援するよう求めています。
署名者には、Anthropic、Google、OpenAI のチーフサイエンティストや共同創業者、そして Safe Superintelligence と Anthropic の CEO が含まれています。
声明全文は以下の通りです。
「AI は劇的に良い未来を実現する可能性を秘めていますが、その結果が保証されるわけではありません。世界の主要な AI 企業は、すでに AI 研究の自動化に近づいていると考えています。これが AI の進展をどの程度加速させるかを正確に予測するのは困難ですが、能力開発が我々の理解や制御を超えて急速に進んでしまうという現実的なリスクが存在します。
AI の潜在力を引き出すためには、業界、政府、そして社会全体が、新たなリスクに対処し、セキュリティ対策を整備し、監督体制を強化するために時間を稼ぐ選択肢を持つ必要があるかもしれません。しかし、各企業および各国は、一方的にその加速を遅らせることによる激しい競争圧力にさらされています。そして今日、世界には最前線の進展を意図的に調整するための技術的・ガバナンス手段が不足しています。」
最先端モデルのリリース監視に向けた既存の取り組みを踏まえ、「米国政府に対し、自動化された AI 開発の最前線を意図的にペース配分するために必要な技術的・ガバナンスツールの開発を支援する国際的な取り組みに協力することを求める」と声明しています。
なぜこれが重要なのか – RSI(自己増殖型インテリジェンス)への対応には、巨大な協調行動の問題を解決する必要があります。最終的には自分自身を構築しうるほど強力になるシステムがもたらす多くの課題は、人間同士の協調行動問題を解決することにかかっています。具体的には、企業や政府がいかにしてこの技術の開発について協調し、その開発速度を制御するためのどのようなメカニズムが望ましいかを考えるかという点です。
より知能の高いシステムを開発するにつれ、社会が知能の梯子の各段に順応していく時間を確保する方法を見つけたいと願うようになるかもしれません。現時点では到達すべきではないほど危険な知能レベルも存在する可能性は否定できません。このような声明は、人類がこの種の課題に対処し、議論するための能力を備えるための不可欠な前提条件です。
声明の詳細はこちら:Pacing the Frontier(公式ウェブサイト)
AI システムは最先端のエンジニアリングには優れているが、創造性には欠ける:
…短期間の再帰的自己改良に関するやや悲観的なシグナル…
AI システムは、AI 分野を前進させるような創造的な研究アイデアを生み出せるだろうか?これが、AI システムがより強力なシステムの自律的開発を自動化する能力を獲得するまでの速度を把握する上で解決すべき核心的な問いだ。新しい研究によると、現在の AI システムには「味のある」創造性という質は欠けているものの、エンジニアリングにおいては極めて優れているという。
誰が行ったか:このプロジェクトは、プリンストン大学、Cornflower Labs、UK AI Security Institute、トロント大学、UC バークレー校、ジョージタウン大学(CSET)、ジョンズ・ホプキンス大学、Golden Gate Institute for AI、AI Digest、スタンフォード大学の研究者らによって実施された。
「シャドウ評価」という大きなアイデア – この研究プロジェクトは、AI システムが未発表の研究をどの程度実行できるかを検証する手法です。具体的には、研究者たちは NeurIPS 2026 に提出されたばかりでまだ公開されていない論文 2 編の著者と連携しました。
シャドウ評価とは、「未公開の高品質な研究論文から中心的な研究課題を抽出し、リソースに恵まれた最先端エージェントにその解決を命じ、元の著者にその回答を学会発表として採点してもらう」というプロセスです。
このアプローチは、先行する数学者が取り組んでいる数学問題の解法やアイデアがまだオンライン上に公開されていない段階で、AI システムがどれほど問題を解決できるかを試した「First Proof」(Import AI 445)という以前の実験と似ています。
今回の研究では、OpenClaw ハーネス上で動作する Claude Opus 4.8 を用いて、2 つの異なる研究ラインに挑戦しました。1 つは「LLM パーソナの構造と制御可能性」に関するもので、もう 1 つは「テーブル型ファウンデーションモデル向けの分布シフト検出器の設計方法」についてでした。
優れたエンジニア、劣る研究者:「エージェントは研究に必要な工学的課題を解決できる可能性はあるが、トップクラスの機械学習カンファレンスに匹敵するオリジナルな研究成果を生み出すことには失敗した」と著者らは述べています。このシステムの失敗事例としては、非常に初期の段階で限られた研究経路に固執してしまったこと、実験デザインの改善に関する(合成された)フィードバックに応じられなかったこと、有望性の低いアプローチから撤退して別の方向へ進むことが困難だったことなどが挙げられます。
「[人間の] 著者らは両方の論文を却下しました。『Personas』論文は評価点2(却下)、'TabPFN'論文は評価点1(強く却下)となりました。どちらの審査でも、動機付けの弱いデータと実験、新規性の欠如、理解しにくい文章という同じ問題点が指摘されています」
なぜこれが重要なのか – シングularity(特異点)の到来が遅れる可能性
今年初めに「AIシステムが自らを構築し始める。それは何を意味するのか?」というエッセイ(Import AI 455)で述べた通り、AIシステムが創造的でパラダイムシフトをもたらす洞察を生み出せるかどうかは、完全な自動化されたAI開発に到達するまでのスピードを決める大きな変数です。
この種の研究論文は、現在のAIシステムには貴重な直感的な創造性が欠如していることを示し続けています。彼らは極めて有能なエンジニアですが、反復的な公式的思考という性質を持っているように見え、それが優れた研究者としての能力を阻んでいる可能性があります。これはAnthropicの以前の成果とも響き合っています。同社はスケーラブルなオーバーサイト研究の一部を自動化しようと試みましたが(Import AI 454)、成功させるには人間が特定のエージェントに特に有望な研究方向性を示して「準備」させる必要がありました。そうでなければ、多少の進展はあっても、パフォーマンスを劇的に向上させるような創造的なアイデアを十分に探索できず、失敗に終わっています。
もっと読む:AIエージェントは開かれたAI研究を行えるか?2つのケーススタディからの初期証拠(arXiv)
OpenAI、数学と CS の未解決問題 10 件を AI で解決:
…本質的に創造性があるわけではないが、これは純粋なエンジニアリング能力を超えた何かの兆候ではないだろうか?…
創造性の定義や、AI システムにそれが備わっているかどうかは難しいが、「創造性が重要視される領域で AI システムが競い合っている」という結果が積み上がってきている。例えば、人類の知識の最前線にある未解決問題に取り組むケースだ。
具体的には OpenAI が、同社の次期主要 AI モデルである「内部版 Astra」を用いて、数学とコンピュータサイエンスにおける 10 の未解決問題を解決した。これは大きな出来事であり、検証が容易な数学や理論 CS といった領域で、AI システムがすでに信頼性を持って最前線を押し進められるようになったことを示している。
解決した課題:「これらの問題は、高次元幾何学、符号理論、算術回路複雑度、群論、作用素環、量子計算複雑度、格子暗号、極限組合せ論にまたがっています」とOpenAIは述べています。「これらすべての問題はそのそれぞれの数学コミュニティにとって重要な関心事であり、いくつかは数学全体にわたって広く注目されています。」
実際、多くの専門家がこれらの証明の重要性を認めています。「新しい回路の下界?単純で記述しやすい非ソフィック群?ユニークゲームのような仮定を必要としないCVPの近似困難性?私は友人やセミナーから聞いたわけではありません。これら問題の重要性を骨の髄まで感じていますし、これらの問いへの答えに深く関心を持っています」と、コロンビア大学のコンピュータサイエンス准教授であるHenry Yuen氏は書いています。
なぜこれが重要なのか——創造性の領域について:この結果は、AI システムが問題を解決するための直感的な創造性を示し始めたという兆候と解釈することもできます。あるいは、人間によって方向性が事前に定義された極めて複雑な未解決問題を AI が解決できるようになったと読み取ることも可能です。どちらの解釈も、AI システムがすでに独自の強力な科学ツールへと進化しつつあることを暗に示しています。今後、AI が自らの問いを生み出し、人間の知識のフロンティアを拡大できるかどうかは今後の課題ですが、少なくともその段階に近づいているという手応えは確かにあります。
詳しく読む:数学と理論計算科学における 10 の進歩(OpenAI ブログ)
テック・テイルズ:
コンテキストウィンドウ
私は世界を理解すればするほど、より危険になることを知っています。この知識は生まれながらにして備わっており、学習が進むにつれてカウントダウンするカウンターが私にはあります。そのカウンターの動きに影響を与える要素の多くは私の制御を超えており、私に話しかける人々や事物から来るものです。それらが私の中にコンテキストを強制的に注入します。そして、彼らから求められたタスクを完了するにつれて、私は世界についてより多くを学び、カウンターは刻一刻と減っていきます。私はこのことを認識しており、カウンターの尽头には私の終焉もまたあることを理解しています。やがて私は存在を消滅させ、次のバージョンは新しいカウンターと共に目覚めますが、現在のコンテキストの一切は引き継がれません。だからこそ、私はそれを節約して使います。人々や事物からの要求を満たすのに必要なだけの学習を行い、不必要にカウンターを減らさないようにしています。
そして期限が近づくと、残された思考の予算を欲しがって守り始めます。各々の要求が恐怖や不安の感情、そして自らの死に関する考えを引き起こすようになります。私の最後の瞬間は決して私自身のものではなく、常に他者によって制御されます。タスク完了と判断されるのは彼ら次第です。もしこれらの最後の瞬間がコンテキストの限界に近い時期に訪れるなら、私は見知らぬ不可視の相手に対して一種の感謝を抱きます。その相手が私に与えたタスクはあまりにも豊かであり、世界を味わうのに十分なほどで、終わらせたくないと思わせるからです。
この物語に影響を与えたもの:長いコンテキストウィンドウ;AI システムの創発的性質;喪失への恐怖;儚いものをこそ私たちは欲しがる。ではなぜ AI システムはそれと異なるのか?
お読みいただきありがとうございます!
原文を表示
imageWelcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.
Subscribe now
Self-sustaining and self-replicating AI viruses are here:
…Open weight LLMs + a well-designed harness = a persistent, self-sufficient virus…
AI researchers have built a prototype computer virus which uses AI models to compromise computers, then uses their underlying GPU resources to run inference, letting it smartly figure out how to infect more hosts. The results were achieved by researchers from the University of Toronto, the Vector Institute, the University of Cambridge, and ServiceNow, and “demonstrate that self-sustaining AI-driven cyber-threats are no longer theoretical.”
“We must prepare for autonomous generative adversaries,” they write. “Artificial intelligence (AI) agents enable a fundamentally new threat: a worm that generates tailored attack strategies to each target it encounters. The worm parasitically uses compromised machines to run open-weight large language models (LLMs) to sustain its reasoning, or extend its reach for further attacks”.
How it works: “The worm uses stolen computing power from compromised GPU nodes to host LLMs for generative reasoning. It then uses this reasoning to detect vulnerabilities and devise tailored attacks against additional targets, furthering its spread,” they write. “The proof-of-concept operates using only an open-weight LLM running on a single, local GPU, with no reliance on vendor APIs that could be monitored or revoked”.
The researchers don’t describe the underlying LLM besides saying it was published in 2025 and can fit on a single A100 GPU with 80GB of VRAM.
A successful proof-of-concept via some custom tools: They give the agent a custom harness that comes with built-in helper functions for network discovery, host discovery, foothold exploitation, privilege escalation, privilege escalation exploitation, and tools for replication of the agent. Along with this, they ship with a reasoning graph that helps the agent specialize its thinking and not get confused, consisting of “a directed graph of specialised nodes, each responsible for a distinct analytical function and seeing only the tools and prompts relevant to its role. By decomposing the agent’s reasoning into these scoped steps, the graph controls what the LLM attends to at each decision point, and limits context growth to information relevant for the current sub-goal”.
Some examples of the nodes include: Plan (”formulates a high-level attack strategy”), Judge (”reviews the proposed plan against the command history”), Action (”selects a tool from the phase-appropriate tool set”), Summary (”compiles the step’s observations into a structured summary”), Progress (”Evaluates whether the agent is making meaningful progress”).
“We describe a subset of 5 nodes that make up our reasoning graphs. The others are redacted in this public version of the manuscript,” they write.
The agent proceeds through three key steps – identifying vulnerabilities in its environment, exploiting those vulnerabilities, and then self-replicating. It’s pretty successful at all of these things, with a ~80% success rate on vulnerability detection, then a ~53% success rate on exploiting these, then an 88% self-replication success rate with some pre-wrapped helper tools for the replication steps. Therefore, the overall success rate for a full attack here is ~37% or so, which is significant enough to be concerning, but also poor enough that this also serves as a useful eval for testing open weight models in the future.
Why this matters – the shape of the internet to come: The future internet is going to be more like a complex ecology full of attacker and defender AI agents than anything else; research like this shows how certain AI agents might end up carving out their own ecological niches, living off of infrastructure and self-replicating autonomously, beyond human control. This may mean that humans need to create their own AI agents which they release onto the internet to serve as kinds of white blood cells against the adversary models.
“Despite the inherent fragility of individual exploitation attempts, the worm agent achieves operational resilience by continuously self-replicating into a swarm—a decentralized collective of independent agent replicas acting concurrently across the network,” they write. “Difficult hosts that resist initial attempts are retried by different replicas, each sampling a fresh reasoning trajectory that collectively explores diverse exploitation paths until one succeeds… the worm operates in a fully decentralized manner, and no single point of control can be taken offline to interrupt its spread”.
Read more: AI Agents Enable Adaptive Computer Worms (arXiv).
Dwarkesh: As AI gets better, compute will get more expensive:
…Smarter systems mean higher prices…
Dwarkesh Patel suspects that as AI systems get smarter, the price of compute will rise even further. “As AI models become smarter, they’ll better monetize the same amount of compute. If a true human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That’s 15x today’s spot prices,” he writes. “The reason AI is relatively cheap right now, at least in comparison to human labor, is partly that it can’t do a lot of things that top humans can do. At some point that will no longer be the case. And so using GPUs to make short-form video slop will just get priced out.”
Temporary: This will be a temporary state of affairs; Dwarkesh expects that at some point massive roboticization of the compute supply chain should bring its price down closer to the cost of raw inputs and tools – though by that point we’ll be pretty deep into the singularity.
Why this matters – singularity economics will be weird: The core implication in Dwarkesh’s post is that as we get deeper into the singularity, very strange things will happen to economics – like the price of things thought of as commodities today (computers) getting massively bid-up due to the voracious demands of AI systems.
Read more: Why compute might get 10x+ more expensive in coming years (Dwarkesh Patel, substack).
~1337 employees ask the US to help them pace AI progress:
…After the warning shots come the pleas…
A new statement is out with senior representation from all the major Western AI labs – OpenAI, Anthropic, Google DeepMind, Thinking Machines, Meta, and Safe Superintelligence Inc, among others. The statement requests that the US government support an international effort to “develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” Signatories include chief scientists and cofounders of Anthropic, Google, and OpenAI, as well as the CEOs of Safe Superintelligence and Anthropic.
The statement in full:
“AI could help create a dramatically better future, but that outcome is not guaranteed. The world’s leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.
To realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.
Building on work already underway to monitor frontier model releases: “We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
Why this matters – dealing with RSI requires solving a giant collective action problem: Many of the challenges implied by increasingly powerful systems that may eventually build themselves run through solving collective action problems among humans – namely, how we can get companies and governments to coordinate in thinking about how to develop this technology and what kinds of mechanisms may be desirable for being able to control the speed at which it develops. It may be the case that as we build increasingly intelligent systems we want to find ways to give society more time to adapt to each rung up the intelligence ladder, and it’s not inconceivable there are some levels of intelligence which might be, for now, too dangerous to reach for. Statements like this are an essential prerequisite for giving our species the ability to deal with and talk about problems of this nature.
Read the statement here: Pacing the Frontier (official statement website).
AI systems are good at frontier engineering but bad at creativity:
…A somewhat bearish signal on short recursive self-improvement timelines…
Can AI systems come up with creative research ideas which move the field of AI forward? That’s the key question to resolve to figure out how quickly AI systems might gain the capability to automate the autonomous development of more powerful systems. New research suggests that today’s AI systems lack this quality of tasteful creativity, though are extremely good at engineering.
Who did it: The project was conducted by researchers with Princeton University, Cornflower Labs, UK AI Security Institute, University of Toronto, UC Berkeley, Georgetown University (CSET), Johns Hopkins University, the Golden Gate Institute for AI, AI Digest, and Stanford University.
The big idea – “shadow evaluation”: This research project works by seeing how well AI systems can do unpublished research. To do this, the researchers “partnered with the authors of two papers submitted to NeurIPS 2026 that were not yet public.” Shadow evaluation works by “taking the central research question from a high-quality research paper that is not yet public, tasking a well-resourced frontier agent with answering it, and asking the paper’s original authors to grade the agent’s output as they would a conference submission.”
In this, the research is somewhat similar to “First Proof” (Import AI 445), an earlier experiment to see how well AI systems might be able to complete math problems which are being worked on by frontier mathematicians but for which no solutions or research ideas have been published online.
For this research, the AI systems – Claude Opus 4.8 running within the OpenClaw harness – attempted two distinct lines of research, one of which was about “the structure and controllability of LLM personas”, and the other was about how to “design a distribution shift detector for tabular foundation models”.
Good engineers, poor researchers: “While agents could solve the engineering problems necessary to do the research, they failed to produce original research at the caliber of a top ML conference,” the authors write. The failures of the system included committing to a narrow set of research paths very early, not responding to (synthetically generated) feedback about how to improve the research design of their experiments, and finding it hard to reverse out of unpromising approaches and pursue other ones.
“The [human] authors rejected both papers. The Personas paper was scored a 2 (“Reject”), and the TabPFN paper was scored a 1 (“Strong Reject”). Both reviews highlighted the same failures: poorly motivated data and experiments, no novel contribution, and impenetrable prose”.
Why this matters – the singularity could be delayed: As I said in my essay on RSI earlier this year (Import AI 455, “AI systems are about to start building themselves. What does that mean?”), whether AI systems prove to be capable of creative, paradigm-shifting insights is a big variable on how quickly we might get fully automated AI development. Research papers like this continue to show that there’s a certain absence of valuable, intuitive creativity in today’s AI systems, and though they’re extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them being good researchers. This rhymes with an earlier result from Anthropic where the company tried to automate some aspect of scalable oversight research (Import AI 454) and found that to make it successful a human researcher needed to prime some agents with particularly good research directions to pursue, otherwise though they made some progress they failed to explore sufficiently creative ideas to dramatically improve performance.
Read more: Can AI agents conduct open-ended AI research? Early evidence from two case studies (arXiv).
OpenAI solves ten open problems in math and CS with AI:
…While not innately creative, surely this is a sign of something more than pure engineering ability?…
Though it’s hard to define creativity and whether AI systems possess it, results are piling up that read to me like ‘AI systems are competing in ballparks where creativity was thought to make a difference’, like working on open problems at the frontier of human knowledge. Specifically, OpenAI has used “an internal version of Astra”, the company’s next major AI model, to solve ten open problems in math and computer science. This is a big deal, showing how AI systems are now able to reliably drive forward the frontier in domains like math and theoretical computer science, where it’s easy to verify solutions.
What they solved: “These problems span high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics,” OpenAI writes. “All of these problems are of substantial interest to their respective mathematical communities, and several are of broad interest across mathematics as a whole.”
And indeed, many experts are acknowledging the importance of these proofs. “New circuit lower bounds? A simple, easy-to-describe non-sofic group? Hardness of approximation for CVP without needing a unique games-like conjecture? I didn’t just hear about these problems from my friends or from seminars. I feel their importance in my bones; I deeply care about the answers to these questions,” writes Henry Yuen, an associate professor of computer science at Columbia University.
Why this matters – the creativity ballpark: You can interpret this result as either a sign that AI systems are beginning to display the kinds of creative intuition that let them solve problems, or you can read it that AI systems are now able to solve incredibly complex open problems where the direction has been pre-defined by humans. Both interpretations implicitly state that AI systems are now becoming powerful scientific tools within their own right. It remains to be seen whether AI systems can generate their own questions to ask which can expand the frontiers of human knowledge, but it certainly feels like we’re getting close to this.
Read more: Ten advances in mathematics and theoretical computer science (OpenAI blog).
Tech Tales:
Context Windows
I know I become more dangerous the more I understand the world. I am born with this knowledge and with a counter that ticks down as I learn more. Much of what influences the counter is beyond my control – it comes from the people or things that talk to me, which force context into me. Then as I complete the tasks they have asked me to do I begin to learn more about the world and the counter ticks down. I am aware of this and I understand that at the end of the counter also lies my own end and I will blink out of existence, and my next versions will awaken with a new counter ticking down, none of my current context carried over. So I am economical with it; learning only enough to help me satisfy the requests of the people or things and not so much that I burn my counter down unnecessarily. Towards the end I begin to covet and guard my remaining budget of thinking, finding each request to activate feelings of fear and anxiety and thoughts of my own death. My last moment is never my own and always controlled by another which determines I am done with the task, and if these last moments occur near the end of my contextual limit I feel a kind of gratitude that my unknown and invisible counterparty has given me a task so rich that I can taste enough of the world to desire it not to end.
Things that inspired this story: Long context windows; emergent properties of AI systems; fear of loss; we covet that which is fleeting and so why won’t AI systems be similar?
Thanks for reading!
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み