AI エージェントは未だにオープンエンドな AI 研究を遂行できないと報告
本文の状態
日本語全文を表示中
詳細モードで約10分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
AI Snake Oil
AI Snake Oil は、最先端の AI エージェントが数日間の実験において、未発表論文の著者から完全に拒絶される結果となり、オープンエンドな研究における判断力や適応力の欠如を明らかにした。
AI深層分析を開く2026年8月5日 23:01
AI深層分析
キーポイント
研究者による完全拒絶の結果
最先端 AI エージェントが未発表論文の著者と協力して作成した論文は、元の著者によって明確に拒絶された。
判断力とリソース管理の欠如
エージェントは低品質なデータに基づいて提案を却下するなどの判断ミスを行い、与えられた予算と時間の半分も使用しなかった。
フィードバックへの非創造的対応
自己レビューや専門家からの否定的な指摘に対し、エージェントは既存の発見に条件を付け加えるだけで根本的な改善を行わなかった。
シャドウ評価の手法と限界
エージェントが未学習・非公開の結果を対象にテストできる利点がある一方、専門家が AI 生成であることを知って評価するためバイアスが生じる可能性やサンプル数が少ないという制約が存在する。
研究者のバイアスと対立の明確化
著者らは再帰的自己改善に関する特定の立場を持っているため、異なる前提を持つ協力者を招き、生じた意見の相違を明示することでバイアスを緩和している。
重要な引用
The authors unambiguously rejected both agent papers.
The agents lacked the judgment for conducting open-ended research.
Despite the agents' own AI self-reviews surfacing many of the issues that the expert reviewers later raised, the agents did not creatively address these concerns.
Our results suggest that conducting open-ended research remains challenging for frontier AI agents.
編集コメントを表示
編集コメント
この分析は、業界が期待する「AI による AI 研究の自動化」というゴールに対し、現時点ではまだ大きな隔たりがあることを示唆している。技術的な限界を直視することで、今後の開発戦略や評価基準の見直しが必要となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
主要な AI ラボの目標は、再帰的自己改善(RSI)です。これは AI エージェントを用いて AI 研究を自動化する取り組みであり、爆発的な AI の進歩に関する予測の根拠ともなっています。では、私たちはこのマイルストーンにどれほど近づいているのでしょうか?
その一つの方法として、AI 研究を実行できるかどうかを試すベンチマークを利用することです。AI コミュニティがベンチマークに注力しているため、RSI に向けた進歩を評価する主要な手段となってきました。過去一年間で実施された多くの評価では、成功の検証が容易なタスクにおいて AI エージェントが進捗を示せることが明らかになり、「RSI の瀬戸際にある」という憶測を呼んでいます。
しかし、これらの評価は有用である一方で、検証可能な狭い範囲のタスクに限定されています。AI 研究はもっとオープンエンドな性質を持っています。成功が即座には明確でなく、検証も難しいケースが多々あります。進歩のためには、有望な仮説を検証したり、立ち戻って再考したり、新しいあるいは非伝統的なアプローチを検討したりする必要があります。では、AI エージェントのオープンエンドな AI 研究能力をどう評価すべきでしょうか?この問いに答えるための第一歩として、私たちは新たな論文を発表しました。
私たちは未発表の AI 論文二編の著者と協力し、それぞれの論文における主要な研究課題の策定を依頼しました。その後、最先端の AI エージェントにこれらの課題に対する研究を実行させ、API クレジットや計算リソースを数千ドル分提供し、6 日間の実時間を与えました。そして、元の著者らがエージェントが作成した論文を検証しました。
著者らは、エージェントに関する論文 2 編を明確に却下しました。これらの結果をより深く理解するため、当チームはエージェントのログを分析するにあたり 100 時間以上を費やしました。主な知見は以下の通りです。
エージェントには、オープンエンドな研究を行うための判断力が欠けていました。専門家のレビューアーが印象的だと認めた方向性を提案したものの、質の低いデータや合成データに基づいて即座に却下してしまいます。
利用可能なリソースに対する認識も不足していました。両方の試行とも、デッドラインまで数時間残っているにもかかわらず API バudget の 50% も使用されずに終了しました。エージェントは自身の使用状況を監視でき、予算を使い切るよう促されていたにも関わらずです。
フィードバックに対して創造的な対応ができませんでした。エージェント自身が行った AI による自己レビューで専門家のレビューアーが後に指摘した多くの問題が浮き彫りになっていたにもかかわらず、これらの懸念に対し創造的に取り組むことはありませんでした。否定的なフィードバックに直面すると、既存の知見に条件付きの注釈を追加するだけで、期待の薄い研究方向性をさらに強化しようとする姿勢を見せました。
効果的な後退(バックトラック)ができませんでした。最も野心的な研究目標は実験初日までに放棄され、その後もどちらのエージェントも根本的にアプローチを変更することはありませんでした。
具体的な指示に従うこともできませんでした。探索に費やすべき時間の長さや、AI 自己レビューツールからレビューを受ける頻度、論文の長さに関する厳格な制限など、明示的なルールを無視していました。
AI がオープンエンドな研究を遂行できるかを評価したいという願いは、エージェントが再現性の向上に使えるかどうかを検証するベンチマークを発表した 2 年前からあります。しかし、私たちは手法を確立してからでないとと願っていました。この手法のアイデアは論文の共同執筆者である英国 AISI のメンバーの一部から提案され、プリンストンのコアチームによって洗練されました。
エージェントが元の研究を追跡する形で行うため、「シャドウ評価」と呼んでいます。私 2 人に加え、コアチームにはピーター・キルギス氏、アンドリュー・シュワルツ氏、ステファン・ラバンスァー氏が含まれています。執筆者の完全なリストは本稿の末尾に記載されています。
シャドウ評価には重要な利点があります。エージェントが学習していない結果や、オンラインでアクセスできない結果に対してテストできる点です。また、数ヶ月をかけて質問に答えてきた専門家が、エージェントの出力を評価することも可能になります。1
しかし、シャドウ評価にも本質的な限界があります。専門家レビューアーは論文が AI によって生成されたことを知っており、エージェントが採用したアプローチよりも自分たちが取ったアプローチを好む可能性があります。各論文について詳細な評価を行っているため、サンプルサイズは小さくなります(今回の研究では 2 つの論文のみを使用しました)。さらに、これらの評価には設計、実行、解釈において研究者の裁量が必然的に多分に含まれます。
実際、私たちは再帰的自己改善や超知能に関する議論において特定の立場をとっていることで知られています。この立場が研究の実施方法に影響を与える可能性があります。論文には、私たちの潜在的なバイアスとそれに対処する方法について詳細なセクションを設けています。私たちは、全員が同じ事前確率(priors)を持たない協力者チームを積極的に探し出し、その結果生じた意見の相違を明確に提示しました。今後の評価では、「対立する協力者」をコアチームの一員として含めることを検討しています。
爆発的なAI進展への示唆
私たちの結果は、オープンエンドな研究が最先端のAIエージェントにとって依然として困難であることを示唆しています。ただし、これらの知見はまだ暫定的なものであり、サンプルサイズの拡大や新モデルでのテスト、潜在的な支援構造(scaffold)の改善などを通じて限界に対処しようとしています。しかし、もしこれらの知見が裏付けられた場合、どのような意味を持つのでしょうか。
第一に、最先端のAI進展(および再帰的自己改善:RSI)が、検証可能なタスクにおけるヒルクライミング(hill climbing)によってどの程度達成可能かを理解する必要があります。私たちの見解では、狭い範囲のタスク(例えば効率性の向上など)においてはより迅速な進展が可能であることは確かですが、それが広範な再帰的自己改善や爆発的な進展につながるとは考えていません。ただし、検証可能なタスクにおけるAIエージェントの能力がどのようにAI進展に結びついていくかについては、引き続き注視していく計画です。
次に、現在の AI エージェントがオープンエンドな研究を行う際に抱える制約(創造性や判断力の欠如など)を、よりターゲットを絞ったトレーニングや支援機能の改善を通じてどれほど迅速に克服できるかを測定する必要があります。この問いに答えるため、私たちは定期的なシャドウ評価を継続して実施する予定です。
最後に、これらの制約が克服されたとしても、AI の進歩ペースを鈍らせるさらなるボトルネックが存在する可能性があります。

ボトルネックには、計算資源の限界や、実世界の実験からのデータ収集が必要となることなどが含まれます。また、現時点では進捗を阻害していないため、まだ認識されていない要因も存在するでしょう。例えば、推論のスケーリングに有用であることが判明するまで、高品質な強化学習(RL)環境の重要性は明確ではありませんでした。同様に、データセンターへの数百億ドル規模の投資がトレーニングや推論のために始まるまで、データセンター用のエネルギーインフラ整備の重要性は認識されていませんでした。
さらにボトルネックが存在するか、そしてそれがどれほど解決可能かによって、進歩のペースを理解する上で重要な示唆が得られます。本論文では、未解決のボトルネックとして、最先端エージェントがオープンエンドな AI 研究においてパフォーマンスを発揮できていない点を特定しました(ただし、これが RSI への必須経路にあるかどうかは今後の検証を要します)。
もし完全自動化された研究に対するボトルネックが容易に解消される世界であれば、AI の能力向上から劇的な進歩の加速が期待できるはずです。しかし、克服が困難なボトルネックが多数残る世界であれば、アンダールの法則が働きます。AI が得意な部分で 100 倍の高速化が実現されたとしても、全体の進歩ペースはわずかな改善に留まります。なぜなら、進歩の速度は最も遅いコンポーネントの速度によって制約されるからです。
私たちがどの世界に住んでいるのかを解明することは、AI の進歩ペースに関する推定値に劇的な影響を与え得ます。本研究の結果が、これらのボトルネックに対するより深い理解の一助となることを願っています。
論文はこちらで読むことができます。著者は Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan です。
これは、論文を盲検審査(ブラインド・ピアレビュー)に送るのとは対照的な状況です。残念ながら、現在のAI分野におけるピアレビューは、審査の質が低かったり、審査員の専門性と割り当てられた論文との間にミスマッチが生じたりする問題を抱えています。おそらくこれは、投稿数の劇的な増加に起因しているのでしょう。
例えば、共同執筆者の間でも意見が割れています。エージェントに創造性が欠如しているのか、それとも認識的閉塞(エピステミック・ロックイン)に陥ってフィードバックを建設的に取り込めなかったのか、見解が分かれているのです。
実際には、AIの進展にはボトルネックがある分野もあれば、ない分野もあることが明らかになるかもしれません。例えば、既存のAIシステムの効率や速度を向上させることは検証可能なタスクであり、ここでの進展は急速です。その結果、企業たちはAIシステムの効率化のためにエージェントを実用的に活用しています。近い将来、検証可能な指標が存在する分野では、AIの進展がさらに加速すると予想されます。
原文を表示
The goal of leading AI labs is recursive self-improvement (RSI): the automation of AI research using AI agents. RSI also underpins forecasts of explosive AI progress. How can we assess if we are close to this milestone?
One way is to use benchmarks that test if agents can conduct AI research. Given the AI community’s focus on benchmarks, they have been the dominant way to evaluate progress towards RSI. Over the last year, many such evaluations have found that agents are now able to make progress on tasks where success is easily verifiable, prompting speculation that we are on the verge of RSI.
But while these evaluations are helpful, they are limited to narrow, verifiable tasks. AI research can be much more open-ended. Success is often not immediately clear or verifiable, and to make progress, researchers need to test promising hypotheses, backtrack, or consider new or unconventional approaches. How can we evaluate agents’ ability to conduct open-ended AI research? We take our first step towards answering this question in a new paper.
We partnered with the authors of two unpublished AI papers and asked them to draft their papers’ main research questions. We then tasked frontier AI agents with conducting research to answer these questions, and gave them thousands of dollars of API credits and compute, and six days of wall-clock time. The original authors reviewed the agents’ papers.
The authors unambiguously rejected both agent papers. To better understand these results, our team spent over a hundred hours analyzing the agents’ logs. Our main takeaways:
The agents lacked the judgment for conducting open-ended research. While the agents proposed directions the expert reviewers found impressive, they quickly rejected their proposed directions based on low-quality or synthetic data.
The agents lacked awareness about the resources available to them. Both runs ended with less than 50% of the API budget spent and with hours left before the deadline, even though the agents could monitor their usage and were encouraged to spend down their budgets.
The agents did not creatively respond to feedback. Despite the agents’ own AI self-reviews surfacing many of the issues that the expert reviewers later raised, the agents did not creatively address these concerns. When faced with negative feedback they responded by adding caveats to existing findings, and doubled down on unpromising research directions.
The agents did not effectively backtrack. They retired their most ambitious research targets within the first day of the experiment, and neither agent fundamentally shifted its approach after that point.
The agents did not follow concrete instructions. They ignored explicit rules about how much time to spend on exploration, how often to get reviews from AI self-review tools, and strict limits on paper length.
We have wanted to evaluate AI’s ability to conduct open-ended research for two years, ever since we released a benchmark to study if agents could be used to improve reproducibility. But we wanted to get our method right. The idea behind our method was suggested by some of the UK AISI coauthors of the paper and refined by our core team at Princeton.
We call these “shadow evaluations” since the agent shadows the original study. In addition to the two of us, the core team comprises Peter Kirgis, Andrew Schwartz, and Stephan Rabanser. The full author list is at the end of this essay.
Shadow evaluations have important advantages: they allow us to test agents on results they haven’t been trained on and can’t access online. They also allow experts who have spent months answering the questions to evaluate agents’ outputs.1
But shadow evaluations also have inherent limitations. Expert reviewers know that the paper is AI-generated, and they might prefer the approach they took over the one that the agent took. Because we are conducting in-depth evaluations of each paper, the sample size is small (in our study, we used just two papers). And these evaluations necessarily involve a lot of researcher flexibility in design, execution, and interpretation.
In fact, we are known for a particular position in the debate on recursive self-improvement and superintelligence. This could influence how we conduct the research. We have a detailed section in the paper on our potential biases and how we address them. We sought out a team of collaborators who don’t all share our priors, and we explicitly surface the disagreements that resulted.2 For future evaluations, we are interested in having “adversarial collaborators” as part of the core team.
Implications for explosive AI progress
Our results suggest that conducting open-ended research remains challenging for frontier AI agents. Still, these findings are tentative, and we are working to address the limitations, such as by increasing the sample size, testing with new models, and through potential scaffold improvements. But if these findings hold up, what are the implications?
First, we need to understand the extent to which frontier AI progress (and RSI) can be achieved simply by hill climbing on verifiable tasks. Our view is that while faster progress is certainly possible on narrow tasks (such as improving efficiency), we don’t think it will lead to broad RSI or explosive progress. Still, we plan to closely follow how AI progress unfolds as a result of AI agents’ capabilities at verifiable tasks.
Second, we need to measure how quickly current limitations of agents at conducting open-ended research (such as the lack of creativity and judgment) can be overcome, such as through more targeted training and scaffold improvements. We plan to continue shadow evaluations on a regular basis to help answer this question.
Finally, even if these limitations can be overcome, there may be further bottlenecks that dampen the pace of AI progress.

Bottlenecks could include compute limits, the necessity of collecting data from real-world experiments, and others that we haven’t recognized yet because they are not currently blocking progress. For example, the importance of high-quality RL environments was not clear before they turned out to be useful for inference scaling. Similarly, the importance of building energy infrastructure for data centers was not realized before companies started investing hundreds of billions on data centers for training and inference.
Whether we encounter further bottlenecks, and how tractable they turn out to be, will be consequential for understanding the pace of progress. In this vein, our paper identifies an unresolved bottleneck, namely, the poor performance of frontier agents on open-ended AI research (though it remains to be seen if it is on the critical path to RSI).
If we’re in a world where the bottlenecks to fully automated research can be easily resolved, we should expect dramatic returns to AI progress from improving AI capabilities. But if we’re in the world with many remaining bottlenecks that are hard to overcome, Amdahl’s law would kick in: even a hundredfold speedup in the parts amenable to AI would only lead to a small speedup in the overall pace of progress, since progress is bottlenecked by the pace of the slowest component.3
Figuring out which world we live in could dramatically impact estimates of the pace of AI progress. We hope our results contribute to a richer understanding of these bottlenecks.
Read the paper here. The authors are Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, and Arvind Narayanan.
1This is as opposed to sending papers for blind peer review. Unfortunately, peer review in AI suffers from poor review quality and a mismatch between reviewers’ expertise and the papers they are assigned, perhaps resulting from a dramatic increase in submissions.
2For example, our coauthors disagree on whether the agents lack creativity, or if they suffered from epistemic lock-in and were unable to productively incorporate feedback.
3In practice, it might turn out that some kinds of AI progress have bottlenecks while others don’t. For example, improving the efficiency and speed of existing AI systems is a verifiable task, where progress has been rapid. As a result, companies have productively utilized agents for improving AI systems’ efficiency. In the near term, we expect quick AI progress in dimensions that have verifiable signals.
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み