ASI-Bench:人工超知能の黎明を測る新たなベンチマーク
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Hugging Face Daily Papers
Hugging Face Daily Papers は、既存の知識の適用に依存する現在の AI を超え、未知への探求と自律的な科学実行を評価する初のベンチマーク「ASI-Bench」を発表し、その結果が示す人間指導への依存度を明確にした。
AI深層分析を開く2026年8月19日 13:35
AI深層分析
キーポイント
ASI-Bench の定義と目的
既存の知識の圧縮や適用をテストする従来のベンチとは異なり、AI が未知を探求し新しいアイデアを検証可能な結果に変える能力を評価するために設計された初のベンチマークである。
段階的指導の撤廃によるテスト手法
1 つの研究プロジェクト内で段階的に人間の方法論的指導を撤回し、AI が独自に方法を選択し研究を実行して検証可能な結果を生み出せるかどうかが試される。
厳格な構築プロセスと評価基準
40 名以上の専門家による 31,000 時間以上の作業で構築され、60 のプロジェクトレベルの課題が 11 の科学分野にわたって設定されている。
現状 AI システムの限界を示す結果
完全な指導下では平均スコア 50.91 を記録したが、方法のみ指定で 29.10、方法を自身で決定する必要がある場合は 26.62 に低下し、自律的な研究実行への依存度が浮き彫りになった。
重要な引用
ASI-Bench is the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains
This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research
編集コメントを表示
編集コメント
このベンチマークは、現在の AI が「学習した知識の応用」から「未知への探求」へと移行する際の障壁を明確に可視化した点で画期的である。開発者は今後のエージェント設計において、人間による細かな指示なしに自律的に課題を解決する能力の向上に注力する必要があるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
人工超知能(ASI)の実現には、AI が既存の知識を習得する段階を超え、未知の領域を探求し、新たな知識を生み出し、そのアイデアを検証可能な成果へと変換できる能力が求められます。しかし現在、AI システムの能力は依然として、人間の既知の知識を学習・圧縮・応用することに依存しています。そのため、既存の評価基準(ベンチマーク)は主に、AI が習得した知識に基づいて正解を導き出せるか、あるいは広範な人間の手引きの下でタスクを完了できるかを問うものになっています。
そこで私たちは、ASI-Bench を導入しました。これは、一般的な研究分野全体にわたって AI システムの革新的な探求能力と自律的な科学実行能力を同時に評価する初のベンチマークであり、また、同じ研究プロジェクト内で段階的に人間による方法論的支援を撤廃し、AI がどこまで自力で進められるかをテストする初の試みです。40 名以上の専門家によって 31,000 時間以上の人件費をかけて構築された ASI-Bench は、11 の科学分野にわたる 60 のプロジェクトレベルの研究タスクを含んでおり、AI が独自に方法を選択し、研究を実行し、検証可能な結果を生成できるかを試すために、段階的に方法論的支援を減らしています。すべてのタスクは専門家のレビュー、AI を活用した監査、サンドボックス環境での実行、そして採点者の検証を経ています。
18 の最先端エージェント・モデル構成において、平均スコアは、完全な方法論的支援がある状態で 50.91 でしたが、方法のみが指定された状態では 29.10 に低下し、エージェント自身が方法を決定しなければならない状態ではさらに 26.62 まで下がりました。
この急激な低下は、現在のシステムがいまだに人間の指導に大きく依存しており、プロジェクトレベルの科学研究を自律的に行うには程遠いことを示しています。ASI-Bench は世界中に向けて公開されています。研究者や開発者の皆様には、新しいタスクの提供、今日時点の AI の限界への挑戦、そして https://asibench.apexin.ai/submit を通じて人類が人工超知能へと向かう集団的な道筋を加速させるお手伝いをお願いいたします。
原文を表示
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み