PreScience:科学の未来をエンドツーエンドで予測する
Allen Institute for AI(Ai2)は、科学の将来を予測する「PreScience」を発表した。この技術レポートとデータセットは、科学分野の進捗をエンドツーエンドで予測する手法を示しており、AIを活用した科学研究支援における注目に値する取り組みである。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
2026 年 2 月 25 日
Ai2
すべての科学論文は、誰と協力するか、どの先行研究を基盤とするか、それらを組み合わせることでどのような科学的貢献が生まれるか、そしてそれをどのように伝えるかという一連の選択から始まります。その後、コミュニティがその結果にどれだけの注目を払うべきかを決定します。これらの選択は数ヶ月から数年にわたって展開され、膨大かつ絶えず進化し続ける研究の蓄積によって形作られます。そしてこれらが、ある分野における科学進歩の方向性を最終的に決定づけるのです。
このプロセスを十分に理解して予測することは、AI システムが科学のダイナミクスをどの程度深く理解しているかを試す意味のあるテストとなります。特定の時点までの科学的記録を基にすれば、彼らは次に来るものを予測できるでしょうか?パズルの一部ではなく、チーム編成から最終的な影響力に至るまで、全体像を予測できるのでしょうか?
これは、PreScience という新しいベンチマークの背景にある問いです。このベンチマークは、シカゴ大学と共同で構築され、米国国立科学財団(NSF)の支援を受けており、NSF グローバル観測所および科学技術進歩のための仮想実験室を通じて、研究ワークフロー全体にわたる科学的予測を評価するために設計されています。PreScience は、科学の進展を 4 つの組み合わせ可能な段階——チーム編成、文献選定、貢献生成、影響予測——に分解し、これらを連結することで、分野が月ごとにどのように進化するかという完全なシミュレーションを実現します。AI 研究を対象とした 12 ヶ月のシミュレーションを実行した結果、驚くべき事実が明らかになりました。つまり、シミュレーションされたコーパスは、人間研究者が実際に生み出したものよりも体系的に多様性に欠け、新規性も低いという点です。ボトルネックはチームの選定や先行研究の選択にあるのではなく、生成ステップそのものに存在していました。
私たちが目指すのは、PreScience が、分野の行方を予測できる AI の開発を加速し、最終的には研究者がより速く目標に到達できるよう支援することです。データセット、評価スイート、そして技術報告書は、すべて現在利用可能です。
なぜ PreScience を構築し、どのように機能するのか
科学的研究の自動化に関する既存の評価の多くは、限定的または合成されたタスクに焦点を当てています。モデルが妥当な抄録を作成できるか、あるいは引用数を予測できるかといった問いです。研究者たちはまた、将来の共同研究の予測、新たなアイデアの組み合わせの予見、続編の研究の予測、出版への影響の見積もりなど、より実質的な問いも研究してきました。しかし、それらは常に個別に扱われてきました。実際には、これらは単一の科学的貢献のライフサイクルにおける相互依存する段階です。チームは共通の関心と補完的な専門知識を基盤として形成され、先行研究を参照し、新たな貢献を行い、コミュニティが時間とともに反応します。各段階は次の段階へとつながっており、それらを別々に研究することは、プロセス全体について得られる学習を制限することになります。
PreScience は、これらを一貫して扱う最初の大規模ベンチマークです。これは arXiv から収集された実際の論文、著者、引用履歴に基づいており、計算言語学、機械学習、コンピュータビジョン、情報検索を含む 7 つの AI サブカテゴリを網羅しています。各論文について、PreScience は予測対象としてタイトルと抄録を含み、さらにその論文の主要な参考文献、著者の出版履歴、およびその他の関連メタデータを含んでいます。
このデータセットは、2023 年 10 月から 2025 年 10 月までに出版された約 10 万本の対象論文をカバーしており、これは 50 万本を超えるより広範なコーパスおよび約 18.3 万人の一意の著者から抽出されたものです。モデルは 2023 年 10 月から 2024 年 10 月までの論文を利用可能であり、翌年のデータに対して評価が行われるため、常に未来への予測(forecasting)を行い、既知の期間内での補間(interpolation)を行うことはありません。また、データの多くが最先端モデルのトレーニングカットオフ日以降に作成されているため、結果に汚染(contamination)が漏れ出すリスクはありません。
いくつかの追加的な設計上の選択により、ノイズや近道ではなく、真の予測能力を反映するようデータセットが十分に清浄化されています。著者の同一性は、クラスタリング品質を向上させる手法によって曖昧さを解消されます。対象論文は、1 から 10 の主要な参考文献を持つものへとフィルタリングされ、予測が極めて容易または困難になり得る外れ値(outliers)は除外されます。さらに、すべてのメタデータ(例:被引用数、h インデックス、出版履歴など)は各論文の出版日に対して時間的に整合性が取られており、モデルが偶然にも未来からの情報を参照してしまうことがありません。これら一連の選択により、ベンチマークテストセットへの情報漏洩がなく、クリーンな評価が可能となります。
この基盤を踏まえ、PreScience は科学上の進展を 4 つの相互依存するタスクに分解し、それぞれが研究が展開される過程における重要な意思決定点を反映しています。
- 共著者予測。ある著者と現在の分野の状況を前提としたとき、次に誰と共同研究を行うでしょうか?これは、研究チームがどのように結成されるかという社会的・トピック的なダイナミクスを反映しています。
- 先行論文選択。あるチームが既存文献からどの論文を基礎として構築するかを予測します。これは、モデルが新しい貢献に対する最も関連性の高い基盤を特定できるかをテストするものです。
- 貢献生成。チームと先行論文が確定した時点で、その論文は実際に何を述べるのでしょうか?モデルは、妥当な新たな貢献を表すタイトルと抄録を生成する必要があります。
- 影響予測。論文が存在したとき、どれほどの注目を集めるでしょうか?モデルはその論文が初年度に獲得する被引用数を予測します。
これらのタスクは組み合わせ可能です。各タスクを個別に研究することもできますし、これらを多段階の「科学シミュレーター」として連鎖させることもできます。チームを予測し、その論文を生成し、それらの論文を文献に組み込み、それを月ごとに繰り返すのです。
生成された貢献の評価における新たなアプローチ:LACERScore
生成された科学的貢献の品質を測定することは、それ自体が課題です。ROUGE や BERTScore といった標準的なテキスト類似度指標は、2 つの抄録が表面的な特徴を共有しているかどうかを示すことはできますが、それだけでは、両者が同じ科学的発見を記述しているかどうかについてはほとんど何も教えてくれません。2 つの抄録は、密接に関連する結果に対して非常に異なる言語を使用することもあれば、根本的に異なる作業を記述しながらも大量の共通語彙を共有することもあります。
これに対処するため、PreScience は「LACERScore」と呼ばれる校正済み評価指標を導入しました。これは言語モデルを判事として用い、生成された抄録と実際の抄録の整合性を 1 から 10 のスケールで評価するものです。この評価は、異なるスコアレベルが何を意味するかを示す基準となる自動構築参照例によって導かれます。実際には、LACERScore は専門家の人間の判断に密接に追従し、人間注釈者同士の合意レベルに近づきつつ、従来の自動指標を大幅に上回っています。
私たちの発見
強力なベースラインや最先端モデルを用いた場合でも、PreScience の 4 つのタスクすべてにおいて改善の余地は依然として大きいです。
共著者予測においては、過去における共著頻度に基づく単純なヒューリスティックが、より複雑な機械学習ベースラインをすべて上回りました。また、2 人の研究者がこれまで一度も共同したことがない「初回」協力を予測する際には、どのベースラインも信頼性を持って対応できません。
先行研究の選択もモデルにとって非常に困難です。最良のベースラインでも nDCG(標準的なランキング指標)は約 0.13 に留まり、著者の完全な出版履歴へのアクセスを許容しても、チームがどの特定の論文を引用するかを特定する点でモデルは苦戦しています。
貢献生成において、最先端のLLMは実際の貢献と中程度にしか類似していない要約を生成します。テストしたモデルの中で最も強力なGPT-5でも、LACERScoreでは平均して10点満点中約5.6点です。より大規模で新しいモデルは、小規模または初期のモデルよりも一般的に改善しますが、その向上は漸進的なものであり劇的ではありません。これを文脈化すると、要約を単純に言い換えただけのもの方がはるかに高いスコアを得るため、モデルが生成するものと研究者が実際に記述したものの間には有意なギャップが存在します。
そして影響予測においては、最も優れた特徴量の組み合わせであっても、依然として大きな予測誤差が残ります。引用数が非常に多い論文——おそらく予測において最も重要な対象——は体系的に正しく予測するのが最も困難です。
より大きな問い:1年間のシミュレーションされた科学
個々のタスクの結果自体が示唆に富んでいますが、PreScienceは究極的に、より野心的な問いのために設計されています。すなわち、4つの段階を組み合わせて完全なシミュレーションを行った場合どうなるかという問いです。
私たちは12ヶ月間のシミュレーションを実行しました。これは各月に研究チームを予測し、彼らの過去の業績を選択し、論文を生成し、その論文を後続の月々の進化する文献に追加していくプロセスです。その結果得られたのは、人間研究者が同じ期間に観察可能に生産したものと比較できる、AI研究に関する合成コーパスです。
主要な発見は、シミュレーションされたコーパスが体系的に多様性に欠け、新規性も低いという点です。個別に生成された論文が完全に的外れであるわけではありません。各論文は、実際の論文と同様に先行研究からおよそ同等の距離にあります。しかし、それらは互いに集約する傾向があります。シミュレーションが進むにつれて、新たに生成される論文は*互いに*ますます類似し、現実の研究コミュニティが探求した範囲よりも狭いアイデアの範囲に収束していきます。
*以下の図:シミュレーション(合成)論文**(A)は、同じ時期に対応する実在(自然)論文と比較して多様性が低く、(B)新規性の低い傾向を示します。新規性を固定された事前シミュレーションコーパスに対して測定した場合(C)**、この傾向は消滅します。H<t:新しい論文を生成する前(コーパスからの合成生成物を含む)。H<t0:テスト期間の前(コーパスからの実在論文のみを含む)*。
PreScience の診断的価値がここではっきりと現れています。シミュレーション中に浮上した著者セットや先行研究のセットは、実際の世界のものよりも*むしろ多様性が高い*です。上流工程(チーム編成と文献選定)がボトルネックになっているわけではありません。多様な入力を与えられた場合でも、言語モデルは依然として、実際の研究者が書くものよりも均質化された出力を生産する傾向があります。
次に何をするか、そしてどのように参加するか
PreScience は、科学的予測におけるいくつかのオープンな課題を浮き彫りにしています。具体的には、初めての共同研究の予測、関連する先行研究の提示、現実世界の新奇性に合致する貢献の生成、そしてどの論文が過大かつ重要な影響を与えるかの予測です。これらすべては、共著者の推薦や文献のナビゲーションから有望な方向性の提案、論文の影響の予測に至るまで、研究者を真に支援できる AI ツールの種類と直接結びついています。科学の実際のあり方に根ざした評価を行うことで初めて、真の発見を支える AI システムを実現できます。
PreScience を生きているベンチマークとして捉えています。科学的記録が成長するにつれて、予測手法を検証する能力も高めていくべきです。今後、機関所属、掲載誌、資金源といったより豊かな文脈信号や、図表などのマルチモーダルな科学アーティファクトの探索にも期待しています。
PreScience には、トレーニング用およびテスト用のコーパス、著者マッピング、ベースライン実装、評価スクリプトが同梱されています。これらすべては、GitHub および Hugging Face リポジトリで確認できます。また、技術報告書については こちら をご覧ください。
*PreScience* は Ai2 で開発され、シカゴ大学と協力して作成されています。この資料は、米国国立科学財団の助成金番号 TF-2404109 の支援に基づいています。本資料に示された意見、発見、結論、または推奨事項はすべて著者のものであり、必ずしも国立科学財団の見解を反映するものではありません。
最新の Ai2 ニュースに関する月次更新を受け取るには、購読してください。
原文を表示
February 25, 2026
Ai2
Every scientific paper starts with a series of choices: who to work with, which prior work to build on, what scientific contribution results from combining them, and how to communicate it. Then the community decides how much attention those results deserve. These choices unfold over months and years, shaped by an enormous and constantly evolving body of research—and they're what ultimately determine the direction of scientific advances in a field.
Understanding this process well enough to forecast it would be a meaningful test of how deeply AI systems grasp the dynamics of science. Can they, given the scientific record up to a fixed point in time, predict what comes next—not just one piece of the puzzle, but the whole picture, from team formation through to eventual impact?
That's the question behind PreScience, a new benchmark we've built with the University of Chicago, supported by the U.S. National Science Foundation (NSF) through NSF Global Observatory and Virtual Laboratory for Science and Technology Advance, to evaluate scientific forecasting across the entire research workflow. PreScience breaks a scientific advance into four composable stages – team formation, literature selection, contribution generation, and impact prediction – that can be chained into a full simulation of how a field evolves month by month. When we ran a 12-month simulation of AI research, the results revealed something striking: the simulated corpus was systematically less diverse and less novel than what human researchers actually produced, and the bottleneck wasn't in selecting teams or prior work—it was in the generation step itself.
Our hope is that PreScience accelerates progress toward AI that can anticipate where a field is heading and eventually help researchers get there faster. The dataset, evaluation suite, and tech report are all available now.
Why we built PreScience and how it works
Most existing evaluations for automating aspects of scientific research focus on narrow or synthetic tasks—can a model write a plausible abstract, or predict a citation count? Researchers have studied more substantive questions too, like predicting future collaborations, anticipating novel idea combinations, forecasting follow-up work, and estimating publication impact—but always in isolation. In practice, these are interdependent stages in the lifecycle of a single scientific contribution. Teams form around shared interests and complementary expertise, draw on prior research, contribute something new, and the community responds over time. Each stage feeds into the next, and studying them separately limits what you can learn about the process as a whole.
PreScience is the first large-scale benchmark to treat them jointly. It's grounded in real papers, real authors, and real citation histories from arXiv, spanning seven AI subcategories including computational linguistics, machine learning, computer vision, and information retrieval. For each paper, PreScience includes the title and abstract as the prediction target, along with the paper's key references, the publication histories of its authors, and other relevant metadata.
The dataset covers ~100K target papers published between October 2023 and October 2025, drawn from a broader corpus of over 500K papers and nearly 183K unique authors. Models can use papers from October 2023 through October 2024 and are evaluated on the following year, so they're always forecasting into the future rather than interpolating within a known period. And because much of the data post-dates the training cutoff of frontier models, there's no risk of contamination leaking into the results.
Several additional design choices ensure the dataset is clean enough to reflect genuine forecasting ability rather than noise or shortcuts. Author identities are disambiguated via a method that improves clustering quality. Target papers are filtered to those with one to ten key references, removing outliers that would be incredibly easy or hard to predict. And all metadata (e.g., citation counts, h-indices, and publication histories) is temporally aligned to each paper's publication date, so models never accidentally see information from the future. Together, these choices ensure clean evaluations with no information leakage into the benchmark test set.
With that foundation in place, PreScience breaks down a scientific advance into four interdependent tasks, each reflecting a key decision point in how research unfolds:
- Collaborator prediction. Given an author and the current state of the field, who will they work with next? This reflects the social and topical dynamics of how research teams come together.
- Prior work selection. Given a team, which papers from the existing literature will they build on? This tests whether models can identify the most relevant foundations for a new contribution.
- Contribution generation. Once a team and the prior work are established, what will the paper actually say? Models must produce a title and abstract that represent a plausible new contribution.
- Impact prediction. Once a paper exists, how much attention will it receive? Models forecast how many citations it’ll accumulate in its first year.
The tasks are composable. You can study each one in isolation, or chain them together into a multi-step "science simulator"—predicting teams, generating their papers, folding those papers back into the literature, and repeating month by month.
A new way to evaluate generated contributions: LACERScore
Measuring the quality of a generated scientific contribution is its own challenge. Standard text-similarity metrics like ROUGE or BERTScore can tell you whether two abstracts share surface-level characteristics, but that alone doesn't tell you much about whether they describe the same scientific finding. Two abstracts might use very different language for closely related results, or share substantial vocabulary while describing fundamentally different work.
To address this, PreScience introduces LACERScore, a calibrated evaluation metric that uses a language model as a judge. Given a generated abstract and the real one, the model rates their alignment on a 1-to-10 scale, guided by automatically constructed reference examples that anchor what different score levels mean. In practice, LACERScore tracks expert human judgments closely—approaching the level of agreement between human annotators themselves and substantially outperforming prior automatic metrics.
What we found
Even with strong baselines and frontier models, there's a lot of room for improvement across all four tasks in PreScience:
In collaborator prediction, a simple heuristic based on how often two authors have co-published in the past outperforms all of the more complex ML baselines. And when it comes to predicting *first-time* collaborations, where two researchers have never worked together before, none of the baselines can do it reliably.
Prior work selection is also quite difficult for models. The best baseline achieves an nDCG (a standard ranking metric) of only about 0.13, meaning models struggle to identify which specific papers from the literature a team will cite even when given access to the authors' full publication histories.
In contribution generation, frontier LLMs produce abstracts that are only moderately similar to the real contributions. GPT-5, the strongest model we tested, averages roughly 5.6 out of 10 on LACERScore. Larger and more recent models generally improve over smaller or earlier ones, but the gains are incremental rather than dramatic. To put that in context, a simple paraphrase of the abstract scores much higher, so there's a meaningful gap between what models generate and what researchers actually wrote.
And in impact prediction, even the best combinations of features leave substantial prediction error. Highly cited papers – arguably the most important ones to forecast – are systematically the hardest to get right.
The bigger question: A full year of simulated science
The individual task results are revealing on their own, but PreScience is ultimately designed for a more ambitious question: what happens when you compose the four stages into a full simulation?
We ran 12-month simulations, predicting research teams each month, selecting their prior work, generating papers, and adding those papers back into the evolving literature for subsequent months. The result is a synthetic corpus of AI research that we can compare against what human researchers observably produced over the same period.
The headline finding is that the simulated corpus is systematically less diverse and less novel than real research. Individual generated papers aren't wildly off-base—each one is roughly as different from prior work as a real paper would be. But they tend to cluster together. As the simulation progresses, newly generated papers become increasingly similar to *each other*, converging on a narrower range of ideas than what the real research community explored.
*Figures below: Simulated (synthetic) papers **(A) are less diverse and (B) trend towards being less novel compared to ground truth (natural) papers that correspond to the same time period. When novelty is measured relative to the fixed pre-simulation corpus (C)**, this trend disappears. H<t: Prior to generating a new paper (includes synthetic generations from the corpus). H<t0: Prior to the test period (includes only real papers from the corpus).*
The diagnostic value of PreScience shows up clearly here. The sets of authors and prior work surfaced during the simulations are actually *more* diverse than their real-world counterparts. The upstream stages (team formation and literature selection) aren't the bottleneck—it's the contribution generation step where diversity collapses. The language model, given a diverse set of inputs, still tends to produce outputs that are more homogeneous than what real researchers would write.
What's next and how to get involved
PreScience highlights several open challenges in scientific forecasting: predicting first-time collaborations, surfacing relevant prior work, generating contributions that match real-world novelty, and anticipating which papers will have outsized impact. Each of these connects directly to the kinds of AI tools that could meaningfully help researchers, from recommending co-authors and navigating the literature to proposing promising directions and anticipating a paper's influence. If we want AI systems that support real discovery, we need evaluations grounded in how science actually happens.
We see PreScience as a living benchmark. As the scientific record grows, so should our ability to test forecasting methods against it. Looking ahead, we're excited to explore richer context signals like institutional affiliations, venues, and funding sources, as well as multimodal scientific artifacts like figures and tables.
PreScience ships with training and test corpora, author mappings, baseline implementations, and evaluation scripts. You can find everything via our GitHub and Hugging Face repos—read our tech report here.
*PreScience is developed at Ai2 in collaboration with the University of Chicago. This material is based upon work supported by the U.S. National Science Foundation under Award No. TF-2404109. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation.*
Subscribe to receive monthly updates about the latest Ai2 news.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み