科学における AI はデータだけでなく推論が不可欠と指摘
本文の状態
日本語全文を表示中
詳細モードで約11分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MIT Technology Review AI
MIT Technology Review は、AlphaFold の成功が特殊なデータ環境に依存する例外的ケースであり、AI が科学を加速させる真の鍵は推論能力を持つ AI エージェントにあると指摘している。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月10日 18:46
AI深層分析
キーポイント
AlphaFold の成功条件の特殊性
タンパク質構造予測における AlphaFold の成果は、国際協力と巨額の投資によって構築された大規模な実験データセットが存在したという稀な条件下でのみ達成されたものであり、この条件を他の分野で再現するのは数十年かかる可能性が高い。
科学データの生成不可能性
タンパク質結晶解析のように結果が安定して再現可能な実験手法は限られており、多くの科学分野では比較可能なデータを生成することが原理的に困難であるという障壁が存在する。
AI エージェントによる加速の提案
大量データに依存しないアプローチとして、推論能力を備えた AI エージェントが科学の加速を実現する次世代のテンプレートになると提唱されている。
比較可能なデータ生成の科学的不可能性
細胞株の変異や実験環境のばらつきにより、現代のニューラルネットワークを訓練するに十分な一貫性・精度を持つ測定データの作成は、多くの分野で現時点では不可能である。
科学における推論と不確実性の克服
実際の科学研究は不完全なデータ下で行われるため、複数の手法の結果を統合し証拠に基づいて結果を修正する人間の推論プロセスが不可欠である。
重要な引用
AlphaFold had shown that the combination of AI and sufficient data could make groundbreaking discoveries
The primary condition for AlphaFold's success was the existence of the Protein Data Bank
the scientific impossibility of generating comparable data
The skill of science is not in any single tool; it is synthesizing what many tools produce, and revising the results as the evidence comes in.
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
数十年に一度、科学が終焉を迎えたという宣言がなされる。1903 年、崇拝されていた物理学者アルバート・ミケルソンは「物理学の事実はすべて発見された」と記した。1980 年代にはスティーブン・ホーキングが、理論物理学はこの世紀末までに完成するだろうと予測していた。人工知能(AI)の爆発的な登場により、再びその空気が漂っている。今回はノーベル賞という重みも伴ってのことだ。
2024 年、Google DeepMind のデミス・ハサビス氏とジョン・ジャッパー氏は、タンパク質の三次元構造を数千もの実験的に測定された形状から学習して予測するニューラルネットワーク「AlphaFold」への貢献により、化学分野でノーベル賞の一部を受賞した。この悪魔的な問題は半世紀にわたり体系的な攻撃に耐え続けてきたが、AlphaFold はそれを一度きりで解決したかのように見えた。世界は同アプローチの可能性に夢中になった。ハサビス氏とチームは AlphaFold を「AI が科学全体をデジタル速度で加速させるためのテンプレート」と呼んだ。生物学、化学、材料探索のための基盤モデルを構築するスタートアップが相次ぎ、DeepMind の成功に後押しされて数十億ドルの資金を集めた。AlphaFold は、AI と十分な量のデータを組み合わせれば画期的な発見が可能になることを示した(その背後にあるメカニズムが必ずしも理解できなくても)。再び、科学の残りの領域を切り開く道筋が目の前に広がっているように思えた。
確かに、AI は科学に並外れた変化をもたらすでしょう。しかし、アルファフォールド(AlphaFold)やそれに類する技術が、その変革のための最良のテンプレートになるとは限りません。これは画期的な成果であることに間違いありませんが、アルファフォールドのような成果を生み出すための条件は極めて稀であり、他の分野でも同様の条件を満たすまでには、年単位ではなく数十年を要すると考えられます。
科学の加速は、別のアプローチによって実現されるでしょう。それが AI エージェントです。
アルファフォールドが成功した主な理由は、タンパク質データバンク(Protein Data Bank)というデータの存在にあります。これは約 17 万もの実験的に検証されたタンパク質構造からなるデータセットで、ディープマインドのチームはこのデータを基にモデルを訓練しました。このタンパク質データバンクの構築は簡単なことではありませんでした。国際的な科学協力が 53 年間にわたり続けられ、最近の推計によると約 210 億ドル規模の実験作業が積み重ねられて初めて完成したのです。
このような大規模な取り組みを資金調達するのは極めて難しく、調整もほぼ不可能に近く、実行には膨大な時間を要します。そのため、こうした試みは往々にして失敗に終わってきました。
しかし、必要な結束力と資源があり、商業的な所有権によってデータへのアクセスが阻まれない分野であっても、もう一つの障壁——比較可能なデータの生成が科学的に不可能であるという事実——はあまり議論されていません。タンパク質構造の例では、主要な実験手法であるタンパク質結晶解析法は非常に再現性が高く信頼性の高いツールであり、25 回以上のノーベル賞がこの技術に依存しています。しかし、実験科学の大部分では、結果がしばしば変動します。細胞株は変異し、化学薬品には微量の不純物が含まれており、実験室の湿度も変化します。現代のニューラルネットワークを生物学や化学の多くで訓練するために、十分に一貫性があり、正確であり、精密であり、スケーラブルな測定データセットを作成するには、新しい種類の測定手法と新たな標準化アプローチが必要ですが、これらがすぐに準備されることはありません。
もちろん、これらの要件が満たされている分野もいくつかあります。気象予報やゲノミクスの多く、そして化学の限られた領域です。すでに実現している場合もありますが、これらは間もなく AlphaFold 型のような画期的な進展を遂げる可能性があります。米国新技術安全保障委員会の主張のように、政府によるデータセットの生産と調整への支援が不可欠です。しかし、科学における多くの未解決の問題については、少なくとも短期的には別の計画が必要です。幸いにも、より静かで控えめな何かが有望な兆しを見せています。
科学者たちは古来より、不確実性の中で推論を行ってきました。新しい創薬標的の特定に取り組む生物学者たちも、完璧なデータセットを手にしたことは一度もありません。彼らはドッキング計算と既知の構造を組み合わせて用い、分子動力学を考慮し、数種類の結合アッセイを実行し、各手法の強みと弱点に応じて自らの判断で重み付けを行います。科学における真の技量は、特定の単一ツールにあるのではなく、複数のツールが生み出した結果を統合し、新たな証拠が得られるたびにその結果を修正していく能力にあります。これが実際の研究の多くが進む道です。しかし、ごく最近まで、これをこなせるソフトウェアは存在しませんでした。
今やエージェントが可能にしています。端的に言えば、エージェントとは、デジタルあるいは物理的なツールへのアクセス権限と、それらを活用する能力を与えられた AI の推論エンジンです。ここ数年で AI における根本的なアーキテクチャの転換が起き、これにより大規模言語モデルを動力とするこれらのプログラムが急速に普及しました。その結果、科学的に特化したデータセットへの依存は劇的に減少しています。科学にとってこの技術的進展は基盤的な変化をもたらします。すなわち、実際の研究プロセスである反復的で極めて状況依存的な過程を模倣できるデジタルツールを創出できるようになったのです。
AlphaFold のようなツールが限定的な問いに対して強力なアプローチを適用するのに対し、エージェントは本質的に一般主義者です。彼らは科学を行う新しい方法を提供するのではなく、発見という人間のプロセスをデジタルモデルとして再現しているのです。
5 月に発表された Google の AI Co-Scientist を考えてみてください。研究者たちは、このシステムに 1 ページの要約と明確な目標を与えました。「細菌種間で抗生物質耐性がどのように広がるか」を解明することです。これは薬剤耐性感染症を引き起こす主要な要因の一つです。
Co-Scientist は複数のサブエージェントを自動起動しました。一つは文献から仮説を起草し、もう一つは査読者のようにそれらを批判的に検証します。三つ目はトーナメント形式で最強の候補をランキング付けし、四つ目は勝った仮説をさらに洗練させました。
その結果、耐性遺伝子が細菌ウイルスに便乗して移動しているという結論に至りました。つまり、新しい宿主へ運ぶことができるウイルスを随時利用しているのです。この仮説は正しかったのです。インペリアル・カレッジ・ロンドンの研究者たちは、10 年間にわたる地道な実験室での研究を通じて同じ結論に達していました。Co-Scientist がその論文を事前に目にしたことはありませんでしたが、その論文はまだ査読中でした。
Co-Scientist のようなエージェントは依然として新しいツールであり、科学プロセスの不可欠な要素となるまでには克服すべき課題がいくつかあります。まだ幻覚(ハルシネーション)を起こす可能性があり、判断に一貫性がなく、自律的に動作できる時間や入力に制約があるためです。
しかし、こうした技術的な障壁はいつか解消されます。その時、科学エージェントが科学の信頼性、一貫性、そして実行速度にもたらす相乗効果に気づき始めることになるでしょう。
最も注目すべき点は、AI エージェントが科学の「再現性危機」に対する構造的な解決策となり得ることです。これは研究者同士が互いの結果を再現できないという広範な問題です。数十年にわたり、科学界は実験プロセスを標準化するために、研究者に対し生データや正確なコードの共有を強く求めてきました。しかし、研究者たちは「面白い研究」が終わった後に発生するこうした退屈な事務作業に対して長年抵抗してきました。
一方、エージェントは自分が行ったすべての動きを自動的にログに記録します。これにより、結果に至るまでの方法論が正確な記録として残され、精密な再現が可能になります。
2 つ目の帰結として、科学の記憶が増幅されることになります。研究者間での知識移転は famously 不透明なプロセスです。数年間の訓練と観察を通じて行われない場合、大学院生たちは数十年にわたる先人たちが残した散らかった実験ノートを読み込み、プロトコルの成否を分ける詳細を探さなければなりません。
しかし、エージェントが科学プロセスのより大きな部分を占めるようになるにつれ、研究室全体の科学的歴史は、組織的知識を記録する中央集権的で標準化されたリポジトリに保存されるようになります。
しかし、エージェントがもたらす最も重要な影響は「速度」です。アイデアの検証に要する時間が会議での議論よりも短くなれば、人々は議論を止めて実際にテストを実行します。1 時間で 1000 本の論文を読み込み、500 種類の分子を設計し、朝まで失敗した実験から学習できるエージェントは、実験コストを劇的に下げ、科学が成し遂げられるペースそのものを根本から変えるでしょう。また、研究者には以前ならリスクが高すぎて手を出さなかった大胆で奇妙な問いに挑戦する自由も与えられ、まだ想像もできない科学的扉が開かれることになります。
AlphaFold のようなテンプレートが驚異的な発見の鍵となることは間違いありませんが、それ単体では科学の終焉をもたらすには不十分です。むしろ、エージェント指向 AI への移行は、より稀有なレベルのブレイクスルーを意味します。それは同時にすべての科学分野を包摂するツールなのです。歴史的に見れば、これほど広範な影響力を持つツールが登場したのはごく限られた時だけです。微積分、統計的推論、分光法、そしてコンピュータがそれです。それぞれが、これまで誰も問題として認識していなかった新たな世界の問題を明らかにし、その結果、各分野自体を再定義しました。エージェントによって、もう一つの同様の転換期が訪れようとしています。
エリック・シュミットは 2001 年から 2011 年まで Google の CEO を務めました。2024 年には妻のウェンディと共に「Schmidt Sciences」を共同設立し、科学と技術における非伝統的な探求分野への資金提供を行う慈善ベンチャーを運営しています。
スーハス・マヘシュは Schmidt Sciences の AI センターで「AI for Science」事業を統括しており、特に材料発見における AI 応用の専門家です。
エリック・シュミット事務所のアソシエイト兼科学部門責任者であるマイヤ・レヴィンによる追加研究。
原文を表示
Every few decades, someone announces that science has reached its end. In 1903, the revered physicist Albert Michelson wrote that the “facts of physical science have all been discovered.” In the 1980s, Stephen Hawking predicted that theoretical physics might be finished by the end of the century. With the explosive arrival of artificial intelligence, the feeling is in the air again—this time accompanied by a Nobel Prize.
In 2024, Demis Hassabis and John Jumper of Google DeepMind were awarded part of the Nobel in chemistry for their neural network AlphaFold, which predicts the three-dimensional structures of proteins by learning from thousands of experimentally measured shapes. This devilish problem had resisted systematic attacks for half a century; AlphaFold seemed to have solved it once and for all, and the world became fixated on the promise of its approach. Hassabis and his team called AlphaFold “the template for how AI can accelerate all of science to digital speed.” A wave of startups building foundation models for biology, chemistry, and materials discovery raised billions of dollars, buoyed by DeepMind’s success. AlphaFold had shown that the combination of AI and sufficient data could make groundbreaking discoveries (even if we did not understand the underlying mechanisms involved), and it seemed, once again, that a path through the rest of science was laid out before us.
To be sure, AI will bring extraordinary changes to science, but it has become increasingly clear that AlphaFold, and things like it, may not be the best template for that metamorphosis. Though it is a profound achievement, the conditions that produced the likes of AlphaFold are rare, and the time it will take to meet those conditions in other fields will be measured in decades, not years. Instead, the acceleration of science will come about thanks to another approach: AI agents.
The primary condition for AlphaFold’s success was the existence of the Protein Data Bank, a data set of roughly 170,000 experimentally validated protein structures on which DeepMind’s team could train its model. The creation of the Protein Data Bank was not simple: It took 53 years of international scientific cooperation and, by a recent estimate, roughly $21 billion worth of experimental work to assemble. Efforts of that scale are infamously difficult to fund, next to impossible to coordinate, and hugely time-consuming to execute; they have often been unsuccessful as a result.
But even in fields with the requisite cohesion and resources, and where the relevant data are not rendered inaccessible by commercial ownership, another barrier is too little discussed: the scientific impossibility of generating comparable data. In the case of protein structures, the key experimental technique—protein crystallography—is an unusually replicable and dependable tool, so much so that over 25 Nobel Prizes have relied on it. But in most of experimental science, results vary more often than not. Cell lines drift. Chemicals have trace contaminants. Lab humidity changes. The creation of measured datasets that will be consistent enough, accurate enough, precise enough, and scalable enough to train a modern neural network in biology or most of chemistry would require new kinds of measurement and new standardized approaches—none of which will be ready anytime soon.
Of course, there are a handful of fields where these requirements are met: weather forecasting, much of genomics, very limited areas of chemistry. These may see AlphaFold-style breakthroughs soon, if they haven’t already. Government support for the production and coordination of those datasets will be critical, as the US National Security Commission on Emerging Biotechnology has argued. But for most open questions in science, we will need a different plan, at least in the short term. Luckily, something quieter and more modest has begun to show promise.
Scientists have always reasoned under uncertainty. Biologists working to identify new drug targets have never had perfect datasets. Instead, they combine docking calculations and known structures, factor in molecular dynamics, run a handful of binding assays, and use their judgment to weigh each method according to its particular strengths and points of failure. The skill of science is not in any single tool; it is synthesizing what many tools produce, and revising the results as the evidence comes in. This is how most working research actually proceeds. But until very recently, no software could do it.
Agents now can. Simply put, an agent is an AI reasoning engine that has been given access to tools—digital or physical—and the capabilities to use them. Over the last few years, a fundamental architectural shift in AI has enabled the rapid proliferation of these programs, which are powered by large language models, dramatically reducing the need for scientifically specialized datasets. For science, this technological advancement represents a foundational change: it has allowed us to create digital tools that can mimic the iterative, highly contingent process of actual research. While tools like AlphaFold apply a powerful approach to a limited question, agents are inherently generalists. They do not represent a new way to do science—instead, they digitally model the human process of discovery.
Consider Google’s AI Co-Scientist, announced in May. Researchers gave it a one-page brief and a goal: Figure out how antibiotic resistance spreads between bacterial species, a key driver of drug-resistant infections. The system spun up sub-agents. One drafted hypotheses from the literature. Another picked them apart like a peer reviewer. A third ran tournaments to rank the strongest candidates. A fourth refined the winning hypothesis. The agent concluded that resistance genes were hitching rides on bacterial viruses, borrowing whichever virus could ferry them into a new host. The hypothesis was correct. Researchers at Imperial College London had spent a decade reaching the same conclusion through painstaking wet-lab work; their paper, previously unseen by Co-Scientist, was still in peer review.
Agents like Co-Scientist are still novel tools, and there are real challenges to overcome before they become a ubiquitous part of the scientific process: They are still liable to hallucinate, their judgment is not consistent, and they have memory and input constraints that limit the time they can run autonomously. But these technical barriers will fall away, and as they do we will begin to notice the compounding effects of scientific agents on the reliability, consistency, and velocity with which science is done.
Perhaps most notably, agents offer a structural fix for science’s “reproducibility crisis,” the widespread problem of researchers’ inability to replicate each other’s results. For decades, the scientific community has begged researchers to share their raw data and exact code in an effort to standardize experimental processes. But researchers have long resisted this tedious administrative work, which happens after the interesting science is already done. Agents, in contrast, automatically log every move they make, creating an exact record of the method that led to their results and allowing for precise replication.
A second consequence will be an amplification of scientific memory. The transfer of knowledge between researchers is a famously murky process; if it isn’t done over years of training and observation, graduate students are left to pore through the messy lab notebooks kept by decades of predecessors, looking for the details that will make or break their protocol. As agents become an increasingly large part of the scientific process, though, a lab’s entire scientific history will be recorded in a central, standardized repository of institutional knowledge.
But the most important impact of agents will be speed. In any field, when testing an idea takes less time than arguing about it in a meeting, people stop debating and just run the test. An agent that can read a thousand papers in an hour, design 500 molecules, and learn from its failed tests by morning will bring down the cost of experimentation and fundamentally change the pace at which science gets done. It will also give researchers the freedom to chase bold, strange questions they never would have risked their time on before, opening scientific doors we have yet to imagine.
While the AlphaFold template will certainly be key to incredible discoveries, it alone will not bring us to the end of science. Instead, the shift toward agentic AI represents a much rarer tier of breakthrough: a tool that envelops every field of science at once. Historically, tools of such scope have arrived just a handful of times: calculus, statistical inference, spectroscopy, the computer. Each revealed a world of problems no one had thought to formulate, and those problems, in turn, defined their fields anew. With agents, another such transformation is upon us.
Eric Schmidt was the CEO of Google from 2001 to 2011. In 2024, with his wife Wendy, he co-founded Schmidt Sciences, a philanthropic venture to fund unconventional areas of exploration in science & tech.
Suhas Mahesh leads AI for Science work at the AI Center of Schmidt Sciences. He is a specialist in AI for materials discovery.
Additional research by Maya Levin, associate and sciences lead, Office of Eric Schmidt.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み