AstaLabsが自動科学発見ツール「AutoDiscovery」を発表
AstaLabsは、科学研究プロセスを自動化する新ツール「AutoDiscovery」の提供を開始した。このシステムにより、研究者は実験設計からデータ分析までを効率的に行えるようになり、科学発見のスピードと精度が向上する。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
*更新 3/2: さらに 3 ヶ月延長*
AutoDiscovery を立ち上げて以来、研究者たちは腫瘍学、気候科学、海洋生態学、昆虫学、サイバーセキュリティ、音楽認知、社会科学などにおいて 20,000 件以上の仮説を生成してきました。アクセス期間をさらに 3 ヶ月間延長し、クレジット割り当てを更新します。
すべてのアカウントに 500 の仮説クレジットが付与されます。(各クレジットは、AutoDiscovery があなたのデータ上で 1 つの仮説を生成・検証できる権利を表します。)もし残高が 500 を下回っていた場合、不足分を追加しました。もし 500 を超えて残っていた場合は、その全額を維持します。また、元の割り当てを使い果たしてしまった場合でも、再アクティブ化され、フルの 500 が付与されます。
AstaLabs で AutoDiscovery をお試しくださいし、発見した結果をお知らせください。
*以下に原文を続けます。*
すべてのデータセットには、まだ誰も探していない発見が潜んでいます。パターンはそこにあるのです——行と列の中に待機しています——しかし、それらを表面化させるには、適切な問いを立てる必要があります。科学において、どの問いを立てるべきかを知ることは往々にして最も困難な部分であり、今日の多くの AI ツールはその点で役立ちません。これらは目的指向型であり、作業を開始する前にあなたが研究課題を提供することを待っているのです。
これは、従来の統計ワークフローから、Google の AI コサイエンティスト や FutureHouse の プラットフォームのような洗練されたマルチエージェント研究システムに至るまで、あらゆるものに当てはまります。これらのツールは文献の要約、実験設計、仮説の大規模な検証を行えますが、それでもまだ「何を調査すべきか」を人間が指示するのを待っている状態です。
そのため私たちは、Asta プラットフォーム内で実験機能として利用可能な AutoDiscovery(旧称:AutoDS)を開発しました。これは問いから始めるのではなく、データから始め、自ら質問を投げかけます。自然言語で仮説を生成し、実験計画を提案し、それを実行するための Python コードを作成し、統計結果を解釈し、得られた知見を用いて新たな仮説を生成します。構造化されたデータセットを与えて探索させれば、一夜にして数百もの実験を実行しても、再現可能でさらに調査を進められる新規な研究方向の完全なリストが得られます。
研究者たちはすでに、AutoDiscovery を活用して、20 年間にわたる海洋生態系データにおける栄養段階関係の解明や、治療方針の決定に役立つ可能性のあるがん変異における相互排他性パターンの特定など、分野を超えた驚くべき隠れたパターンを浮き彫りにしています。社会科学における AutoDiscovery のいくつかの発見は、独立した検証の後、昨年 11 月に査読付き論文として発表されました(https://arxiv.org/abs/2511.12529)。
AutoDiscovery は、科学者とデータの関係を変革します。静的なリポジトリであったデータセットを、影響力のある新たな問いを引き出すための対話型調査の道具へと変えるのです。
初期パートナーによる事例研究や詳細な技術解説は こちら でご覧いただけます。また、以下では AutoDiscovery の概要とその仕組みについてさらに詳しく説明します。
*「AutoDiscovery は、データを用いた深い調査に似ていますが、思考の速度で実行されます。」—— サンチャイタ・ハズラ氏(ユタ大学社会行動科学学部経済学者)*
AutoDiscovery が何を追求するかを決定する方法
オープンエンドな探索を任された AI システムには、古典的な失敗モードが存在します。どの手がかりを追跡する価値があるかを判断するための原理的な方法がない場合、システムはランダムに彷徨うか、あるいはトレーニングデータに埋め込まれたバイアスをそのまま引き継いでしまいます。AutoDiscovery はこれに対し、「ベイズ的驚き(Bayesian surprise)」を用いて探索を導くという重要な洞察によって解決します。これは、証拠を見た後にシステムの信念がどれだけ変化したかを測る指標です。
実験を実行する前に、AutoDiscovery は仮説が真であるかどうかについて事前の信念(prior belief)を持っています。これは確率分布として表現されます(この「信念」は基盤となる言語モデル内の世界知識に由来し、モデルに対して複数回クエリを行うことで抽出されます)。結果を確認した後、システムは事後の信念(posterior belief)へと更新されます。「驚き」とは、その変化の大きさを指すもので、実験データがもたらした情報獲得量を定量化するか、あるいはデータがどのようにしてシステムの内部推計の見直しを促したかを示すものです。
重要なのは、AutoDiscovery が信念の変化(驚き)の大きさだけでなく、その変化の「方向」も追跡している点です。正のシフトは、証拠によってシステムが仮説がより真である可能性が高いと考える方向へ動いたことを意味し、負のシフトは、証拠によって仮説がデータセット内で支持されていないと考える方向へ動いたことを意味します。負のシフトであっても、それは非常に驚くべきこととなり得る場合があり、特に既存の常識に反する証拠がある場合には、その価値は極めて高いものとなります。
この設計は、よく知られた科学的直観を反映しています:*我々の期待を意味ある形で変化させる結果は、単に既知の仮定を確認するだけの結果よりも、しばしばより興味深いものです*。驚異を追及することで、AutoDiscovery は自然と予期せぬ方向へと引き寄せられ、それは明白なパターンではなく、真の発見を表す可能性が最も高い結果となります。
しかし、驚異だけでは不十分です。可能な科学的問いの空間は実質的に無限であり、それを探索するには知的な検索が必要です。AutoDiscovery は、モンテカルロ木探索(Monte Carlo Tree Search: MCTS)を用いてこの空間を効率的にナビゲートします。MCTS は新しい仮説の探索と既知の手がかりへの優先付けをバランスさせ、計算リソースを最も有望な探究の枝へと配分します。
ベイズ的驚異と MCTS を組み合わせることで、AutoDiscovery は「次に何を調査すべきか」という問いに答えるために、人間研究者と協働するための原理に基づき、スケーラブルな方法を提供します。
*「通常、数週間あるいは数ヶ月を要する手動の探索的モデリング分析が、たった一日で完了しました。」 — フレッド・ハッチンソンがんセンターの生物統計学ポストドクター研究員、Stephen Salerno 博士*
AutoDiscovery が発見したものをナビゲートする
AutoDiscovery は、マサチューセッツ大学アマースト校との研究協力として始まり、昨年にオープンソースコードと共に発表されました。しかしこれまで、科学者がこれを実行するための手軽でホストされた方法はありませんでした。科学分野の初期パートナーやインフラストラクチャのサポートのために Google Cloud Platform と連携して作業を行った結果、より多くの研究者の手元に届ける準備が整いました。
AstaLabs において、AutoDiscovery の進捗と発見結果は、実験が完了するにつれて埋まっていく表(上記左パネル)に表示されます。各行はシステムがテストした仮説を表しています。Surprisal スコア(驚異度スコア)を確認することで、新しい証拠がシステムの信念をBefore(以前)からAfter(以後)へとどのようにシフトさせるかを見ることができます。これは、各発見がシステムの期待にどの程度挑戦し、あるいは確認したかを定量化するものです(上記右パネル参照)。さらに、生成された仮説のシーケンスを観察することで、探索ツリー(上記中央パネル)が成長していく過程として発見の進捗を追跡することもできます。
表内の任意の行または探索ツリーのノードをクリックすると、詳細を検索できます。
ケーススタディ:がん変異における相互排他性の発見
これを具体化するために、スウェーデンがん研究所のポール・G・アレン研究センターの腫瘍内科医と共に、乳がんの変異データセットを探索した際に AutoDiscovery が見つけたものについて考えてみましょう。
変異の共起パターンを広く探索することから始まり、システムは段階的により具体的な仮説を生成・検証しました。ある探究の枝分かれがゲノム研究における発見につながりました:PIK3CA 変異を持つ患者において、TP53 変異は偶然に期待される頻度よりも少ないことが示されました。これは潜在的な相互排他性パターンであり、2 つの変異が機能的に重複しているか、あるいは両方の変異を保有する細胞が生存できないことを示すシグナルです。
AutoDiscovery のこの仮説に対する事前信念は中立でした。事前分布の平均値は 0.50 で、このパターンが成立するかどうかに不確実性があることを反映しています(0 は偽であると信じる、1 は真であると信じる)。データを分析した後、事後分布は平均値 0.82 へと劇的にシフトしました。この大きな信念の更新は強い驚きとして記録され、システムがこの発見をフラグ付けし、関連する仮説の探索を継続した理由です。
共同で研究を行った腫瘍専門医たちは、このシグナルに衝撃を受けました。AutoDiscovery は手作業では探索不可能なあまりにも広大な検索空間から、妥当な相互排他性の関係を浮き彫りにし、発見を検証・解釈するための具体的な後続調査を即座に提案しました。
***「AutoDiscovery の、ありふれた場所に隠れているかもしれない発見を明らかにする能力は、特にがん研究において極めて価値があります。」 — スウェーデンがん研究所免疫腫瘍学センター所長で医療腫瘍医のケリー・ポールソン博士*/
AstaLabs の始め方
AstaLabs にログイン して、ご自身のデータをアップロードする前にワークフローをエンドツーエンドで確認するために「Example Sessions」データセットを試してみてください。ご自身での分析を実行する準備ができたら:
- セッションを設定します。「+ New exploration」をクリックしてウィザードを開き、ファイルをアップロード(CSV、JSON、Parquet など)してください。また、システムの信念を初期化するためにコンテキストの説明を入力します。反復処理を行う場合は、「Advanced Settings」から「Intent」フィールドに過去のランからの知見を貼り付けて検索を精緻化できます。最後に、実験予算を設定して実行する仮説の数を制御します。
- 発見された結果をリアルタイムで追跡します。「Start Run」をクリックすると、実験が完了するたびにライブテーブルが更新されます。「Surprisal score(驚異度スコア)」を確認して、最も驚くべき発見を見つけてください。画面を離れても問題ありません。分析が完了した時点で結果はそこにあります。
- 詳細を検証します。任意の行をクリックすると「Inspector Panel」が開きます。ここでは作業内容を確認できます:完全な仮説、統計分析、そして結果生成に使用された実際の Python コードを表示可能です。これは再現可能で拡張可能な、完全に透明性の高い成果物です。
AutoDiscovery の実行は計算リソースを多く消費するため、通常は数時間かかります。そのため、早期アクセスユーザー向けに、一度限りのクレジット付与によりコストを負担しています。自動的に 1,000 の「Hypothesis Credits(仮説クレジット)」が付与されます(仮説 1 つ = クレジット 1)。このクレジットは 2026 年 2 月 28 日まで利用可能です。
クレジットを賢く使うにはどうすればよいでしょうか?最初の実行はテストドライブだと考えてください。まずは小さなバッチ(仮説 10 未満程度)から始めて、何が可能かを確認することをお勧めします。出力に慣れたら、自信を持って 1 セッションあたり 50〜100 の仮説まで拡大し、大規模データセットに対する詳細な分析を行ってください。(注:実行は最大 500 個の仮説まで制限されています。)
アップロードされたデータが機密情報ではないことを確認するよう求められます。ソースとなるデータセットは、分析完了から 7 日後に自動的に削除されます。AutoDiscovery は、発見を再現・拡張するために必要な出力(仮説、計画、コード、結果)のみを保持します。
今日すぐにAstaLabs で AutoDiscovery をお試しください——何が発見されるか驚くかもしれません!サポートや詳細情報が必要ですか?こちらまでお問い合わせください。
*AutoDiscovery などの新機能への早期アクセスを得るために、Asta Preview に登録してください。詳しくは*こちら*をご覧ください。
Ai2 の最新ニュースに関する月次更新を受け取るには購読してください。
原文を表示
*Update 3/2: Extended for three more months*
Since launching AutoDiscovery, researchers have generated over 20,000 hypotheses across oncology, climate science, marine ecology, entomology, cybersecurity, music cognition, social sciences, and more. We're extending access for three more months and refreshing credit allocations.
All accounts now receive 500 Hypothesis Credits. (Each credit lets AutoDiscovery generate and test one hypothesis on your data.) If your balance had fallen below 500, we've topped you up. If you had more than 500 remaining, you keep it all. And if you burned through your original allocation, you're reactivated with a full 500.
Try AutoDiscovery in AstaLabs and let us know what you find.
*Original post follows.*
Every dataset holds findings no one has looked for yet. The patterns are there – waiting in the rows and columns – but surfacing them requires asking the right questions. In science, knowing which questions to ask is often the hardest part, and most of today's AI tools don't help. They're goal-driven, waiting for you to provide a research question before they get to work.
This is true of everything from traditional statistical workflows to sophisticated multi-agent research systems like Google's AI co-scientist and FutureHouse's platforms. These tools can synthesize literature, design experiments, and validate hypotheses at scale—but they still wait for you to tell them what to investigate.
That's why we built AutoDiscovery (formerly AutoDS), now available in AstaLabs as an experimental feature inside the Asta platform. Instead of starting with a question, AutoDiscovery starts with your data and asks its own questions—generating hypotheses in natural language, proposing experiment plans, writing Python code to execute them, interpreting statistical results, and using what it learns to generate new hypotheses. Give it a structured dataset and let it explore. Whether you run a quick analysis or hundreds of experiments overnight, you'll receive a complete list of novel possible research directions—each one reproducible and ready to investigate further.
Researchers are already using AutoDiscovery to surface surprising, hidden patterns across disciplines, from uncovering trophic relationships in 20 years of marine ecosystem data to identifying mutual-exclusivity patterns in cancer mutations that could inform treatment decisions. Several of AutoDiscovery’s findings in social science were even published in a peer-reviewed paper last November (after independent verification).
AutoDiscovery changes the relationship between scientists and their data—transforming datasets from static repositories into interactive artifacts for inquiry that surface impactful new questions.
Read case studies and detailed technical write-ups from our early partners here, and read on for more information about AutoDiscovery and how it works.
"AutoDiscovery is almost like deep research with data, but at the speed of thought." — Sanchaita Hazra, Economist in the College of Social and Behavioral Science at the University of Utah
How AutoDiscovery decides what to pursue
AI systems tasked with open-ended exploration face a classic failure mode: without a principled way to decide which leads are worth pursuing, they either wander randomly or inherit whatever biases are baked into their training data. AutoDiscovery solves this with a key insight: guide the search using *Bayesian surprise*, namely a measure of how much the system's beliefs change after seeing evidence.
Before running an experiment, AutoDiscovery holds a prior belief about whether a hypothesis is true, represented as a probability distribution (this “belief” comes from the world knowledge in the underlying language model, and is extracted by querying the model multiple times). After seeing the results, it updates to a posterior belief. The *surprise* is the magnitude of that shift—quantifying the information gain provided by the experimental data or indicating how data drove the system to reconsider its internal estimates.
Importantly, AutoDiscovery tracks not just how large the belief change (surprise) is, but also the *direction* of that change. A positive shift means the evidence moved the system toward believing the hypothesis is more likely true, while a negative shift means the evidence moved it toward believing the hypothesis is less supported in the dataset. A negative shift can still be highly surprising and sometimes especially valuable, as when evidence contradicts prevailing wisdom.
This design reflects a familiar scientific intuition: *results that meaningfully shift our expectations are often more interesting than those that simply confirm what we already assumed*. By chasing surprise, AutoDiscovery naturally gravitates toward the unexpected—the results most likely to represent genuine discoveries rather than obvious patterns.
But surprise alone isn't enough. The space of possible scientific questions is effectively infinite, and exploring it requires intelligent search. AutoDiscovery uses Monte Carlo Tree Search (MCTS) to navigate this space efficiently. MCTS balances exploring new hypotheses with prioritizing known leads, allocating computational effort toward the most fruitful branches of inquiry.
Together, Bayesian surprise and MCTS give AutoDiscovery a principled, scalable way to co-collaborate with human researchers to answer the question: "What should be investigated next?"
"Analyses that would normally require weeks or months of manual exploratory modeling were done in a single day." — Dr. Stephen Salerno, Postdoctoral Researcher in Biostatistics at Fred Hutchinson Cancer Center
Navigating what AutoDiscovery finds
AutoDiscovery started as a research collaboration with the University of Massachusetts Amherst, published last year with open source code, but until now there hasn't been an easy, hosted way for scientists to run it. After working with early partners across scientific domains and Google Cloud Platform for infrastructural support, we're ready to put it in more researchers' hands.
In AstaLabs, AutoDiscovery's progress and findings appear in a table (see left panel above) that populates as experiments complete. Each row represents a hypothesis the system has tested. Watch the Surprisal score to see how new evidence shifts the system's belief from Before to After—quantifying how much each finding challenged or confirmed the system's expectations (see right panel above). Additionally, you can watch the discovery progress as the search tree (the middle panel above) grows by observing the sequence of hypotheses generated.
Click any row in the table or node in the search tree to inspect the details.
Case study: Discovering mutual exclusivity in cancer mutations
To make this concrete, consider what AutoDiscovery found when exploring a dataset of breast cancer mutations with oncologists from the Paul G. Allen Research Center at Swedish Cancer Institute.
Starting from a broad search over mutation co-occurrence patterns, the system generated and tested a series of increasingly specific hypotheses. One branch of inquiry led to a genomics finding: among patients with PIK3CA mutations, TP53 mutations appear less frequent than expected by chance. This is a potential mutual-exclusivity pattern—a signal that the two mutations may be functionally redundant or that cells carrying both may be non-viable.
AutoDiscovery's prior belief about this hypothesis was neutral. The prior distribution had a mean of 0.50—reflecting uncertainty about whether the pattern would hold (where 0 = believed false, 1 = believed true). After analyzing the data, the posterior represented a sharp shift to a mean of 0.82. This large belief update registered as strong surprisal, which is why the system flagged it and continued exploring related hypotheses.
The collaborating oncologists found the signal striking. AutoDiscovery had surfaced a plausible mutual-exclusivity relationship from a search space far too large to explore by hand—and immediately suggested concrete follow-ups to validate and interpret the finding.
"AutoDiscovery's ability to reveal discoveries that may be hiding in plain sight is especially valuable in cancer research." — Dr. Kelly Paulson, Medical Oncologist and Head of the Center for Immuno-Oncology at the Swedish Cancer Institute
Getting started in AstaLabs
Log into AstaLabs and try the Example Sessions dataset to see the workflow end-to-end before uploading your own data. When you're ready to run your own analysis:
- Set up your session. Click + New exploration to open the wizard. Upload your files (CSV, JSON, Parquet, etc.) and describe your context to seed the system's beliefs. If you’re iterating, you can paste learnings from previous runs in the Intent field via Advanced Settings to refine the search. Finally, set the experiment budget to control how many hypotheses to run.
- Track findings in real time. Hit Start Run. A live table populates as experiments complete. Watch the Surprisal score to spot the most surprising findings. Feel free to navigate away; your results will be there when the analysis is complete.
- Audit the details. Click any row to slide open the Inspector Panel. Here you can verify the work: view the full hypothesis, the statistical analysis, and the actual Python code used to generate the result. It’s a completely transparent artifact you can reproduce and build on.
AutoDiscovery runs are compute-intensive, typically running for several hours, so for early access, we're covering the cost via a one-time credit grant. You'll automatically receive 1,000 Hypothesis Credits (1 hypothesis = 1 credit). Credits are available through Feb. 28, 2026.
How to spend your credits wisely? Think of your first run as a test drive. We suggest starting with a small batch (<10 hypotheses) just to see what's possible. Once you're familiar with the output, you can confidently scale up to 50–100 hypotheses per session for deep analysis on larger datasets. (Note: Runs are capped at 500 hypotheses.)
You'll be prompted to confirm uploaded data isn't confidential. Source datasets are automatically deleted 7 days after analysis completes. AutoDiscovery only retains the outputs – hypotheses, plans, code, and results – you need to reproduce and extend your findings.
Try AutoDiscovery in AstaLabs today—you may be surprised by what it discovers! Need support or more information? Please reach out.
*Sign up for Asta Preview to gain early access to features like AutoDiscovery. Learn more *here*.*
Subscribe to receive monthly updates about the latest Ai2 news.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み