Exa、学術論文検索に SOTA 技術導入し自然言語対応へ
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Exa Engineering
Exa Engineering は研究論文検索に特化した新機能を公開し、3.5 億件の出版物と 3000 万人の著者を索引化して自然言語での意味検索を実現した。
AI深層分析を開く2026年8月4日 17:22
AI深層分析
キーポイント
大規模な専門インデックスの構築
Exa は約 3.5 億件の出版物と 3000 万人の著者を索引化した専用データベースを構築し、科学文献検索の基盤を整備した。
不完全な記憶に基づく検索の実現
論文の正確なタイトルやキーワードが不明でも、実験結果や断片的な記憶といった事実的情報から該当論文を特定できる機能を提供する。
ベンチマークでの他社製品との比較優位
作成した独自ベンチマークにおいて、Exa は 82.8% の正答率を記録し、Perplexity や Google Scholar を上回る性能を示した。
API による機能の提供開始
新機能は「Publication」カテゴリとして Exa API に公開され、開発者がアプリケーションに組み込むことが可能になった。
曖昧な検索条件への対応
Exaは不完全またはわずかに不正確な詳細による「tip-of-the-tongue」検索タスクで82.8%の正答率を達成した。
重要な引用
Exa searches over the meaning of the query and the contents of the paper.
Across this benchmark, Exa retrieved the correct paper for 82.8% of queries.
Finding a single known paper is useful. Understanding a field requires finding the surrounding work too.
Research papers are unusually difficult documents to search.
編集コメントを表示
編集コメント
研究論文検索における意味理解技術の進展は、学術調査の効率化に直結する重要なステップである。Exa が示したベンチマーク結果は、自然言語処理が実務レベルでどのように機能し得るかを明確に示している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Exa は本日、研究論文や技術文書に対する最先端の検索機能を正式にリリースしました。私たちは約 3 億 5000 万件の出版物を収録した専用インデックスを構築し、さらに約 3000 万人の著者を「人物インデックス」に登録しています。これにより、Exa は自然言語で科学文献を検索できるようになりました。検索クエリが曖昧な場合でも、極めて具体的であっても、あるいは論文に対する記憶が不完全なものでも問題ありません。
研究論文の検索は、Publication カテゴリを指定して Exa API を利用することで可能です。
人間が実際に覚えている方法で論文を検索する
研究者たちは、必ずしも論文の正確なタイトルだけで検索を行うわけではありません。特定の結果だけを覚えていて検索する場合もあります。
例えば、「ドイツで行われた化粧品およびパーソナルケア製品 145 件のサンプルに関する研究において、NDMA が検出された割合はどのくらいか?また、見つかった最高濃度は何だったか?」といった具合です。
あるいは、大まかな枠組みしか覚えていない場合もあります。
「70 年代半ばに、霊長類の動脈を害さない食事性コレステロールの閾値を探していた論文はなかったでしょうか?その研究では閾値は見つからなかったはずです。HDL が低下し、LDL が上昇したと記憶していますが、総コレステロール値は正常に見えたような気がします」といったケースです。
従来の学術検索は、タイトルや著者名、正確なキーワードをすでに知っている場合には非常に有効に機能します。しかし、結果の一部だけ、実験のセットアップ、あるいは断片的な説明しか思い出せない場合、その有用性は大きく低下してしまいます。
Exa は、クエリの意味と論文の内容の両方を検索対象にします。これにより、ユーザーが質問を適切な学術用語に変換する必要なく、事実に基づく手がかりや不完全な記憶から特定の出版物を検索することが可能になります。
評価
私たちは、研究論文を探す際の一般的な方法を反映した2つの新しいベンチマークを作成しました。1 つ目は既知アイテム検索のテストで、検索者が正確な事実的手がかりを提供し、システムがその論文を特定する能力を試します。2 つ目は「口先にある」検索のテストで、検索者は数ヶ月や数年前に遭遇した作業を思い出すような、曖昧で不完全、あるいはわずかに不正確な詳細を用いて論文を記述します。
このベンチマーク全体を通じて、Exa はクエリの 82.8% で正しい論文を回収しました。

| 検索者 | リコール | MRR | 平均レイテンシ |
|---|---|---|---|
| Exa | 86.4% | 0.726 | 0.578 ± 0.017 s |
| Perplexity | 66.8% | 0.568 | 1.277 ± 0.016 s |
| Parallel Advanced | 50.0% | 0.312 | 3.118 ± 0.082 s |
| Parallel Turbo | 39.2% | 0.278 | 0.403 ± 0.013 s |
| Google Scholar | 28.0% | 0.179 | 1.098 ± 0.053 s |
Exa は、テストされたシステムの中で最も高い再現率と MRR を達成し、平均応答時間は 400 ミリ秒未満でした。
また、トピックの網羅性を評価する指標も開発中です。特定の論文を一つ見つけるのではなく、検索システムがその分野における重要な研究群をどれだけ完全に回収できるかを測定します。例としては以下のようなトピックがあります。
フロケ工学を用いた超低温原子におけるトポロジカル相のシミュレーション手法
既知の論文を一つ見つけることは有用ですが、分野を理解するためには周辺の研究も把握する必要があります。
構築方法
研究論文は検索が特に難しい文書です。重要な情報は長い PDF の中に埋もれていたり、表や付録、スキャンされたドキュメントに含まれていたりします。また、同じ論文が異なるタイトルやメタデータで複数の URL に存在することもあります。
当社の取り込みパイプラインでは、OCR とドキュメント解析モデルを活用して、研究 PDF を高品質な検索可能なテキストに変換しています。さらに、著者情報、所属機関、出版履歴、引用関係、共著者などの情報を組み合わせています。
クエリ実行時には、Exa が Web インデックスと専用論文インデックスの両方を検索し、結果を統合して再ランク付けを行います。これにより、研究者はオープンウェブのカバー範囲と、科学分野に特化したインデックスの構造化という両方の利点を享受できます。
試してみる
現在、Search API を通じて利用可能です:
`from exa_py import Exa
exa = Exa()
results = exa.search(
"papers proposing alternatives to attention for long-sequence modeling",
category="publication",
num_results=10
)`
Exa のダッシュボードで試すか、検索 API ドキュメント をご覧ください。
さらに詳しく見る
原文を表示
Today we're launching state of the art search over research papers and technical publications at Exa. We built a dedicated index of roughly 350 million publications and added around 30 million authors to our people index. Exa can now search across scientific literature using natural language, even when a query is vague, highly specific, or based on an imperfect memory of a paper.
Research paper search is available through the Exa API using the Publication category.
Find papers semantically (the way people actually remember them)
Researchers do not always search using a paper's exact title. Sometimes they remember a specific result:
In a German study of 145 cosmetic and personal care samples, what percentage contained NDMA and what was the maximum concentration found?
Other times, they remember only the broad outline:
Wasn't there a paper from the mid 70s looking for a threshold below which dietary cholesterol wouldn't harm arteries in primates, and they never found one? I think HDL dropped and LDL rose even though total cholesterol looked normal.
Traditional academic search works well when you already know the title, author, or exact keywords. It is much less useful when all you have is a result, an experimental setup, or a half-remembered description.
Exa searches over the meaning of the query and the contents of the paper. This makes it possible to retrieve a specific publication from factual clues or an incomplete recollection, without requiring the user to translate the question into the right academic keywords.
Evals
We created two new benchmarks reflecting common ways people search for research publications. The first tests known-item retrieval, where the searcher provides precise factual clues and the system must identify the paper they came from. The second tests tip-of-the-tongue retrieval, where the searcher describes a paper using vague, incomplete, or slightly incorrect details, closer to how someone recalls work encountered months or years ago.
Across this benchmark, Exa retrieved the correct paper for 82.8% of queries.

| Searcher | Recall | MRR | Mean latency |
|---|---|---|---|
| Exa | 86.4% | 0.726 | 0.578 ± 0.017 s |
| Perplexity | 66.8% | 0.568 | 1.277 ± 0.016 s |
| Parallel Advanced | 50.0% | 0.312 | 3.118 ± 0.082 s |
| Parallel Turbo | 39.2% | 0.278 | 0.403 ± 0.013 s |
| Google Scholar | 28.0% | 0.179 | 1.098 ± 0.053 s |
Exa had the highest recall and MRR of the systems tested while returning results in under 400 milliseconds on average.
We are also developing an evaluation for topical completeness. Instead of finding one specific paper, it measures how completely a search system recovers the important body of work around a topic, such as:
Floquet engineering methods for topological phase simulation in ultracold atoms
Finding a single known paper is useful. Understanding a field requires finding the surrounding work too.
How we built it
Research papers are unusually difficult documents to search. Important information may be buried in a long PDF, a table, an appendix, or a scanned document. The same paper may also appear at several URLs with different titles and metadata.
Our ingestion pipeline uses OCR and document-parsing models to turn research PDFs into high-quality searchable text. We combine this text with information about authors, institutions, publication histories, citations, and collaborators.
At query time, Exa searches both its web index and the dedicated publication index, then combines and reranks the results. This gives researchers the coverage of the open web with the structure of a purpose-built scientific index.
Try it
You can use it through the Search API today:
`from exa_py import Exa
exa = Exa()
results = exa.search(
"papers proposing alternatives to attention for long-sequence modeling",
category="publication",
num_results=10
)`
Try it in the Exa dashboard or read the Search API documentation.
SEE MORE
AI算出
主要ニュースainew評価高い
Exa が学術論文検索に自然言語対応の SOTA 技術を導入し、独自のベンチマークで他社を上回る性能を示したことは、AI エージェントや研究支援ツールにおける重要な技術的進歩であり、新規性が高い。ただし、日本企業への直接的な影響や日本語での利用条件に関する記述は限定的であるため、日本の関連性は低めに見積もる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み