Microsoft Research、推論中に自己経験から学習する「EvoLib」を発表
Microsoft Research は、推論時に自己経験から学習する「EvoLib」というフレームワークを発表し、モデルの更新なしにブラックボックスLLMが過去の実績を蓄積・洗練された知識として活用できる仕組みを提示した。
AI深層分析を開く2026年7月31日 01:21
AI深層分析
キーポイント
推論時の自己学習機能
EvoLib は、正解ラベルや外部フィードバックを必要とせず、モデルが推論中に自身の経験から直接学習する仕組みを実現する。
経験の知識への変換
単なる過去の記録の蓄積ではなく、成功した解決策からのスキルや失敗からの洞察を抽出し、再利用可能な形式へと変換する。
モデル更新不要な適用性
EvoLib はモデル自体の再学習や更新を必要としないため、API経由で展開された既存のブラックボックス言語モデルに即座に適用可能である。
知識の進化と洗練
新しい経験が流入するたびにスキルは一般化され、洞察は精度が高まり、重み付けが再調整されることで、時間とともに性能が向上する。
知識の進化と統合
EvoLib は単なる記憶の蓄積ではなく、新しい経験に基づいて既存の知識を継続的に統合・再評価する仕組みを持つ。これにより、個々の事例を超えた汎用的なスキルや洞察へと知識を進化させる。
重要な引用
EvoLib enables large language models to learn from their own experience during inference, without requiring ground-truth labels or external feedback.
Rather than treating memory as a growing archive of past experiences, EvoLib extracts reusable knowledge from those experiences and continually refines it as new experiences arrive.
By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.
Instead, the greatest gains come from transforming experience into reusable knowledge that can be continually refined and applied across tasks.
編集コメントを表示
編集コメント
モデルの再学習を伴わない推論時学習の実現は、リソース制約のある環境や迅速な適応が求められる現場において大きな意義を持つ。既存のAPIベースのAIシステムにこのフレームワークを組み込むことで、静的なモデルから動的に進化するエージェントへの転換が現実味を帯びてくる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

一目でわかるポイント
自己教師あり学習。EvoLib は、正解ラベルや外部フィードバックを必要とせず、推論中に大規模言語モデルが自身の経験から学べるようにします。
経験から知識へ。過去の試行錯誤を、将来のタスクに活用できる再利用可能なスキルや省察的な洞察へと変換します。
進化する知識。有用なスキルや洞察は継続的に洗練・統合・重み付けされ、特定の事象に関する観察が時間の経過とともにより一般的な知識へと進化していきます。
タスク横断的な学習転移。経験を活用可能な知識に変えることで、EvoLib は AI モデルに過去の成功と失敗から学ばせ、将来のパフォーマンス向上に最も高い可能性を持つ知識を進化させます。
現在の AI モデル向けに設計。モデルの更新を必要としないため、API を通じて展開されたブラックボックス型の言語モデルや AI システムにも適用可能です。
記憶は重要な AI エージェント機能となりました:過去の経験を保存し、検索する能力です。しかし、記憶があるだけでは学習にはなりません。過去の会話、推論の痕跡、行動履歴を集めただけでは、膨大な経験アーカイブが形成されるだけで、新しいタスクに最も関連性の高い知識を特定することは難しくなります。ましてや、時間をかけてパフォーマンスを向上させるためにこの知識を洗練・進化させることはなおさら困難です。
人間は異なる方法で学びます。過去の経験のすべての詳細を記憶するわけではありません。重要なのは、機能する戦略や避けるべきミス、そして状況を超えて応用できるスキルです。時間が経つにつれて、これらの教訓はより一般的で再利用可能な知識へと洗練されていきます。このように経験を転送可能で進化する知識へ変換する能力こそが、人間学習の基盤の一つです。
最近発表した論文「Test-Time Learning with an Evolving Library」では、AI システムがいかにして人間と同様に経験から学べるかを探究しました。そこで提案したのが EvoLib です。これは生きた経験を、進化し続ける知識のライブラリへと変換するフレームワークです。EvoLib は記憶を単なる過去の蓄積として扱うのではなく、そこから再利用可能な知識を抽出し、新たな経験が入ってくるたびにそれを継続的に洗練させます。ライブラリの進化を通じて、スキルはより汎用性を帯び、洞察は精度を増し、下流タスクのパフォーマンスは時間とともに一貫して向上します。このようにして、AI エージェントは基盤モデルを更新することなく、蓄積される経験から絶えず学習することが可能になります。
How EvoLib Works
従来の AI メモリシステムが生の経験を静的な情報として保存するのに対し、EvoLib は「知識が進化する」という考え方を基盤に構築されています。EvoLib における知識の単位は、成功した解決策から抽出された再利用可能なスキルや、失敗から得られた反省的な洞察といった形をとります。単に時間をかけて記憶を蓄積していくのではなく、新しい経験が訪れるたびに既存の知識を継続的に精緻化し、統合し、重み付けを更新します。具体的には、知識の進化のために以下のメカニズムを設計しました。
統合(Consolidation):最近の経験から新たな知識が抽出されると、EvoLib はライブラリ内から類似した知識を検索し、それを新しい知識と統合して、より汎用的で再利用可能な形へと高めます。これにより、知識は個々の経験に限定されず、さまざまなタスクに応用可能になります。
重み付けメカニズム:EvoLib は、各知識単位の重要性を、現在のタスクにおける即時的な有用性だけでなく、将来のタスクで有益な知識を生み出すための貢献度に基づいて継続的に更新します。時間の経過とともに、長期的な影響力が最も大きい知識が自然とライブラリの中で目立つようになります。
image図 1. EvoLib は生の経験を再利用可能なスキルや洞察に変換し、統合と動的な重み付けを通じて継続的に進化させます。
主要な結果
EvoLib の評価では、異なる種類の経験と学習要件を持つ多様な難易度のタスクでテストを行いました。
数学的推論問題の解決
効率制約の下で与えられたタスクを実行するコードの記述
環境を探索・操作して長期にわたるタスクを遂行するための意思決定
これらのタスク全体において、EvoLib はトークン使用量をより効率的に抑えながら、トップクラスの検索ベースの記憶手法や他の抽象的な記憶機構を上回る性能を発揮しました。
また、継続的に進化する知識を通じて、テスト時の計算リソースをいかに効果的にパフォーマンス向上に変換できるかも評価しました。図 2 は、各タスクを孤立して処理する計算スケーリング手法と、強力な記憶ベースの学習アプローチに対する EvoLib の比較を示しています。各曲線は、テスト時の計算量が増加するにつれてどのように性能が向上するかを表しています。
3 つのベンチマークすべてで、EvoLib は計算量の大部分の範囲において高い性能を達成し、計算量が増えるほどパフォーマンスが急速に改善されました。
これらの結果から、より良い学習のための鍵は、単に記憶を増やすことや計算リソースを投入することだけではないことが示唆されます。むしろ、最大の効果を得るには、経験を再利用可能な知識へと変換し、それを継続的に洗練させてさまざまなタスクに応用できる仕組みが重要なのです。
image図 2:あらゆるタスクにおいて、EvoLib は既存手法よりも効率的にテスト時の計算リソースを性能向上に変換します。
ランダムなタスク順序に対する頑健性
自然な疑問として、このような学習がタスクの遭遇順に強く依存するかどうかがあります。現実世界では、AI システムは多様な種類のタスクを任意の順序で直面することになり、有用な学習フレームワークはタスク順序のランダムさに対して頑健であるべきです。
これを評価するため、同じ異種タスクセットを用いながら、異なるタスク順序での性能を測定しました。その結果、EvoLib は既存のメモリベース学習手法を一貫して上回り、異なる順序付けにおいても安定したパフォーマンスを維持することがわかりました。これは、エージェントが構造化されたカリキュラムに依存せず、多様なユーザー要求の混合ストリームを処理・学習する必要がある現実的なシナリオにおいて、EvoLib が実用的な優位性を有していることを示唆しています。
AI システムがより長時間実行され複雑化するタスクを引き受けるにつれ、経験からの学習はますます重要になっていきます。AI の未来は、単にモデルの大型化や計算リソースの増強だけでなく、システムが知識を継続的に蓄積・洗練・再利用できるメカニズムにも依存するでしょう。
EvoLib は、そのビジョンに向けた一歩です。経験を進化していく知識へと変換することで、AI システムは展開後も継続的に改善し、適応できるようになります。毎回ゼロからやり直すのではなく、将来の AI システムは、人間のように再利用可能なスキルや洞察が蓄積されたライブラリを基盤に成長していくことが可能になるでしょう。
コードと実験結果は GitHub に公開されており(別タブで開く)、AI システムにおける記憶と知識の進化に関する今後の研究をサポートします。
本記事「EvoLib: Turning experience into evolving knowledge」は、Microsoft Research の投稿として初めて発表されました。
原文を表示

At a glance
Self-supervised. EvoLib enables large language models to learn from their own experience during inference, without requiring ground-truth labels or external feedback.
From experience to knowledge. EvoLib transforms past attempts into reusable skills and reflective insights that can be applied to future tasks.
Knowledge that evolves. Useful skills and insights are continually refined, consolidated, and reweighted, turning instance-specific observations into increasingly general knowledge over time.
Learning that transfers across tasks. By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.
Built for today’s AI models. As EvoLib does not require model updates, it can be applied to any black-box language models and AI systems deployed through APIs.
Memory has become an important AI agent capability: the ability to store and retrieve past experiences. But memory alone is not learning. A collection of past conversations, reasoning traces, or action histories can quickly grow into a vast archive of experiences, making it difficult to identify the most relevant knowledge for a new task—let alone refine and evolve this knowledge to improve performance over time.
Humans learn differently. We do not remember every detail of our past experiences. Instead, we remember what matters: strategies that work, mistakes to avoid, and skills that transfer across situations. Over time, these lessons are refined into increasingly general and reusable knowledge. This ability to transform experience into transferable, evolving knowledge is one of the foundations of human learning.
In our recent paper, Test-Time Learning with an Evolving Library, we explore how AI systems can learn from experience in a similar way. We introduce EvoLib, a framework that transforms raw experience into an evolving library of knowledge. Rather than treating memory as a growing archive of past experiences, EvoLib extracts reusable knowledge from those experiences and continually refines it as new experiences arrive. Through the evolution of library, skills become more general, insights become more accurate, and downstream performance gets improved consistently over time. In this way, AI agents can continually learn from accumulating experience without updating the underlying model.
How EvoLib Works
Unlike traditional AI memory systems that store raw experiences as static information, EvoLib is built around the idea of evolving knowledge. In EvoLib, a unit of knowledge can take the form of a reusable skill distilled from a successful solution or a reflective insight learned from mistakes. Rather than simply accumulating more memories over time, EvoLib continually refines, consolidates and reweights existing knowledge as new experiences arrive. Concretely, we design the following mechanisms around knowledge evolution:
Consolidation. As new knowledge is extracted from recent experience, EvoLib retrieves similar knowledge from the library and tries to consolidate it with the new knowledge into a more general and reusable one. This allows knowledge to move beyond individual experiences and become applicable across tasks.
Weighting mechanism. EvoLib continually updates the importance of each knowledge unit based not only on its immediate utility on the current task, but also on how much it contributes to generating useful knowledge on future tasks. Over time, knowledge with the greatest long-term impact naturally becomes more prominent in the library.
imageFigure 1. EvoLib transforms raw experiences into reusable skills and insights, then continually evolves them through consolidation and dynamic weighting.
Key Results
To evaluate EvoLib, we tested it across a diverse set of challenging tasks with different types of experiences and demands for learning:
Solving mathematical reasoning problems
Writing code to perform the given tasks under efficiency constraints
Making decisions to explore and interact with an environment to perform long-horizon tasks
Across these tasks, EvoLib consistently outperforms the top retrieval-based memory approaches and other abstract memory mechanisms with more efficient token usage.
We also evaluated how effectively EvoLib converts test-time compute into performance gains through continually evolving knowledge. Figure 2 compares EvoLib against both compute scaling methods that perform each task in isolation and strong memory-based learning approaches. Each curve shows how performance improves as the amount of test-time compute increases.
Across all three benchmarks, EvoLib achieves higher performance throughout most of the compute range and improves performance more rapidly with increasing compute.
These results suggest that the key to better learning may not simply be storing more memories or spending more compute. Instead, the greatest gains come from transforming experience into reusable knowledge that can be continually refined and applied across tasks.
imageFigure 2. Across all tasks, EvoLib converts test-time compute into performance gains more effectively than existing methods.
Robustness to random task order
A natural question is whether such learning depends heavily on the order in which tasks are encountered. In the real world, an AI system may face diverse types of tasks in arbitrary order, and a useful learning framework should be robust to the randomness in task order. To evaluate this, we measured the task performance on the same set of heterogeneous tasks but with different task orders. We found that EvoLib consistently improves over existing memory-based learning approaches and maintains stable performance across different orderings. This indicates that EvoLib can continually learn from diverse tasks even when they are interleaved, suggesting its practical advantage in real-world scenarios where an agent must handle and learn from a mixed stream of heterogeneous user requests without relying on a structured curriculum.
As AI systems take on longer-running and more complex tasks, learning from experience will become increasingly important. The future of AI may depend not only on larger models and more computation, but also on mechanisms that allow systems to continually accumulate, refine, and reuse knowledge.
EvoLib is one step toward that vision. By transforming experience into evolving knowledge, it enables AI systems to continually improve and adapt after deployment. Rather than repeatedly starting from scratch, future AI systems may be able to build upon an evolving library of reusable skills and insights, much like humans do.
Code and experiment results are available on GitHub (opens in new tab) to support future research on memory and knowledge evolution in AI systems.
Opens in a new tabThe post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.
AI算出
主要ニュースainew評価高い
AI モデルの学習パラダイムそのものを変える新しい技術(EvoLib)が発表されており、具体的なメカニズム(統合・重み付け)と評価結果が示されているため新規性は高い。ただし、日本企業や日本固有の適用事例に関する言及はないため、日本の関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み