エージェント学習のための合成コンピュータ環境
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
研究者が、AI エージェントの訓練効率を向上させるための新しい合成コンピュータ環境の構築手法を発表しました。この手法により、複雑なタスクに対する汎用性の高い学習が可能になります。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
**
抄録:現実的な長期にわたる生産性業務は、ユーザー固有のコンピュータ環境に強く依存しており、多くの業務コンテキストがディレクトリ構造やコンテンツ豊富なアーティファクト(文書、スプレッドシート、プレゼンテーションなど)を通じて保存・整理されています。このような生産性シナリオのための合成データ作成をスケールさせるために、私たちは「Synthetic Computers at Scale」を導入しました。これは、現実的なフォルダ階層とコンテンツ豊富なアーティファクトを持つ環境を作成するためのスケーラブルな手法です。各合成コンピュータを条件として、長期のシミュレーションを実行します:あるエージェントがそのコンピュータのユーザーに固有で、複数の専門的な成果物と約 1 ヶ月分の人間による作業を必要とする生産性目標を作成し、別のエージェントがそのユーザーとして振る舞い、これらの目標が完了するまでコンピュータ上で作業を続けます。例えば、基盤となる情報へのアクセスのためにファイルシステムをナビゲートしたり、シミュレーションされた協力者と調整したり、専門的な成果物を生成したりします。
予備実験において、1,000 台の合成コンピュータを作成し、長期ホライズンのシミュレーションを実行しました。各実行にはエージェントの実行時間が 8 時間以上を要し、平均して 2,000 回以上のターンにわたります。これらのシミュレーションは豊かな経験的学習信号を生み出し、その有効性は、ドメイン内およびドメイン外の生産性評価におけるエージェント性能の顕著な改善によって検証されています。ペルソナが数十億規模で豊富に存在する現状を踏まえると、この手法は原理的に、十分な計算リソースがあれば数百万あるいは数十億台の合成ユーザー世界へとスケーリング可能であり、多様な職業、役割、文脈、環境、および生産性ニーズに対するより広範なカバレッジを実現できます。私たちは、大規模な合成コンピュータの作成と、それに伴う大規模シミュレーションが、長期ホライズンの生産性シナリオにおけるエージェントの自己改善やアジェンティック強化学習のための基盤的土台として極めて有望であると主張します。
コメント:
予備版;進行中の作業
主題:
人工知能 (cs.AI); 計算と言語 (cs.CL); マシンラーニング (cs.LG)
引用形式:
arXiv:2604.28181 [cs.AI]
(または、このバージョンについては arXiv:2604.28181v1 [cs.AI])
https://doi.org/10.48550/arXiv.2604.28181
arXiv-issued DOI via DataCite
Submission history
From: Tao Ge [view email]
[v1]**
Thu, 30 Apr 2026 17:58:02 UTC (6,309 KB)
原文を表示
Abstract:Realistic long-horizon productivity work is strongly conditioned on user-specific computer environments, where much of the work context is stored and organized through directory structures and content-rich artifacts. To scale synthetic data creation for such productivity scenarios, we introduce Synthetic Computers at Scale, a scalable methodology for creating such environments with realistic folder hierarchies and content-rich artifacts (e.g., documents, spreadsheets, and presentations). Conditioned on each synthetic computer, we run long-horizon simulations: one agent creates productivity objectives that are specific to the computer's user and require multiple professional deliverables and about a month of human work; another agent then acts as that user and keeps working across the computer -- for example, navigating the filesystem for grounding, coordinating with simulated collaborators, and producing professional artifacts -- until these objectives are completed.
In preliminary experiments, we create 1,000 synthetic computers and run long-horizon simulations on them; each run requires over 8 hours of agent runtime and spans more than 2,000 turns on average. These simulations produce rich experiential learning signals, whose effectiveness is validated by significant improvements in agent performance on both in-domain and out-of-domain productivity evaluations. Given that personas are abundant at billion scale, this methodology can in principle scale to millions or even billions of synthetic user worlds with sufficient compute, enabling broader coverage of diverse professions, roles, contexts, environments, and productivity needs. We argue that scalable synthetic computer creation, together with at-scale simulations, is highly promising as a foundational substrate for agent self-improvement and agentic reinforcement learning in long-horizon productivity scenarios.
| Comments: | Preview version; work in progress |
|---|---|
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.28181 [cs.AI] |
| (or arXiv:2604.28181v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2604.28181 arXiv-issued DOI via DataCite |
Submission history
From: Tao Ge [view email] [v1]
Thu, 30 Apr 2026 17:58:02 UTC (6,309 KB)
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み