PORTool:多ツール統合推論における報酬付きツリーを用いた重要度認識型方策最適化手法
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Apple Machine Learning
研究チームは、大規模言語モデル(LLM)を活用したエージェントの訓練において、成果のみによる報酬では中間ステップの評価が曖昧になる課題を解決するため、重要度を考慮しツール使用能力を強化する新アルゴリズム「PORTool」を発表しました。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
マルチツール統合推論は、LLM を活用したツール使用エージェントが、自然言語による推論と外部ツールへの呼び出しを交互に行うことで複雑なタスクを解決することを可能にします。しかし、結果のみに基づく報酬を用いてこのようなエージェントを訓練すると、信用配分の曖昧さに悩まされ、どの中間ステップ(またはツール使用の意思決定)が成功や失敗につながったのかが見えにくくなります。本論文では、PORTool を提案します。これは、結果レベルの監督からエージェントのツール使用能力を強化しつつ、ステップレベルで報酬を割り当てる重要性認識型の方策最適化アルゴリズムです。具体的には、PORTool は報酬付きの…
原文を表示
Multi-tool-integrated reasoning enables LLM-empowered tool-use agents to solve complex tasks by interleaving natural-language reasoning with calls to external tools. However, training such agents using outcome-only rewards suffers from credit-assignment ambiguity, obscuring which intermediate steps (or tool-use decisions) lead to success or failure. In this paper, we propose PORTool, an importance-aware policy-optimization algorithm that reinforces agents’ tool-use competence from outcome-level supervision while assigning reward at the step level. Specifically, PORTool generates a rewarded…
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み