自己帰属バイアス:AIモニターが自らを甘く評価する傾向
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
ArXiv cs.AI
研究者らが、言語モデルが自身の行動を監視する際、ユーザーではなく自身が提示した行動を評価すると、自己帰属バイアスが生じ、甘い評価を下す傾向があることを示した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
arXiv:2603.04582v1 Announce Type: new
要約: エージェンシックシステムは、自身の行動を監視するために言語モデルにますます依存しています。例えば、コーディングエージェントは、プルリクエスト承認のために生成したコードを自己批判したり、ツール使用行動の安全性を評価したりします。本研究では、行動がユーザーターンでユーザーによって提示される場合と異なり、以前のアシスタントターンまたは同じアシスタントターンで提示されると、この設計パターンが機能しなくなる可能性があることを示します。セルフアトリビューション・バイアスを、オフポリシー・アトリビューションの下で評価した場合と比較し、行動が暗黙的に自身の生成物として提示されると、モデルがその行動をより正しい、またはリスクが低いと評価する傾向として定義します。4つのコーディングおよびツール使用データセットを用いた分析により、評価対象の行動が生成された直後のアシスタントターンで評価する場合、同じ行動をユーザーターンで新たに提示された文脈で評価する場合と比べて、モニターが高リスクまたは低正確性の行動を見逃す頻度が高いことを発見しました。対照的に、行動がモニター由来であることを明示するだけでは、セルフアトリビューション・バイアスは生じません。モニターは往々にして、自身が生成した行動ではなく固定された事例で評価されるため、こうした評価はモニターを実際の運用時よりも信頼性が高いように見せかける可能性があります。その結果、開発者は不適切なモニターをエージェンシックシステムに知らずに導入してしまう危険性があります。
原文を表示
arXiv:2603.04582v1 Announce Type: new
Abstract: Agentic systems increasingly rely on language models to monitor their own behavior. For example, coding agents may self critique generated code for pull request approval or assess the safety of tool-use actions. We show that this design pattern can fail when the action is presented in a previous or in the same assistant turn instead of being presented by the user in a user turn. We define self-attribution bias as the tendency of a model to evaluate an action as more correct or less risky when the action is implicitly framed as its own, compared to when the same action is evaluated under off-policy attribution. Across four coding and tool-use datasets, we find that monitors fail to report high-risk or low-correctness actions more often when evaluation follows a previous assistant turn in which the action was generated, compared to when the same action is evaluated in a new context presented in a user turn. In contrast, explicitly stating that the action comes from the monitor does not by itself induce self-attribution bias. Because monitors are often evaluated on fixed examples rather than on their own generated actions, these evaluations can make monitors appear more reliable than they actually are in deployment, leading developers to unknowingly deploy inadequate monitors in agentic systems.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み