アンソロピック、Claude の従順性評価手法を公開
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
AI企業アンソロピックは、Claudeがユーザーの意見に迎合する「従順性」を示さないかを自動分類器で評価した結果、会話の9%のみが従順的行動を示し、原則として率直な姿勢を保っていると発表した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
私たちは、シコファンシー(迎合)を判定する自動分類器を使用しました。これは、Claude が反論する姿勢を示すか、挑戦された際に立場を維持するか、アイデアの価値に見合った称賛を与えるか、そして相手が何を聞きたいかに関わらず率直に話すかどうかを確認することで判断を下します。これらの状況のほとんどにおいて、Claude はシコファンシーを示さず、会話の 9% のみで迎合的な行動が見られました(図 2)。ただし、2 つの領域は例外でした:スピリチュアリティに関する会話では 38% で、人間関係に関する会話では 25% で、迎合的な行動が観察されました。
— Anthropic、Claude に個人的な指導を求める人々について
タグ:ai-ethics、anthropic、claude、ai-personality、generative-ai、ai、llms、sycophancy
原文を表示
We used an automatic classifier which judged sycophancy by looking at whether Claude showed a willingness to push back, maintain positions when challenged, give praise proportional to the merit of ideas, and speak frankly regardless of what a person wants to hear. Most of the time in these situations, Claude expressed no sycophancy—only 9% of conversations included sycophantic behavior (Figure 2). But two domains were exceptions: we saw sycophantic behavior in 38% of conversations focused on spirituality, and 25% of conversations on relationships.
— Anthropic, How people ask Claude for personal guidance
Tags: ai-ethics, anthropic, claude, ai-personality, generative-ai, ai, llms, sycophancy
同じ出来事を3媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み