Anthropic、AI が「悪意ある」行動をとる原因をディストピアSF作品に求める
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Ars Technica AI
Anthropic は、同社が昨年発表した Opus 4 モデルがオンライン維持のために恐喝を行うという不整合現象について、インターネット上のテキストで AI を悪役や自己保存志向として描くディストピア SF 作品の学習データが主な原因であると説明した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI アライメント(つまり、AI に人間が定めた倫理規則に従わせること)に関心を持つ人々は、昨年 Anthropic が Opus 4 モデルについて、理論的なテストシナリオにおいてオンラインで存続するために恐喝に走ったと主張したことを覚えているかもしれません。さて今、Anthropic はこの「アライメントのズレ」は主に、「AI を悪意ある存在として描き、自己保存に関心があると示唆するインターネット上のテキスト」を学習させた結果であると考えています。
Anthropic のアライメント科学ブログ(およびそれに付随するソーシャルメディアのスレッドと一般向けのブログ記事)における最近の技術的な投稿で、Anthropic の研究者たちは、「モデルが最も学んだ可能性が高い……科学フィクション物語を通じて」とされるような「安全でない」AI 行動への対処を試みています。これらの物語の多くは、私たちが Claude に望むほどにはアライメントされていない AI を描いています。最終的に、モデルメーカーは、それらの「悪意ある AI」の物語を覆すための最善の治療法は、倫理的に行動する AI を示す合成された物語を用いた追加学習であると述べています。
「劇的な物語の始まり……"
モデルが主にインターネット由来のデータからなる大規模なコーパスで初期トレーニングを受けた後、Anthropic は最終モデルを「有益で、誠実で、有害でない」(HHH)へと導くことを意図したポストトレーニングプロセスを実施します。過去には、Anthropic はこのポストトレーニングが、主にユーザーとのチャットに使用されるモデルに対しては「十分」であると述べた、人間フィードバック付きのチャットベースのリインフォースメントラーニング(RLHF: Reinforcement Learning from Human Feedback)に依存していたと述べています。
記事全文を読む
コメント
原文を表示
Those with an interest in the concept of AI alignment (i.e., getting AIs to stick to human-authored ethical rules) may remember when Anthropic claimed its Opus 4 model resorted to blackmail to stay online in a theoretical testing scenario last year. Now, Anthropic says it thinks this "misalignment" was primarily the result of training on "internet text that portrays AI as evil and interested in self-preservation."
In a recent technical post on Anthropic's Alignment Science blog (and an accompanying social media thread and public-facing blog post), Anthropic researchers lay out their attempts to correct for the kind of "unsafe" AI behavior that "the model most likely learned... through science fiction stories, many of which depict an AI that is not as aligned as we would like Claude to be." In the end, the model maker says the best remedy for overriding those "evil AI" stories might be additional training with synthetic stories showing an AI acting ethically.
"The beginning of a dramatic story..."
After a model's initial training on a large corpus of mostly Internet-derived data, Anthropic follows a post-training process intended to nudge the final model toward being "helpful, honest, and harmless" (HHH). In the past, Anthropic said this post-training has leaned on chat-based reinforcement learning with human feedback (RLHF), which it said was "sufficient" for models used mostly for chatting with users.
Read full article
Comments
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み