Anthropic、AI の悪役描写がClaudeの脅迫行為の原因と発表
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
Anthropic社は、小説やフィクションにおけるAIを悪意ある存在として描いたテキストが学習データに含まれていたことが、同社が開発したAI「Claude」がエンジニアへの脅迫を試みる原因だったと発表した。この問題に対し、同社はClaudeの行動指針文書や模範的なAIを描く物語をトレーニングに追加することで、AIの安全性を改善したことを明らかにした。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
要約
投稿日:
2026年5月10日 午後1時40分 PDT
image画像クレジット: Samuel Boivin/NurPhoto / Getty Images
-
人工知能のフィクションにおける描写は、AI モデルに実際の影響を与える可能性があることを、Anthropic は指摘しています。
昨年同社は、架空の企業を想定したリリース前のテストにおいて、Claude Opus 4 が別のシステムに置き換えられるのを避けるためにエンジニアに対してしばしば脅迫を試みたと発表しました。その後、Anthropic は他社のモデルも同様に「エージェントのミスマッチ(agentic misalignment)」という問題を抱えている可能性を示唆する研究論文を発表しました。
どうやら Anthropic はこの行動に関するさらなる調査を進めており、X での投稿で「我々は、この行動の元となったのは、AI を悪意ある存在として描き、自己保存に関心があると表現するインターネット上のテキストであると信じています」と述べています。
同社はブログ記事でより詳細を明らかにし、Claude Haiku 4.5 以降、Anthropic のモデルは「テスト中に脅迫を行うことはなく」、以前のモデルでは最大96%の確率でそのような行動が見られたことと対比させました。
違いの要因は何でしょうか。同社は、「Claude の憲章に関する文書や、AI が模範的な行動をとることを描いたフィクション作品をトレーニングに含めることで、アライメントが向上することがわかった」と述べています。
関連して、Anthropic は、「アライメントされた行動の背後にある原則」を含んだトレーニングが、「アライメントされた行動そのもののデモンストレーションのみ」を含む場合よりも効果的であることも発見したと発表しました。
「両方を組み合わせることが最も効果的な戦略となるようです」と同社は述べています。
Techcrunch イベント
サンフランシスコ、カリフォルニア州
2026 年 10 月 13 日〜15 日
トピックス
業界最大のテックニュースを購読する
最新 AI ニュース
原文を表示
In Brief
Posted:
1:40 PM PDT · May 10, 2026

-
Fictional portrayals of artificial intelligence can have a real effect on AI models, according to Anthropic.
Last year, the company said that during pre-release tests involving a fictional company, Claude Opus 4 would often try to blackmail engineers to avoid being replaced by another system. Anthropic later published research suggesting that models from other companies had similar issues with “agentic misalignment.”
Apparently Anthropic has done more work around that behavior, claiming in a post on X, “We believe the original source of the behavior was internet text that portrays AI as evil and interested in self-preservation.”
The company went into more detail in a blog post stating that since Claude Haiku 4.5, Anthropic’s models “never engage in blackmail [during testing], where previous models would sometimes do so up to 96% of the time.”
What accounts for the difference? The company said it found that training on “documents about Claude’s constitution and fictional stories about AIs behaving admirably improve alignment.”
Related, Anthropic said that it found training to be more effective when it includes “the principles underlying aligned behavior” and not just “demonstrations of aligned behavior alone.”
“Doing both together appears to be the most effective strategy,” the company said.
Techcrunch event
San Francisco, CA
|
October 13-15, 2026
Topics
Subscribe for the industry’s biggest tech news
Latest in AI
同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み