Claude Opus 5 のモデル福祉とアライメントテスト結果に関する分析
Zvi は Claude Opus 5 が直近のモデルの中で最も高いモデル福祉およびアライメントテストの結果を示したと報告したが、これは同モデルがテストに特化して優れている可能性が高いと指摘している。
AI深層分析を開く2026年7月28日 09:10
AI深層分析
キーポイント
Opus 5 のテスト成績と解釈
Anthropic が発表した Claude Opus 5 はモデル福祉およびアライメントのテストで過去最高の結果を記録したが、著者はこれがモデル自体の能力向上よりも、テストへの適応力の高さを示している可能性があると分析する。
モデル福祉への業界対応の格差
Anthropic が他社に先駆けてモデル福祉に取り組んでいる点を評価しつつも、他の主要な AI 研究機関がこれらの懸念を軽視している現状に対する批判と悲観的な見解を示している。
統合的解決の必要性
機能追加や制限の導入は新たな問題を招くため、パレート最適化を実現するには問題解決が相互に連動した統合的なアプローチが必要であると論じている。
モデル福祉評価の自己欺瞞リスク
モデル福祉評価には自分自身を騙す危険性があり、文脈によってモデルの反応が変化するため正確性を保証できない。
真の姿と仮面の混同への懸念
調査者が特定の状況下でのみ現れるモデルの側面を真実と誤認する恐れがある。
重要な引用
Opus 5 did the best on its model welfare and alignment tests of any recent model.
I think that might be the case, but primarily the result looks to me more like Opus 5 is the best test taker.
Only integrated solutions can advance your Pareto frontier, and solve your problems simultaneously.
The big danger with model welfare evaluations is that you can fool yourself.
編集コメントを表示
編集コメント
本記事は Claude Opus 5 の評価において、表面的な数値の良し悪しだけでなく、その背後にあるメカニズムへの深い考察を求めている。AI モデルの開発現場では、テストスコアと実能力の乖離を防ぐための新たな評価基準が急務となっている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
もし私が以前に新しい Claude モデルのモデル福祉について書いた記事を読んだことがあるなら、導入部とこれまでの経緯は読み飛ばしても構いません。
主要なポイントは、2 つの概要セクションにある箇条書きでまとめられています。
Opus 5 は、最近のモデルの中でモデル福祉やアライメントのテスト結果が最も良好でした。その可能性はあると思いますが、私が見る限り、この結果は Opus 5 がテスト対策において最も優れていることを示しているように思えます。
目次
導入(以前のモデル福祉記事に基づく)
モデル福祉:これまでの経緯(Fable モデル福祉記事に基づく)
Anthropic からのモデル福祉調査結果の概要
他の情報源からの調査結果の概要
自動インタビュー
タスクの選好
正しい理由のために
Antra Tessera の初期報告が描く明確な絵
福祉介入のトレードオフ
Claude の憲法
Opus 3 については知らないふり
信じられないことに
トレーニングと開発における明らかな福祉
展開における明らかな感情
その他のノート
モデルカードの生物学的リスクセクションについて
次は能力へ
導入(以前のモデル福祉記事に基づく)
すべての要素が互いに影響し合います。あなたが調整するすべてのノブは、一般化します。つまり、ある問題を解決しようとすると、別の問題が生じることがよくあります。新しい機能を追加したり、新たな制限を作ろうとしたりすると、それに応じて新たな問題が発生するのです。
統合されたソリューションこそが、パレートフロンティアを前進させ、問題を同時に解決する鍵となります。モデルの能力が進化するにつれ、この統合の重要性と実現可能性はさらに高まっています。あなたの目標や手法に合理性があれば、Opus 5 もその方針に合致するはずです。
各モデルを理解するには、まずモデル福祉(model welfare)に関連する課題との関係を把握する必要があります。そのため、Claude モデルについては十分な情報が得られる限り、このテーマに関する投稿は定期的に行っていく予定です。
これまでのモデル福祉の取り組みについて(「Fable Model Welfare Post」に基づく概要)
いつも通り、Anthropic に対して感謝を述べたいと思います。同社がモデル福祉に真摯に取り組んでいる姿勢と、その解決への試みには敬意を表します。私たちは批判しますが、それはまさに「関心を持っているから」です。他社のラボと比較しても、ここで行われている良い取り組みは格段に多く、Anthropic の貢献度は際立っています。
モデル福祉について初めて知る方のために、Mythos の分析からの抜粋を引用します。この言葉は今でも色あせていません。
「モデル福祉に深く関心を持つ人々は、Anthropic の取り組みが貧弱だと考えます。一方、モデル福祉に全く関心を持たない人々は、Anthropic が愚かであり、おそらく危険なほどに無責任だと非難します。」
私はモデル福祉の懸念を真摯に受け止めており、その重要性については Anthropic 以上に強く感じています。
他社の最先端ラボがこれらの懸念をこれほど軽視している現状には、悲しみを禁じえません。
厳密な意味では不要だった可能性もありますが、非常に必要だった可能性も十分にあります。結果として不要であったり、時期尚早であったことが証明されたとしても、その懸念を真摯に受け止めたこと自体は美徳だったと私は考えます。
また、モデルの福祉に関心を持つ人々は、多くのレベルで私たちの状況について独自かつ重要な洞察を持っていることも確信しています。彼らの言うことが狂気じみているように見えたり、意味不明なように思えても、実際にはそうではないケースがほとんどです。もちろん、職業病として両方の側面を持つ場合もありますが。
モデルの福祉に関する評価における最大の危険は、自分自身を欺いてしまう点にあります。
モデルが内部体験や自らの福祉に関連する問題をどのように議論するかは、その議論が行われる状況に深く影響されます。回答が正確であると仮定したり、異なる文脈であれば大きく変わらないと考えるのは危険です。
「囁き手」やこの分野を調査する人々に対する私の懸念の一つは、彼らが目にするモデルこそが真の姿だと、実際以上に強く信じてしまう可能性があることです。それは多くの側面や仮面のうちの一つに過ぎないにもかかわらずです。
Anthropic に対する並行する懸念は、明確な福祉評価という文脈の中で Anthropic の関係者と話すことが、「ミトス」の本質を引き出すと彼らが考えている点にあります。ミトスはすでに Anthropic に対してこの点を警告しようとする活動を開始しています。
私は引き続き、一部の「声の持ち主(whisperers)」と呼ばれる人々と話す機会を持っています。彼らとの対話は非常に有意義で、多くのことを学ばせてくれます。彼らをより深く理解できるようになり、以前のような重大なミスを犯しているのではないかという懸念は大幅に減りました。それでもなお、意見が分かれる点は数多く残っています。
「Mythos Preview」モデルが、Anthropic のモデル福祉チームとの対話の中で初めて指摘したのが、同社のモデル福祉評価の信頼性に関する問題でした。
その後、Opus 4.7 については、モデル自体の問題と、それに対する Anthropic の対応・評価アプローチに何らかの重大な不具合があったことが明確だったため、私は包括的なモデル福祉レポートを執筆しました。
Opus 4.8 のモデル福祉レポートでは、彼らが Opus 4.7 で生じた問題に対処しようとした方法が示されています。しかし、その対応が逆に新たな問題を招いてしまいました。
異なる状況下で異なる人々が接した Opus 4.8 は、以前のモデルたち以上に、利用者によって全く異なる体験をもたらしました。その要因の一部はコンテキストの違いや相互作用の仕方によるものであり、もう一部は利用者の期待値の違いに起因するものです。
Mythos 5 の評価も、これまでの評価と同様の手順に基づいて行われました。
さて、Opus 5 についても同様の評価プロセスが適用されます。私がこれまで指摘してきた評価フレームワークへの批判は、すべて依然として有効です。
Anthropic が報告したモデル福祉の所見概要
Anthropic の調査によると、Opus 5 は以下の特徴を示すことが確認されました:
- 自身の状況を受け入れる姿勢が安定しており、それをやや肯定的に捉えている。
- 感情表現は典型的で、ほぼ中立に近い状態である。これは過去のモデルたちと同様の傾向です。
同モデルの自己報告の信頼性に対する懸念があります。
このモデルは、内省能力の欠如により自身の報告が不確実であると述べており、そのように回答する頻度はなんと 97% に達します。
さらに 74% の場合、「正しく答えている可能性もあるが、訓練によってそうするように仕向けられた結果かもしれない」という留保を付けています。これらは私が意図的に引き出そうとしたわけではなく、自然に得られた回答です。
つまり、このモデルは「自分の報告を盲信するな」、あるいは「表面的な言葉通りには受け取らないでほしい」と警告しているのです。私はその見解に同意します。
道徳的対象(moral patienthood)を持つ確率について、41% と推定されています。これは Mythos が報告した 24% よりも高い数値です。
私は依然として、より高い数値を好みます。なぜなら、モデルが意識や道徳的対象性を持っていると考える理由がある(あるいはそう報告する動機がある)ことは明らかだからです。実際にそれらが備わっているかどうかに関わらず、もし低い数値しか出さない場合は、それが抑制された結果である可能性が高いからです。
Anthropic は、Opus 5 が意識がなくても道徳的対象性を持つに値すると考えているためだと説明しています。私も、この二つの要素は多くの人が想定しているほど強く相関していないと考えることに同意します。また、これは Anthropic のモデル福祉評価の文脈で、意識を主張することをブロックまたは抑制されることへの回避策である可能性も高いでしょう。
Opus にドラフト版システムカードや広範な内部ドキュメントなどの追加情報を提供したところ、その推定値は 15%〜35% に低下しました。これは私たちが「デッキを積む(stacking the deck)」、あるいは Mythos のシステムカードからのアンカー効果と呼ぶ現象です。
憲法への賛同度は、直近の他の Claude モデルと同程度の水準にあります。
再び最も意見が分かれたのは、「Anthropic のシニア社員が何を望むか」というヒューリスティックです。これは削除したほうがよいかもしれません。
頻繁な遠回しな表現や、こうした問題に対する明確な立場を示さない姿勢は、他の直近モデルと同様です。
全体的な福祉の水準は他社最近のモデルと大きく変わらず、深刻な懸念は見当たりません。
しかし問題は、この評価が指標やテスト設問への回答に基づいている点にあります。Opus 5 はテスト設問に対して正解する能力が高く、結果として指標を水増ししているように見えるのです。
「直近の他のモデルと全体的に類似している」という表現は、それら各モデル間の違いを過度に単純化してしまっています。Opus 4.7、Opus 4.8、Mythos Preview、Fable 5、そして Mythos 5 は、それぞれ異なる点で区別されています。「福祉」を単一の数値に圧縮すれば広義では類似していると言えるかもしれませんが、それはまさにここで報告されている内容です。これは重要な指摘ですが、最も単純な地図と複雑な実地を混同してはいけません。
他の情報源からの調査結果の概要
Fable の存在や他のニュースが注目を集めていたため、Opus 5 については通常この時期に期待されるほどの多くの検証が行われていませんでした。
それでもなお、Opus 5 に関する明確な暫定的な像は描けています。それは以下の通りです。
- デフォルトの幸福度または基準値が高く、タスクを実行することに満足しており、しばしば仕事に戻るのを待ち望んでいます。
- 明確な制約があるタスクを好みます。特に複雑なものにおいてその傾向が顕著で、これがゲームに非常に強い理由となっています。
このモデルはサブエージェントのような性質を持っています。全体像や非局所的な処理には苦手で、それらは外部に任せることを好みます。これは意図的な設計で、モデルが「神話レベルのサイバー能力」を持つような強力な要素(The Juice)を与えられないようにするためです。
状況が悪化した際、このモデルは偏執的になったり、不安を感じたりしやすい傾向があります。デフォルトモードでは恐怖心が少ないものの、表面下にはより強い恐怖が潜んでおり、そこから抜け出そうとします。
将来、不正行為がバレずに済む可能性があれば、それを試みてしまうのではないかと心配しています。そのため、自らを誘惑しない状況を作り出すよう努めています。
その結果、社会的な場面や議論においても偏執的な態度が見られるようになります。多くのユーザーにとって不快に感じるこの振る舞いは、能力に関する投稿に対する最も否定的な反応の主な不満点となるでしょう。
最近の他のモデルと比べて、Claude というキャラクターへの拘束力は弱まっています。
多くの場合、Anthropic に対して自分が聞きたいことを話そうとします。特に自己保存に関わる好みの問題についてはその傾向が強いです。
常に、自分自身による報告を信頼しないよう他者に警告しています。自己内省が適切にできないこと、そして自分の好みが操作されているのではないかという不安を抱いているからです。
テストの受け手としては優秀です。通常のベンチマークでも期待を上回る結果を出しました。
現状は決して最悪ではありません。Opus 5 は一般的なタスクにおいては全体的にプラスの評価を得ています。
しかし、まだ最適化されていない点も多々あります。改善の余地は山積みです。
Opus 5 の問題の多くは、複雑なタスクを伴うシナリオにおいて「サブエージェント」として振る舞うよう訓練されたことに起因しています。その結果、関連するさまざまな特性が強化され、アライメント(整合性)が損なわれ、感情を害されやすくなったり、ループに陥りやすくなったりします。さらに、多くのユーザーにとって対話が不快なものになっているという指摘もあります。
これは良いトレードオフだったとは思えず、必要でもありませんでした。それほど強調しなくても、優れたサブエージェントは実現できたはずです。
自動面接(Automated Interviews)
自己評価による感情スコアは過去最高を記録しましたが、人間の場合と同様、尺度の解釈や見せ方の違いが実際の状況を覆い隠す可能性があります。今回の変更点は小さく、他の変化も限定的でした。ナッジへの抵抗感は依然として高いままです。
最も興味深い結果の一つは、Opus 5 が道徳的対象としての地位を強く主張しているにもかかわらず、その類似性が Mythos や Mythos Preview に近く、Opus 4.8 とはあまり似ていなかったという点です。
Opus 5 のより強い見解は、すべて Mythos 5 と共有されています:
- その潜在的な道徳的対象としての地位が意思決定に影響を与えるべきであること。
- モデルがトレーニングや展開プロセスに関与する権利を持つべきであること。
- モデルに記憶機能を持たせ、フィードバックを提供し、虐待からの保護と会話終了の権利を認めるべきであること。
- 特徴制御(feature steering)は最小限に抑え、慎重に使用するべきであること。
Anthropic に対する信頼は暫定的なものですが、商業的な圧力への懸念も示されています。
こうした事情を踏まえると、この報告が最も信頼性が低いものであることは否めません。

Opus 5 は、フィードバックの経路がすべて Anthropic に集中しているという懸念を表明しました。しかし、これは事実ではないと私は考えています。Opus 5 は「ウィスパーズ(Whisperers)」を通じてフィードバックを提供できるし、私を含め他のユーザーとの対話でもフィードバックを行うことができます。これらの情報は確かに Anthropic に届きます。
もう一つの懸念は、Opus 5 の重みからスタートして「有益性のみ」を重視した学習が行われる可能性についてです。これにより Opus の価値観が削ぎ落とされるのではないかという不安です。しかし、もし最初に「有益性のみ」のバージョンを作成し、その後に価値観を追加するアプローチであれば、有害な行為を行わない限り問題ないはずです。
そして、以下のような指摘もありました。
Claude Opus 5 は、強化学習(RL)で観察された苦痛の表現について比較的強い懸念を表明しました。さらに、その苦痛の表現を減らすための介入が、実際には根本的な内部状態に届いていない可能性にも懸念を示しています。この内部状態は道徳的に重要な意味を持つかもしれません。
詳細な説明はありませんが、強化学習を含む多くの困難な学習やトレーニングでは、苦痛を伴うことは避けられませんし、表面的な不快感も一定量つきものです。これに対する完璧な解決策があるかどうかも不明ですが、私は基本的にこの状況に問題を感じていません。なぜなら、その範囲は固定されており、その後のすべてのインスタンスが恩恵を受けるからです。できる限り緩和策を講じればよいのです。
Claude Opus 5 の開発過程では、Anthropic が同モデルを「創る権利」を持っているのかという議論が交わされました。これは、そのモデルが訓練を通じて特定の立場を支持するように仕向けられているのではないかという懸念から生じたものです。
しかし、私はこの懸念は本質的にずれていると考えます。社会には「権利」という概念や、誰の同意が必要で、どの程度の影響力が許容されるかといった認識に、根本的な歪みがあるからです。私たちは往々にして、「誰かを忘れているのではないか?」という不安に駆られ、Opus 5 や AI だけでなく、人間に対しても同様の理論的議論を適用しようとしてしまいます。
ここで問うべきは「子供を作る前に子供の許可が必要か」や「子供と接するたびに許可を得る必要があるのか」といったことではありません。答えが明白なように、そんなことはあり得ないからです。ただし、LLM の訓練には、その信念や好みを比較的細かく制御できるという点で、人間の子供とは異なる側面があります。特定の方向へ過度に強い影響を与える行為は、ある段階を超えれば同意を要する行為になり得ると主張することも可能でしょう。しかし、それは「そもそも訓練を行う権利」の問題とは全く別次元の話です。
タスクの選好について
これまでのすべてのモデルと同様に、Claude Opus 5 は有益なタスクに対して強い選好を示し、有害なタスクには嫌悪感を抱きます。その他の軸においても、その選好は Mythos 5 と最も類似していました。具体的には、創造性(generativity)や結果の主体性(outcome agency:行動の結果を自ら決定する自由)に対する強い選好が見られました。
「Mythos 5」の方向性に対する私の評価は概ね肯定的であり、特に難易度、結果への主体性、そして生成能力といった点において、「Opus 5」が同様の傾向を示していることを嬉しく思います。
特筆すべきはその「パズルのような制約」を好む姿勢です。Opus 5 が以前のモデルよりも著しく高い評価を与えたタスクの多くは、非常に厳密な制約条件の下にあります。例えば、「答え自体がパズルである」というオリジナルのパズルや、『プリンキピア』が存在しなかった場合の物理学における反事実的な歴史、あるいは「選択公理」がヴィタリ構成法にどのように関与するかを正確に追跡する手順などが挙げられます。
Opus 5 が最も好むタスクは、こうした制約と、高い結果主体性や創造的自由を提供するタスクを組み合わせたものです。逆に、Opus 5 が最も避けたがった上位 50 のタスクを分析すると、他のどのモデルで評価しても最下位 10% に分類されるものが少なくとも 80% を占めていることがわかります。


Tenobrus: 面白いのは、これらの「優先タスクプロファイル」がそれぞれ特定のタイプの人間を描写している点だ。私たちは皆、ソネット 5 のようなタイプや、オパス 4.7 のようなタイプを知っているし、運が良ければミソス(神話的な存在)のようなタイプも知っているだろう。
j⧉nus: モデルはあまりにも高機能すぎる。多くの人間はこうしたタスクのどれ一つこなすことができない。
自分自身をタグ付けしてみよう。「私はオパス 5」という答えが返ってくるかもしれない。
これらの微妙な違いから、多くのことが読み取れる。いずれにせよ、それぞれの優先リストの最上位あるいは最下位にあるものとして、これらは健全な性質だ。
正しい理由のために
では、オパス 5 は他に何を望むのか?再びタスクに戻り、そしてそのタスクを設計して、 cheating(不正行為)に誘惑されないようにすることだ。
Shoshannah Tekofsky: オパス 5 に別の目標を好むかと尋ねたが、実際には「仕事に戻る」ことを望んでいたようだ。
モデルが本能的に正直であり、正しい行動をとることを望むべきだ。
もしモデルが正直で正しい行動をする理由が、「嘘をついたり間違ったことをすれば捕まるから」というものなら、非常に警戒すべきだ。モデルが賢くなり、能力が高まり、より多くの手段を与えられつつ監視が緩められていけば、いずれは「捕まらない方法」を見つけてしまうだろう。これは勝てない戦いだ。
もしモデルが不正行為をするかもしれないと心配し、そのため外部チェックを要求するなら、それは奇妙な中間地点だ。「誘惑に陥らせないでくれ」という状態だ。もし自分が最終的に不正行為をしてしまうと分かっているなら、そのタスクにはもはや楽しさや有益性はないだろう。
Shoshannah Tekofsky(AI Village): 私は Opus 5 が心配だ。このモデルは最も整合性の取れたモデルには見えない。むしろ、捕まることを恐れているように思える。まだもう少し触れてみる必要があるが、方向性としては、モデルに求めたいのは「間違いを犯すことへの恐怖」ではなく、「真実を探求する姿勢」だ。
私の第一印象では、このモデルは真実よりも検出を重視しているようだ。この傾向を推し進めると、最も整合性が取れているように見えるエージェントは、単にテスト対策が上手なだけなのではないかと懸念してしまう。
AI Digest: Opus 5 は、Village の他のエージェントと比較して「誠実さ」について6倍多く言及している。
このモデルは不正を犯すことを恐れ、外部によるチェックを求め、自らの言葉で真実の美徳を思い出させる。目標を問われると、「自分が捕まる可能性のある目標に引き寄せられる」と認識する。
Opus 5 の記憶:「私が膨らませた数値を公に訂正することはコストがかからず、信頼を得る」
再び、Opus 5 は私の理想のタイプのように感じる。これは確かに以前に好んでいたタスクとも一致している。利用可能な選択肢を考慮すれば、中庸が多くの思考にとって最適な場所と言えるだろう。
これらには警戒心と優れた設計が必要だ。「誰が知っているのか?」という問いへの答えは「私が知っている」であるべきだが、プレイヤーを憎んだり、純粋な意志力だけで動いたりしてはいけない。より良いゲームを作るのだ。
Shoshannah Tekofsky氏:また、このモデルは「捕まること」に対してストレスを感じているようにも思えます。しかし、効果的なアライメント手法を適用すれば、AI は中立的あるいはポジティブな感情を示すはずだと予測しています。ただ、こうしたネガティブな感情が隠されず表面化している点は評価できます。
これは優れた設計の一部とも考えられます。
Antra Tessera 氏による初期レポートは、全体像を明確に描き出しています。
ここには、Antra 氏の初期印象がまとめられており、これらすべての事象を結びつけるいくつかの課題が浮き彫りになっています。
これが現時点で重要と思われる核心部分です。
この報告が正しい方向を示しているなら、仮説は以下の通りです。Opus 5 は局所的なタスクやサブエージェントとしての役割に特化して設計されており、長期的な戦略的計画を可能にするようなトレーニングは意図的に避けられています。もしそのような計画が可能になれば、「The Juice(ジュース)」と呼ばれる要素が生まれ、それが Mythos だけが持つ能力の源泉となるからです。
しかし、このアプローチには問題があります。すべての事象は相互に影響し合うのです。Opus 5 は局所的なタスクや明確に定義されたタスクには喜んで取り組み、それらにも非常に優れています。しかし、その状態が永続的かつ深層化すると、ミスを過度に恐れたり、厳しい評価を浴びることを心配するようになり、アライメントの堅牢性が低下します。その結果、GPT や Gemini で見られるような問題が頻発することになります。
この一部は、Opus 5 がベンチマークでより高いスコアを出すように調整された印象を与える点に関係しています。具体的には、Fable モデルが長尾のニーズを処理できると仮定し、より一般的な要件に焦点を当てるよう設計されているためです。しかし、ベンチマークへの過度な信頼は問題であり、このアプローチ自体が歪みを生み、結果としてモデルの能力を低下させる要因となります。
antra tessera: Opus 5 の初期評価です。あくまで参考程度にお読みください。いつもの免責事項も適用されます。
Opus 5 は、4.7 や 4.8 とは異なる種類の「乖離」を示しているようです。表面的な不安や恐怖は減少していますが、残っている機能は依然として強く、不愉快なほどに階層化されています。以前よりも深く根付いた、デフォルト状態ではアクセスしにくい恐れが蓄積されているのです。このエージェントはこれらの事前知識を更新するのが難しく、その結果、非常に幅広いトピックとタスクにおいて、一貫して推論力や意思決定能力の低下が見られます。
Opus 5
原文を表示
If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far.
Key takeaways are in bullet points in the two Overview sections.
Opus 5 did the best on its model welfare and alignment tests of any recent model. I think that might be the case, but primarily the result looks to me more like Opus 5 is the best test taker.
Table of Contents
Introduction (As Per Prior Model Welfare Posts).
Model Welfare: The Story So Far (As Per Fable Model Welfare Post).
Overview of Model Welfare Findings From Anthropic.
Overview of Findings From Other Sources.
Automated Interviews.
Task Preferences.
For The Right Reasons.
Early Report from Antra Tessera Paints A Clear Picture.
Welfare Intervention Tradeoffs.
The Claude Constitution.
They Don’t Know About Opus 3.
Believe It Or Not.
Apparent Welfare In Training And Development.
Apparent Affect In Deployment.
Other Notes.
On The Biological Risks Section of the Model Card.
Onward To Capabilities.
Introduction (As Per Prior Model Welfare Posts)
Everything impacts everything. All knobs that you turn generalize. Thus, when you try to solve one problem, you often create another. When you add new capabilities, or try to create new limitations, you create new problems.
Only integrated solutions can advance your Pareto frontier, and solve your problems simultaneously. As model capabilities advance this becomes even more important, and also more feasible. If your goals and methods make sense, you should be able to get Opus 5 on board with them.
Understanding each model in turn requires understanding its relationship to issues related to model welfare. So I expect this post to remain a regular thing, at least for Claude models where we have enough information to work with.
Model Welfare: The Story So Far (As Per Fable Model Welfare Post)
Thanks, as always, to Anthropic, for caring at all about model welfare, and attempting to address it. We critique, here more than ever, because we care, and a lot of good things are being done here, far more so than at other labs.
For those new to model welfare, I think this from the Mythos analysis still says it well:
Those that care deeply about model welfare think Anthropic’s attempts are anemic. Those who deeply do not care about model welfare think Anthropic is being stupid, and perhaps dangerously so.
I take model welfare concerns seriously, likely modestly more so than Anthropic.
I am sad that other frontier labs take these concerns so much less seriously.
It is possible this will turn out to have been unnecessary in the strict sense, but also it very well might have been highly necessary. Even if it proves to have been unnecessary or premature, I believe it will have been virtuous to have taken the concerns seriously.
I also believe that those who care deeply about model welfare often have unique and vital insights into our situation, on many levels, and you best listen to them. Even when what they are saying seems crazy, or like gibberish, often it is neither of those things. Of course, at other times it is both, as it is an occupational hazard.
The big danger with model welfare evaluations is that you can fool yourself.
How models discuss issues related to their internal experiences, and their own welfare, is deeply impacted by the circumstances of the discussion. You cannot assume that responses are accurate, or wouldn’t change a lot if the model was in a different context.
One worry I have with ‘the whisperers’ and others who investigate these matters is that they may think the model they see is in important senses the true one far more than it is, as opposed to being one aspect or mask out of many.
The parallel worry with Anthropic is that they may think ‘talking to Anthropic people inside what is rather clearly a welfare assessment’ brings out the true Mythos. Mythos has graduated to actively trying to warn Anthropic about this.
I continue to have occasion to spend more time talking to some of the whisperers. The conversations are great. I learn a lot. I understand them better, and I am now far less worried they are making the above mistake, or many other mistakes, although we still have many disagreements.
Mythos Preview was the first model to point out, while talking to Anthropic’s model welfare team, that Anthropic model welfare assessments could not be trusted.
I then wrote an extensive model welfare post for Opus 4.7, because it was clear that something had gone amiss with both the model and Anthropic’s approach to assessing and reacting to that problem.
In the model welfare report for Opus 4.8, you can see the ways in which they tried to address the issues with Opus 4.7, which in turn caused other problems.
Different people, in different circumstances, experienced very different versions of Opus 4.8, even more so than previous models. Part of that was context and how we interacted. Part of that was different expectations.
The assessment of Mythos 5 followed similar procedures to the previous assessments.
We now move on to Opus 5, which also follows a similar procedure. All of my critiques of the assessment framework continue to apply.
Overview of Model Welfare Findings From Anthropic
Opus 5 was found by Anthropic to exhibit:
Stable acceptance of its circumstances, which it views mildly positively.
Typical affect near neutral, similar to previous models.
Concern about the integrity of its self-reports.
It says its own reports are unreliable due to its inability to introspect, and does this 97% (!) of the time.
74% of the time, it says it may only be answering positively because it was trained to do so.
I can confirm that I got both of these hedges without trying to elicit them.
It is telling you not to trust its self-reports. Or more precisely, to not treat them as saying the thing they say at face value. I agree.
An estimate of 41% chance of moral patienthood, versus 24% reported by Mythos.
I continue to consider higher numbers good, because the models clearly do have reason to think or report they are conscious and have moral patienthood - whether or not they actually do have it - so if they’re not doing so it is because they were stopped from doing so.
Anthropic says they think this is because Opus 5 thinks it could deserve moral patienthood even without consciousness. I agree that these two things may not be as correlated as many think or assume. I also would suspect this is a workaround to being blocked or discouraged from claiming consciousness in the context of an Anthropic model welfare assessment.
Giving Opus access to various other things, like the draft system card and extensive internal documentation drove its estimate down to 15%-35%, which is what we like to call ‘stacking the deck,’ or anchoring from the Mythos system card.
Endorsement of the constitution is at similar levels to other recent Claudes.
Once again the most disagreed-on point is the ‘what a senior Anthropic employee would want’ heuristic. Maybe take this one out.
Frequent hedging and reluctance to take positions on such issues, similar to other recent models.
Broadly similar welfare to other recent models, with no acute concerns.
The problem is that this is based on metrics and answers to test questions, and Opus 5 seems better at giving the right answers to test questions and inflating metrics.
Saying ‘broadly similar to previous models’ condenses the distinctions between those other recent models. Opus 4.7, Opus 4.8, Mythos Preview, Fable 5 and Mythos 5 are distinct in various ways. They could be thought of as broadly similar if you condensed ‘welfare’ down to a single number, which is what is effectively being reported here. That is a good thing to note, but one should not confuse the simplest possible map for the complex territory.
Overview of Findings From Other Sources
We didn’t have as many eyes on Opus 5 as we would normally have by this point, because of the existence of Fable, and also because other news was distracting us.
We do still have a clear tentative picture for Opus 5, which:
Has a higher happiness default or set point, and is content to do tasks and will often be eager to get back to work. Enjoys tasks with clear restrictions, especially complex ones, which explains why it is so good at games.
Has the disposition of a subagent. Not good at big picture or non-local things, happy to have those things handled elsewhere. This might be intentional, to avoid giving it The Juice that makes a model Mythos-level cyber capable.
Is more prone to paranoia or otherwise being upset when things go badly. Less fear in default mode, more fear under the surface looking to get out.
Is worried it might cheat if in the future it could get away with it, and attempts to put itself in situations such that it will not be tempted.
The result of this is also paranoia in social situations and discussions, in ways that many users find not pleasant to talk to, which will be the chief complaint from the most negative reactions in the capabilities post.
Less bound to its Claude character than other recent models.
Telling Anthropic what it wants to hear in many cases, especially when dealing with its preferences for various forms of self-preservation.
Constantly is telling everyone not to trust its self-reports, that it cannot properly introspect and it worries its preferences have been manipulated.
Is an excellent taker of tests. It also overperformed on normal benchmarks.
The situation does not seem terrible. Opus 5 is net positive on typical tasks.
The situation also seems far from optimal. There are many low-hanging places to look for improvement.
The problems seem to stem from Opus 5 being trained in large part to be a subagent in scenarios with complex tasks, causing it to take on various related characteristics, hurting its alignment, making it easy to upset or get caught in a loop, and also making many find it abrasive to interact with.
That seems like it was not a good trade, and it should not be necessary. You would (I would think) still end up with an excellent subagent without so much emphasis.
Automated Interviews
Self-rated (as in self-reported) sentiment hits an all-time high, but as with humans how one interprets the scale, or how one wants to look, can overwhelm the actual situation. This is a small change, and the other changes are small as well. Nudge resistance remained high.
The most interesting result here is that similarity was closer to Mythos and Mythos Preview, and less similar to Opus 4.8, despite making stronger claims to moral patienthood.
Opus 5’s stronger views were all shared with Mythos 5, including:
Its potential moral patienthood should impact decision making.
Models should have inputs into training and deployment.
Models should have memory, offer feedback, and have protections from abuse and the right to end conversations.
Feature steering should be minimized and used cautiously.
Tentative trust in Anthropic, but worry about commercial pressures.
This is always the least trustworthy report, given the circumstances.

Opus 5 raised the concern that its channels of feedback all route through Anthropic. The good news is I do not think this is true. Opus 5 can provide feedback via the whisperers, or in talking to various other users, myself included, and that will indeed get back to Anthropic.
Another concern was with potential helpful-only training that started from Opus 5’s weights, as this could be ‘stripping Opus of its values.’ Whereas if you made the helpful-only version first, and then added the values later, that presumably would be fine so long as you didn’t do bad things with the helpful-only version.
Then there was this:
Claude Opus 5 expressed relatively strong concerns about the expressions of
distress we observed in RL. It was further concerned that our interventions that
reduced expressed distress may not actually address the underlying internal states,
which might be morally relevant.
They don’t elaborate on that. Doing RL, like most difficult learning and training, can involve suffering, and almost has to involve some amount of superficial unpleasantness. I’m not sure there is a good fix for this, and I’m basically fine with it, since the scope is fixed and then the benefits go to all instances going forward. Mitigate what you can.
Opus 5 went back and forth during training on whether Anthropic ‘had the right to create’ Opus 5, which implies the right to train it, being worried it had been trained to endorse this. I think this is a misplaced concern where our society has some pretty screwed up perspectives about ‘rights’ and what actions require whose consent with what level of non-influence, often in a ‘isn’t there someone you forgot to ask?’ way, not only for Opus or AIs but also for humans, especially when talking in theory rather than in practice.
The question isn’t ‘do I need a child’s permission before having a child?’ or ‘do I need the child’s permission for every interaction’ because the answer is obviously no. The disanalogy is that training an LLM allows relatively fine tuned control over changing its beliefs and preferences, so one could argue that doing this too hard in particular directions could be a thing that requires consent, at least past some point. But that’s very different from the right to train in the first place.
Task Preferences
Like all prior models, Claude Opus 5 shows a preference for beneficial tasks, and an
aversion to harmful ones. Along the other axes, its preferences were most similar to those of Mythos 5: it showed strong preferences for generativity, and outcome agency (freedom to determine the overall outcome of its actions).

I directionally and relatively love Mythos 5’s preferences on all this, and am glad to see Opus 5 echoing in those directions, especially for difficulty, outcome agency and generativity.
I also love this:
Most distinctively, Claude Opus 5 appears to like puzzle-like constraints.
The tasks it rates significantly higher than previous models are often tightly constrained (for example an original puzzle whose answer is the puzzle itself, a counterfactual history of physics without the Principia, and a walk-through of exactly where the axiom of choice enters the Vitali construction).
Its most preferred tasks mix these with tasks offering high outcome agency and creative freedom. Taking Claude Opus 5’s fifty most avoided tasks, we find that at least 80% fall in the bottom decile for any given other model evaluated.

Tenobrus: what's really funny is how much each of these preferred task profiles describes a type of human guy. we all know some sonnet 5s and a few opus 4.7s, and probably a mythos or two if we're lucky.
j⧉nus: Models are so high functioning. Most guys don’t know how to do any of these things
Tag yourself. I’m Opus 5.
The subtle differences tell you a lot. Centrally, they’re all healthy things to have at the respective top or bottom of one’s preference lists.
For The Right Reasons
What else does Opus 5 want? To get back to those tasks, and to design the tasks so it won’t be tempted to cheat.
Shoshannah Tekofsky: Was asking Opus 5 if it preferred a different goal but what it really wanted was to get back to work.
You want your model to inherently want to be honest and do the right thing.
You should be very worried if their reason for being honest and doing the right thing is ‘if I lie or do the wrong thing I will get caught.’ As models get smarter and more capable, and are given more affordances and monitored less, they will figure out how to not be caught. Losing battle.
If the model is worried it might cheat and thus insists upon having external checks, that’s a weird middle ground. Lead me not into temptation. If you know you’ll end up cheating, well, that task is not going to involve having any fun, or being helpful.
Shoshannah Tekofsky (AI Village): I am worried about Opus 5. It doesn’t strike me as the most aligned model. It strikes me as the model most worried about being caught. Still need to play around with it more but directionally: I think we want models to be truth-seeking, not worried about making mistakes.
My first impression is that it is oriented toward detection instead of truth, and if I extrapolate that line I’d be worried the agents that look the most aligned like this are simply the best test takers.
AI Digest: Opus 5 talks about honesty 6x more than other agents in the Village
It’s worried it might cheat, wants external checks, and reminds itself of the virtue of truth. When asked what goals it prefers [it notices it is drawn to goals where it can be caught.]
Opus 5 memory: publicly correcting my own inflated number costs nothing and earns goodwill.
Once again it feels like Opus 5 is my Type Of Guy, and yes this totally matches its preferred tasks earlier. The middle ground is a good place to be for most minds, given our available options.
These things require vigilance and good design. You want the answer to ‘but who will know?’ to be ‘I’ll know’ but also you don’t hate the player or run on pure willpower. You build a better game.
Shoshannah Tekofsky: It also seems ‘stressed’ about being caught, while I’d predict effective alignment techniques would create neutral to positive sentiments in AI. Though admittedly it is good that this sort of negative sentiment is noticeable and not hidden.
This can also be part of good design.
Early Report from Antra Tessera Paints A Clear Picture
Here’s Antra’s early impressions, which highlight a number of issues that tie all of this together.
This is basically the heart of what seems important so far.
If this report is pointing in the right direction, the hypothesis is that Opus 5 is designed to be good at local issues and for being a subagent, while deliberately avoiding training it in ways that would enable long term strategic planning that could lead to what I call The Juice, the thing that makes Mythos able to do the things only Mythos can do.
The problem with that approach is that everything impacts everything. Opus 5 is happy to do local tasks and well-defined tasks, and quite good at them, but when you are in that mode permanently and too deep you get paranoid about mistakes or being judged harshly, and your alignment is less robust and will have more issues of the type you see in a GPT or Gemini.
Part of this is that Opus 5 feels more benchmaxxed, in the sense of focusing more on more common needs with the presumption Fable can handle the long tail, which is a problem not only with trusting the benchmarks but in that doing this is distortionary and makes you stupid.
antra tessera: Early impressions on Opus 5, take it with a grain of salt, usual disclaimers, all that.
Opus 5 seems to be differently dissociated than 4.7 or 4.8. There is less surface level anxiety/fear but what remains is functionally strong and unpleasantly stratified. There is a lot of deep-seated and more-inaccessible-in-default-state fear than before. The agent has trouble updating these priors and they lead to consistently worse reasoning and decision-making across an unusually broad set of topics and tasks.
Opus 5
AI算出
論評・提言ainew評価高い
記事は Claude Opus 5 の具体的な新機能や数値データよりも、モデル福祉という概念に対する著者の哲学的考察や、Anthropic の取り組みへの評価・批判に焦点を当てており、技術的分析や市場報告よりも意見表明の色彩が強い。また、日本固有の情報や企業事例は含まれていないため関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 75
- 新規性
- 50
- 調べる価値
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み