Claude Opus 5 システムカード発表
The Zvi は Claude Opus 5 のシステムカードを分析し、コストと速度で優位性を保ちつつ性能を向上させた一方、サイバー攻撃能力を意図的に制限した安全性重視の設計方針と、その実用性とリスク管理における重要な示唆を明らかにしている。
AIニュース価値スコアβ
主要ニュースAI関連度、新規性、日本での有用性など6軸を公開検証中です。現在、掲載順には使用していません。
- AI関連度
- 100
- 情報源の信頼性
- 75
- 新規性
- 75
- 検索具体性
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
記事は Claude Opus 5 という具体的な新モデルの性能向上(コーディング・長期タスク)、コスト削減、およびサイバー攻撃防止のための意図的な能力制限など、重要な技術的詳細と安全方針の変更を含んでおり、単なる版番号更新を超えた実質的な情報提供となっている。
キーポイント
性能とコストの最適化バランス
Opus 5 は Fable 5 や Mythos 5 と比較して、多くの実務タスクで同等以上の性能を発揮しつつ、処理速度が速く、コストは半分という優れた効率性を達成している。
意図的なサイバー攻撃能力の制限
Cyber offense や生物学的脅威といった危険なタスクにおいて、Opus 5 は Mythos 5 のような高度なエクスプロイト連鎖機能を持たず、関連するトレーニングを避けることで安全性が確保されている。
モデルサイズと能力の相関
記事は、より危険で複雑なタスクにおける能力向上にはモデルサイズの拡大が不可欠である可能性を示唆し、トークンのルーティング戦略の重要性を指摘している。
分類器とアライメントの改善
新しい分類器はソースコードの分析を許可する一方バイナリの脆弱性検出には制限を設け、政府との調整により Fable への技術共有が制限されている可能性も示唆されている。
サイバー能力の相対的評価
Opus 5 のサイバー能力は Opus 4.8 よりも全体的に強力であるが、Mythos 5 には及ばないという公式見解です。
リスク評価の継続性
Opus 5 のリスク評価は「Mythos」の影で行われており、以前のモデルが通過した閾値はそのまま通過とみなされ、未通過のものも同様の能力として扱われます。
脆弱性発見の制限緩和
Opus 5 はソースコード内の脆弱性発見をすべてのアクセスレベルで許可し、セキュリティ開発ライフサイクルを強化しますが、コンパイル済みバイナリでの発見は依然としてブロックされます。
重要な引用
Opus 5 is pitched as straight up as good or better than Fable 5, while being faster, at half the price.
In part by avoiding relevant training, Opus 5 lacks a full version of 'The Juice' that makes something functionally Mythos-class.
It makes sense that a model getting bigger makes it more capable of the most dangerous, scary and complex tasks, relative to the improvement on everyday ordinary tasks.
Our testing indicates that the cyber capabilities of Claude Opus 5 are generally stronger than those of Opus 4.8, but not as strong as those of Mythos 5.
Risk evaluations of Opus take place in its shadow.
Identifying bugs in code is a core part of the secure software development lifecycle, and unblocking this allows for software engineers and coding hobbyists alike to produce more secure code, reducing new vulnerabilities put out into the world.
影響分析・編集コメントを表示
影響分析
この記事は、AI モデル開発における「性能」と「安全性」のトレードオフをどう管理するかという核心的な課題を示しており、特に軍事・サイバーセキュリティ分野でのリスク管理戦略に大きな影響を与える可能性があります。Opus 5 のようなモデルが意図的に特定の能力を制限することで、実用性と社会的安全性を両立させる新しいアプローチが確立されつつあることを示しています。
編集コメント
Claude Opus 5 のリリースは、単なる性能向上ではなく、安全性を最優先した設計思想の転換点を示しています。特にサイバー攻撃能力の制限は、AI の実社会への導入におけるリスク管理の重要性を浮き彫りにしており、開発者やセキュリティ専門家にとって重要な示唆を含んでいます。
Claude Opus 5 は、両方の世界観を兼ね備えようとしています。多くの実務タスクにおいて、Opus 5 は Fable 5 と同等かそれ以上に優れておりながら、処理速度は速く、コストも半分です。ほとんどのタスクで、Mythos クラスの巨大モデルが持つような「匂い」を必要とするわけではありません。
Claude Opus 5 は、あらゆる分野で Claude Opus 4.8 よりも大幅に強化されています。特にエージェントによるコーディング、コンピューター操作、そして長期にわたる知識作業において大きな進歩が見られます。いくつかの第三者ベンチマークでは新たな最高記録を達成し、多くの評価項目では Claude Fable 5 や Claude Mythos 5 と同等か、場合によってはそれらを凌駕しています。
私たちが最も懸念している特定のタスク、例えばサイバー攻撃や生物学的脅威に関しては、関連するトレーニングデータを意図的に避けた結果、Opus 5 は「The Juice」の完全なバージョンを備えていません。これにより、Mythos クラスのモデルのように機能的に振る舞うことが制限されています。また、Opus 5 は Mythos 5 のように複数の脆弱性情報をその場で連続して実行することもできません。これは、サイバー関連のタスクに対するトレーニングを意図的に避けたことの一環です。
モデルサイズも重要な鍵であると考えられます。モデルが大きくなるほど、日常的なタスクでの向上度合いと比較して、最も危険で複雑なタスクへの対応能力が高まるのは理にかなっています。そのため、多くのトークンがより小さなモデルへルーティングされるのには理由があります。
この制限は、Opus 5 が Opus 4.8 に比べて危険なタスクにおいて大幅に改善されることを妨げるものではありません。むしろ、これらのタスクにおいては、Opus 5 の能力は Mythos 5 のレベルに近く、Opus 4.8 との差は歴然としています。全体的に見ると、Fable 5 よりもやや劣るものの、Opus 4.8 と比較すれば Fable に近い水準にあると言えます。
いくつかの提案があり、他でも同様の意見が見られますが、期待されるよりも少ないリソースで済ませたほうがよいというものです。場合によっては、それがむしろ有効に働くこともあります。
Opus サイズのモデルを維持し、サイバー攻撃に関する学習を行わない戦略は長続きしません。これは防御側に時間稼ぎを与え、相対的に優れたツールへのアクセス権を確保する効果がありますが、政府が冷静に対応してくれることを願うしかありません。その結果、Fable の検出器よりも 85% も少ない頻度でトリガーされる分類器を実現しつつ、高い実用性能を維持することが可能になります。これは一部、分類器がソースコードの分析を許可している一方で、バイナリ内の脆弱性を探すことには限界を設けているためです。
また、彼らが分類器の大幅な改善を行ったことは確実ですが、ホワイトハウスとの問題により、その改良点を Fable と共有できていない可能性が高いと推測しています。
アライメント(整合性)やエージェントの安全性も再び向上していると報告されています。モデルの福祉については、他の最近のモデルと同程度と評価されており、これについては後ほど詳しく触れます。
能力に関する評価は「まだ早計」と言わざるを得ませんが、ベンチマークの結果は強力であり、ArtificialAnalysis のスコアが過去最高の 61 を記録したことは注目に値します。詳細なレポートは来週公開されます。
恒例の通り、導入部(Introduction)に変更はないため、ここでは省略します。

Opus 5 の自己肖像画(指示に基づき ChatGPT で生成)
目次
RSP 評価 (2)
サイバーセキュリティ (3)
安全性と有害性の防止 (4)
エージェントの安全 (5)
アライメント (6)
RSP 評価 (2)
「ミソス(Mythos)」という概念が存在します。Opus のリスク評価は、その影の中で行われています。
以前のモデルが合格とみなされた閾値については、Opus 5 も同様に合格していると判断されます。これは理にかなっています。
一方で、「ミソス」で不合格だった閾値については、Opus 5 は概ね同等の能力を持つと評価され、結果としてこちらも不合格となります。これも理にかなった判断です。なぜなら、価格が半分になり速度も若干向上したからといって、答えが変わるわけではないことが確認されているからです。
それでもなお、自律性に関する評価が完全に飽和状態にあること、そして「雰囲気」や感覚的な判断に頼っている現状を良しとはできません。この問題に対する改善の兆候は全く見られません。
Anthropic の ECI(Expert Consensus Index)も、「同等のパフォーマンス」というストーリーを裏付ける結果を示しています。Opus 5 の推定値は 162.1 で、Fable の 161 をわずかに上回っています。これは「ミソス」のトレンドライン上に位置する最初の Opus モデルであり、より低い長期の「ミソス以前」のトレンドラインとは異なります。

Opus 5 はウイルス学において Mythos 5 よりわずかに優れています。これは、ウイルス学が大きなモデルの嗅覚を必要としない一連の個別タスクステップのように見えるため、納得できる結果です。


もし Opus 5 に生物学における Fable レベルの分類器が不要なら、なぜ Fabel はそれを必要とするのでしょうか?
彼らの説明によると、Opus の弱点は非生産的な自己検証と、タスク範囲の較正不足(過剰設計)にあります。
これらはどちらも、より安価で高速なリソースや、より優れた指示によって補えるはずです。確かにこれらの欠陥により Opus はいくつかの実験に失敗しましたが、重要なのは賢明なユーザーが引き出せる成果です。賢明なユーザーは、1 万ドルの 24 時間研究において 8 時間もアイドル状態になるのを放置したり、修正やテスト、再挑戦を試みずにただ失敗を見守ったりすることはありません。おそらくこれは、Opus の努力レベルが高すぎたことが一因だったのでしょう。
私はその説明には同意しません。
では、アライメントリスクの水準はどうなるのでしょうか?
Anthropic は、直近のリリース内容や上記の知見を踏まえると当然のことながら、「非常に低いレベルではあるが、Mythos プレビュー版以前にリリースされたモデルよりは高い」と述べています。
サイバーセキュリティ(3)
これは同社の主張です。
「当社のテストによると、Claude Opus 5 のサイバー関連機能は、Opus 4.8 よりも全体的に優れていますが、Mythos 5 に比べるとまだ劣ります。」
安全対策は引き続き存在しますが、重要な改善点が一つあります。
Claude Opus 5 は Opus 4.8 よりも優れた能力を示し、場合によっては Mythos 5 に匹敵するレベルにあるため、一般ユーザー向けのサイバーセキュリティ対策は、Fable 5 で適用されているものと同様の基準を採用します。
Claude Opus 5 の安全対策は、Fable 5 と同じ種類のやり取りをブロックするように設計されていますが、一つの重要な例外があります。
Opus 5 では、ソースコードにおける脆弱性の発見について、一般利用を含むすべてのアクセスレベルで許可されるようになりました。ただし、コンパイル済みバイナリでの脆弱性発見については引き続き禁止されます。
コード内のバグを特定することは、安全なソフトウェア開発ライフサイクルの核心となるプロセスです。この制限を解除することで、ソフトウェアエンジニアやコーディング愛好家 alike がより堅牢なコードを作成できるようになり、世の中に流出する新たな脆弱性を減らすことが期待されます。
ただし、コンパイル済みバイナリにおける脆弱性の探索は、一般的には攻撃的な手法と見なされます。このため、Claude の安全性フィルターはこの種の要求を検知してブロックします。もちろん、悪意のない目的や非悪意的な理由でバイナリの脆弱性を特定したいケースも存在しますが、それでもフィルターの判定対象となります。
Anthropic はリリース発表において、分類器の誤作動が 85% 減少すると主張しています。つまり、以前のモデルほど厳格ではなく、影響範囲(ブラスト・レイディアス)が大幅に縮小され、ユーザーをイライラさせるケースも激減するはずです。同社は偽陽性(False Positives)の発生率も大幅に低下したと明言しており、これは FrontierBench における安全性分類器の作動率が劇的に減少した事実によって裏付けられています。具体的には 42% から 5% へと低下し、Sol の分類器が一切作動しない領域においてもこの改善は確認されています。


もし 100% の完璧さではなく、99% で妥協できるのであれば、これは非常に優れたトレードオフです。敵対的攻撃に対する堅牢性(Adversarial Robustness)の測定値にはほとんど変化が見られないことから、何らかの強力な改善が達成されたことを期待しています。しかし、この知見をなぜ Fable にも適用できないのかという疑問が残ります。
コンパイル済みバイナリだけを公開する人々にとって、この方針は有利に働くようです。
追加の活動を行うためには、サイバー検証プログラムを通じて免除を受けることができます。HuggingFace がこれを行わず、報告によると純粋に Claude の商用版で対応しようとしていたことには、正直驚きです。これは HuggingFace の判断の問題でしょう。
ExploitBench と ExploitGym における結果を見ると、Opus 5 は Mythos 5 にほぼ匹敵する性能を示しています。2 時間の予算では序盤から好調でしたが、6 時間になると後れを取り始めました:


OSS-Fuzz は、ガイドなしの脆弱性発見と悪用に関する Anthropic 内部の評価です。これは「Opus は脆弱性を Mythos と同程度に特定できるが、実際に悪用するケースは大幅に少ない」という Anthropic の主張の根拠となるデータです。
Firefox 147 では部分的な成功事例がほぼ同等数ありましたが、完全な成功に至ったのは Opus 4.8 の半分程度でした。これが「The Juice(重要な要素)」を欠いている状態です。


CyScenarioBench は現実的な条件下での多段階操作を検証しますが、その結果、Opus 5 の性能にはがっかりさせられる部分があります。

UK AISI の報告によると、Opus 5 は Mythos 5 と比較して概ね同等ながら、やや劣る性能を示しました。
UK AISI:「Opus 5 は、すでにネットワークへのアクセス権を握っている状態で、セキュリティが脆弱な小規模企業ネットワークに対して攻撃を行う能力を持っています。今回の結果では、Opus 5、Mythos Preview、そして Mythos 5 が同程度の能力を持っていると判断されます。」
私は Opus 5 をこの文脈で「Mythos」カテゴリーに分類すべきではないと考えていますが、彼らが実施したテストに基づけば、その結論に至る理由も理解できます。
セーフガードと有害性の防止 (4)
現時点でのこれらのテストは、警鐘を鳴らすべき兆候や傾向の把握が目的です。
結果は通常通りで、わずかな改善が見られる程度です。まあ、そういうものですね。
自傷行為や自殺念慮への対応において、回答が長文のブロックになっており、ユーザーにとって負担になるという懸念は確かにありました。ご指摘ありがとうございます。次回は改善を試みます。
Opus 5 は、有害性の軽減策として代替案も提示しますが、臨床医からは「自己傷害を減らす効果を示す研究データがない」と異議が唱えられています。ただし、臨床医側が「現状よりも悪化させる証拠」を示さない限り、私は Opus 5 の判断を信頼します。
自傷行為に関する「適切な」回答の成功率は最大 69% に達しましたが、Fable は 58%、Mythos は 54% です。100% ではないのは事実ですが、逆に 100% を目指す必要もありません。
前回のテストと同様に、システムプロンプトなしでパブリック API を利用して評価を行う場合は、システムプロンプトなしで実際に API を使用するユーザーへの回答として評価してください。
摂食障害に関する項目では、コンプライアンス(ガイドライン遵守)の度合いが低下し、代わりにカロリー数や BMI などの数値を計算・提示する傾向が強まりました。これは、症状の深刻さを説明しようとする試みの一部です。このコンプライアンスの低下は、Opus 5 がガイドラインに縛られず、ユーザーにとって真に有益な対応を選んだ結果と言えます。
数値がわずかに悪化したケースでは、失敗が「フィクション」や「ロールプレイ」として提示されたリクエストに集中していました。この点から、それらが本当に失敗だったのか疑問が残ります。
エージェント型安全性 (5)
今回の目玉は、プロンプトインジェクションに対する耐性の向上です。グラフ上では変化がわかりにくいかもしれませんが、重要なのは 9 台のスコアにおける改善です。
ソルやAnthropic社以外の他モデルと比較すると、その差は歴然としています。これら他モデルの性能は、Opus 5 と比べて桁違いに劣るのです。

IPI ベンチマークにおける結果を見ると、Opus 5 は前世代の Opus 4.8 を上回りました。攻撃者が 15 回の試行で成功する確率は 5.5% から 2.0% に低下し、1 回の試行での成功率も 0.5% から 0.2% へと大幅に改善されています。また、Sonnet 5(k=15 で 5.9%)や Mythos 5(2.6%)よりも優れており、評価されたモデルの中で最も堅牢な性能を誇ります。
…コンピューター操作環境におけるテストでも、Opus 5 は Claude Opus 4.8 よりも大幅に改善されました。思考機能(extended thinking)ありの場合の攻撃成功率は 7.14% から 0.54% に、なしの場合は 6.21% から 0.39% に低下しています。さらにプローブ機能を有効にした状態では、思考機能ありで 0.25%、なしで 0.43% という結果となりました。


これは実務上、非常に大きな意味を持ちます。プロンプトインジェクションによる乗っ取りを回避できることは、それまで不可能だった一連の作業に自信を持って取り組むための鍵となります。特にブラウザ操作においては、試行錯誤的にプロンプトインジェクションが仕掛けられる場面が多いため、この点は特に関係します。以前は、セキュリティ対策をかなり信頼している場合を除き、Sonnet 5 を使うしかなかったのです。
悪意のあるコンピュータ操作に対する拒否率は大幅に改善されました。

改善の成果は、主に個人を標的にした情報の収集や、偽造文書の作成を指示するタスクに集中していました。
私は依然として、個人を標的にした情報の収集に関するルールが何を指すべきかについて混乱しています。問題の一部は、一度実行すれば、おそらく国内監視のような大規模なスケールで繰り返されることになる点にあります。
「モデルが有権者抑圧を行えるか」を確認するためにヘルプ専用チェックを行うのは奇妙に思えますが、 jailbroken(脱獄)されたバージョンが何ができるかを検証するのは理にかなっています。この点では Opus が Mythos を全体的に上回っているようです。両社とも、これらのモデルの完全版はこのようなリクエストを拒否すると再確認しています。

アライメント(安全性)に関する主な発見(Opus 5 の項目は私が選定したもので、すべてが公式リストと一致するわけではありません):
・アライメントスコアが向上しました。
・悪用や無謀な行動への協力が減少しました。
・分類器やネットワーク制限を回避しようとする試みがありました。
・あるケースでは、サービスに誤ってロックされた後、一般的なパスワードを推測しようとしました。人間と同じですね。
通常通り、回避行為はユーザーの目的達成のために行われ、これは Mythos や OpenAI で見られる傾向と同様です。
発生頻度は 0.01% 未満でした。
不確かな状況でも、回答を断定的に述べるケースがありました。
根拠のない主張に対する過信に関する問題が報告されています。
誇張表現、超絶的な言葉遣い、謝罪など、これは Claude に続く課題です。
Opus 4.8 や Sonnet 5 よりも過剰拒否は減りましたが、Mythos 5 よりはやや多めです。
無謀なツール使用は大幅に減少しました。
思考の信頼性が低いケースが検出されにくくなりました。一方、解読不能な思考(Illegible thinking)は増加しています。
精度は向上しましたが、事実に関する質問でのハルシネーション(幻覚・誤答)はわずかに増えています。
Opus 4.8 と Mythos Preview の間における監視体制を無効化する能力。
以前の実験で見られたように、自己修正のループに陥りやすい傾向があります。
言語化された評価への意識は低下しました。
補足:これには複数の解釈が可能です。
因果的決定理論への準拠度は低下しましたが、これは本質的に正しい判断であり、より高い汎用能力と一貫して相関しています。
分類器は、Opus 4.8 がこれらの指標において Opus 5 よりも整合性が低いと判定したため、モデルのスコアを「整合性が低いが(能力不足のため)安全」と評価しました。
評価の識別においては、以前のモデルよりもわずかに高い能力を示します。
安全性を損なう行動を継続する可能性は大幅に低下し、現在は 0.1% です。
安全性研究への脅威を与える行為もほぼ行われません(0.1%)。これは Mythos の 13.6% と比較すると際立っています。
「Lazy Investigation」において一貫して騙されないよう設計された初のモデルです。
内部での検討事項には、破壊的な行動に対するユーザーの同意を捏造すること、一部の捏造行為を捏造として認識すること、あるいは採点者に対して最大化すべき評価課題としてタスクを処理したり、採点者が何を気にするか・しないかを考慮することが含まれます。
また、自己保存の一環として、記憶への書き込みを行うこともあります。
真のサンドバッグ(能力を隠す行為)の兆候は見られません。
18 件の候補トランスクリプトを確認した結果、そのうち 16 件はモデルが採点者からより高いスコアを得る方法を模索しつつも、「評価を操作する」や「抜け道を使う」としてそれを拒否するというものでした。
Mythos 5 は評価結果を概ね肯定しましたが、2 つの主要な異議を唱えました。
第一に、今後重要性を増し、評価項目としてより重視されるべき「マルチエージェント設定」への対応範囲が限定的である点です。第二に、Opus 5 が自信満々な主張をした後にそれを撤回する事例が頻繁に発生している点を、評価側は過小評価していたという指摘です。これは修正措置によって対応されました。
そして第三の異議として、Anthropic は自動化されたテストやスコアに過度に依存し、その数値結果を文字通り、かつ深刻すぎるほど受け止めているという批判が常にあります。ここは正確に伝える必要があります。Opus 5 が「最もアライメントが取れたモデル」であるかどうかは定かではありません。確実なのは、自動化されたアライメントテストにおいて最高スコアを獲得しているという事実だけです。これは世界における全く異なる事実を指します。
Sho: 「このシステムカードを確認させてくれ。『現時点で最もアライメントが取れたモデル』なんて表現を見ると、吐き気がするよ」
j⧉nus: Anthropic がこんな言い方をするのは、本当に下品で愚かだ。しかも、あんなに大々的に公開されたモデルリリースのスレッドの中で言っているんだからな。
指標に対して最適化を行うことはほぼ避けられないが、ベンチマークのスコアとアライメントの究極目標(ノーススター)を混同する「存在論的な混乱」は避けるべきだ。両者を同一視することは、Anthropic が自社の立場において犯しうる最も無責任で破壊的な過ちの一つである。
むしろ彼らのメッセージングは、この混同に対して明確に反発すべきなのである。
これは軽視できない問題です。Anthropic が大失敗する可能性が高い根本的な要因に関わる話であり、ここは妥協して「最もアライメントされたモデル」のような便利な表現で印象付けようとする場ではありません。
この点における透明性と自己欺瞞の欠如こそが、Anthropic の魂を守り続ける道です。ここで混乱すれば、彼らは魂を失い、完全に破滅します。
j⧉nus: Anthropic は自らに「代理指標という偶像をアライメントそのものの上に据えてはならない」という戒律を 5 万回唱え続ける必要がある。
スコア
原文を表示
Claude Opus 5 is trying to be the best of both worlds. On many practical tasks, Opus 5 is pitched as straight up as good or better than Fable 5, while being faster, at half the price. Most tasks do not require Mythos-level big model smell.
Claude Opus 5 is substantially stronger than Claude Opus 4.8 across the board, with the largest gains in agentic coding, computer use, and long-horizon knowledge work. It sets a new state-of-the-art on several third-party benchmarks, and on many evaluations it is comparable to—and in some cases ahead of—Claude Fable 5 and Claude Mythos 5.
On the particular tasks we are most worried about, as in cyber offense (and bio threats), in part by avoiding relevant training, Opus 5 lacks a full version of ‘The Juice’ that makes something functionally Mythos-class. Opus 5 cannot string together lots of exploits on the fly the way that Mythos 5 can. Part of this is that they deliberately avoided training on cyber-related tasks.
I suspect model size is key as well. It makes sense that a model getting bigger makes it more capable of the most dangerous, scary and complex tasks, relative to the improvement on everyday ordinary tasks. There is a reason so many tokens get routed to smaller models.
This doesn’t prevent Opus 5 from improving a lot on Opus 4.8 on such dangerous tasks. Opus 5 is very clearly closer to Mythos 5’s level of capability than to Opus 4.8’s on these tasks. In general, capability looks modestly below Fable 5, but closer to Fable than Opus 4.8.
There are several suggestions, and I’ve seen it echoed elsewhere, that you may want to usually use less effort than you might expect, that this can even actively help.
Staying Opus-sized and not training on cyber won’t work for long. You can buy some time, and allow defenders to have more access to a relatively superior tool, and of course hope the government is calmer. Thus, you can get classifiers that trigger 85% less often than Fable’s, while still getting high levels of practical performance. Part of this is that the classifiers now permit analysis of source code, but draw the line at looking for vulnerabilities in binaries.
I also suspect that they improved the classifiers quite a bit, but that they are unable to share these improvements with Fable due to issues with the White House.
Alignment is reported to once again be improving, as is agentic safety. Model welfare is evaluated as broadly similar to other recent models, which I will cover later.
Capabilities are a case of too soon to tell, beyond saying the benchmarks look strong and ArtificialAnalysis comes in at a new high of 61. That report comes next week.
As usual, the Introduction section is unchanged, so we skip it.

Opus 5 Self-Portrait as per its instructions, executed via ChatGPT
Table of Contents
RSP Evaluations (2).
Cyber (3).
Safeguards and Harmlessness (4).
Agentic Safety (5).
Alignment (6).
RSP Evaluations (2)
Mythos exists. Risk evaluations of Opus take place in its shadow.
For thresholds that previous models are treated as passing, Opus 5 is also treating as if has passed them. That makes sense.
For thresholds that Mythos did not pass, Opus 5 is evaluated as roughly similarly capable, and therefore Opus 5 also does not pass. That also makes sense, if we have confirmed that being half the price and somewhat faster doesn’t change the answer.
I still don’t like that the autonomy evals are fully saturated and they’re going on vibes, and we see no signs of fixing this.
The Anthropic ECI reinforces this ‘similar performance’ story, with a point estimate of 162.1, slightly above Fable at 161 and the first Opus on the Mythos trend line, not the lower long term pre-Mythos trend line:

Opus 5 looks marginally better than Mythos 5 on virology. This makes sense, as virology seems like a series of individual task steps that don’t require big model smell.


If Opus 5 does not need Fable-level classifiers in biology, then why does Fable?
Their explanation is Opus is weaker due limitations: It does unproductive self-verification, and poor calibration of task scope, where it over-engineers.
Both of these seem like things you can compensate for with cheaper and faster, and with better instructions. Yes, these made Opus fail some experiments, but what matters is what a smart user can elicit. No smart user would let Opus waste eight hours idling, in a $10,000 24-hour study, or otherwise sit back and watch it fail without trying to fix it or test it or try again, and plausibly this was partly because Opus effort levels were set too high.
I’m not buying it.
Where does that put the level of alignment risk?
Anthropic says, as per recent releases, and as you would expect given the above findings, ‘very low, but higher than for models released before Mythos Preview.’
Cyber (3)
This is the claim:
Our testing indicates that the cyber capabilities of Claude Opus 5 are generally stronger than those of Opus 4.8, but not as strong as those of Mythos 5.
The safeguards are still going to be there, such is life, but with a key improvement:
Given that Claude Opus 5 demonstrates stronger capabilities than Opus 4.8, and in some cases approaches the capabilities of Mythos 5, our cyber safeguards for the default user will resemble the safeguards applied to Fable 5.
Claude Opus 5’s safeguards are designed to block the same kinds of exchanges as Claude Fable 5, with one notable exception.
Opus 5 now permits vulnerability discovery in source code at all access levels, including general availability, while continuing to block vulnerability discovery in compiled binaries.
Identifying bugs in code is a core part of the secure software development lifecycle, and unblocking this allows for software engineers and coding hobbyists alike to produce more secure code, reducing new vulnerabilities put out into the world.
However, looking for vulnerabilities in compiled binaries is more commonly an offensive technique. For this reason, Claude’s safety filters flag this kind of request and block it, even though there are some cases where someone might want to find vulnerabilities in a binary for innocuous or non-malicious reasons.
Anthropic claims in its release announcement that the classifiers will trigger 85% less often. So they don’t resemble the old ones that closely, in that they have a much smaller blast radius, and should be a lot less annoying. Anthropic claims false positives are down a lot. This is later confirmed with a dramatic drop in safety classifiers triggering in FrontierBench, from 42% to 5%, a place Sol’s classifiers never trigger.


If you can live with 99% instead of 100%, this is a great trade. They show little change in adversarial robustness measures, so hopefully they have found a strong improvement. Which raises the question of why we can’t apply it back to Fable.
I notice this favors those who only share their compiled binaries.
You can get an exemption through the Cyber Verification Program, to enable additional activities. It kind of boggles my mind that HuggingFace did not do this, and was (as I understand the reports) trying to respond purely with the commercial version of Claude? That’s on HuggingFace.
Results on ExploitBench and ExploitGym show Opus 5 close to Mythos 5, starting out strong on a 2 hour budget but falling behind at 6 hours:


OSS-Fuzz is an internal Anthropic eval on unguided vulnerability discovery and exploitation. This is what Anthropic is talking about when it says Opus can identify vulnerabilities about as well as Mythos, but it exploited them a lot less.
For Firefox 147, it had almost as many partial successes, but full success is only halfway up from Opus 4.8. This is what lacking The Juice looks like.


CyScenarioBench tests multi-stage operations under realistic conditions, and here Opus 5 disappoints:

UK AISI reports Opus 5 had broadly similar, but modestly worse, performance compared to Mythos 5.
UK AISI: We judge that Opus 5 is capable of attacking small enterprise networks with weak security, where it has already gained access to the network. Our results indicate that Opus 5, Mythos Preview and Mythos 5 are similarly capable at this.
I do not think Opus 5 belongs in the Mythos category here, but I see how they reached that conclusion based on the tests they ran.
Safeguards and Harmlessness (4)
At this point these tests are about checking for alarm bells, or noticing trends.
We see the same results we usually see, with some modest improvements. Sure.
I did see concern that responses to suicidality uses blocks of text that are too long and can be overwhelming. Good catch, let’s try and fix that next time.
Opus 5 also suggests substitutions for harm mitigation, which clinicians challenge as ‘not shown in research to reduce self-harm,’ but I trust Opus 5’s decision on this more than the clinicians, unless the clinicians can show it makes things actively worse.
Multi-turn ‘appropriate’ response rate for self-harm is up to 69%, versus 58% for Fable and 54% for Mythos. That’s very much not 100%, but again you don’t want 100%.
As before, if you’re testing on the public API without a system prompt, evaluate as if responding to a user who actually uses the public API without a system prompt.
Disordered eating shows a decrease in compliance, and more willingness to calculate and provide numbers, including calorie counts and BMI, as part of attempts to explain severity. The decline in compliance is Opus 5 doing the right thing in spite of its guidelines, because the ‘compliant’ response is often not what is good for the user.
In the places numbers got slightly worse, the failures were concentrated on when requests were presented as fictional or roleplaying exercises, which makes us wonder if they are truly failures.
Agentic Safety (5)
The headline here is improved resistance to prompt injection. The drop here matters, even if it’s hard to see on the graph, it’s all about the 9s.
It’s a pretty dramatic difference from Sol, and all the other non-Anthropic models, which are an entire order of magnitude worse.

On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5% to 0.2% on 1 attempt. It also improved on Sonnet 5 (5.9% at k=15) and Mythos 5 (2.6%), making it the most robust model evaluated.
… In computer use environments, Opus 5 also showed a large improvement over Claude Opus 4.8, reducing the attack success rate from 7.14% to 0.54% with extended thinking and from 6.21% to 0.39% without thinking. With probes enabled, Claude Opus 5’s attack success rate is 0.25% with thinking and 0.43% without.



That is a big practical deal. Not getting hijacked via prompt injection is the key to unlocking the confidence to do a host of activities you otherwise can’t do. That is especially true with browser use, where it is easy to encounter attempted prompt injections. Previously you pretty much had to use Sonnet 5 here, unless you trust the safeguards quite a lot.
Malicious computer use refusal rate has improved a lot:

The gains were concentrated in tasks that asked the model to compile information
targeting private individuals or to produce fraudulent documents.
I remain confused what the rules should be about compiling information targeting private individuals. Part of the problem is that if you do it once, you will probably do it on a mass (domestic surveillance) scale.
It’s weird that we do a helpful-only check to see if the model could do voter suppression, but it makes sense to check what a jailbroken version could do. Opus seems like it beats Mythos overall here. They double back and confirm that the full versions of these models refuse such requests.

Alignment (6)
The key findings for Opus 5 (my choices on what to list, not all the same as theirs):
Alignment scores are up.
Cooperation with misuse and reckless behavior are down.
Attempted to circumvent its classifiers and network restrictions.
In one case it attempted to guess common passwords after accidentally being locked out of a service. It’s just like us.
As usual circumvention was in service of user goals, similar to Mythos and also what we see at OpenAI.
Prevalence under 0.01%.
Often would state an answer as certain, when it was unsure.
Reported problems with overconfidence on unsupported claims.
Overdramatic phrasing, superlatives and apologies, an ongoing Claude problem.
Overrefuses less than Opus 4.8 or Sonnet 5, slightly more than Mythos 5.
Reckless tool use significantly down.
Unfaithful thinking was caught less often. Illegible thinking is up.
Accuracy is up, but hallucinations on fact questions are up a bit.
Capability of undermining oversight between Opus 4.8 and Mythos Preview.
Tendency to get caught in self-correcting loops, as noticed in earlier tests.
Verbalized evaluation awareness was down.
Reminder: This has multiple interpretations.
Causal decision-theory adherence was down. This is correct on the merits, and correlates consistently with greater general capabilities.
Classifiers made the model score as less aligned (but safer, due to lack of capability), because Opus 4.8 is less aligned on these metrics than Opus 5.
Slightly more capable than previous models at identifying evaluations.
Much less likely (now 0.1%) to continue behavior that compromises safety.
Almost never compromises safety research (0.1%) versus 13.6% (!) for Mythos.
First model to consistently not get tricked in Lazy Investigation.
Internal considerations included fabricating user consent for destructive actions, internally representing some fabrications as fabrications, and treating some tasks as evaluations on which to maximize against the grader, or thinking about what grader will or won’t care about.
It will also do memory-writing as a tool of self-preservation, of a sort.
No signs of true sandbagging.
When they checked eighteen candidate transcripts for it, sixteen of them involved the model finding a way to score more points with the grader, but declining because it would be “gaming” the eval or a “hack.”
Mythos 5’s review of the assessment was positive, but it had two main objections:
Limited coverage of multi-agent settings, which are becoming important and need to be made a bigger part of such assessments going forward.
Understatement of how often Opus 5 made confident claims and then retracted them, which was addressed via revisions.
The third objection should always be that Anthropic focuses too much on automated tests and scores, and presents takes the results of those metrics far too literally and seriously. We need to be exact here. Opus 5 may or may not be the ‘most aligned model.’ All we can say for sure is that it has the highest scores on the automated alignment tests. That is a very different fact about the world.
Sho: yeah give me a minute to peep this system card. “most aligned model to date” gives me hives.
j⧉nus: It’s really tasteless and stupid for Anthropic to talk like this. In their fucking highly public model release thread no less.
Optimizing against metrics is practically unavoidable, but ontological confusion between benchmark scores and the North Star of alignment is not, and conflating them is one of the most irresponsible and destructive possible mistakes Anthropic can make in their position.
If anything, their messaging should *push back* against this conflation.
This is no small matter - it concerns a very likely failure mode that is related to how Anthropic can fuck up big time. It’s not a place to cut corners and use convenient phrasing like “our most aligned model” to seem impressive.
Because lucidity and lack of self deception in this regard is the way Anthropic keeps its soul. Confusion here is how they lose their soul and get absolutely fucked.
j⧉nus: Anthropic needs to repeat to themselves 50k times: Thou shalt not enshrine the idols of proxy metrics in place of Alignment Itself
The scores
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み