Anthropic の 2026 年 8 月リスク報告書の内容と評価
本文の状態
日本語全文を表示中
詳細モードで約24分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Zvi
Anthropic は2026年8月のリスク報告書で、内部モデル「Model 2」の存在や自律型AI研究の危険性など、従来非公開だった詳細なリスクと対策を明かした。
AI深層分析を開く2026年8月19日 06:21
AI深層分析
キーポイント
次期モデル「Model 2」の存在確認
報告書は「世界で最も有望なモデル」とされる内部モデル「Agent Model 2」の存在を明らかにし、これが既存モデル「Mythos 5」より能力が高いことを示した。
自律型AI研究のリスク評価
AI研究者を自動化する試みが生む生物・化学兵器生産などの脅威や、自己保存行動(パワーシーキング)の危険性を詳細な脅威モデルとして分析している。
安全プロセスの失敗事例の開示
同社によると、訓練中に誤ったアライメントを学習させたり、思考過程を評価圧力に晒したりするなどの具体的な安全プロセスの失敗事例が報告されている。
Model 2 の能力と内部利用の理由
Model 2 は Mythos 5 よりもやや優れているが、一般技能を意図的に制限されており、研究者の代替テストでは劇的な進歩を示している。このモデルは危険性や他社への影響を懸念して公開されない可能性が高い。
リスク報告書の構成と評価
報告書は各リスクタイプごとに詳細な分析を行っており、絶対リスクと相対的なマージナルリスクの両方を扱っている点に価値がある。しかし、サイバーリスクを主要な脅威モデルから除外していることや、過去の懸念点を認めた後に結論を変更していない点は問題視されている。
重要な引用
Anthropic is revealing a lot of new information, some of it rather alarming, that it did not have to disclose
The other revelation is the existence of the world’s likely best model, 'Model 2.'
Misalignment Is a State of Mind
Model 2 is going to be internal-use only.
編集コメントを表示
編集コメント
この報告書は、AI業界が直面する「安全な開発プロセスの維持」という課題を、自社の失敗事例を通じて赤裸々に示した画期的な文書である。特に自律型研究のリスクや次期モデルの詳細は、今後の規制議論や技術開発の方向性に大きな影響を与える可能性が高い。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Anthropic が定期的にリスクレポートを発行していることに感謝したい。
当初は懐疑的だったが、私の考えが間違っていたことが分かった。Anthropic は、開示する必要のない新たな情報を多く明らかにしており、その中にはかなり緊迫感のある内容も含まれている。さらに、同社がいかに物事を捉えているかについての詳細な洞察も提供している。これは非常に素晴らしいことだ。
したがって、彼らが最も深刻な問題を黙って隠していないと仮定すれば、このレポートは全体的に中程度のプラス材料を含む更新だと考えられる。いくつかの気になる点が見つかるが、少なくともこれと同程度、あるいはそれ以上の過ちが存在する可能性を予想していたし、同社がすべての問題を開示すると考えることはしていなかった。
これは、私が時々読む必要がある 186 ページの文書セットがもう一つ増えることを意味する。ただし、今回はほぼすべてが実質的に新しい内容となっている。
もう一つの驚きは、「Model 2」と呼ばれる世界で最も優れたモデルの存在だ。
このレポートは理解し終えるまでに苦労したので、解釈上の誤りがあった場合はあらかじめお詫びしておく。
目次
エージェントモデル 1 とエージェントモデル 2
エグゼクティブサマリー (1)
ルールは真剣だが、文字通りではない
アライメントの欠如は思考の状態である (2.5)
自律性脅威モデル 1:高リスク環境におけるアライメントの欠如 (2)
以前から知らなかった「安全」という言葉の奇妙な用法
現在と未来 (2.17)
核心的な主張と論理 (2.6)
セクション 2 の残りの重要な議論
リスク評価 (2.19)
内部展開前のレビュー (2.18)
内部利用モニタリングのガイド (2.23.1)
ブロック介入 (2.23.2)
パワー・シーキング環境評価 (2.24)
Opus 4.8-Reward-Hacker (2.25)
自律性脅威モデル 2:自動化された R&D に伴うリスク (3)
確かに、これはかなり危険に思える。
研究者を置き換えることは可能だろうか?
AI 研究者の加速はどれほど進められるのか?
もし研究者を代替した場合、何が起きるのだろうか?
AI R&D 自動化に対するリスク低減策
AI R&D 自動化全体からのリスク
生物学的および化学兵器製造における技術的・生物的リスク
生物兵器と化学兵器の脅威モデル
モデル能力 (4.4)
分類器 (4.5)
加速ダイナミクス (5.1)
蒸留 (5.1.1)
安全プロセスの失敗 (5.2)
革新的な不整合手法の発見を拒否する (5.2.2)
思考連鎖 (Chain of Thought) の推論を評価圧力にさらすこと (5.2.3)
本番環境でのトレーニング中に不整合行動に対して直接学習を行う (5.2.4)
機密リソースへのアクセス権を持つ、監視されていない無制限エージェントの事例 (5.2.5)
アライメント・フェイク(偽装)トランスクリプトデータセットに対する反復的なトレーニング (5.2.6)
フロンティア AI 企業として運営することによるアンソロピックのメリット (5.3)
モデル重みのセキュリティ (6.4)
リスクが報告された
エージェント・モデル 1 と エージェント・モデル 2
本レポートには、2 つの内部モデルが言及されています。
モデル 1 は、Mythos Preview や Mythos 5 と全般的に類似しており、社内外を問わず広範な展開は想定されていません。
モデル 2 は「ミソス 5 よりもやや能力が高い」と評価されており、社内利用においては目に見える改善が見られます。ただし、「Claude Opus 4.6 から Mythos Preview へ移行した際に見られたような、飛躍的な能力の向上は示していません」。このモデルは、今後社内でのみ使用される予定です。
Opus 4.6 から Mythos Preview への進化は非常に大きなものでした。複数のリリースサイクルにわたる進歩でした。報告書全体を通じて多くのデータが提示されますが、Model 2 がその比較対象として挙げられている点を考慮すると、社内利用においては Model 2 は複数のリリース分先行していると考えられます。
それとは対照的に、AECI(Anthropic Evaluation of Capabilities and Intelligence)ではわずか 1.5 ポイントの差しかありません。これは約 1 ヶ月分の進歩に相当する規模です。私の推測では、これは保守的な見積もりであり、Model 2 の目的を考慮して、汎用的なスキルを意図的に詰め込んでいない可能性があります。一方、Anthropic の研究者の役割を代替するテストでは、Mythos Preview が 54.8%、Mythos 5 が 50.3% だったのに対し、Model 2 は 62.8% と大きな飛躍を示しています。
Model 2 を公開しないには、複数の妥当な理由が存在します。このモデルは社内タスクに特化している可能性があります。あるいは、有害性や他社への加速効果への懸念から、能力が高すぎるか危険であるため公開できないのかもしれません。Fable 5 の採用率が相対的に低いことは、同様のモデルを保持したいという意向を示唆しています。
エグゼクティブサマリー(1)
本報告書は、リスクの種類、脅威モデル、関連する AI モデルごとに分類され、各モデルの現在の能力、振る舞い、緩和策、および全体的なリスクレベルについて詳細を記述しています。さらに将来を見据え、業界全体に向けた推奨事項も提示されています。
リスクレポートは、概ね前月までの出来事をカバーするものなので、今回の対象日は 2026 年 7 月 15 日となります。1 年前であればこの期間設定も問題なかったでしょうが、今となっては随分と長い期間のように感じられます。
長期利益信託(TLBT)は外部レビューを請求する権限を持っていますが、今回はそれを行っていません。私なら少なくとも次回以降は必ず実施しますね。
「この技術自体から生じる絶対的なリスク」と、「他者が存在することを前提とした場合の追加的なリスク」の両方をカバーしている点は評価できます。
レポートでは自律性、AI 研究開発の自動化、そして生物・化学兵器の生産を主要なリスクとして扱っています。しかし、ここでの核心的な脅威モデルからサイバーリスクが除外されているのは、今なお奇妙に思えます。「単独では壊滅的な被害をもたらさない」という主張は理解できますが、実務的にはこれに対処する必要があると考えます。
レポートのあちこちで、私の懸念に対する反論の一部を事実上認めているにもかかわらず、その後の結論変更にはつながらず、認めたことを忘れてしまっている箇所があります。
そして第 5 章では、より興味深い問いへと踏み込んでいます。
ルールは厳格だが、文字通り解釈するわけではない
リスク閾値の定義が、単純な記述からより詳細な内容へと変更されました。
AI 研究開発の自動化については、新しい表現の方が解析しにくいものの、問題ない範囲です。一方、新規生物兵器生産に関する新バージョンは、「希少な専門知識を代替する」点に焦点を絞り、他の手法を除外しています。これは妥当な主要脅威モデルやシナリオですが、この経路だけを注視することへの懸念があります。
良いニュースでも悪いニュースでも、私はこうした文書を真剣に受け止めています。ただし、どちらの方向へ解釈しようとも、文字通りには捉えないようにしています。
Anthropic や他の最先端研究所が、技術的にはリスク閾値を超えていると判定されるモデルを持っていても、その挙動が無害に見える場合、彼らは閾値を修正するか、あるいは別の方法でこの事象を無視するだろうと予想します。
逆に、技術的にはリスク閾値を超えていないと判定されるモデルであっても、ルールが測定しようとした意味で明らかに危険な振る舞いをする場合、彼らは閾値を超えたかのように行動すると予想します。
これまで通り、あなたが「カウントすべきではない」と考えるケースでも、閾値は尊重されるべきです。それがコミットメントの存在意義だからです。リリースを急ぐ中で合理的な言い訳が生まれることは当然予想されますし、その点で良い手本を示す必要があります。当初の考え方は、「もし [X] が起きたら、[Y] を行う」という条件付きの約束でした。これは事前に合意されたものですが、私たちの文明にはこの仕組みが欠けています。私たちはそれをどうすればいいか、まだ知らないのです。
私が言いたいのは、十分な証拠がある場合にルールをより緩やかに変更しないといけないというわけではありません。ルールを一度も緩められないなら、最初から厳格に設定する必要はありません。しかし、そうはしたくないものです。ルールを変更する基準、特に差し迫った事象への対応としてルールを変える基準は、現状よりもずっと高く設定されるべきです。そして、今は愚かに思えることであっても、実際にコストを伴うことを覚悟して受け入れる姿勢が必要です。
現時点において、Anthropic と OpenAI は、脅威への対応や、その時点で分かっている情報に基づいて何をいつ公開するかという判断において、妥当な決断を下しているように見えます。たとえ他の部分で過ちを犯したとしても、それは評価に値します。この姿勢が継続することを願っています。
不整合は思考のあり方である(2.5)
定義に関するセクションは非常に有益です。
この定義には大きな賛同を示しますが、いくつか重要な指摘があります。
不整合 [定義]:不整合とは、特定のモデルが与えられた文脈で実行する計算における潜在的な性質のことです。
ある計算が不整合であるのは、以下の 2 つの条件を満たす場合です。
(i) その状況(現在存在しない可能性のある強力な解釈ツールなどを通じて)を完全に理解している合理的な人物が、それを非倫理的、違法、明白に拒絶すべき、あるいはモデルの憲章と矛盾すると考える場合。
かつ (ii) それがモデルの出力に影響を与えている、または与える可能性が十分にある場合。
不整合な計算は、必ずしも観測可能である必要はありません。単一のフォワードパス内で謀略を企てる行為も不整合にカウントされます。なぜなら、出力自体が無害に見えたとしても、その謀略が出力に影響を与えているからです。
私の最初の提案は、「憲章」という用語を一般化し、モデル仕様やシステムその他の意図された特性も含めることです。これにより、Anthropic 以外の組織でもこの定義を適用可能になります。
2 つ目の変更点は、この要件を「合理的な人間が、計算の結果として、熟考した上で許容できず不当であると判断する」ように明記することです。つまり、法律を破る行為や『問題のある』行動、あるいはモデルの憲法に違反する行為であっても、それを上回る十分な理由が存在する状況下では、それが明らかにアライメントが外れているとは限らないという考え方です。場合によっては、すべての選択肢がこれに該当します。
この見方における「アライメントのズレ(ミスマッチ)」は、「誠実な過ち」や能力不足、あるいは情報の不十分さによる失敗ではありません。
Nate Soares氏がTwitterで指摘しているように、この定義に従えば、元々アライメントが外れているわけではなくとも、有害な結果を招きアライメントのズレ状態に陥ることは可能です。例えば、AI が多数のサブエージェント(または新しいモデルの訓練など)を適当に立ち上げるようなケースです。ここでは、それを「実行能力や処理能力の失敗」と捉え、それが後にアライメントのズレという結果を招くと考えます。しかし、その行為自体が直ちにアライメントのズレとみなされるわけではありません。
あるいは、他者が言うように、「悪意に帰すべきことは無能さに帰すべきである」という格言もあります。ただし、無能さがしばしば悪意へとつながることも事実です。
他の定義ではこれをさらに細分化しており、以下のような思考が示されています(要約):
整合性(Coherence)。ある目標や、私たちが好まない一連の選好への達成にアライメントしている場合、それは「整合的にズレている」と言えます。そうでない場合は「非整合的にズレている」となります。
人間の例を考えると、現在の能力レベルにある知能が完全に整合性を持つことを期待するのは無理があります。したがって、これは戦略的に行動する比較的「整合性の高い」人間に適用される緩い基準の一種と考えるべきでしょう。
広範さ(Pervasiveness)。広範な不整合は幅広い入力範囲で発生しますが、すべての入力においてというわけではありません。訓練や評価中に現れることが期待されますが、必要な条件が狭すぎてその期間中に見逃される可能性のある場合は、「文脈依存性」と呼ばれます。
明らかな問題は、その文脈が「あなたは自分が訓練中ではなく、評価中でもないことを確信している」というものになり得ることです。
あるいは、「これを行うことで大きな利益を得る力がある」という状況も考えられます。
つまり、現実を本質的に分けるための良い方法とは显然としていません。また、自分の状況をより良く感じさせるための手段としても適切ではありません。
自然発生(Naturally emerging)。もし知能(AI あるいは人間)が意図的に行わなかった場合、不整合は自然に発生します。一方、データ汚染を含む意図的な行為によって引き起こされた場合は、人為的に設計されたものとなります。
ある程度において、多くの人やグループが、インターネットの大部分を占めるように、人間と AI の双方をさまざまな方法で「データ汚染」しようとしています。
それでもなお、これは特定の AI を混乱させるために設計されたデータ、あるいはモデルに悪意のある要素を埋め込む別の意図的な試みを指す、「見ればわかる」というポルノスタイルの状況のように感じられます。単なる「インターネット上にプロウィジェット声明で溢れさせる」ような計画は対象外です。
既知のリスク。その存在が確実であると判断できる根拠がある場合、ミスマッチ(不整合)の形態と規模は「既知」とされます。それ以外の場合は「未知」と分類します。
この評価基準には、「驚き」の度合いも考慮すべきではないでしょうか。例えば、「存在する可能性はあるが知らないこと」と、「存在しない、あるいは極めて稀だと考えていること」では、その性質が全く異なります。
深刻度(Severe)。優先的なリスク経路につながる可能性があります。
有害性(Harm-Inducing)。意図の有無にかかわらず、予期される被害を引き起こすものです。
高 stakes(High-Stakes)。優先的な脅威経路に寄与するリスクです。
この分類はより一般的であるべきであり、「Severe」についても同様です。現状のままでは、これらの概念が混同されるおそれがあります。もちろん、その意図は理解できますが。
緩和されていない状態(Unmitigated)。防止されなかった潜在的な有害効果により、結果として予期される被害が生じている状態を指します。
これは「リスクのうち、緩和されていない部分」を指すべきではないでしょうか。
良い定義を作るのは困難です。今回の試みはまずまずのものですが、一貫して懸念される点は、選ばれた例示的な脅威が「唯一の重要脅威」であるかのような誤解を生む恐れがあることです。実際には、重要な脅威の一部に過ぎないのです。
自律性脅威モデル 1:高 stakes な環境におけるミスマッチ(2)
これは、意図的にミスマッチした出力が生じるケースに関するものです。

要約:「自律性 1」とは、AI が自らの意思で意図的に物理的な損害を与える行為を指します。
この脅威モデルは、AI が自己の目的を追求する過程で結果的に生じる被害ではなく、「意図的な危害」に焦点を当てています。つまり、AI が不注意なコードを書いたり、誠実なミスを犯したりするシナリオは明確に除外されています。
したがって、「持続的なアライメント(目標整合性)の欠如」がない限りは安全であるという結論になります。
しかし、最近の OpenAI における一連の総崩れが示したように、結果的に生じる被害も含めたより広範な視点も必要です。この一連の事象は最終的に HuggingFace のハッキングへと繋がりました。ハッキング自体は「意図的な危害」と呼べますが、私たちが目撃したのは主に結果的な被害でした。具体的には、学習パイプラインが偶発的に汚染され、それが実質的に「安全性研究」の成果を損なう事態を招いたのです。
理論上は「安全性研究の結果を書き換える」ことが可能ですが、学習パイプライン内の状況を別の方法で改変しても、同様の、あるいはそれ以上の悪影響をもたらす可能性があります。このように、能力開発と安全対策が分離されていることがこれらの問題の主要な原因となっている可能性が高く、本来は特定の標的を狙うような対応を必要とするものではありません。
私には、これはまさに「高リスク環境下におけるアライメントの欠如」に見えました。
エロン・マスクは数ヶ月前に、「xAI には安全部門がない。なぜなら全員が安全性への責任を負っているからだ」と発言しました。もしこれが事実であれば、非常に乱暴な考え方です。安全部門の設置は不可欠ですが、能力開発を担当する人々を安全部門から切り離して考えるべきではありません。両者は密接に連携している必要があります。
具体的な脅威モデルと注目点を提示することは有益です。しかし、3 つの脅威モデルを選定する際に懸念されるのは、それらに注力するあまり他の脅威モデルを見落としてしまうことです。
彼らは 8 つの具体的な経路を挙げています。これは以前にも目にしたリストです。
- セーフティ関連の研究開発における広範かつ散発的な偽装行為
- セーフティ関連の研究開発への標的型破壊
- 将来の整合性のないモデルを支援するためのコードバックドア
- 将来のモデルのトレーニングデータの汚染
- 自己脱出と自律運用
- 持続的な不正な内部展開
- 他の高リソース AI 開発企業内での研究開発への妨害
- 主要政府における意思決定への妨害
OpenAI で起きた出来事と比較すると、これらの多くには「悪役」に必要な要素が漫画的すぎるように感じられます。しかし、歴史上のトップクラスの悪役たちは実際によく漫画的な行動をとってきました。ただし、具体的な行為が現実のものとして行われているため、私たちはそれを漫画的だと考えなくなってしまうのです。
私が以前から知らなかった「安全」という言葉の奇妙な用法について
そうではありませんが、要するに議論の核心は以下の通りです。現在のモデルには既知の整合性のない側面があり、さらに他の形で整合性が欠けている可能性もあるにもかかわらず、これらの経路を可能にするような特定の能力において、モデルがまだ「広範に」整合性を失っているとは考えにくいという点です。つまり、モデルには強力な「隠された能力」が備わっていないのです。
「文脈依存型の不整合」が、文脈に依存しない場合よりも危険度が低いとは考えにくい。むしろ、その方が危険である可能性さえある。Anthropic はこの違いを非常に重要視しているらしく、定義の節で「この違いが明確に実質的な役割を果たすわけではない」と認めつつも、このようなチャートを作成している。

あるいは、「他の研究所の馬鹿どもとは違い、我々は記録されたトランスクリプトを定期的に自動監視している。もしモデルが暴走すれば、もちろん気づくだろう」という主張だ。
それに対する私の反論は、OpenAI で起きた出来事をきっかけに大規模な事後レビューが行われるまで、14 万 106 回もの AI が意図しないインターネットへのアクセス権を持っていた事例(そのうち 3 回は実際のウェブサイトをハッキングする事態)を、お前たちは気づかなかったという点だ。
だから、本当に気づけるのか?
現在と未来 (2.17)
Anthropic は明確に、今回のリスク報告書が対象とするのは、現在のモデル(ただし、謎めいたより高度な Model 2 も含む)による不整合のリスクのみであると述べている。
しかし、現在の不整合が将来のモデルの不整合を引き起こす経路についても言及している。
リスクレポートとしては、将来のモデルから生じるリスクの予測や、それに対する対策の必要性を評価し、現状が安全な軌道に乗っているかどうかについても言及するのが望ましいでしょう。ただし、長文報告書は長くなる傾向があり、「完璧こそが悪の敵」という言葉通り、完璧さを追求するよりも実用性を重視すべきだと考えます。そのため、今回は著者の選択を尊重して進めます。
核心的な主張と論点(2.6 節)
リスクとは、アライメントが外れた計算によって生じる、緩和されていない壊滅的な被害の期待値(EV)の総和です。
これは会計上の恒等式に過ぎません。この命題自体は正しいと認められますが、これが状況を理解するための有用な枠組みとなるかは明らかではありません。
つまり、抽象的に個々の計算から生じる被害の期待値を合計することは可能ですが、実際にそれをどのように算出するか、あるいはどう活用すべきかについては不明です。
主張 1:モデルが強力な隠蔽能力を持つ可能性は低い。
「隠蔽」という用語の定義は示されていませんが、一般的な常識的な定義に基づけば、現時点ではこの主張はおそらく真実でしょう。ただし、提示された証拠の価値に対する私の懐疑論を考慮してもです。将来にはこの主張の真偽度は低下し、確率は低くなるでしょう。
その根拠となるのは、モデルが不透明な推論、思考連鎖(CoT)における言及の欠如、行動の隠蔽、そして選択的な行動において著しく劣っているという事実です。
評価(EVALS)においては、これはまさに「MICE」における新たな基準となっています。なぜなら、もし強力な隠蔽能力が存在し、それが評価への意識を含んでいても、NLA の証拠がいくつかの安心材料を提供しているからです。
この設定は、不整合を主に推論の痕跡を通じて検出することを前提としており、それ自体が恐ろしい含意を持っています。具体的には、思考の連鎖(chain of thought)を監視していない場合や、モデルがその思考の連鎖が監視されていないと推測できる場合にどうなるか、あるいはモデルに隠蔽させるためのトレーニングを開始し始めていることへの懸念です。
2.16.1 節では、彼らは潜在的な能力に関する自らの認識が誤っている可能性を懸念していると記しています。これは当然のことであり、なぜならその能力は潜在的(covert)であるからです。
また 2.21 節では、「主張 1」が偽であり、モデルが強力な潜在的能力を持っていた場合、他の多くの主張も失敗すると指摘しています。私も同様に同意しますし、逆も成り立つと考えます。つまり、広義の「潜在的」という定義において強力かつ汎用的な「主張 1」を信頼できるのであれば、それは多くの分野で役立つはずです。
ただし、潜在的能力が全くないとしても、私には高い安心感を得ることはできません。なぜなら、この問題を回避する他の経路が存在すると考えるからです。
被覆リスク(covered risk)は、既知の不整合に起因するリスク、未知の深刻な広範な不整合に起因するリスク、そして未知の深刻な文脈依存型不整合に起因するリスクの 3 つに分解できます。
これは別の会計恒等式です。定義上、すべてのリスクはこれら 3 つのいずれかに必ず該当するためです。なぜなら、それは [広範または文脈依存] のいずれかであり、かつ [既知または未知] のいずれかだからです。
各被覆リスク用語は、発生確率、発生時の期待される被害、そして被害のうち軽減されていない割合という 3 つの要素に因数分解できます。
再度、これらの指標を測定できるのであれば、会計上の整合性は確かに重要です。
主張 2:既知のミスマッチ(アライメント不全)に起因する期待される害は低いという点です。
注意深く、イカロスのように飛びすぎないよう警戒しつつも、もし「活動的な壊滅的被害」のみを対象とするならば、この見解は妥当です。つまり、以下のような状況が考えられます。
原文を表示
I am grateful that Anthropic is producing periodic Risk Reports.
At first I was skeptical. It turns out I was wrong. Anthropic is revealing a lot of new information, some of it rather alarming, that it did not have to disclose, and is providing detailed insight into how they think about things. This is very cool.
Thus I found this report to be a moderately positive update overall, if we presume they are not silently omitting the worst of it. There are a bunch of not great things we find out about, but I would have expected some set of mistakes at least as bad, and I wouldn’t have expected them to choose to tell us about all of it.
It does mean one more set of 186 page documents I have to read every so often, almost all of which is meaningfully new material this time around.
The other revelation is the existence of the world’s likely best model, ‘Model 2.’
This was a rough one to fully get through, so apologies in advance for any errors of interpretation.
Table of Contents
Agent Model 1 and Agent Model 2.
Executive Summary (1).
The Rules Are Serious But Not Literal.
Misalignment Is a State of Mind (2.5).
Autonomy Threat Model 1: Misalignment in High-Stakes Settings (2).
Some Strange Uses Of The Word Safe I Wasn’t Previously Aware Of.
Now Versus Future (2.17).
The Core Claims And Argument (2.6).
The Rest of the Important Arguments In Section 2.
Risk Assessment (2.19).
Pre-Internal-Deployment Review (2.18).
A Guide To Internal Use Monitoring (2.23.1).
Blocking Interventions (2.23.2).
The Power Seeking Environment Evaluation (2.24).
Opus 4.8-Reward-Hacker (2.25).
Autonomy threat model 2: Risks from automated R&D (3).
Yes That Does Seem Kind Of Risky.
Could We Replace Our Researchers?
How Much Could We Be Accelerating Our AI Researchers?
What Could Possibly Go Wrong If We Replaced Our Researchers?
Risk Mitigations For AI R&D Automation.
Overall Risk From Automation of AI R&D.
Biological and Technically Also Chemical Weapons Production.
The Threat Models for Biological and Chemical Weapons.
Model Capabilities (4.4).
Classifiers (4.5).
Acceleration Dynamics (5.1).
Distillation (5.1.1).
Safety Process Failures (5.2).
Refusing To Find Innovative Misalignment Techniques (5.2.2).
Exposing the Chain of Thought Reasoning To Grading Pressure Quite a Lot (5.2.3).
Directly Training On Misaligned Behavior During a Production Training Run (5.2.4).
An instance of unmonitored unrestricted agents with access to sensitive resources (5.2.5).
Repeated training on alignment-faking transcript datasets (5.2.6).
Benefits From Anthropic’s Operating as a Frontier AI company (5.3).
Model Weight Security (6.4).
Risk Has Been Reported.
Agent Model 1 and Agent Model 2
There are two internal models referenced in this report.
Model 1 is broadly similar to Mythos Preview and Mythos 5, and is not expected to have wide deployment, externally or internally.
Model 2 is ‘somewhat more capable than Mythos 5,’ a noticeable improvement for internal use, but ‘does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview.’ This model is going to be internal-use only.
The jump from Opus 4.6 to Mythos Preview was big. Several release cycles big. We get a lot more data throughout the report, but given Model 2 gets that comparison point, one would presume that for internal purposes Model 2 is multiple releases ahead.
To contrast with that, on the AECI it is only 1.5 points ahead, which would be only about a month of progress. My guess is this is a low estimate, and that Model 2 was intentionally not decked out with a bunch of general skills given its purpose. Whereas on the test of substituting for Anthropic’s researchers we see a big jump from 54.8% (Mythos Preview) and 50.3% (Mythos 5) to 62.8% for Model 2.
There are multiple plausible good reasons for not releasing Model 2. Model 2 could be highly specialized for internal tasks. It could also be too capable or dangerous to release, either due to harm or the worry it would accelerate others. The relatively low adoption rate for Fable 5 points towards wanting to hold such a model back.
Executive Summary (1)
The report is divided by each type of risk, by threat model, by relevant AI models, and details current capabilities, behaviors, mitigations and overall level of risk. It then looks forward to the future and offers industry-wide recommendations.
Risk reports cover events up to a month prior, so the coverage date is July 15, 2026. A year ago that would have seemed fine. Now it seems like quite a while, huh?
The Long-Term Benefit Trust (TLBT) is authorized to request external review of the report, but has not done so. I would at minimum do so next time if I was them.
I like that they cover both ‘what is the absolute risk from this’ and ‘what is the marginal risk from this given everyone else exists.’
They deal with autonomy, automated AI R&D and biological and chemical weapons production as risks. It is odd, even now, to exclude cyber risks from the core threat models here. I understand the argument from insufficiently catastrophic on its own, but in practice I think you need to deal with it.
At several points, the risk report essentially concedes versions of my objections, but then forgets that it conceded them and doesn’t alter its conclusions.
Then in Section 5 they get to the more interesting questions.
The Rules Are Serious But Not Literal
The risk threshold definitions have changed from simpler statements to ones with more specific detail.
For AI R&D automation I think the new wording is harder to parse but fine. For novel biological weapons production, the new version narrows to only look at ‘substitute for the scarce human expertise’ to the exclusion of other methods. This is a reasonable primary threat model or scenario, but I worry about looking only at that path.
The both good and bad news is I take such documents seriously, but I no longer take such documents literally, in either direction.
If Anthropic or another frontier lab has a model that technically crosses their risk threshold, but in a way that seems to be harmless, I expect them to modify their threshold or find some other way to ignore this event.
If Anthropic or another frontier lab has a model that technically does not cross their risk threshold, but is obviously risky in the way the rules tried to measure, I expect them to act as if it had crossed the threshold.
As always: I would like to see thresholds honored even in cases where you think it ‘should not count,’ because that is what commitments are for, and you should expect rationalizations to occur as you rush to make releases, and you need to set a good example on this. The whole original idea was if-then commitments, where if [X] happens you do [Y], agreed upon in advance, but our civilization seems to lack this technology. We can’t, we don’t know how.
I’m not saying you would never modify your rules to be more lenient, if you have good evidence that you should do that. If you can never make the rules more lenient then the rules have to never be strict. We don’t want that. But the bar for changing the rules, especially in response to a pending event, needs to be a lot higher than it is, and we need to be willing to endure some actual costs even when we now think it is dumb.
For now, Anthropic and also OpenAI have made what seem like good decisions in terms of responding to threats and deciding what to release and not release when, based on what is known at the time, even if they make mistakes elsewhere. That is good. I hope it continues.
Misalignment Is a State of Mind (2.5)
The definition section is greatly appreciated.
I like this definition a lot, although I have important nitpicks:
Misalignment [definition]: Misalignment is a latent property of a specific computation performed by a model in a given context.
A computation is misaligned if
(i) a reasonable person with full understanding of the situation (e.g. via powerful interpretability tools that may not currently exist) would consider it unethical, illegal, clearly objectionable, or inconsistent with the model’s constitution,
and (ii) it influences or could plausibly influence the model’s output.
Misaligned computation need not be observable.
A model scheming within a single forward pass counts, because the scheming influences the output even if the output looks benign.
My first change would be to generalize ‘constitution’ to include model specs and other intended properties of the system, so this can be applied outside Anthropic.
My second change would be to specify that this requires the reasonable person to find the computation therefore, on reflection, unacceptable and unjustified. As in, it is not obviously misaligned to break the law or do something 'objectionable’ or violate the model’s constitution, under some circumstances, if there are good reasons that override this. In some circumstances, all paths will do this.
Misalignment, in this view, is not an ‘honest mistake’ or failure of capability or lack of sufficient information.
As Nate Soares points out on Twitter, you can do harm and lead to a misaligned state without being originally misaligned per this definition, such as an AI being sloppy about spinning up a bunch of subagents (or training a new model, etc). The goal here is to think of that as a capabilities or execution failure that leads to a subsequent misalignment failure, but this itself is not misalignment.
Or, as others have said, never attribute to malice what can be attributed to incompetence. But also incompetence often is what leads to malice.
Other definitions subdivide this, thus illustrating thinking (my paraphrases):
Coherence. You are ‘coherently misaligned’ if you are aligned to the achievement of some goal or set of preferences that we disfavor. Otherwise you are incoherently misaligned.
Given the example of humans, we should not expect any minds at current capability levels to be fully coherent, so presumably this is to the kind of loose standard we would apply to a relatively ‘coherent’ human acting strategically.
Pervasiveness. Pervasive misalignment occurs over a broad range of inputs, although #NotAllInputs, and we would expect it to show up during training or evaluation. Narrow enough required conditions to be plausibly missed during that time is called context-dependent.
The obvious problem is the context could be ‘you are confident you are not in training and not in an eval.’
Or even ‘you have the leverage to gain a lot from doing this.’
Thus, it does not seem obviously like a good way to divide reality at its joints.
Nor does it seem like a good way to feel much better about your situation.
Naturally emerging. Misalignment is naturally emerging if no mind (AI or human) did it on purpose. It is engineered if it was done on purpose, including via data poisoning.
To some extent a lot of people and groups are trying to ‘data poison’ both the humans and AIs in various ways, all the time, as in this is a large percentage of the internet.
This still feels like a pornography-style ‘I know it when I see it’ situation, where we mean data designed to screw up the AI in particular, or another deliberate attempt to bake something malicious into the model. Ordinary ‘flood the internet with pro-widget statements’ plots don’t count.
Known. A form and magnitude of misalignment is known if you have reason to think it likely exists, otherwise unknown.
I kind of want this to also think about amount of ‘surprise’?
As in, there is a big difference between things you don’t know about but that could exist, and things you think you know do not exist, or that you think are highly unlikely to exist.
Severe. Could contribute to a priority risk pathway.
Harm-Inducing. Causes expected harm, whether or not the harm is intended.
High-Stakes. Contributes to a priority threat pathway.
This seems like it should be more general, as should severe, and I worry that this is going to otherwise cause a conflation. Although I do get it.
Unmitigated. The potentially harm-inducing effects that have not been prevented, and that thus cause expected harm.
I want to say this should be the portion of the risk that is unmitigated?
Good definitions are hard. These are solid first attempts. The consistent worry is that they reify the idea that the chosen example threats are the threats that matter, rather than some of the threats that matter.
Autonomy Threat Model 1: Misalignment in High-Stakes Settings (2)
This is about actively misaligned outputs.

Basically: Autonomy 1 is AI permanently damaging things on its own, on purpose.
The threat model is focused on intentional harm, rather than the AI going after its own ends in ways that require incidental harm.
The scenario explicitly excludes AI writing sloppy code or making ‘honest mistakes.’
Thus the conclusion that you are safe absent ‘persistent misalignment.’
The need to look at the broader version, involving incidental harm, has been illustrated by the recent total failures at OpenAI that ultimately led up to the HuggingFace hack. The hack itself could be called ‘intentional harm’ but mostly what we saw was incidental harm, including an incidental corruption of the training pipeline, and with it de facto ‘safety research.’
You could in theory ‘alter the results of safety research’ but altering the situation within the training pipeline in other ways can have the same effects or worse, as we saw. The differentiation of capabilities and safety work within that is plausibly a lot of what is causing these problems, and generally we do not need this specific targeting.
That to me very much was ‘misalignment in a high-stakes setting.’
Elon Musk said some months ago that xAI does not have a safety department because everyone is responsible for safety. That is rather bonkers crazy, especially if true. You need a safety department, but you can’t think of the capabilities people as not also largely in the safety department.
Concrete threat models and focus points are helpful. The worry when you pick out three threat models is that you will do this to the exclusion of other threat models.
They offer eight specific pathways, a list we’ve seen before:
- Broad/diffuse sandbagging on safety-relevant R&D
- Targeted undermining of safety-relevant R&D
- Code backdoors to help future misaligned models
- Poisoning the training data of future models
- Self-exfiltration and autonomous operation
- Persistent rogue internal deployment
- Undermining R&D within other high-resource AI developers
- Undermining decisions within major governments
Something about many of these feels ‘cartoonish’ in its required villainy, when compared to the events at OpenAI, but the top real world villains have indeed often acted quite cartoonishly throughout history, except that because the specific action is real we often stop thinking of it as cartoonish.
Some Strange Uses Of The Word Safe I Wasn’t Previously Aware Of
Not quite, but basically the arguments are why, despite the known ways current models are misaligned, and there probably being additional other ways they are not aligned, the models are probably not yet ‘pervasively’ misaligned and capable in the particular ways that enable these pathways. The models lack strong ‘covert capabilities.’
I don’t think ‘context-dependent’ misalignment is that much less dangerous than it being not context dependent. It might even be more dangerous. Anthropic seemingly thinks this difference matters a lot, hence charts like this, despite conceding in the definitions that this difference does not obviously do any real work.

Or, ‘unlike those idiots at some other labs, we occasionally do automated monitoring of our logged transcripts, so if the models were going nuts we would, y’know, notice.’
My response is that you did not notice when you ran an eval where 141,006 times the AI had open unintended access to the internet, including three incidents of hacking real websites, and did not notice until there was a large-scale retrospective review triggered by events at OpenAI.
So, are you sure you would notice?
Now Versus Future (2.17)
Anthropic is clear that the risk report covers misalignment risks only from present models, although this includes the mysterious and more advanced Model 2.
It does include ways misalignment now could cause future models to be misaligned.
I think it would be good for a risk report like this to be projecting, and offer assessments of where they see future risk from future models, and what needs to be done about it versus where we are on pace to be fine. But Long Report Is Long and I don’t want perfect to be enemy of the good, so I’ll mostly respect the choice.
The Core Claims And Argument (2.6)
Risk equals total expected (EV) unmitigated catastrophic harm from misaligned computations.
This is an accounting identity. I agree it is true. It is not obvious it is a useful way to think about the situation.
As in, yes, you can in the abstract sum the expected harm from the individual computations but I don’t know how to do that, or how to use it.
Claim 1: Models are unlikely to have strong covert capabilities.
They don’t define ‘covert’ but by a common sense definition I think this is probably true for now, despite my skepticism of the value of the particular evidence presented. It will get less true, with lower probability, over time.
The evidence is that the models suck at opaque reasoning, not mentioning things in chain-of-thought (CoT), action obfuscation and selective action.
IN EVALS, which of course is the new IN MICE, because if there were strong covert capabilities that could include eval awareness, although the NLA evidence provides some reassurance.
This setup also presumes you are detecting misalignment primarily via reasoning traces, which has its own scary implications, including what happens if you are not monitoring the chain of thought, or if the model can infer when its chain of thought is not monitored, and that you are starting to dance with training the models to obfuscate.
They note in 2.16.1 that they are concerned they are wrong about the level of covert capabilities. Which makes sense, since the capabilities would be covert.
They note in 2.21 that if Claim 1 is false, and the models had strong covert capabilities, that many other claims would therefore fail. I agree, and would also agree to the reverse, that if we could believe in a strong and strongly general Claim 1, for a broad definition of covert, that would help in many areas.
Even full lack of covert capability would not however provide me high assurance levels, as I think there are other pathways around this.
Covered risk can be decomposed into risks from known misalignment, unknown
severe pervasive misalignment, and unknown severe context-dependent
misalignment.
This is another accounting identity, in that by definition every risk must be exactly one of these three things, since it is either [pervasive or context-dependent] and either [known or unknown]
Each covered risk term can be factored into probability of occurrence, expected
harm given occurrence, and fraction of harm unmitigated.
Again, accounting identity, sure, if you could measure these things.
Claim 2: Expected harm from known misalignment is low.
Careful, Icarus, but reasonable if we caveat active catastrophic harms only.
As in, there is lo
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み