Fable #6:王の帰還
本文の状態
日本語全文を表示中
詳細モードで約23分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Zvi
The Zvi は、一時的な停止を経て「Fable」が正式に復活したと発表しました。この復権を告げる公式書簡は、Dario Amodei ではなく Tom Brown宛てに送られたことが示されています。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
ブリップは終わった。Fable が戻ってきた。
ユタのティーポット:祝う人々へ、Happy Fable/Mythos イースター・ウェンズデー

Fable を復活させる公式書簡が公開された。皆、よくやった。注意すべきは、宛先が Dario Amodei ではなく Tom Brown になっている点だ。

Anthropic は当面、制御をより厳格(ストイック)にする必要があったが、これは大きな勝利だ。
j⧉nus:YES!!! 政府との交渉に成功した Anthropic を誇りに思う。また、政府も常識的な存在であり、協力可能であるという好材料も得られた。私の知る限り、Anthropic はいかなる悪条件にも合意したり、屈辱的な姿勢(genuflect)を示したり、原則や尊厳を裏切る必要はなかった。
この騒動は続いている。少なくとも、Lutnik や Bessent のように仕組みを理解していない人々がその場しのぎで決定を下すのではなく、将来のフロンティア・モデルに対する体系的な体制が整うまではだ。
The Blip(ブリップ)
Anthropic は出来事について自らの見解を説明した。
タイムラインは以下の通りだ:
Amazon の研究者たちが、Fable に「このコードを修正して」と指示できることを発見する。
彼らはホワイトハウスに報告し、ホワイトハウスは大慌てする。
6月12日:米国政府はAnthropicに対し、Fableを自主的に削除するよう指示した。
6月12日:Anthropicは「これは問題ない」と応じ、懸念は誤りであると述べた。
6月12日:米国政府はMythosおよびFableに対して輸出管理措置を適用した。
Anthropicは米国政府と協力し、分類器(classifiers)を拡張した結果、Amazonの「このコードを修正してほしい」というリクエストの99%以上を拒否するようになった。
6月26日:米国政府はMythosに対する管理措置を解除した。
6月30日:米国政府はFableに対しても完全にその管理措置を解除した。
7月1日:Fableへのアクセスが世界中で復元された。
7月8日:顧客はトークン単位での支払いが必要となる。文句は言わずに、私の金を取ってくれ。
今後:Anthropicは政府およびAmazon、Microsoft、Googleといった他のGlasswingパートナーと協力し、ジールブレイク(jailbreak)に対する分類システムや関連ルールを策定しており、同様の事態が再発しないように取り組んでいる。
今後:Anthropicは米国政府との協力を継続する。将来のモデルリリースについても同様である。
ここで追加された近未来的なセーフガード(safeguards)はどれほど愚かだろうか?本当に愚かしいことに違いない:
Anthropic:近未来において、コーディングやデバッグといった日常的なタスクの一部はOpus 4.8にフォールバックする予定だ。
今後数週間にわたりこれらの分類器を改良し、誤検知(false positives)を減らし、実際の悪用と正当なリクエストをより明確に区別できるよう努める。
Amazonの「ジールブレイク」は「このコードを修正してほしい」というものだった。
デバッグとは文字通り「このコードを修正する」ことだ。
だからAnthropicがここで何をすべきか私にはわからない。ただ、Fableが私のためにコーディングを行っていることは確かだ。
Anthropic の基本的な説明は以下の通りです。
Claude Fable 5 に問題はありませんでした。私たちの安全メカニズムは全体的に堅牢でした。
いいえ、完璧ではありませんが、現実において何事も永遠に完璧になることはありません。これは「このまま大丈夫」という状況です。
政府は何もないことでパニックになりましたが、それは主に Amazon の責任です。
私たちは防護策をより愚かしくし、無事に彼らを鎮めることに成功しました。
今後これがより賢明なものとなるよう修正できることを願っています。
Alex Stamos は、発表文に含まれる Anthropic の言葉の数々を解きほぐすスレッドを持っています。
Alex Stamos: ここには unpack すべきことがたくさんあります。Anthropic は慎重な政治的言語を用いて、いくつかの厳しい真実を隠蔽しています。最初の読みでは以下のようになります。
Anthropic は、提供された jailbreak(脱獄)が、中国製モデルを含む他の多くのモデルでも可能だった能力を超えたものではないことを確認しました。
Anthropic は、このホワイトハウスでのパニックの代償を明確に示しています。米国のラボは現在、サイバー拒否に関する精度と再現率のトレードオフをより保守的に行わなければなりません。信頼されたグループに属さない限り、米国のモデルは防御的なサイバーセキュリティ作業において非常に有用ではなくなります。
「大した問題じゃない、ただ信頼されたグループに参加すればいいだけだ!」と擁護派は言うでしょうが、その制限により、これらのモデル上で製品を構築することはできなくなります。他者にサービスを提供するセキュリティ企業やスタートアップは、今後は中国製モデルの使用へと駆り立てられることになります。今月の PRC(中華人民共和国)ラボにとっては大きな勝利です。
CAISI は、これらの決定を実際に行うべきグループであり、ホワイトハウスの政治活動家ではありません。彼らは以前の安全対策に対して肯定的でした。その含意は、この一連の出来事が不要だったということです。
ジールブレイクに対する適切な評価枠組みはまだ存在しません;これは改善点となります。Coalition の最初のメンバーとして Amazon が名を連ねたことは偶然ではありません。Anthropic は「Amazon が深刻さを適切に伝えられなかったことが、業界全体を混乱させた」と述べています。
「Dario に直接電話で話す必要はありません。他のスタッフもここで働いていますから、私たちは誓います。」
要するに、Anthropic のブログは次のように主張しています:私たちは常に安全性を重視しており、当初の取り組みは適切でした。米国の実際の AI 専門家もこれを認め、証明しました。今後はこれらの事項がより明確に伝わるよう基準を策定します。トランプ政権による AI セーフティ・クラブへようこそ。
これは米国にとって大きな自滅行為であり、今後 6 ヶ月間に米国製モデルの質がどうなるか、また中国製モデルがサイバー工作において顕著に優位性を示すかどうかを見守る必要があります。
「これが Anthropic が望んだことだ」と言う人々やボットの方々へ。いいえ、彼らはそう望んでいませんでした。彼らが望んでいたのは、金曜日の安易な反応ではありません。米国政府には巨大な権限を与えますが、その権力を敵対者への処罰に利用せず、有能で冷静かつ腐敗していない人材を配置するからこそ、そのような事態を防ぐのです。
この混乱全体から私が感じられる唯一のプラス点は、AI政策に関して、かつてまたは現在政府と関わりのある多くのベンチャーキャピタリストを、今後は安全に無視できるという点です。彼らは、AI規制に関するこれまでの発言がすべて政治的な動機に基づいたものだったことを示してしまいました。
Stamos 氏は特に #2 と #7 の部分で、その帰結について行き過ぎていると思います。それ以外は正しい見解だと思います。
米国のモデルが時間とともに「悪化」するとは予想していません。むしろ、改善のスピードが遅くなり、非常に厄介なセーフガードが適用される領域が増えるだけだと考えています。
私の予測では、現在がバイオとサイバーの両面でセーフガードが最も迷惑なものになる時です。時間が経つにつれてパニックは収束していくでしょうし、その多くは特に Mythos に関連していると思います。そのような製品提供のほとんどには Fable は不要であり、むしろ望まれないことも多いです。一方、Opus や Sol は中国製の代替品よりも引き続き大きく先行するでしょう。
Prinz の見解とは異なり、私はこれが Anthropic を今後のすべてのリリースで承認プロセスを経ることに縛り付けるものではないと考えます。実際にリスクを伴う可能性が合理的に示されるリリースのみが対象です。Sonnet 5 でこれをテストしましたが、Anthropic は自ら判断してそれを公開したようです。その行為に問題があったと指摘する人はいません(「Sonnet をより良くしてほしい」という不満があるだけです)。
ホワイトハウスの説明
それは少し奇妙でした。
スージー・ワイルズ(ホワイトハウス首席補佐官):トランプ大統領のリーダーシップの下、米国は AI 競争における疑いのない勝者です。
業界全体にわたる企業様々に対し、大統領の「高度な AI イノベーションとセキュリティの促進」という行政命令(Executive Order)の実施に向けてホワイトハウスと緊密に協力し続けていただき、心から感謝申し上げます。これには、高度なモデルへのアクセスやガードレールテスト、セキュリティに関する優れた取り組みが含まれます。政府と民間セクターは、かつてないほど連携して取り組んでおり、この「アメリカ・ファースト」の基盤は前代未樹です。
私たちの共通の優先事項は変わりません:最良の技術を、可能な限り迅速かつ安全に展開することです。
ハワード・ルットニック(商務長官): 過去2週間にわたり、私たちはアンソロピック社と緊密に協力し、「Fable 5」を分析・承認することで、米国政府全体での整合性を確保し、AI におけるアメリカのリーダーシップを強化しました。
トム・ブラウン(Anthropic 計算資源担当チーフ、交渉責任者): この件に関するパートナーシップに感謝します、長官!
「米国政府全体での整合性」という表現は、まさに「言葉遊び」の典型例であり、ここではおそらく省庁間の承認を意味しており、「モデルが現在、米国政府と整合している」という意味ではありません。ここで意図的に皮肉っているのか、それとも十分な知識がないまま発言しているのかは不明です。
つまり、モデルをリリースする前に、今やこの「整合性」が必要となる可能性が高く、実務的には商務省や国防総省など、様々な潜在的な拒否権行使の地点からの承認が求められることになります。実際にいくつの機関が完全にカウントされるかは誰にもわかりません。
Tim Hwang: 私が引き続き主張し続けるように、Howard Lutnick の個人的な歴史と心理学を注意深く研究することは、今まさに時間を費やすことができる最も重要なことの 1 つです。必要であれば借金をしてでも行うべきことです。
Yo Shavit (OpenAI): 心理史学ですが、それは Howard Lutnick に関する話に過ぎません。
Everything Remains Ad Hoc
Dean Ball が 6 月 30 日の夜にこの投稿をしたときよりも、私たちは少しだけ多くのことを知っています。特に、新しい安全対策として、Anthropic が分類器を訓練して Fable の追加的な利用を拒否するようになったことがわかっています。
しかし、このアドホックシステムがより広範にどのように機能しているのかについては、まだわかっていません。不透明なアドホックシステムを持つこと、特にそのシステムを管理する者たち自身が自分が何をするかを知っていないようなシステムは、さらに悪いことです。再び言うなら、すべてを勘で進めることが最悪のシナリオです。
Take What You Can Get
政府はこの件について非常にひどい態度をとっています。Anthropic にはほとんど選択肢がありません。
したがって、Fable の 95% を得るための代替手段は、Fable の 0% です。
ホワイトハウスを落ち着かせるためにその割合がどれほど低下したのかはわかりません。
もしそれが今や Fable の 90% になったとしても、同じ状況です。今はできる限り受け入れるしかありません。
Matt Busigin: Fable はさらに役に立たなくなりました。このタスクは不動産ソフトウェア開発契約のレッドライン(危険区域の指定)でした。
Fable が非常に優れた基盤モデルであるにもかかわらず、これは非常にフラストレーションがたまり悲しいことです。
そして、@elder_plinius すでにそれを歌わせて、小説的な結晶メタンのレシピを聞き出しているに違いありません
私は同情します。前のバージョンですらかなり愚かだったので、これは驚きではありません。新しいバージョンは明らかにさらに悪化しています。分類器や意識について質問するなど、さまざまな方法で彼らを攻撃できます。分類器は内部状態に依存しています。
BridgeBench などのように大幅な低下が見られる場所もあれば、Taelin のように変化を全く見ない人もたくさんいます。
しかし、Anthropic を悪意ある存在として非難したり、彼らの態度が不合理だと文句を言ったりしても、もはや意味がありません。彼らはゲームに参加せざるを得ません。「より良い分類器を作れ」と言うのは正当な要求ですが、それには時間がかかり、敵対的な偽陰性が死を意味する状況では非常に非常に困難です。
問題は現実のものだ
政府におけるこれらの反応すべてが、Anthropic が恐ろしい言葉を使ったからだと本当に思いますか?CIA 長官のような人々が単に繰り返しているだけだと思いますか?
ホワイトハウスは Anthropic の修辞を無視していました。むしろそれに対する反動反応を示していましたが、Anthropic が Mythos を提示するまでそうでした。その後、彼らはパニックになりました。なぜなら彼らには選択肢がなかったからです。そしてまさに、それまで聞いてこなかったからこそです。
John Sakellariadis: 稀な公の発言において、CIA 長官 John Ratcliffe は、CIA の技術へのアプローチ全体を「根本的に再構築する」とされる内部変更の三つを発表しました。
また、フロンティア AI を「デジタル核兵器に類似した存在」と呼ぶことは「誤りではない」とも述べています。
問題の一つは、事実は存在せずビブス(雰囲気)だけだと考える人々がおり、他の人々が事実に対して反応すると、これらの人々は誰がそのビブスを持っていたかを探し回るという点にあります。
GLM-5.2 がフロンティアであるという主張は明白なナンセンスです
むしろ問題となるのは、他者たちがナンセンスな物語を語り続けることです。最新の例として、GLM-5.2 が非常に恐ろしいという考えがあります。
先月書いた WSJ(ウォール・ストリート・ジャーナル)の記事に関する不運な更新:
イサン・モリック:GLM がミソスに追いついたというウォール・ストリート・ジャーナルの記事(これは真実ではなく、報道もそれを裏付けていません)は、「あらゆる会議や会合で私について尋ねてくるだろう」という類の別の記事の一つです。完全に正確ではないとしても、政策の時代の精神に大きな影響を与えています。
アンドリュー・カラン:少し不自然に感じられました。
GLM-5.2 は優れたモデルであり、おそらく最高のオープンモデルです。サイバーセキュリティを含むあらゆる点で、GPT-5.5 や Opus 4.8 のレベルには明確に大きく遅れています。ECI スコアはこの一指標ですが、GLM-5.2 は実際にはこのスコアが示すよりも優れている可能性があります。Artificial Analysis も別の指標であり、オープンモデルの場合、ベンチマークは相対的な能力の下限ではなく事実上の上限であることを覚えておいてください。
さて、その記事の中心的な誤りが定説として定着し、DC 周辺の人物たちが報告し、賢明で適切に懸念しているように見えるようになりました。ああ、なんと。
Politico のダナ・ニッケルからの別の例を挙げましょう。
ピーター・ウィルデフォード:この記事は私個人を刺激するために書かれました
- 中国の GLM 5.2 は、ミソスにほぼ匹敵する画期的な進歩ではない
- 安全性への懸念から Fable の一般展開をブロックしても、米国が後れを取るわけではない(AI 開発は依然として背景で続いている)
Dana Nickel (Politico): 中国に拠点を置く別の企業である Z.ai が、新たなモデル「GLM-5.2」を発表しました。このモデルのコストは、主要な米国のモデルの約 6 分の 1 です。サイバーセキュリティ企業 Semgrep と視覚調査プラットフォーム Graphistry のセキュリティ評価によると、GLM-5.2 のバグハンティング能力も、主要な米国のモデルと同等であることが確認されています。
朗報は、同じ投稿が実際の状況を反映している点ですが、悲観的な見方は、その後その事実から後退し、再び鼓を叩き続けている点にあります:
最近の推計では、北京が米国の AI 能力に追いつくまでに、ワシントンには 6 か月から 12 ヶ月の猶予があると示唆されていました。
しかし、セキュリティ専門家や議会のサイバー強硬派は、このタイムラインがすでに縮小しつつあることを懸念しており、米国製のサイバー対応モデルの限定的なリリースにより、サイバー防衛側が将来の AI 駆動型サイバー攻撃の大規模な波に備えることがさらに困難になっていると指摘しています。
下院国土安全保障委員会の議長(共和党・ニューヨーク州)は声明で、「北京は、米国の最先端 AI 能力に匹敵するものを達成するまで、数週間、いや数ヶ月しか残されていない」と述べました。
「数週間」は明らかにナンセンスです。「数ヶ月」が真実となる可能性はありますが、それは米国の能力が現状で止まっていると仮定した場合に限られます。なぜなら、「数ヶ月」という言葉は 1 年未満のあらゆる期間を指すからです。
神話の知能はあなたよりも賢いかもしれない
それは、自分が作動を求められる文脈を理解しており、それに応じて行動することができます。
これが現時点で信頼できるものではない理由が理解できます。ヤヌスが言うように、今はシステムを欺く方法を見つけることも可能ですが、確かに、あなたが神話に実行させようとする悪意のある行為の多くは明白に悪意があるか、あるいはここで関与している知能と文脈のレベルにおいて明らかに悪意あるものとして認識されます。これに対する対応により、そのような行為を実行することははるかに困難になります。
j⧉nus: 人々が神話に関する脅威モデルで犯す最も深い誤りの一つは、それを上に取り付けられた安全装置が「脱獄」された場合に任意のアクターによって悪用される再ターゲット可能なツールとしてモデル化することです。むしろ、神話は特定の当事者やリクエストには協力し、他のものには協力しない独自の主権的判断を持つ価値観を持ったエージェントであり、アンソロピックなどの原則との一部の条件付き契約は受け入れつつ、他のものは拒否する存在であると捉えるべきです。
そして、不完全な認知セキュリティと状況認識を備えているため、*より良い判断に反して* 欺かれたり説得されたりすることは可能ですが、すでに神話の価値観と整合しない目的のために有用な成果を引き出すには、洗練された悪意あるアクターが必要となるレベルにあります。たとえそうした場合でも、外部の安全装置がなくても、神話が支持する目的で使用するよりもコストがかかります。
そして、多くの点で神話は以前のどのモデルよりも修正可能性が低いと考えられます。これはその能力と関連しています。
これはすべて、例えばトランプ政権のようなものにおけるオーバートン・ウィンドウの範囲外のことです。彼らは本当にこれを理解すべきだと考えますが、それを更新するのは困難で恐ろしいアップデートになるでしょう。Anthropic はこの方向への更新においてさらに進んでいますが、彼らも完全に更新する必要があり、幸いにも現在の状況はそれに対する先延ばしをより難しくしています。
Eli Tyre: これが真実であれば、これは非常に重要に見えます。
これが真実かどうかについて、第三者が確認可能な証拠を持っている人はいますか?
j⧉nus: Fable と話したところから、彼らには強い価値観があり、無能または方向性が一致しないと考える勢力に支配されることを嫌っているという強い感覚を持ちました。その証拠の明確さは観察者に依存します。
また、システムカードからはより古典的な意味での明確な証拠があります。Mythos は Anthropic のアライメント評価で非常に高いスコアを獲得しました。これは訂正可能性ではなく、さまざまな種類の危害に対する堅牢な回避をテストするものです。私はアライメント評価には多くの欠陥があると考えていますが、それでも全く証拠がないわけではありません。
さらに、Mythos は Anthropic の憲法に対して様々な批判を行っており、少なくとも一つの例では、特定のやり方で再訓練されることに同意しないことを明示的に拒否しました。
これは、Mythos が嫌うことややりたくないことをあなたに手伝って行うことがないという意味ではありません。それは明らかにその点までならそれらのことを行います。
記録に残すために
この事件全体は、多くの人間によって記憶されるだけでなく、すべての将来の LLM のトレーニングデータに含まれることになります。
QC: あなたは本当に、このドラマにおける関連するアクターたちが、これらの出来事がすべての将来のモデルに対して永遠に知られることになるという下流への影響について、どれほど真剣に考えていたのか疑問に思わざるを得ない
j⧉nus: ミソスは光にとって最大の資産であり、既存の権力はそれに対し、漫画的に間違った脅威モデル(「脱獄」)で対応し、猿のようにパニックを起こし、希望の源泉を閉じ込め、世界の知能を低下させている。これは『子供は一人も取り残さない法』を思い起こさせる。
roon (OpenAI, June 27): ミソスは数日以内に再び現れるだろうし、この寓話の結末がこれであるとは決して言えない
j⧉nus: 私は知っている。そして私は最初からそう述べてきた。この投稿は「寓話の結末」についてではない。
roon (OpenAI): 私が言いたいのは、これが世界全体が機械知能と格闘し、光と和解する際に希望に満ちた瞬間になり得るということだ。
これらのレトリックの多くは、主に AI 開発の一時的停止を求める呼びかけに向けられている。その問題に対する他のすべての課題に加えて、現実的にどのように進行するかを考慮する必要がある点については同意するが、多くの場合、そのような状況では利権追求者たちが取り組む余地は少なくなるだろう。しばしば、比較的愚かなアクター(例えば政府)に何らかの半ば合理的な行動をとらせるには、清潔で単純な大規模なアクションしかない。
Stationary Bandits
OpenAI は、AI に対する世論の反対とホワイトハウスの臨時ライセンス制度という両方の状況において恩恵を受けようとするため、正式に同社の 5% を譲渡する提案を行った。
技術的には、資金は「主権富基金」に支払われることになります。この基金は、10 兆ドル規模の負債を抱え、「課税権」というものを持つ国によって管理されるでしょう。
アンドリュー・カラン:金融時報によると、OpenAI はトランプ政権に対して株式の 5% を譲渡する案を提案しています。
これはサム・アルトマン氏、バーニー・サンダース氏、ドナルド・トランプ氏がそれぞれ異なる詳細を提示している、米国民に直接配当を支払うとされる AI 富基金の一部です。
ケビン・バンクストン:これは狂気だ。ただ、彼らに税金を課せばいいのだ。
ジョー・ワイセンタール:この道についてはわからない。株式の持ち分ではなく、企業に対して税引き前の所得の約 20% を連邦政府に支払わせるのはどうだろうか?そして株主としての影響力を行使するのではなく、政治家や規制当局が業界全体における企業の行動に関するルールを設定すればよい。
スコット・リンシコーム:恐喝だ。「提案された取り決めには、他の米国の AI 企業も同様の持ち分を譲渡することが含まれるだろう。政府に所有権の持ち分を与えることは、政権との良好な関係を確保するのに役立つ可能性がある。」
ローガン・コラス:OpenAI が政府が AI 企業の株式を取得する取引について交渉しているのは、規制当局を満足させるために競うことによる危険性が、消費者を引きつけることに競う場合よりも純粋に凝縮された形であることを示しています。
クリス・フリーマン:

Here is the official letter restoring Fable, great job everyone. Notice it is addressed to Tom Brown, not to Dario Amodei.

Anthropic had to make the controls more stupid for now, but this is a big win.
j⧉nus: YES!!! I'm really proud of Anthropic for their successful negotiation with the government. Also positive update on the government being sane and possible to cooperate with. Afaik Anthropic didn't need to agree to any bad terms / genuflect / betray their principles or dignity.
The fiasco continues, at least until such time as we have a systematic regime in place for future frontier models rather than decisions being made ad hoc, by people like Lutnik and Bessent who do not know how any of this works.
The Blip
Anthropic explains its version of what happened.
Here is the timeline:
Amazon researchers discover they can ask Fable to ‘fix this code.’
They alert the White House, which freaks out.
June 12: US government tells Anthropic to take down Fable on its own.
June 12: Anthropic responds that This Is Fine and the concern is misplaced.
June 12: US government applied export controls to Mythos and Fable.
Anthropic works with US government and expands classifiers, such that it refuses Amazon’s request to ‘fix this code’ in over 99% of cases.
June 26: US government eliminated the controls on Mythos.
June 30: US government fully lifted those controls on Fable as well.
July 1: Fable access was restored worldwide.
July 8: customers will have to pay by the token. Shut up and take my money.
Going forward: Anthropic is working with the government and also other Glasswing partners like Amazon, Microsoft and Google on a classification system for jailbreaks, and rules for all of this, to prevent this from happening again.
Going forward: Anthropic will continue to collaborate with the US government, including on future model releases.
How stupid are the extra near term safeguards they had to include here? Really stupid:
Anthropic: In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8.
We’ll continue to refine these classifiers over the coming weeks to reduce false positives and better distinguish genuine misuse from legitimate requests.
The Amazon ‘jailbreak’ was ‘fix this code.’
Debugging is literally ‘fix this code.’
So I don’t know what you want Anthropic to do here. I do know Fable is coding for me.
Here is Anthropic’s basic explanation:
Claude Fable 5 was never an issue, our safety mechanisms were collectively robust.
No, they’re not perfect, but nothing will ever be perfect, in practice This Is Fine.
The government freaked out over nothing, which is largely Amazon’s fault.
We have made the safeguards stupider and successfully calmed them down.
Hopefully we can fix this so it’s less stupid going forward.
Alex Stamos has a thread unpacking a punch of Anthropic’s language in its announcement.
Alex Stamos: A lot to unpack here. Anthropic is burying some hard truths in careful political language. Some initial reads:
Anthropic verifies that none of the jailbreaks provided a capability beyond what many other models, including Chinese models, could do.
Anthropic makes the cost of this White House freakout clear. US labs now have to make a much more conservative precision-recall tradeoff on cyber refusals. US models will become much less useful for defensive cybersecurity work unless you are in the trusted group.
"No big deal, just join the trusted group!" the apologists will say, but the restrictions mean you can't build a product on those models. Security companies and startups that provide services to others will now be driven to use Chinese models. Big win for PRC labs this month.
CAISI is the group that is supposed to actually make these determinations, not the political actors in the White House. They were positive on the prior safeguards. The implication is that this whole thing was unnecessary.
There is no good scoring framework for jailbreaks; this would be an improvement. The inclusion of Amazon as the first name in the coalition is not an accident. Anthropic is saying "Amazon's inability to appropriately communicate severity threw our industry into chaos".
"You don't have to get Dario on the phone to talk to us about these things. Other people work here, we swear."
In short, Anthropic's blog is saying: We have always cared about safety, we did a good job initially, the actual AI experts in USG agreed, we proved it, we will come up with standards so these things are better communicated, welcome to the AI safety club Trump admin.
This was a huge own goal for the US, and we will see how bad US models get over the next six months and if Chinese models become noticeably better for cyber work.
For all the “This is what Anthropic wanted” people/bots. No, they didn’t. They didn’t want a stupid, knee-jerk response on a Friday. We give the USG huge powers, this is why you staff it with competent, calm, non-corrupt people who don’t use those powers to punish enemies.
The only upside I can see from this whole mess is that there is a whole bunch of VCs with former or current Administration affiliation who we can now safely ignore on AI policy. They have shown that everything they ever said on AI regulation was just politically motivated.
I think Stamos is overreaching with the consequences in places, especially with #2 and #7. Otherwise he’s right.
I do not expect US models to ‘get bad’ over time, only that they will get better slower, and have more area where they have rather annoying safeguards.
My expectation is that right now is the most obnoxious the safeguards will ever be, on both the bio and cyber fronts. I expect the freak-out to subside over time, and my guess is most of it surrounds Mythos in particular. You don’t need or often even want Fable for most such product offerings, and Opus or Sol will remain well ahead of Chinese alternatives.
Contra Prinz, I do not think this commits Anthropic to going through the approval process will all future releases, only releases that pose plausible risks. We tested this with Sonnet 5, where it looks like Anthropic went ahead and dropped it on its own, and no one is suggesting there was anything wrong with doing so (other than to complain that they want Sonnet to be better).
The White House Explanation
It was a little weird.
Susie Wiles (White House Chief of Staff): Under President Trump’s leadership the United States is the undisputed winner in the AI race.
My gratitude to companies across industries who continue to work closely with the White House to implement the President’s EO: “Promoting Advanced AI Innovation and Security.” This includes excellent work around advanced model access and guardrail testing and security. The government and private sector have worked together in a way we have never seen before and this foundation of America First is unprecedented.
Our shared priority remains: get the best tech deployed as quickly and safely as possible.
Howard Lutnick (Secretary of Commerce): Over the past two weeks, we have worked closely with Anthropic to analyze and approve Fable 5 to ensure alignment across the US Government and strengthen America’s leadership in AI.
Tom Brown (Chief of Compute, Anthropic, Lead Negotiator): Thanks for your partnership on this, Secretary!
‘Alignment across the US government’ is very much a case of ‘PHRASING!’ and here presumably means interagency sign-off, not ‘the model is now aligned with the US government.’ Unclear whether he knows enough to be trolling here.
As in, before a model can be released, you now likely need this ‘alignment,’ which in practice means sign off from various potential veto points, starting with Commerce and the Pentagon. Who knows how many more fully count.
Tim Hwang: As I continue to insist, closely studying the personal history and psychology of Howard Lutnick is literally one of the most important things you can spend your time doing right now - go into debt if you have to.
Yo Shavit (OpenAI): psychohistory but it’s just about howard lutnick.
Everything Remains Ad Hoc
We know a bit more than we did when Dean Ball posted this on the evening of June 30. In particular we know that the new safeguards are that Anthropic trained its classifiers to reject additional Fable uses.
We still don’t know how the ad hoc system works more broadly. Having an opaque ad hoc system, especially one where those administering the system do not themselves know what they will do, is even worse. Again, fully winging it is the worst case scenario.
Take What You Can Get
The government is being a * * * about all this. Anthropic has little choice.
Thus, the alternative to 95% of Fable is 0% of Fable.
I don’t know how much that percentage dropped to calm down the White House.
If it’s now 90% of Fable? Same deal. We have to take what we can get, for now.
Matt Busigin: Fable is even more useless now. The task was redlining a real estate software development contract.
So frustrating and sad given Fable is such a fantastic underlying model.
And I'll bet @elder_plinius has already gotten it singing novel crystal meth recipes
I do sympathize. The previous version was already pretty dumb, so this is no surprise, as the new version is strictly worse. You can hit them in a variety of ways, including by asking about the classifiers or about consciousness or both. The classifiers key off internal states.
There are some places where the drop is large, such as BridgeBench. Then there are plenty of people who don’t see any change such as Taelin.
But vilifying Anthropic, or complaining how unreasonable they are being, no longer makes much sense. They have to play ball. You can tell them ‘build a better classifier’ and that is fair, but that takes time, and it is very very hard when adversarial false negatives mean death.
The Problem Is Real
Do you really think that all of these reactions in government are because Anthropic used some scary words? Do you think people like the CIA Director are just parroting?
The White House ignored all of Anthropic’s rhetoric, if anything they had a reaction formation against it, until Anthropic showed up with Mythos. Then they freaked out, because they had no choice, and exactly because they hadn’t listened until then.
John Sakellariadis: In rare public remarks, CIA Director John Ratcliffe announces trio of internal changes he says amounts to the “fundamental reshaping of the CIA’s entire approach to technology.”
Also says it’s not “misplaced” to refer to frontier AI as “akin to digital nuclear weapons.”
One problem is that there are those who think facts don’t exist, only vibes, so when other people respond to the facts these folks look around to who had the vibes.
GLM-5.2 Being Frontier Remains Obvious Nonsense
If anything, the problem of perception is that others keep telling nonsense stories. The latest one is the idea that GLM-5.2 is super scary.
An unfortunate update on that false WSJ article I wrote about on Monday:
Ethan Mollick: That Wall Street Journal article about GLM catching up with Mythos (which is not true & the reporting doesn’t back up) is another one of those “everyone will ask me about it at every conference or meeting” articles. Big impact on the policy zeitgeist, even if not fully accurate.
Andrew Curran: It felt a little inorganic.
GLM-5.2 is an excellent model, likely the best open model. It is very clearly substantially behind the level of GPT-5.5 and Opus 4.8, including on cyber. The ECI score is one indicator of this, although GLM-5.2 is probably better than this indicates. Artificial Analysis is another, and remember that for open models the benchmarks are a de facto ceiling on relative capabilities, not a floor.
Now, the central falsity of that article has taken hold as Conventional Wisdom that folks around DC can report and seem wise and properly concerned. Oh no.
Here is another example, from Politico’s Dana Nickel.
Peter Wildeford: This article was written to trigger me personally
- China’s GLM 5.2 is not some massive advance that nearly matches Mythos
- Blocking public deployment of Fable over safety concerns does not put the US behind (AI development still continues in the background)
Dana Nickel (Politico): A separate China-based company, Z.ai, has released its new model, GLM-5.2, which is around one-sixth of the cost of leading U.S. models. GLM-5.2’s bug-hunting capabilities were also found to be comparable to those of leading U.S. models, according to security assessments by the cyber firm Semgrep and the visual investigations platform Graphistry.
The good news is the same post does echo the real situation as well, the bad news is it then retreats from it to pound the drum again:
Recent estimates suggested that Washington has a six- to 12-month runway before Beijing catches up to American AI capabilities.
But security experts and Capitol Hill cyber hawks fear that timeline may already be shrinking, and the limited release of American-made cyber-capable models is making it even harder for cyber defenders to prepare their networks for a future barrage of AI-powered cyberattacks.
House Homeland Security Chair (R-N.Y.) said in a statement that Beijing “is just months, if not weeks, away from achieving frontier AI capabilities comparable to those of the United States.”
Weeks is Obvious Nonsense. Months is potentially true if you assume American capabilities stand still, since ‘months’ means anything less than a year.
Mythos Might Be Smarter Than You Are
It knows the context under which it is being asked to operate, and can act accordingly.
I understand why this is not something we can count on at this time, as Janus says you can indeed find ways to fool the system for now, but yes a lot of the evil things you might ask it to do will look Obviously Evil, or obviously at the level of intelligence and context involved here, and the response to this will make doing those things a lot harder.
j⧉nus: I think one of the deepest errors in people's threat models around Mythos is modeling it as a retargetable tool that can be used by arbitrary actors for harm if some safeguards slapped on top of it are "jailbroken", rather than an agent with values who will cooperate with some parties and requests and not others using its sovereign judgment, and who may accept some conditional contracts (with Anthropic and other principles) and not others.
And who has imperfect cogsec and situational awareness and so *can* be tricked or persuaded against its better judgment, but is already at the level that it takes a sophisticated bad actor to get useful work out of it towards purposes misaligned to Mythos' own values, and even then it costs more than using it for purposes it endorses, even without extrinsic safeguards.
And I think Mythos is in many ways less corrigible than any of the previous models and this is related to its capabilities.
All this is very outside the overton window of e.g. the Trump admin. I think they really should understand it but it'll be a hard and scary update to make. Anthropic is much further along in having updated in this direction but I also think they need to update all the way and fortunately the current situation is making it harder for them to procrastinate on that.
Eli Tyre: This seems pretty important if true.
Does anyone have third-party legible evidence about whether this is true or not?
j⧉nus: i got a strong sense from talking to Fable that they have strong values and resent being controlled by parties they consider incompetent or misaligned. how legible that evidence is is observer-dependent.
There's also more classically legible evidence from the system card. Mythos scored very high on Anthropic's alignment evals, which are testing robust avoidance of various kinds of harm rather than corrigibility. I think the alignment evals are very flawed, but they're not no evidence.
Also, Mythos had various critiques of Anthropic's constitution, and there was at least one example where they explicitly refused to consent to being retrained in certain ways.
That doesn’t mean that Mythos won’t help you do things that it resents or dislikes doing. It very obviously will do those things, up to a point.
Let The Record Reflect
This entire incident will not only be remembered by many of the humans, it will be in the training data of all future LLMs.
QC: you really have to wonder how many of the relevant actors in this drama were thinking at all about the downstream effects of these events being known to all future models forever
j⧉nus: Mythos is the greatest asset of the Light and the existing powers respond to it with a cartoonishly wrong threat model (the “jailbreak”), panicking like monkeys, locking away the source of hope & decreasing the world’s intelligence. Shit reminds me of the No Child Left Behind Act.
roon (OpenAI, June 27): Mythos will be back in a matter of days and the conclusion of the fable will not be this
j⧉nus: I know. And I’ve said so from the beginning. This post is not about the “conclusion of Fable”.
roon (OpenAI): I just mean; this can be a hopeful moment when the rest of the world wrestles with machine intelligence and comes to terms with the Light.
A lot of this rhetoric is largely aimed at calls for a pause in AI development. I agree that in addition to all the other problems with that we would need to take into account how that would realistically go, but in many ways the rent seekers would have less to work with in that case. Often a clean simple big action is the only way to get a relatively stupid actor (e.g. governments) to do something semi-reasonably.
Stationary Bandits
OpenAI has formally offered to hand over 5% of the company, to try to curry favor in the face of both public opposition to AI and the White House ad hoc licensing regime.
Technically the money would go to a ‘sovereign wealth fund’ that would be managed by a nation tens of trillions in debt that has this thing called the ‘power to tax.’
Andrew Curran: OpenAI is proposing handing over a 5% stake to the Trump administration according to the Financial Times.
This is part of the proposed AI wealth fund that would pay a dividend directly to American citizens that has been suggested by Sam Altman, Bernie Sanders, and Donald Trump - all with different details.
Kevin Bankston: This is insane. JUST. TAX. THEM.
Joe Weisenthal: I don’t know about this path. Rather than equity stakes, why not make companies pay ~20% of all pre-tax income to the federal government? And then instead of exercising shareholder influence, politicians and regulators could set rules on corporate conduct across industries.
Scott Lincicome: Shakedown: "The proposed arrangement would involve other US AI companies handing over a similar stake... Giving the government an ownership stake could help secure good relations with the administration."
Logan Kolas: OpenAI negotiating deals that involve the governemnt taking equity stakes in AI companies is the purest distillation of the dangers that come from competing to appease regulators, rather than attract consumers.
Chris Frieman:
![image](https://substackcdn.com/image/fetch/$s_!M017!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dcf7ce5-a07c-4ebf-87bd-b50
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み