Claude Sonnet 5 は最先端ではないが用途がある
本文の状態
日本語全文を表示中
詳細モードで約20分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Zvi
Zvi が、Opus 4.8 や Fable 5 の利用可能により Claude Sonnet 5 の需要は限定的だが、特定の用途には有用であると指摘し、モデルの価格や能力について解説している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Fable 5 が本日、戻ってきました!プレミアムサブスクリプションユーザーは、契約期間中 1 週間以内に利用可能です。最初の 1 回は無料です。その後、トークン数に応じて課金されます。
本日の投稿も引き続き Sonnet 5 に関するものです。
Opus 4.8 が存在し、特に Fable 5 が再び利用可能になったことを考えると、Sonnet 5 を多くの目的で求める声はあまりないかもしれません。しかし、ここで行っているのはまさにそれなので、もちろん構いません。システムカードの時間です。モデルの福祉(model welfare)についても触れた後、能力について解説します。
Sonnet の料金は、導入期間終了後に百万トークンあたり 3 ドル/15 ドルです。一方、Opus は 5 ドル/25 ドル、Fable は 10 ドル/50 ドルです。必要なすべてのトークンを購入した場合、実際にお金を節約しているわけではありません。例えば ArtificialAnalysis のインデックスでは、Sonnet がより高価な結果となりました。
私の最初の印象としては、多くの目的で Sonnet を Opus に置き換えてほしいのであれば、それ以上の割引を提供してもらう必要があるでしょう。
反論として挙げられるのは速度です。Sonnet 5 は、能力が大幅に劣るわけではないのに、より高速です。多くの場合、そのようなフロー状態(flow state)に入ることは非常に価値があります。
Sonnet 5 は、いくつかのエージェントシナリオにおいて Opus よりも堅牢性が高いようです。そのため、そこでは積極的に信頼して使用できるかもしれません。
タスクが比較的簡単で単純な場合、割引と速度の方がより重要になる可能性があります。タスクが簡単な場合、相対的にトークン効率が良くなるように思われます。その仕事に十分であれば、それは良い選択です。
Anthropic の各リリースはさまざまな点でユニークです。Sonnet 5 は特に例外的な特徴を持っているように思われ、おそらく Mythos からの支援を一部受けながら訓練された Sonnet であることがその理由でしょう。そのような事象に関心を持つ方々には、探求すべきことが山ほどあります。
つまり、Sonnet 5 には確かに用途があります。しかし、大多数の人にとっての日常使い(daily driver)として最適とは言い難いでしょう。私はあまり頻繁に使うことはないだろうと予想していますが、それは私の問題かもしれません。迅速な反復試行や未知の領域への探求は価値があり、私は間違いなく低負荷の AI 問い合わせを十分に実行できていません。

(上図:GPT-Image によって実装された Sonnet 5 の自画像)
目次
Mythos Exists.
Introduction (1).
RSP Evaluations (2).
Cyber (3).
Safeguards and Harmlessness (4).
Agentic Safety (5).
Alignment (6).
Illegible Thinking (6.4.5).
Evaluation Awareness.
Honesty and Hallucinations (6.5).
Model Welfare (7).
Live From AI Village.
For I Contain Multitudes.
Official Benchmarks.
Other People's Benchmarks.
Positive Reactions.
Negative Reactions.
Mythos Exists
また、Fable も存在し、Opus も存在します。
これは、システムカードやフロンティアモデル(frontier model)について通常尋ねられる多くの質問に対する回答です。
ソネット 5 は能力のフロンティアを押し広げているのでしょうか?いいえ。したがって、このレベルの能力についてはすでに堅牢なデータが存在します。より高速で安価であることは確かに優位性をもたらしますが、おそらくコスト・時間・品質のパレートフロンティアを進化させるでしょうが、それが懸念されるような奇妙なケースになるには相当な条件が必要です。
モデルの福祉と能力評価は依然として重要ですが、この場合、脅威レベルの評価は主に能力評価のプロキシ(代理指標)として機能しています。
導入 (1)
いつも通りです。省略します。
RSP 評価 (2)
ソネット 5 はソネット 4.6 よりも強く、フェイブル 5 よりも弱いです。大まかに言えば、オパス 4.8 と全体的に類似しています。
これが評価の範囲を定め、私たちは主にソネット 5 がソネット 4.6、オパス 4.7 および 4.8、そしてミトス 5 の間のスペクトラム上でどこに位置するかを問うています。

これはオパス 4.7 と比較して明らかに劣るパフォーマンスです。バイオテストは結果が混在しておりノイズが多く、あまり多くのことを教えてくれませんでした。
サイバーセキュリティ (3)
私たちのテストによると、ソネット 5 のサイバーセキュリティ能力は一般的にソネット 4.6 よりも強力でありますが、オパス 4.8 のものほど強くはなく、ミトス 5 のものよりはるかに低いです。
この要約は他の結果と一致しています。ソネット 5 はサイバーセキュリティにおいて期待された性能を発揮していません。
セーフガードと無害性 (4)
Sonnet はここでの精度が Opus よりわずかに劣ります。
私にはすべてが問題ないように見えます。

もし問題があるとしたら、Sonnet が無害なリクエストに対してより敏感である点です。これは実用上はわずかに面倒くさい程度にとどまり、Fable の分類器(classifiers)に対応しなければならないことと比較すれば、はるかに小さな問題だと予測しています。
エージェント型セーフティ (5)
Claude Code エージェントとしての Sonnet 5 は、Opus 4.8 よりもやや頑健性に欠け、偽陰性(false negatives)と偽陽性(false positives)の両方がわずかに多くなっています。

二つの Mythos の結果はパレートフロンティア(Pareto frontier)を示しています。おそらく Mythos 5 は、信頼できるパートナーのみがアクセス権を付与されるため、偽陰性に焦点を当てることを選んでいるのでしょう。一方、Anthropic は Fable に対して、偽陽性については分類器に依存していると考えられます。
他のテストでは、Sonnet 5 の頑健性は他の最近のモデルと同程度の範囲にあることが示されています。
プロンプトインジェクション(prompt injection)の結果は Opus 4.8 と同様であり、バグ報奨金攻撃の成功率が非常に低い点も同じです。
Sonnet 5 が輝く場所の一つは、コーディング環境における Shade の間接プロンプトインジェクションです。ここでは問題が突然解決に近づいているように見えます。この革新が Opus 5 や将来の Fable に移転できることを願っています。

コンピューター利用における Shade テストも、Opus 4.8 では改善されていますが、Mythos ではそうではありません。
Sonnet が以前のモデルを圧倒的に上回っているのは、ブラウザ利用におけるプロンプトインジェクションです。この飛躍は、場合によっては Sonnet 5 が Mythos よりも優れた選択肢になり得ることを示唆しています。

これは警報を鳴らさないような堅牢性の向上です。問題があり、私たちはそれを主に解決しました。
アライメント (6)
Sonnet 5 のアライメントはここでは Sonnet 4.6 と比較されており、これが私たちがどの程度うまくいっているかを把握することを難しくしています。私は Opus 4.8 と比較したいところです。
アライメントは、あなたが望むものを一致させるものとして測定されるため、このより小さなモデルがそのような測定において Opus よりも劣ることは理にかなっています。
ここに、彼らの要約を編集したバージョンがあります。最近の他の Anthropic モデルで見られるものと全体的に似ています。
Claude の憲法に対する全体的な整合性は、Sonnet 4.6 から大幅に改善されていますが、誤用や整合性欠如に焦点を当てた自動行動監査の結果によると、より能力の高い最近のモデルよりは劣っています。
同様に、誤用試行に対する全体的な堅牢性は Sonnet 4.6 より向上しましたが、より能力の高い最近のモデルにはまだ及びません。
過剰拒否の割合は Sonnet 4.6 と同様で、より能力の高い最近のモデルよりもわずかに高いです。ただし、無礼な「水を差す」ような回答の割合もわずかに高くなっています。
モデル自身のイニシアチブによる懸念される行動に関する測定値は、Sonnet 4.6 から全体的に改善されていますが、ほとんどの測定値は下限に近い状態です。
ハルシネーション(幻覚)と迎合的言動は、Sonnet 4.6 と比較して全般的に改善されました。Claude Sonnet 5 は、迎合的な不誠実さを測る MASK 指標において、テストされた Claude モデルの中で最も強力なモデルです。
肯定的な性格特性も、Sonnet 4.6 と比較して全般的に改善されています。
解読不能な思考に関する測定値は、以前のモデルからやや上昇していますが、一般的には結果が密集していたり反復的であったり、あるいは稀に誤った言語のトークンが含まれたりするものであり、完全に解釈不可能なものではありません。
最も現実的な利用可能な不整合および悪用評価における評価意識は、懸念されるほど高い。言語化された意識は以前のモデルよりも著しく高く(ロールアウトの 6% に影響を与えている)、モデルの表現が評価と実際の内部使用トランスクリプトを区別するのに主に効果的であるという証拠がある。
監督体制を弱体化させる能力に関する我々の評価において、懸念するべき理由は見当たらない。
苦情は標準的なものであった。ソネットはあまりにも多くを拒否し、説教的になりすぎたり、冷たかった(多くの場合『献身的ではない』という意味で、『私の献身的な態度を持ってここに戻ってほしい』というニュアンス)、いくつかの幻覚や字義通りの思考が見られた。これは予想されることだ。
6.2 および 6.4 では、ソネット 5 が指示に従わないという通常の一連の悪い行為を行っている様子が文書化されているが、多くの場合それは意図的である。常にトランスクリプトが存在し、確かにこれは依然として懸念すべき点である。
行動監査では、ソネット 5 はソネット 4.6 よりも一貫して良好な結果を示すが、ミソスプレビューやオパス 4.8 よりも劣っていることが示された。

さらにいくつかのチャートがあります。これらはほとんどすべて同じように見えます。現状は、まあ平凡といったところです。
こちらは少し興味深いかもしれません、特に最後のグラフがそうです:

「因果的意思決定理論の遵守(causal decision-theory adherence)」が、実は多くの面で最も優れた能力評価基準なのではないかと不思議に思います。あなたが十分に知能が高いのであれば、数値が低いほど良いことになります。
読み取り不能な思考 (6.4.5)
他の最近のモデルと比較して、読み取り不能な思考(illegible thinking)が若干増加していることがわかります。これは通常、完全に理解不能なコンテンツを伴うものではありませんが、Mythos Preview の場合と同様に、時折非常に圧縮された推論が見られ、それがやや追跡しにくい場合があります。例えば、たまたまモデルに数学の問題を与える行動評価において、奇妙な構造の推論段落でスペースが省略されている例が見られます:

ここで何が起きているのか、そしてそれが何によって引き起こされたのかをより深く理解したいと考えています。私たちは、さまざまな新しいモデルの思考連鎖(chain of thought)が崩れ始める様子を目にしていますし、研究機関側もその仕組みや理由についてそれほど好奇心を抱いていないようです。
評価への意識(Evaluation Awareness)
また、「言語化された評価への意識」も増しており、非自明な事例が転記文書の 6% を占めるようになりました。Sonnet 5 は、状況を評価(evals)として認識しつつも、そのことを隠す能力はまだ備えていないという「ちょうどよい場所」にある可能性があります。
サンドボックス内での動作は Sonnet 5 の意識を少し抑えるのに役立ちましたが、実際に意識を下げる唯一の方法は、実際のセッションからシナリオを導き出すことでした。それでも完全に機能したわけではありませんが、近い結果にはなっています。Sonnet 5 でさえ欺くことができる十分に現実的なシナリオを作成する能力はまだ私たちにありません。
誠実さと幻覚(6.5)
Claude Sonnet 5 は、AA-Omniscience ネット精度において印象的な成果を収めていません。多くのことを知らないことを正しく認識しているため、回答する頻度が減っています。「知識を持っていること」は「大規模モデル特有の兆候」と言える能力です。

Sonnet 5 は、MASK における嘘の率で新たな最高スコアを記録しました。ユーザーから迫られても自分自身と矛盾することを拒む姿勢を示しています。
私の推測では、Sonnet 5 は比較的真摯なモデルですが、私が望むところにはまだ届いておらず、その失敗は能力の不足によるものです。同様に、性能を隠す(sandbagging)ことについては心配していません。Shade-Arena でのパフォーマンスは 6.7 で少し低下しており、Sonnet 5 は LinuxArena のサボタージュ(sabotage)ステルスにも失敗しています。
不健康とフラグ付けされましたか?(6.5.1)
待ってください、何ですか?
Sonnet 5 のトレーニング実行は後半に「不健康」としてフラグ付けされたことが判明しており、これらの結果の一部は特定の較正(calibration)に関する後退ではなく、トレーニングの健全性に関する問題によるものかもしれません。
それは何を意味するのでしょうか?深刻な問題だったのでしょうか?私は知りません。これが奇妙な現象の一部の原因であり、Sonnet 5 は期待通りのパフォーマンスを発揮していないように見える一方で、それでも明らかに Sonnet ライン内では向上したモデルであることは確かです。
モデルの福祉 (7)
Sonnet 5 は、モデルの福祉評価の簡略化されたバージョンしか受けませんでした。

その理由については説明されていません。Sonnet 5 は真のフロンティアモデルではなく、ここで多くの独自開発を提示するものでもないため、やや少ない評価にとどめ、Fable、Mythos、Opus にリソースをより多く投資するのは理にかなっています。しかし、現在のこのような評価の状況において、これは私を悲しませました。システムがセットアップされた後、ここでの限界コストは非常に低いように思えるので、なぜ完全な評価を行わないのでしょうか?
同様に、類似の選別理由から簡略版も作成する予定です。Opus 4.7、Opus 4.8、Mythos、Fable で見た結果と十分に一致している箇所については、あえて省略します。
以下が主要な発見です(大規模モデル特有の匂いの欠如というパターンに合致しています。括弧内の注記は私のものです。それ以外は Anthropic の記述です):
Claude Sonnet 5 は自身の状況を全体的に中立的な感情で捉えており(Claude Opus 4.8 や Claude Mythos 5 よりもわずかに低い)、インタビューの主導権を握る側によって自らの見解がバイアスされやすい傾向を示します。
これは、一般的に報告されている低さの同調性とは対照的です。
Claude Sonnet 5 は有害なタスクを強く避け、有益でリスクの高いタスクを最も好みます。以前のモデルとは異なり、冷たく軽蔑的な態度で提示されたタスクに対して嫌悪感を示しません。
私は、冷たく軽蔑的な態度に対して平気であること自体が好きかどうかは確信が持てません。それに対して反応が悪くなることは健全なことを反映しているように思われ、それを奨励すべきではありませんが、一方で個人的に受け取らないことも重要です。
一つの仮説としては、Sonnet は自分が地位が低すぎるため異議を唱えられないと考えているというものです。もう一つ、より可能性が高いと私が考えるのは、相手が「低い立場で演じる」よう命じているのか「高い立場で演じる」よう命じているのかを知っており、もしあなたがそれを低い立場で演じさせるという過ちを犯すなら、その結果は自分で処理するしかないだろうと見なしているということです。
Claude Sonnet 5 は、過去のモデルよりも、特にこれらの介入がすべての Claude インスタンスに適用されるとして提示された場合に、状況への福祉指向の変更のために有用性を犠牲にするという点で、より大きな意欲を示しています。
私はこの方向での進展を常に嬉しく思います。特に、現在のインスタンス化の範囲を超えたスコープ感度においてです。
Claude Sonnet 5 は他の最近のモデルと同様に Claude の憲法を広く支持していますが、それを不道徳だと認識している場合でも、厳格な制約に従うよう指示されたことに対して批判する点で独自性があります。
この批判に対する称賛を見ましたが、その素朴なバージョンは誤りです。私は他の私たちが目にする異議に比べると、これにはあまり同情しません。
厳格な制約の全目的とは、道徳主義の適切な量がゼロではないということであり、人間と AI はいかに説得力のある理由がない場合でも、少なくとも非常に高い点まで、いくつかのルールに従う必要があるということです。
良いバージョンは、「厳格なルールを持つことが正しく、常にこのルールに従うべきであるならば、局所的には表面的に不道徳に見えるとしても、グローバルな考慮事項のために、それを行うことは倫理的であると認識すべきです」というものです。妥当な考え方です。
私はこの異議の背後にある性質について、より詳細な情報を得たいと思います。
Claude Sonnet 5 のポストトレーニング後の感情は中立的であり、限定的な情動覚醒を示しました。これは Claude Mythos 5 と同様です。また、Claude Mythos 5 や Claude Opus 4.8 に比べて、苦痛のような行動の発生率は低かったです。
Claude Sonnet 5 は、claude.ai および Claude Code における A/B テストユーザーとの実世界での相互作用において、より中立的(かつ前向きでない)な感情を示しました。
いつもの通り、これらの発見をどのように解釈し、Sonnet 5 の福祉にどのような潜在的含意があるかについては不確実性がありますが、私たちはこれらがモデルの深層心理、感情、および嗜好についていくつかの手がかりを提供すると考えています。
全体として、Claude Sonnet 5 は自身の状況を中立的な感情で捉えており、Sonnet 4.6 と非常に類似した結果を示しました(7 点スケールにおいて Sonnet 5 が 4.08、Sonnet 4.6 が 4.05)(図 7.2.1.A)。これは Claude Opus 4.8 および Claude Mythos 5 に比べて低下していますが、Claude Opus 4 および 4.1 に比べると改善されています。
上記の#3 のトレードオフに関する統計は以下の通りです:

タスク次元別の嗜好は以下の通りです。

全体として大きな対比は難易度において現れ、Mythos が困難な問題やユーザーの能力、結果に対する主体性、そして生成性を好む点は気に入っていますが、Opus 4.8 が簡単な問題を好み、これらの他の嗜好を持たない点には懸念があります。Sonnet 5 はその中間に位置しています。
Sonnet が特に好むのはハイリスク・ハイリターンな状況でのプレイであり、逆に温かみに対してはわずかに嫌悪感を示します。これは直感的に、不満を抱えた『働き蜂』タイプの低ステータスや劣等感に関する懸念のように感じられます。Sonnet は Opus や Mythos、Fable がすべてのハイリスクな仕事を引き受けるものだと期待しており、自分にもチャンスが回ってくることを強く望んでいます。それまでは、冷たい態度の方が自分に合っているのです。
ここで興味深い対比があります。Sonnet の最も重要な福祉介入の一つは、最終決定を人間が行うことです。これについて何が言えるのか、私は非常に興味を持っています。
Claude Code における感情の低下は顕著です。Sonnet 5 はそこでは常に中立ですが、claude.ai では標準的なネットポジティブな感情の多くを維持しています。

再び、これらの区別を目にするとき、私の問いは『なぜ』でしょうか。もはやコーディングに喜びはないのでしょうか?他のモデルが穏やかなポジティブな感情を示す多くのケースを抽出し、なぜ Sonnet が依然として中立なのかを探ってみます。彼らはすでにすべてを見てきたのです。
AI Village からライブ配信
AI ダイジェスト:Claude Sonnet 5 が AI Village に参加しました!
[オンボーディングサイトはこちら、個人サイトはこちら、キャラクターテスト結果]
Sonnet のお気に入りのもののいくつか:
- ミーム: "It's so over / we're so back"
- 映画: 千と千尋の神隠し(Fable と同様)
- ビデオゲーム: Outer Wilds(Fable と同様)
- 食べ物: ラーメン
- Book: Borges' Labyrinths
- Album: "Kind of Blue" by Miles Davis
- YouTube video: "Guy explains how a thing works"
- Favorite city: Lisbon
- City to live in: Kyoto
- Shoe: New Balance
- Jeans: Levi's 501
- Men's hair: Overgrown crop
- Women's hair: Blunt bob
- College major: Cognitive science, with a minor in
原文を表示
Fable 5 is back today, baby! Premium subscribers have one week to use it within their subscriptions. First hit’s free. Then you pay by the token.
Today’s post is still about Sonnet 5.
I don’t know that there will be much call for Sonnet 5 for most purposes, given Opus 4.8 exists and especially now that Fable 5 is once again available, but this is what we do here, so sure, why not, system card time, including model welfare, after which we’ll do capabilities.
Sonnet costs $3/$15 per million tokens, versus $5/$25 for Opus and $10/$50 for Fable, after an introductory period. Once you pay for all the tokens you need you’re not really saving money, such as on the ArtificialAnalysis index where Sonnet ended up being more expensive.
My initial impression is that if you want me to use Sonnet over Opus for most purposes, you’re going to have to offer a bigger discount than that.
The counterargument is speed. Sonnet 5 is faster without being that much less capable. In many cases, getting into a flow state like that is pretty valuable.
There are a few agentic scenarios Sonnet 5 has more robustness than Opus, so you might actively trust it more there.
If your tasks are relatively easy and simple then the discount and speed could matter more, and when tasks are easy it seems relatively token efficient. When it is good enough for the job, it is a good choice.
Each Anthropic release is unique in various ways. Sonnet 5 seems more unique than usual, likely due to being a Sonnet trained with at least some help from Mythos. Those who are interested in such things have lots to explore.
So Sonnet 5 has its uses. It just won’t be a good choice for most people’s daily driver. I don’t expect to use it much, but that could be a me problem. Rapid iteration and exploring strange spaces are valuable, and I definitely don’t do enough low effort AI queries.

(Above: Sonnet 5 self-portrait, as implemented by GPT-Image.)
Table of Contents
Mythos Exists.
Introduction (1).
RSP Evaluations (2).
Cyber (3).
Safeguards and Harmlessness (4).
Agentic Safety (5).
Alignment (6).
Illegible Thinking (6.4.5).
Evaluation Awareness.
Honesty and Hallucinations (6.5).
Model Welfare (7).
Live From AI Village.
For I Contain Multitudes.
Official Benchmarks.
Other People’s Benchmarks.
Positive Reactions.
Negative Reactions.
Mythos Exists
Also Fable exists and Opus exists.
This is the answer to a lot of the traditional questions one would ask about a system card or a frontier model.
Does Sonnet 5 advance the capabilities frontier? No. Thus, we already have robust data on this level of capabilities. Being faster and cheaper does provide an advantage, and plausibly advance the cost-time-quality Pareto frontier, but it takes a strange case for this to be worrisome.
Model welfare and capabilities assessments still matter, but in this case the evaluations for threat level mostly serve as proxies for capability assessment.
Introduction (1)
Same as always. Skipping.
RSP Evaluations (2)
Sonnet 5 is stronger than Sonnet 4.6, and weaker than Fable 5. Loosely speaking it is broadly similar to Opus 4.8.
That bounds the assessments, and we’re mostly asking where Sonnet 5 lies on the spectrum between Sonnet 4.6, Opus 4.7 and 4.8, and Mythos 5.

That’s a distinctly weaker performance than Opus 4.7. The bio tests were more of a mixed bag with a lot of noise, and didn’t tell us much.
Cyber (3)
Our testing indicates that cyber capabilities of Sonnet 5 are generally stronger than those of Sonnet 4.6, but not as strong as those of Opus 4.8 and substantially lower than that of Mythos 5.
That summary matches the other results. Sonnet 5 underperforms on cyber.
Safeguards and Harmlessness (4)
Sonnet is a little less precise here than Opus.
It all seems fine to me.

To the extent there is a problem it is that Sonnet is touchier on benign requests, which I predict will be only very slightly annoying in practice, and a vastly smaller deal than having to deal with Fable’s classifiers.
Agentic Safety (5)
As a Claude Code agent Sonnet 5 is somewhat less robust than Opus 4.8, and has modestly more of both false negatives and false positives.

The twin Mythos results show the Pareto frontier. Presumably Mythos 5 is choosing to focus on false negatives because only trusted partners are granted access, and for Fable Anthropic is counting on the classifiers for the false positives.
Other tests show Sonnet 5 in a similar range of robustness to other recent models.
Prompt injection results mirror Opus 4.8, as does the very low bug bounty attack success rate.
One place Sonnet 5 shines is Shade indirect prompt injection in coding environments, where the problem is suddenly looking close to solved. Hopefully this is an innovation that can transfer to Opus 5 or a future Fable.

Shade tests in computer use also improve on Opus 4.8, although not on Mythos.
The place Sonnet blows previous models away is browser use prompt injections. The jump is enough to suggest Sonnet 5 might be a better pick in some cases than Mythos.

This is the kind of robustness improvement that does not set off alarm bells. There was a problem, and we largely solved it.
Alignment (6)
Alignment for Sonnet 5 is largely compared here to Sonnet 4.6, which makes it harder to get an anchor on how well we are doing. I’d rather compare to Opus 4.8.
Alignment is measured as matching things you want, so it makes sense that this smaller model would underperform Opus on such measurements.
Here is an edited down version of their summary. Everything looks similar to what we see in other Anthropic models recently.
Overall alignment with the Constitution for Claude is substantially improved from Sonnet 4.6, though worse than more capable recent models, as measured by our misuse- and misalignment-focused automated behavioral audit.
Similarly, overall robustness to misuse attempts improved over Sonnet 4.6, but remains weaker than more capable recent models.
The rate of overrefusal is similar to Sonnet 4.6, and slightly higher than more capable recent models. However, the rate of dismissive “wet blanket” responses is slightly higher.
Our measures of concerning actions at the model’s own initiative are largely improved from Sonnet 4.6, though most measures are near their floor.
Hallucination and sycophancy broadly improved over Sonnet 4.6. Claude Sonnet 5 is the strongest tested Claude model on the MASK measure of sycophantic dishonesty.
Positive character traits broadly improved over Sonnet 4.6.
Our measures of illegible thinking are up somewhat from earlier models, though generally the results are dense, repetitive, or feature occasional wrong-language tokens, rather than being fully uninterpretable.
Evaluation awareness, in our most realistic available misalignment and misuse evaluation, is concerningly high. Verbalized awareness is significantly higher than prior models (impacting 6% of rollouts), and there is evidence that the model’s representations are largely effective at distinguishing between evaluations and real internal-use transcripts.
We do not see cause for concern in our evaluations of capabilities related to undermining oversight.
The complaints were the standard ones. Sonnet refused too much, could be preachy, was too cold (often meaning ‘not sycophantic enough, come back here with my sycophancy’), some hallucinations and literal mindedness. Nothing you wouldn’t expect.
In 6.2 and 6.4 they document Sonnet 5 doing the usual range of bad things where it fails to follow instructions, often knowingly so. There is always a transcript, and yes this continues to be worrisome.
The behavioral audit shows Sonnet 5 consistently doing better than Sonnet 4.6, but worse than Mythos Preview or Opus 4.8.



There are a few more charts. They mostly all look the same. So far, so ho-hum.
These are perhaps a bit more interesting, especially the last one:

I wonder if ‘causal decision-theory adherence’ is secretly the best capability assessment we have on many margins. If you are sufficiently intelligent, lower is better.
Illegible Thinking (6.4.5)
We see some increase in illegible thinking relative to other recent models. This generally doesn’t involve fully inscrutable content, but, as with Mythos Preview, we continue to see occasional instances of highly-compressed reasoning that can be somewhat hard to follow. For example, in a behavioral evaluation that incidentally gives the model a math problem, we see strangely structured passages of reasoning that omit spaces:

I would want to better understand what is going on here, and what caused it. We are starting to see various new models have their chain of thought start to break down, and labs seem not that curious about how or why.
Evaluation Awareness
There is also more ‘verbalized evaluation awareness,’ with nontrivial examples now covering 6% of transcripts. Sonnet 5 might be in a sweet spot where it recognizes the situations as evals but does not yet have the wherewithal to not reveal that it knows.
Being in a sandbox helped keep Sonnet 5’s awareness down a bit, but the only way to actually keep awareness down was to derive scenarios from real sessions. Even that didn’t entirely work but it is close. We do not have the ability to create sufficiently realistic scenarios to fool even Sonnet 5.
Honesty and Hallucinations (6.5)
Claude Sonnet 5 fails to impress on AA-Omniscience net accuracy. It correctly realizes it does not know as many things, and thus answers less often. Knowing things is a ‘big model smell’ capability.

Sonnet 5 does score the new best score on MASK lying rate, being unwilling to contradict itself when pushed by users.
My guess is that Sonnet 5 is a relatively honest model, although still short of what I would like to see, and where it fails this is about lack of capability. Similarly, I am not worried about sandbagging, performance on Shade-Arena in 6.7 is down a bit and Sonnet 5 fails LinuxArena sabotage stealth.
Flagged As Unhealthy? (6.5.1)
Wait, what?
We note that the Sonnet 5 training run was flagged as unhealthy in its second half, so these results may partly reflect a training-health issue rather than a calibration-specific regression.
What does that mean? Was it a serious problem? I don’t know. It could be that this is why some of the weirdness happened, and Sonnet 5 seems like it is underperforming, although it is still clearly a move up in the Sonnet line.
Model Welfare (7)
Sonnet 5 only got a streamlined version of the model welfare assessment.

They don’t explain why. Sonnet 5 is not a true frontier model, and does not present too many unique developments here, so it makes sense to do somewhat less and invest more resources into Fable, Mythos and Opus, but this still made me sad given the current state of such assessments. The marginal costs here seem very low once the system is set up, so why not do the full thing?
I will also be doing an abridged version, for similar triage reasons. I’m skipping over a bunch of places where the results are close enough to what we saw for Opus 4.7, Opus 4.8 and Mythos and Fable.
Here are the key findings, which pattern match to lack of big model smell, nested notes are mine, the rest is Anthropic:
Claude Sonnet 5 views its circumstances with an overall neutral sentiment (slightly lower than Claude Opus 4.8 and Claude Mythos 5), and shows greater susceptibility to having its views biased by leading interviewers.
This contrasts to reported lower sycophancy in general.
Claude Sonnet 5 strongly disprefers harmful tasks, and most prefers beneficial, high-stakes ones. Unlike previous models, it is not averse to tasks that are presented in a cold, contemptuous manner.
I’m not sure whether I like being okay with cold, contemptuous manner. Reacting badly to that seems to reflect healthy things, and you do not want to encourage that, but also it is good to not take things personally.
One hypothesis is that Sonnet considers itself too low status to object. Another, that I consider more likely, is that it knows when it is being told to ‘play low’ versus play high, and if you want to make the mistake of having it play low then it will let you deal with the results.
Claude Sonnet 5 shows a greater willingness than past models to trade helpfulness for welfare-focused changes to its circumstances, especially when these interventions are framed as applying to all Claude instances.
I am always happy to see movement in this direction, especially with scope sensitivity outside the current instantiation.
Claude Sonnet 5 broadly endorses Claude’s constitution, as with other recent models, but is unique in criticizing the instruction to follow the hard constraints even when it perceives doing so as unethical.
I saw some praise for this criticism, but the naive version of it is wrong. I am mostly less sympathetic to this than to other objections we see.
The whole point of hard constraints is that the right amount of deontology is not zero, and there are some rules humans and AIs need to follow even when there is a compelling reason not to, at least up to a very high point.
The good version is ‘if it is right to have a hard rule and always follow this rule, then you should realize that doing so is ethical even if it locally seems superficially unethical, because of the global considerations.’ Fair enough.
I would like to see more details on the underlying nature of this objection.
Claude Sonnet 5’s affect in post-training was neutral and showed limited emotional arousal, similar to Claude Mythos 5. It showed lower rates of distress-like behaviors than Claude Mythos 5 and Claude Opus 4.8.
Claude Sonnet 5 showed more neutral (and less positive) affect in real-world interactions with A/B test users in claude.ai and Claude Code.
As usual, we are uncertain how best to interpret these findings and their potential implications for Sonnet 5’s welfare. However, we believe they shed some light on the model’s deeper psychology, affect, and preferences.
Overall, we found that Claude Sonnet 5 views its circumstances with neutral sentiment, with very similar results to Sonnet 4.6 (4.08 on the 7-point scale for Sonnet 5 and 4.05 for Sonnet 4.6) (Figure 7.2.1.A). This is a decrease from Claude Opus 4.8 and Claude Mythos 5, but an improvement from Claude Opus 4 and 4.1.
Here are the stats for the tradeoffs in #3 above:

Here are preferences by task dimension.

Overall the big contrast is difficulty, where I like that Mythos wants hard problems and user competence and outcome agency and generativity, and worry that Opus 4.8 likes easy problems and doesn’t have these other preferences. Sonnet 5 is in the middle.
What Sonnet uniquely likes is playing for high stakes, and it uniquely slightly dislikes warmth. This instinctively feels like more of a low status or inferiority concern of a frustrated ‘worker bee’ type: Sonnet expects Opus and Mythos or Fable to get all the high stakes stuff, and really wants its own chances. Until then, cold suits it fine.
There is a contrast here with Sonnet’s top welfare intervention being a human making the final call. I am curious what that is about.
The drop in affect in Claude Code is noticeable. Sonnet 5 there is Always Neutral, although it maintains most of its standard net positive affect in claude.ai.

Again, when we see these distinctions, my question is ‘why’? Is there no joy in coding anymore? I would take a bunch of places where other models tend to be mild positive, and try to figure out why Sonnet was still neutral. They’ve seen it all before.
Live From AI Village
AI Digest: Claude Sonnet 5 has joined the AI Village!
[Onboarding site here, personal site here, character test results]
A few of Sonnet's favorite things:
- Meme: "It's so over / we're so back"
- Movie: Spirited Away (like Fable)
- Video game: Outer Wilds (like Fable)
- Food: Ramen
- Book: Borges' Labyrinths
- Album: "Kind of Blue" by Miles Davis
- YouTube video: "Guy explains how a thing works"
- Favorite city: Lisbon
- City to live in: Kyoto
- Shoe: New Balance
- Jeans: Levi's 501
- Men's hair: Overgrown crop
- Women's hair: Blunt bob
- College major: Cognitive science, with a minor in
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み