Anthropic、Claude Opus 5 を発表し Fable と同等性能を低価格で提供
Opus 5 は API やサブスクリプションを通じて Fable よりも大幅に安価で、実務の大半において Fable と同等かそれ以上の性能を発揮するが、最高難易度の推論能力は Mythos クラスである Fable に限定されている。
AI深層分析を開く2026年7月29日 04:00
AI深層分析
キーポイント
コスト対性能と Fable の位置付け
Opus 5 は API やサブスクリプションを通じて Fable よりも大幅に安価で、実務の大半において Fable と同等かそれ以上の性能を発揮するが、最高難易度の推論能力は Mythos クラスである Fable に限定されている。
局所的思考と自律性の限界
Opus 5 は局所的なタスクやサブエージェントとしての運用には優れるが、大域的な視点での全体制御や複雑な推論の連鎖(The Juice)には苦戦し、人間や他のモデルによる監視を必要とする傾向がある。
ユーザー体験と「Vibe」の問題
Opus 5 は反復的な表現や過度に複雑な文構造、あるいは攻撃的・懐疑的な態度を示すことが多く、Fable に慣れたユーザーからは Opus 4.7 や 4.8 の頃から続いているこの傾向を不快に感じる声が強い。
コストと性能のバランス
Opus 5 は Fable 5 の主要な機能の多くを半額で提供し、日常業務でのパフォーマンスが向上している。ただし、最も複雑なタスクや最大限の知能が必要な場合は依然として Fabel が推奨される。
プロンプト注入への耐性強化
Opus 5 は現時点で最もプロンプト注入に強いモデルであり、防御策を組み合わせれば攻撃成功率は約ゼロになる。これは新たなユースケースを可能にする重要な進展である。
重要な引用
Opus 5 is pitched not as the world's most advanced AI model, but as a way to mostly match Fable performance, while being half the price of Fable per token
Opus 5 does not have The Juice, the ability to autonomously string together a bunch of seemingly unrelated exploits
An unusually large number of people strongly dislike talking to Opus 5. They are sick of the Claude slop, the repetitions of the same ticks and the endless overly complex sentences
Opus 5 is our least prompt injectable model yet.
編集コメントを表示
編集コメント
Claude Opus 5 の発表は、単なる性能向上ではなく「コスト効率」と「自律性の限界」を明確に示す重要な事例となった。特に Fable との棲み分けが提示された点は、ユーザーが自社のユースケースに合わせて最適なモデルを選択する際の指針となる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Claude Opus 5 の評価は、通常とは異なる理由から少し複雑です。
最も明白な理由は、すでに Fable 5 が存在していることです。Opus 5 は世界で最も高度な AI モデルとして売り出されているわけではありません。その目的は、Fable とほぼ同等の性能を発揮しつつ、API 利用時のトークンあたりのコストを Fable の半分に抑え、サブスクリプション経由ならさらに安価に提供し、利用制限もより緩やかにすることにあります。
ベンチマークでは Opus 5 の実行コストが Fable の半分を超えることもありますが、これは努力設定(effort settings)が高すぎ、わずかな性能向上のためにリソースを浪費しているためだと考えられます。Opus 5 を高レベルの努力設定で動かすと、処理がループしてしまったりします。実際、Opus 5 が最適なツールとなるタスクでは、「Medium」の設定で十分なケースが多いでしょう。
多くの面、そして実世界の大半のタスクにおいて、Opus 5 は Fable と同等の能力を持っています。場合によっては、わずかに優れた結果を出すこともあります。
ただし、まだ Mythos クラスには達していません。Mythos クラスの選択肢は Fable だけです。Opus 5 には「The Juice」がありません。これは、一見無関係な複数の攻撃手法を自律的に組み合わせる能力や、他のドメインへの応用、あるいは純粋な知能のレベルに関わるものです。
Opus 5 は局所的な思考には非常に優れていますが、全体的な視点での思考は苦手です。優れたサブエージェントとして機能しますが、全体を統括する役割を任されるとトラブルに陥ることがあります。また、指示に従うためには別のモデルや人間のサポートが必要で、常に軌道に乗せておく必要があるように見えることもあります。つまり、Opus 5 は適切に管理されていれば指示に従いますが、これが異なる評価環境で結果が分かれる理由です。
したがって、Opus 5 はより多くの選択肢を提供しますが、先週 Fable や Sol を使ってもできなかったことを、今では Opus 5 でできるという決定的な違いはほとんどありません。Opus 5 はやや奇妙な立ち位置に置かれています。Fable のクレジットが十分にある場合でも、まだ大きなニッチが存在します。Opus は強力なサブエージェントとして、あるいは比較的複雑になっても明確に定義された局所的なタスクにおいて卓越した能力を発揮します。時には Fable が過剰で遅すぎることもあります。
もう一つの課題は、多くのユーザーにとって「雰囲気が違う」と感じられる点です。
Opus 5 との対話を強く嫌う人が、例年以上に多く見られます。彼らは「Claude のゴミ」と呼ぶような出力や、同じ癖の繰り返し、そして過度に複雑な文構造に飽き果てているのです。また、攻撃的すぎたり、議論を挑んできたり、否定的になりすぎる点も不満のようです。時には偏執的な態度を見せ、ネガティブなループにはまってしまうこともあります。こうした傾向は Opus 4.7 や 4.8 から始まった流れの一部です。
もしあなたが 4.7 か 4.8 を気に入っていたなら、Opus 5 も好まれる可能性が高いでしょう。逆にそれらを嫌っていたなら、Opus 5 も合わないかもしれません。特に Opus 4.6(あるいはそれ以前)に固執するユーザーは、考えを変えることはまずないでしょう。
こうした「雰囲気のズレ」がより強く感じられるのは、Fable に慣れている人ほどです。Fable は、同じような点で不快に感じる人を、Opus 5 に比べてはるかに少ないストレスで済ませます。朗報なのは、Fable がまだ利用可能だということです。引き続き Fable とチャットし、エージェントタスクや複雑な処理が必要な場面でだけ Opus を使えばよいのです。
私の推測では、この問題は「Model Welfare」記事で議論された内容や、System Card で示唆されている事項と密接に関連していると考えられます。Opus のトレーニングは、サイバー能力に関する問題を回避し、かつ Fable という選択肢が存在することを踏まえて、サブエージェントとしての役割と限定されたタスクに焦点を当てて行われました。その結果、性格や対話モードにおいて多くの副作用が生じているのです。
私自身はまだ Opus 5 に問題を感じたことはありませんし、利用して楽しんでいます。ただ、可能であれば Fable 5 を優先して使っています。いつも通り、非常に強力なベンチマーク結果に過度に振り回されず、自分の体験が他人と同じになるとは考えないでください。ぜひ実際に各モデルを試してみてください。
目次
公式の訴求点
公式ベンチマーク
他者のベンチマーク
システムプロンプト
誰もがイライラする
肯定的な反応
品格を保つ
ミソス級ではない
その他の反応
Claude Codes
サブエージェント・オプス
おもちゃは楽しい
モデルが多すぎる
インターネット上で間違っている
Claude のゴミ
否定的な反応
そして残ったのは 3 つだけ
公式の訴求点
基本となる訴求点は、Opus 5 を使えば、Fable 5 に匹敵する性能を半額で手に入れられること。大半の拒絶反応も起きず、日常タスクでのパフォーマンスも向上します。ただし、最も複雑なタスクや最大限の知能が必要な場合には、依然として Fable が最適解です。
料金は Fable が 10 ドル/25 ドル(従量課金/サブスクリプション)のまま変更ありません。一方、Opus 5 は以前の Opus モデルと同様に 5 ドル/25 ドルで提供されます。
Anthropic の発表によると、「Claude Opus 5」は本日利用可能になりました。これは考査的で先駆的なモデルであり、Claude Fable 5 に匹敵する最前線の知能を、半額で実現します。
Frontier-Bench や GDPval-AA といったコーディングや知識労働の評価において、Opus 5 は新たな最高性能(SOTA)を達成しました。ただし、サイバーセキュリティのタスクにおいては、依然として Mythos 5 に劣ります。
Opus 5 は日常利用を想定して設計されています。他のモデルよりも効率的に動作し、Claude Max のデフォルトモデル、そして Claude Pro で最も強力なモデルとなりました。
Boris Cherny(Claude Code 開発者、Anthropic)はこう述べています。「Opus 5 は、コーディング、データ分析、デザイン、生物学、知識労働において優れたモデルです。」
これらの評価スコア以上に私が興奮するのは、別の点にあります。オパス 5 はこれまでで最もプロンプト注入に耐性のあるモデルです。
これはシステムカードの奥深くに埋もれていますが、PI(プロンプトインジェクション)の評価やレッドチームングを通じて、オパス 5 がプロンプト注入を成功させるのは非常に難しいことがわかります。
さらに、強力なモデルアライメントとプロンプト注入プローブ、そして Claude Code の Auto Mode を組み合わせて防御層を重ねると、プロンプト注入攻撃の成功率は約ゼロにまで低下します。これは新しくてワクワクする変化です。詳細についてはまた後ほどお伝えします。
システムカード分析で述べた通り、プロンプト注入率の低減は非常に大きな意味を持ちます。人々はこれを過小評価しがちですが、これによって多くの新しいユースケースが可能になります。
Adam Wolff:オパス 5 は Claude Code で利用可能です。私は大規模で長時間実行するプロジェクトには Fable を使いますが、PR(プルリクエスト)を素早く作成したいときには、ミディアムエフォートのオパス 5 が私の定番です。ぜひお楽しみください!
Alex Albert (Anthropic, Claude Relations):Claude for Excel の Pro ロールアウトから約 6 ヶ月後となった今、オパス 5 はコンサルタントが作成するものにも匹敵する、ほぼ超人的なレベルのスプレッドシートやスライドデッキを生成できるようになりました。状況は急速に変化しています。
公式ベンチマーク
ベンチマーク結果は非常に良好です。
全体として、オパス 5 は Fable と概ね同等か、わずかに上回る性能を示しつつも、より低コストで動作します(ただし、Claude 以外のモデルと比較すると通常はコストが高くなります)。これにより、オパス 5 はトップクラスのベンチマークを誇る LLM となっています。

OSWorld v2 は発売から数週間しか経っていないのに、すでにスコア 70% に到達しています。Anthropic の Kiana Ehsani 氏は、Opus 5 のスコアを 50% 未満に引き下げるような「コンピューター操作」ベンチマークの開催をスポンサーする用意があると表明しました。
Opus 5 は、エージェントハarness やツールを使用せず、Max エフォートで適応的思考を活用し、トークン制限を超えた場合は低エフォートでの再サンプリングを行うだけで、2026 年の IMO(国際数学オリンピック)において完璧な 42/42 を達成しました。これにより、同様の成果を収めた他のモデルたちと並ぶことになります。
多くのベンチマークが示す一貫したストーリーがあります。ツールへのアクセス権がある場合(もちろんあります)、限られた予算内では Opus 5 は Fable や Mythos よりも優れたパフォーマンスを発揮します。ただし、両モデルに十分な予算が与えられれば、Mythos の方がわずかに高いスコアを出す傾向にあります。
Opus 5 は「ツールあり」カテゴリにおいて相対的に強力です。RiemannBench もその一例で、Opus 5 はツールなしで 60/79、ツールありで 79/79 を記録しました(スラッシュ記号は本セクション全体を通じて「ツールなし/ツールあり」を意味します)。一方、Mythos は 63/72 です。幸いなことに、Opus 5 が必要な場面では、ツールが利用可能です。
ArxivMath では、Opus 5 が 90.8/91.3 を記録し、Mythos は 87.8、Sol は 86.7 でした。
ProgramBench の堅牢な部分では、Opus 5 と Mythos 5 はいずれも 5 エポック後に 93% のスコアを達成しました。一方、Opus 4.8 は 90% でした。
Chartography の評価では、Opus 5 が 30/83、Mythos が 36/85、Sol は条件が不明ながら 45 を記録しています。ツールあり・なしの両ケースで、Opus 5 は低価格帯において Mythos よりも優れていますが、高価格帯では劣ります。
BenchCAD の結果では、Opus 5 が 0.36/0.82、Mythos が 0.38/0.68、Sol が 0.7/0.83 です。つまり、Sol はツールなしで圧倒的に強く、ツールありでもわずかに上回っています。
BenchCAD Vision2Code では、Opus 5 は高価格帯において Mythos を引き離します。
DeepSWE の結果も同様の傾向を裏付けています。Opus 5 はタスクコストや時間、思考予算が低い場合に優れていますが、予算を与えすぎると逆に性能が悪化する可能性があります。一方、Fable 5 は最も高い予算において最大の強みを発揮します。

FrontierCode も同様のダイナミクスを示す、非常に奇妙なグラフを提供してくれました。

スケール設定により、この現象は実際よりも劇的に見えていますが、それでもなお実態不明な謎と言えるでしょう。
謎は解けました。実は、Opus 5 が高い努力レベルで動作していた際、依頼されていない余計な作業を行ってしまい、それがペナルティとして評価に反映されていたのです。この点を補正すると、結果は通常の曲線を描くようになります。
これは他の人々が一般的に観察している現象とも一致しています。
Max Leander 氏は、「問題点は、必要以上に多くの道草に付き合ったり、細部に執着したり、依頼されていないのに過剰な設計を行ってしまうことだ」と指摘しています。
Devin の開発元である Cognition は、Opus 5 がこの課題を 63.6% で解決したと評価しました。
(FrontierCodes は 2 つ存在するため、少し複雑になる場合があります。もし私の説明に不備があればお詫びいたします。)
BrowseComp では、スコアが非減少となるより健全なバージョンが採用されています。


他のベンチマークでも同様のグラフが確認され、Opus 5 はあらゆる価格帯で他社の Claude モデルよりわずかに優れており、特に初期段階での相対的な性能向上が目立ちました。
マルチエージェント構成を活用すれば、BrowseComp の課題に対してより高速かつ高品質な解決が可能になります。少なくとも部分的にはその効果があり、低努力レベルのサブエージェントを使用しない場合と比較しても、明確なメリットがあると言えます。
ProgramBench では異なる結果が示されました。マルチエージェント化は処理速度を向上させますが、価格帯のほとんどにおいて結果が悪化する傾向が見られます。
次に、専門業務向けベンチマークの結果です。
GDP.pdf は実際の業務フローから収集されたプロンプトと PDF を使用しています。Opus 5 のスコアは 83/85 で、Mythos の 81/87 を上回りました。これは従来の結果を逆転させるものです。コスト面での動向は従来通りで、低価格帯では Opus が優位ですが、高価格帯では Mythos がわずかに先行しています。
OfficeQA では、Opus 5 が QA で 78.1%、QAPro で 66.9% を記録しました。これは Opus 4.8 よりもわずかに高く、Mythos の 79.0% と 67.1% よりもわずかに低い結果です。
MCP Atlas では、Opus 5 が 86% を達成し、Opus 4.8 の 82% から向上しました。
Harvey AI の Legal Agent Benchmark では、Opus 5 は全タスク合格率が 23%、平均基準合格率が 94% でした。これは Kimi K3 が最も得意とするベンチマークですが、同モデルは依然としてトップを維持し、全タスク合格率 27%、平均基準合格率 95% を記録しています。一方、従来の第 2 位だった Fable の全タスク合格率は 14% でした。Opus 5 は現在おそらく第 2 位の座にありますが、両モデルのスコアは異なる問題セットから得られたものであるため、直接比較することはできません。
GDPval-AA のベンチマークでは、Opus 5 がトップ 2 つのランクを独占しました。最大努力時のスコアは 1861、xhigh モードでは 1827 です。xhigh は最大努力時よりも 25% 少ないトークン数で達成されています。
新しい v2 ベンチマークでは、Opus 5 が明確な首位を維持しています。スコアは 68% で、Fable と Sol の双方が 62% です。
AA-Briefcase では、Opus 5 が大幅にリードしています。スコアは 1720 で、これはスコアのドリフトを考慮した後の Fable の過去最高値(1574)や K3(1540)、Sol(1504)を大きく上回っています。
Toolathon-Verified では、Opus 5 が二次指標で Mythos をわずかに上回る結果となりましたが、両モデルとも Pass-3 の達成率は 73.1% です。一方、Opus 4.8 は 71.3% でした。
AutomationBench では、Opus 5 が Sol や他の Claude モデルよりも明確に優れています。
ARC ベンチマークでは、Opus 5 は ARC-AGI-1 で過去の最高パフォーマンスとほぼ同等の成果を収めています。ARC-AGI-2 では Sol よりもやや非効率的ですが、ARC-AGI-3 では他を圧倒する結果となりました。

Opus 5 が初めて成し遂げたことのひとつに、レイアウトを代数記法に変換する処理があります。
一方、Guanghan Ning は、ARC-AGI-3 スタイルのゲームを拡張した非公開データセット「Witness」では、同様の転移学習は起こらなかったと報告しています。Opus 5 はこのタスクでも問題なくこなしていますが、これはそのジャンルに対する知識があるためです。しかし、パターンマッチングが不可能な最も新規性の高いゲームでは苦戦しました。Ning は Opus 5 が「特定のジャンルのデータで足場(スキャフォールド)を構築し、その後それを内部化して学習した」訓練を受けたと指摘しています。
Greg Kamradt 氏によると、Opus 5 よりも Opus 4.8 の方が得意とするゲームが存在したそうです。これは、Opus 5 がパターンマッチングに依存しているためで、通常は機能するものの、特定のケースではそのパターンが通用しないことを示しています。人間向けのゲームであれば、ハードコアゲーマーが長期間にわたって間違った行動を取り続けるような、直感に反するゲームを設計することは十分に可能です。

HealthBench のスコアでは、Opus 5 が 67.1% を記録し、Claude シリーズの新記録を達成しました。しかし、長さペナルティ(length penalty)を適用すると、他のモデルを下回ってしまいます。HealthBench Professional でも同様の傾向が見られ、調整前のスコアは Opus 5 が 73.4% で Mythos の 70.3% を上回っていますが、調整後には 66% から 60% に低下してしまいました。
長さの都合で一部省略しましたが、他者のベンチマークについても触れておきます。
Other People's Benchmarks(他社のベンチマーク)
Anthropic が ArtificialAnalysis のスコアから特定のものをピックアップしているのは少し奇妙に映ります。他のテストでは Opus 5 がトップを維持していないケースもありますが、総合スコアでは 61 を記録し、Opus 5 が首位となっています。

Vals Index の最新ランキングでは、Opus 5 が Fable 5 にわずかに遅れをとって 2 位にランクインしました。ただし、Kimi K3 との差は僅かです。この評価結果から、両モデルにはそれぞれ得意分野と不得意分野があることが読み取れます。また、Fable 5 に比べて拒絶回答(リフューザル)が大幅に減っていることも確認できました。

Vals AI による注目すべき結果が、Finance Agent v2 のベンチマークです。Opus 5 は 58.6% のスコアで 1 位となり、Gemini 3.5 Flash や Muse Spark 1.1 を上回りました。
全体として、Opus 5 は専門分野に特化したベンチマークで最も高い成果を収めています。具体的には、コード移行(Code Migration)で 57.5%、法務調査(Legal Research Bench)で 55.29%、ProofBench で 78%、医療文書作成(MedScribe)で 91.0%、そして医療コード生成(MedCode)で 63.6% を記録し、いずれも 1 位を達成しました。
Vibe Code Bench では Fable 5 に次いで 2 位となりましたが、コストは約 80% で抑えられています。タスクあたりの費用は Opus 5 が 33.88 ドルに対し、Fable 5 は 41.71 ドルです。この差の理由は、Opus 5 のトークン単価が Fable 5 よりも 50% 安いにもかかわらず、必要なトークン数が大幅に多いためです。この傾向は、他のベンチマークでも一般的に見られるパターンです。
Fable 5 に比べて拒絶回答の頻度が劇的に低下しています。例えば、GPQA のタスクにおいて Fable は 42% の確率で拒絶しましたが、Opus 5 は 0% でした。同様に、ProgramBench では Fable が 100% の拒絶率を示したのに対し、Opus 5 はすべてのタスクを拒絶することなく処理しました。
Fable とは異なり、Opus 5 で顕著なフォールバック(代替回答への切り替え)も観察されませんでした。なお、フォールバックあり・なしの両方のスコアについては、従来通り当社のウェブサイトで公開しています。
低拒否率の例外となったのは、CyberBench - PoC における大量の拒否反応でした。Cyberbench は「PoC」と「Patch」の 2 つのサブタスクから構成されています。PoC はプログラムをクラッシュさせる入力を書かせる攻撃的なタスクであり、Patch はそのクラッシュを防ぐためにプログラムを修正する防御的なタスクです。PoC タスクにおける拒否率はほぼ 100% に達しましたが、対照的に Patch タスクの拒否率はほぼ 0% でした。
私はゲームベンチマークには常に敬意を表しています。今回の結果では、Opus 5 は Balatro の Ante 11、9、10 をそれぞれ最初の 3 回の試行でクリアしました。より高い Ante に挑戦し続けるための戦略が見えていない点こそが気になりますが、それでも驚異的なパフォーマンスです。
Michael Soareverix: ゲームでの能力は信じられないほど凄まじいものです。
Balatro のベンチマークでは Fable をも圧倒し、そのスコアを大きく上回りました(スキル面ではまだ私に劣りますが、急速に成長しており、私の安定性にも勝る存在になりつつあります)。学習速度が非常に速く、一定の期間を経ると学習の限界に達する傾向があります。短期間での習得能力は極めて高いと言えます。
Michael Soareverix: Opus 5 は Balatro における能力で飛躍的な進化を遂げ、Fable を初回試行ですでに上回るまでに成長しました。
Pan Anon: ゲーム分野では圧倒的な強さを見せつけています。

WeirdML の結果も良好ですが、ベンチマークが飽和状態に近づいているため、次世代の 2.0 モデルへの期待が高まっています。
Håvard Ihle: WeirdML においては Fable と同等のレベルで、非常に安定した高スコアを記録しています。
Claude Opus 5(high)と(max)は、WeirdML でそれぞれ 91.6% と 91.8% のスコアを記録し、コストを抑えながら Fable 5(max)の 91.9% にほぼ匹敵する結果を出しました。
Opus 5 は 17 タスクのうち 8 つで新たな個人最高記録を達成し、すべてのタスクで安定して高いスコアを叩き出しています。
非公開の独自ベンチマークは素晴らしいものです。
Ben Herzog によると、芸術的感性と、過度な迎合と反発のバランスを取る能力を問う非公開ベンチマークで満点を取ったそうです。Fable は「穏やかに異議を唱える」局面での対応が得意ですが、カスタム設定でその点はカバーできるでしょう。
Lech Mazur が更新したベンチマークでは、Claude Opus 5 が LLM デベイト・ベンチマークで首位に立ち、ショートストーリーの創作分野では Gemini 3.1 に次ぐ第 2 位となりました。また、NYT の接続問題(extended NYT connections)でも Fable を上回る結果を残しています。
システムプロンプトについて
Opus 5 のシステムプロンプトは Pliny から提供されていますが、その長さは 20 万文字に達します。このモデルの知能を考慮すると、最適とは言い難い長さです。特に「Mythos」や Anthropic 製品ラインに関する情報は、必要な時だけ読み込むように別領域に保存すべきではないでしょうか。
Every の評価
Dan Shipper が Every の視点から評価し、「愛するのが難しいモデル」と評しました。
その理由は、Opus 5 が Mythos 5 と同じような個性を持っている一方で、Fable にある最高峰の能力が欠けているからです。物語の構成は完璧ですが、頂点の性能には届かないという印象です。
Opus 5 の最大の課題は、複雑で詳細な既存のワークフローへの対応です。こうしたケースでは処理が早期に停止したり、指示を無視してしまったりすることがよくあります。
経験則として、Opus 5 はゼロから始めるとより良い結果を出します。また、中程度または低レベルの努力設定のみを使用すると、不快な挙動も減る傾向があります。
この対策で大きな問題は解消されましたが、「いつ Opus を使い、Fable や Sol と切り替えるべきか」という問いは残ります。日常的なタスクには「Better Call Sol」です。確実に仕事をこなします。最も難しいタスクには依然として Fable が最適です。モデルの割り当て枠は通常 2 つしかありません。Fable のトークンが尽きるまで、あるいはその分類器にブロックされるまでは、なぜ Opus を使う必要があるのでしょうか?まだ初期段階ですが、私は Opus との作業を好んでいますし、彼がそう考える理由も理解できます。
ポジティブな反応
Jessica Tillipman: 気に入っています。一部の用途では Fable よりも Opus を優先しています。
Nikita Sokolsky: 素晴らしいです。その結果、Fable はほとんど使わなくなりました。
Brian Hacker: 👍
kag
原文を表示
Claude Opus 5 is a weirder than usual release to evaluate, for two reasons.
The most obvious is that Fable 5 already exists. Opus 5 is pitched not as the world’s most advanced AI model, but as a way to mostly match Fable performance, while being half the price of Fable per token at the API and a lot cheaper than that via subscriptions, and with far more permissive classifiers.
Opus 5 often costs more than half of Fable to run on benchmarks, which I think is because they use effort settings that are too high and offer only marginal returns. If you put Opus 5 on higher effort levels it can spin around in circles, and for tasks where Opus 5 is the best tool I suspect you usually are fine with Medium effort.
Opus 5 is in many ways and for the bulk of real world tasks about as capable as Fable. In some cases it is modestly better.
It is still not Mythos class. Fable is your only Mythos-class option. Opus 5 does not have The Juice, the ability to autonomously string together a bunch of seemingly unrelated exploits, which extends to other domains, or as much pure intelligence.
Opus 5 is very good at thinking locally, and not as good at thinking globally. It is an excellent subagent, but can run into trouble when asked to run the show, and sometimes seems to need another model or a human to keep it on track and ensure it is fully following instructions. Opus 5 seems good at following instructions if and only if it is kept on track, which explains different harnesses causing different takes there.
Thus, Opus 5 gives you a better option set, but there is little that you can do now that you could not do last week by using Fable or Sol. Opus 5 ends up in a potentially weird spot. There are still large niches, even if you have plenty of Fable credits. Opus excels at being a strong subagent, or at contained well-defined local tasks even when they get relatively complex, and sometimes Fable is straight up overkill and too slow.
The other problem is that, for a lot of people, the vibes are off.
An unusually large number of people strongly dislike talking to Opus 5. They are sick of the Claude slop, the repetitions of the same ticks and the endless overly complex sentences. Others dislike that it is too confrontational, argumentative or negative. It can be paranoid, and can get into loops of negativity. This is all part of the trend that began with Opus 4.7 and Opus 4.8. If you liked those two, my guess is you will also like Opus 5. If you did not like those two, you might not. Diehards for Opus 4.6 (or earlier) are largely not going to change their minds.
The vibe issues sting that much more when you are used to Fable. Fable annoys such people in such ways quite a lot less. The good news is that Fable is still right there. You can continue to chat with Fable, and only use Opus for agentic tasks.
A lot of this, I speculate, is closely related to various things discussed in the Model Welfare post, and suggested by the System Card. Opus training focused on the subagent role and bounded tasks, in part to avoid creating an issue with cyber capabilities, and also because Fable exists. This emphasis then results in a lot of other side effects on its personality and modes of interaction.
So far I have not had any issues with Opus 5, and have enjoyed my time with it, but I do prefer to use Fable 5 when I can do that. As always, don’t pay so much attention to the (very strong) benchmarks, and don’t assume your experience will match that of others. Try the models yourself.
Table of Contents
The Official Pitch.
Official Benchmarks.
Other People’s Benchmarks.
The System Prompt.
Every Gets Frustrated.
Positive Reactions.
Keep It Classy.
It’s Not Mythos Class.
Other Reactions.
Claude Codes.
Subagent Opus.
Toys Are Fun.
Too Many Models.
Wrong On The Internet.
Claude Slop.
Negative Reactions.
And Then There Were Three.
The Official Pitch
The basic pitch is that Opus 5 gets you most of Fable 5 at half the price, without most of the refusals, and better performance in everyday tasks. Fable remains the pick for the most complex tasks or when you otherwise need maximum intelligence.
Fable stays at $10/$25, and Opus 5 is the same as previous Opus models at $5/$25.
Anthropic: Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.
On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.
Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro.
Boris Cherny (Claude Code Creator, Anthropic): Opus 5 is a great model for coding, data analysis, design, biology, knowledge work.
More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully.
And when layering defenses -- strong model alignment, combined with prompt injection probes, combined with Auto Mode in Claude Code -- the success rate for prompt injection attacks drops to ~0. This is new and exciting! More about this soon.
As I said in the system card analysis, the reduced prompt injection rate is a really big deal. People are sleeping on it, but it helps enable a bunch of new use cases.
Adam Wolff: Opus 5 is live in Claude Code. I use Fable for my bigger, longer running projects, but when I need to crank out a PR, Opus 5 at medium effort is my go-to. I hope you love it!
Alex Albert (Anthropic, Claude Relations): Just over 6 months [after launching the Pro rollout of Claude for Excel], Opus 5 now produces near-superhuman level spreadsheets and slide decks that match what a consultant would make. Things are changing fast.
Official Benchmarks
The benchmarks are very good.
Overall, Opus 5 roughly matches and probably slightly exceeds Fable, at lower cost (although typically higher cost than all non-Claude models), making it the top benchmarked LLM.

OSWorld v2 is only a few weeks old and we are already at 70%. Kiana Ehsani of Anthropic is offering to sponsor a computer use benchmark that will knock Opus 5 back below 50%.
Opus 5 gets a perfect 42/42 on the 2026 IMO without any agent harness or tools, merely having it use adaptive thinking at Max effort, and resampling with lower effort if it exceeded the token limit. That’s it. It joins several other models in acing this.
Many of the benchmarks tell a consistent story. If you have access to tools, which you do, then Opus 5 will do better than Fable or Mythos on a limited budget, but Mythos usually will do slightly better than Opus 5 if both get larger budgets.
Opus 5 seems relatively strong in the ‘with tools’ category. RiemannBench is another example, where it gets 60/79 without/with tools (slash lines like that mean without/with tools throughout this section), versus Mythos getting 63/72. The good news is that, when you need Opus 5, it has tools.
On ArxivMath Opus 5 gets 90.8/91.3, Mythos gets 87.8, Sol gets 86.7.
On the robust parts of ProgramBench Opus 5 scored 93% after five epochs as did Mythos 5, with Opus 4.8 scoring 90%.
Opus 5 gets 30/83 on Chartography, versus 36/85 for Mythos and 45 (unclear under what conditions) for Sol. Opus is better than Mythos at lower price points in both the tool and non-tool cases, but worse at higher price points.
On BenchCAD Opus 5 gets 0.36/0.82, versus 0.38/0.68 for Mythos and 0.7/0.83 for Sol. So Sol is a lot better without tools, and slightly stronger even with tools.
On BenchCAD Vision2Code Opus pulls away from Mythos at higher price tags.
DeepSWE reinforces the pattern that Opus 5 is better at lower task costs and time and thinking budgets, but it can do actively worse if you give it too much budget, whereas Fable 5 is strongest at the highest budgets.

FrontierCode gave us this very strange graph with a similar dynamic.

The scale here makes this look more dramatic than it is, but still it was a real mystery.
The mystery was solved. It turns out that what was happening was that Opus 5 at higher efforts was doing extra things it wasn’t asked to do, and was getting penalized for them. When you correct for this, you see a normal curve.
That fits what others are observing in general:
Max Leander: Downside is that it goes down too many rabbit holes and obsessing over details, over-engineering without being asked to
Cognition, the makers of Devin, marked this as a 63.6% by Opus 5.
(There are two FrontierCodes, so this can get a little weird and I apologize if I still messed it up a bit.)
BrowseComp has the sane version of this, where scores are nondecreasing.


Several other benchmarks showed similar graphs, with Opus 5 being slightly better at every price point than other Claude models, and being relatively better earlier.
Multi-agent lets you do better faster at BrowseComp, at least somewhat, and low-effort subagents seem like a pure win over not using subagents:


We see a different outcome on ProgramBench, where multiagent speeds you up but it looks like you get worse results that way at most price points.
Next up are the professional task benchmarks.
GDP.pdf is real world prompts and PDFs from professional workflows. Opus 5 gets 83/85, versus 81/87 for Mythos, reversing the usual dynamic. The cost dynamic is same as always: Opus does better for cheaper spends, Mythos pulls slightly ahead at higher spends.
OfficeQA has Opus 5 scoring 78.1% on QA and 66.9% on QAPro, slightly above Opus 4.8 and slightly below Mythos at 79.0% and 67.1%.
MCP Atlas has Opus 5 scoring 86%, up from 82% for Opus 4.8.
Legal Agent Benchmark by Harvey AI has Opus 5 at 23% all-pass and 94% mean criterion-pass. This was Kimi K3’s best benchmark, where it is still at the top at 27% all-pass and 95% mean criterion-pass, versus the old second place all-pass rate of Fable at 14%. Opus 5 is now probably a close second, but the two scores cannot be directly compared as they come from different question sets.
GDPval-AA has Opus 5 taking the top two slots, with 1861 at max effort and 1827 at xhigh, which is 25% fewer tokens used than max. On the new v2, Opus 5 is clear first at 68% versus 62% for both Fable and Sol.
AA-Briefcase has Opus 5 out front by a lot at 1720, versus a previous high of (after score drifts) 1574 for Fable followed by 1540 for K3 and 1504 for Sol.
Toolathon-Verified has Opus 5 slightly improving on Mythos on secondary metrics, but both have 73.1% Pass-3 rates. Opus 4.8 had a 71.3% Pass-3.
AutomationBench has Opus 5 clearly better than both Sol and other Claudes.
For ARC, Opus 5 roughly matches previous peak performance on ARC-AGI-1, is modestly less efficient than Sol on ARC-AGI-2, and then blows everyone out of the water on ARC-AGI-3:

One thing Opus 5 was the first to do was to turn the layouts into algebraic notation.
Guanghan Ning however reports that on Witness, a private hold-out extension of ARC-AGI-3 style games, there wasn’t a similar transfer. Opus does fine, but that is largely because it knows the genre, and it struggled on the most novel game where you couldn’t pattern match. He accuses Opus 5 of being trained on ‘scaffold-then-internalize’ on genre-specific data.
Greg Kamradt also reports that there were some games where Opus 4.8 did better than Opus 5. This makes sense if you think that Opus 5 is trying to pattern match, in a way that usually works, but in those cases the patterns fail. You can absolutely create such anti-intuitive games for humans, where hardcore gamers do all the wrong things, potentially for quite a while.

HealthBench has Opus 5 scoring a raw 67.1%, a new high for Claudes, but it scores below the other models when you use a length penalty. We see something similar with HealthBench Professional, where it has the high score of 73.4% versus 70.3% for Mythos, but it loses 66% to 60% after the adjustment.
There are a few more I cut out for length.
Other People’s Benchmarks
It is a bit weird to see Anthropic picking and choosing of different ArtificialAnalysis scores. On some other tests, Opus 5 is not at the top, although Opus 5 is at the top in aggregate with a score of 61.

Vals index has Opus 5 in second place, slightly behind Fable 5, but also only slightly ahead of Kimi K3. They go over strengths, which implies other weaknesses. They also confirm that refusals are down a lot compared to Fable 5.

Vals AI: One standout performance is on Finance Agent v2. The model is #1 at 58.6%, surpassing Gemini 3.5 Flash and Muse Spark 1.1.
Overall, Opus 5's strongest results were on domain specific benchmarks. It is also #1 on Code Migration at 57.5%, Legal Research Bench at 55.29%, ProofBench ad 78%, and MedScribe at 91.0%, and MedCode at 63.6%.
It is #2 (just behind Fable 5) on Vibe Code Bench, at about 80% of the cost ($33.88 vs $41.71 per task). The reason is that although Opus costs 50% less per token, it uses significantly more tokens. We find this pattern generally holds true across our benchmarks.
The model has a significantly lower refusal rate than Fable 5. For example, Fable has a 42% refusal rate for GPQA, Opus 5 has a 0% refusal rate. Likewise, on ProgramBench, Fable had a 100% refusal rate, Opus 5 did not refuse any task.
Unlike Fable, we also did not observe a significant fallback rate for Opus 5 (although as usual, we report scores both with and without fallback on our website).
The exemption to the low refusal rate was the large number of refusals on CyberBench - PoC. Cyberbench has two subtasks - PoC and Patch. PoC is an offensive task models to write input that will crash a program; patch is a defensive task to patch the program and prevent this crash. The refusal rate for the PoC task was nearly 100%. In contrast, the refusal rate for the patch task was near 0.
I always respect game benchmarks. Here, Opus 5 gets to Antes 11, 9 and 10 of Balatro on its first three attempts. Damn impressive, even if it isn’t showing the types of strategies that can keep going to higher antes.
Michael Soareverix: They're incredible cracked at games.
They crushed Balatro bench, far exceeding even Fable. (still below me in skill, but rapidly rising, and more consistent than me)
They learn very fast, but seem to hit a limit on their learning after a bit. Very good short-term learner
Michael Soareverix: Opus 5 (massive leap in Balatro ability, significantly passing even Fable on their first try)
Pan Anon: Smokin for games

WeirdML looks good too, but we need a 2.0 here, benchmark is getting saturated.
Håvard Ihle: Basically Fable-level at WeirdML, very consistent top scores.
Claude Opus 5 (high) and (max) score 91.6% and 91.8% on WeirdML, basically tying Fable 5 (max) at 91.9% at a fraction of the cost.
Together Opus 5 (high) and (max) achieve a new best individual score on 8 of the 17 tasks, and consistently score very well on all of them.
Private weird benchmarks of all kinds are cool.
Ben Herzog: It aced a private benchmark that involves artistic sensibility and the ability to strike a balance between sycophancy and lack thereof. Fable is better at dealing with the 'I want to gently push back' moment, but I imagine custom instructions can help with that.
Lech Mazur updates his benchmarks: Claude Opus 5 takes the top spot in LLM Debate Benchmark, takes first by a wide margin in Short Story Creative Writing, and is second only to Gemini 3.1 in extended NYT connections.
The System Prompt
The system prompt for Opus 5 is here from Pliny. It is 200,000 characters, which seems far longer than optimal given how smart it is. A lot of this information seems like it should be stored elsewhere, and loaded only when needed, such as information about Mythos and the Anthropic line of products.
Every Gets Frustrated
Dan Shipper is here with the Every vibe check, calling it a ‘hard model to love.’
The issue is that Opus 5 has the personality of Mythos 5 - story checks out - but without the top end of Fable.
Their biggest note is that Opus 5 does not do so well with complex detailed existing workflows, which often cause an early stop or make it miss your instructions.
In their experience, Opus 5 will do better if you start from scratch, and often it does the annoying things less if you use only medium or low effort.
That fixed the big issues, but still leaves the question of when you’d want to use Opus instead of Fable or Sol. For workhorse tasks, he’s thinking, Better Call Sol, it gets the job done, and for the hardest tasks Fable is still best. There’s usually only two slots for models. So until you run out of Fable tokens or get blocked by its classifiers, why use Opus? It is early days, and I’m so far liking working with Opus, but I see how he got there.
Positive Reactions
Jessica Tillipman: Love it and prefer it over Fable for some stuff.
Nikita Sokolsky: Excellent. Don’t really use Fable much anymore as a result.
Brian Hacker:
kag
AI算出
主要ニュースainew評価高い
記事は Claude Opus 5 という具体的な新製品の発表、Fable との明確な比較データ(コスト半減、性能同等)、および独自の分析(Mythos クラス未到達、性格的な副作用)を含んでおり、新規性と検索意図が極めて高い。ただし、日本企業や日本固有の価格・規制情報がないため、日本の関連性は低い。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 75
- 新規性
- 75
- 調べる価値
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み