Anthropic、Claude Opus 5 の詳細を公開(3 分読了)
Anthropic は新モデル Claude Opus 5 を発表し、同社によると Frontier-Bench や GDPval-AA などで業界最高性能を達成したが、サイバーセキュリティ分野では Mythos 5 に劣る。
AI深層分析を開く2026年7月28日 00:08
AI深層分析
キーポイント
コスト効率と性能の両立
Anthropic によると、Claude Opus 5 は前作 Opus 4.8 と同等のコストで大幅な性能向上を達成し、特にソフトウェアエンジニアリングタスクでは他モデルを凌駕する。
主要ベンチマークでの首位
Frontier-Bench v0.1 や ARC-AGI 3 などの評価において、Opus 5 は他のモデルを大きく引き離し、特に問題解決タスクでは次点の3倍のスコアを記録した。
特定分野での競合優位性
コーディングや知識処理においては Fable 5 に匹敵する性能を発揮するが、サイバーセキュリティタスクに限っては Mythos 5 の後塵を拝している。
サービスへの統合と設定
Opus 5 は Claude Max のデフォルトモデルとなり、Claude Pro では最強のモデルとして提供されるが、コストと速度のバランスを取るための努力設定(effort setting)も用意されている。
重要な引用
Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.
On Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8's performance at a lower cost per task.
Opus 5 is the new default model on Claude Max, and the strongest model on Claude Pro.
編集コメントを表示
編集コメント
Claude Opus 5 の登場は、高性能モデルが低コストで利用可能になるという業界のトレンドを加速させる。特にサイバーセキュリティ分野での競合優位性については、用途に応じたモデル選定がさらに重要となることを示唆している。
Claude Opus 5 が本日利用可能になりました。このモデルは思考深く、先回りする能力を持ち、価格が Claude Fable 5 の半分でも、最前線の知能に迫る性能を発揮します。
Frontier-Bench や GDPval-AA といったコーディングや知識作業の評価において、Opus 5 は新たな最高記録を達成しました。ただし、サイバーセキュリティのタスクにおいては Mythos 5 にまだ及びません。
Opus 5 は日常使いを想定して設計されています。他のモデルよりも効率的に動作し、Claude Max のデフォルトモデルとなり、Claude Pro では最も強力なモデルとして採用されました。

パフォーマンスとコストパフォーマンス
Claude Opus 5 は、前作の Opus 4.8 と同じコストで、大幅に向上したパフォーマンスを提供します。このセクションのグラフは、モデルの努力設定(effort setting)に応じてパフォーマンスがどのように変化するかを示しています。顧客はこの設定を使って、知能を最大化するか、トークンを節約してより高速かつ低コストな結果を得るかを最適化できます。
Opus 5 は、価値の高いソフトウェアエンジニアリングタスクにおいて卓越した性能を発揮します。例えば、Frontier-Bench v0.1では他のすべてのモデルを凌駕し、タスクあたりのコストを抑えながら Opus 4.8 のパフォーマンスを倍以上に引き上げました。CursorBench 3.2でも最大努力設定で Fable 5 の最高スコアと 0.5% 以内の差しか出ず、かつタスクあたりのコストは半分です。さらに、高・超高・最大努力設定において、同じコストであれば他のすべてのモデルよりも高い性能を達成しています。
知識労働や問題解決に関するタスクでも同様の結果が得られています。具体的には以下の通りです。
- ARC-AGI 3(新規の問題解決を評価するベンチマーク)では、Opus 5 のスコアは次点のモデルの 3 倍に達します。
- Zapier AutomationBench(ビジネスタスクを最初から最後まで完遂できるかを測定するベンチ)では、同じコストで Opus 5 のパス率は次点のモデルの約 1.5 倍です。最低努力設定であっても、Opus 5 は他のどのモデルよりも多くのタスクを成功させます。
- OSWorld 2.0(コンピュータ操作を評価するベンチ)では、あらゆるコスト設定において他モデルを上回り、Fable 5 の最良結果をコストの約 3 分の 1 で達成しています。
これら関連する評価項目においても、Opus 5 は最も高性能かつコスト効率に優れたモデルです。
Opus 5 は、科学的研究において Opus 4.8 よりも有意な進歩です。構造生物学、有機化学、バイオインフォマティクスなどを含む生命科学研究の評価項目すべてで、Opus 4.8 を上回る性能を示しています。
その改善点は特に顕著です。分光データから分子構造を推論するといった有機化学のタスクでは、内部ベンチマークで Opus 4.8 よりも 10.2 ポイント高いスコアを記録しました。また、タンパク質の配列変化が機能にどう影響するかを予測するタンパク質関連のタスクでも、7.7 ポイント高いスコアを達成しています。
さらに、Opus 5 ははるかに強力な視覚出力も生成できるようになりました:
Claude Opus 5 の活用方法
Claude Opus 5 は、自身の作業を検証し、成功するまで慎重に反復処理を行う能力が大幅に向上しました。評価や早期アクセステストにおいて、私たちが発見した事例やユーザーからの報告を通じて、Opus 5 の自律性と徹底性が数多く確認されています:
Frontier-Bench のあるタスクでは、Opus 5 に機械部品の図面が提示され、それを 3D FreeCAD モデルとして再構築するコードを書くよう求められました。ただしこのタスクでは、モデルが図面を直接見られるような仕組みは意図的に用意されていませんでした。それに対し Opus 5 は、自らコンピュータビジョンパイプラインを記述して生ピクセルから幾何情報を抽出し、機械部品全体を再構築しました。これを何度も成功させ、同じ設定で競合モデルが 5 回試しても解けた例はありません。
人気のあるオープンソースパッケージマネージャーに実際に存在するバグに対し、Opus 5 は根本原因を特定し、コミュニティの修正パッチで見落とされていたエッジケースも正しく修正しました。一方、競合モデルは表面症状のみに対処して根本原因には手を付けず、「バグは解決した」と報告するにとどまりました。
ある取引所のエンジニアが、Opus 5 を使って新しい取引所向けの市場データフィードを単一のセッションで構築しました。以前のモデルではこのタスク自体完了できず、エンジニアから詳細な計画を与えられても無理でした。検証用の生データフィードが存在しない状況下でも、Opus 5 は自らのコードが取引所のデータを正しくパースしているか確認するためのテストハッチを自ら構築しました。
以下は、早期アクセス顧客が Opus 5 との連携を通じて得た体験に関する追加報告です。
「FrontierCode 1.1」における Claude Opus 5 は、Fable レベルの性能を半分のコストで達成しました。また Devin 内では、困難なデバッグや根本原因分析タスクにおいて特に高い強みを示しています。
Claude Opus 5 は、Opus の速度とコストを実現しながら、Fable 5 に匹敵する知能レベルを提供します。CursorBench では Fable 5 にわずかに及ばないものの、多くの点で同様の振る舞いを示しました。開発者が Cursor でこのモデルをどのように活用するか、楽しみにしています。
Claude Opus 5 は、トークン消費量を以前の Claude モデルよりも増やさずに、Zapier の AutomationBench リーダーボードで首位を獲得しました。生のアカウント健康度ワークシートを受け取り、リスクのあるアカウントの特定から適切な担当者のアラート通知、リテンションオペレーション向けの要約まで、チャーン防止シーケンスをエンドツーエンドで実行しました。以前のモデルはこれを突破できませんでしたが、Opus 5 は見事に 100% の成功率を達成しています。
ゲノム解析の作業においては、Claude Opus 5 はこれまで試してきたどのモデルよりも慎重な科学者のように振る舞います。交絡因子を排除するために適切な統計検定を選び取り、独立した手法で自身の結果をクロスチェックし、長期的な多段階分析においても軌道から外れることなく遂行します。
Claude Opus 5 は、同シリーズのすべてのモデルを社内評価で上回りました。単に最も困難なエージェント型コーディングタスクにおいて Opus 4.7 よりも 22% 向上しただけでなく、実行ごとのばらつきが大幅に減少し、安定性も高まっています。Lovable で活動する数百万人のビルダーにとって、この一貫性が何よりも重要です。信頼できる結果を、ビルドごとに提供します。
Claude Opus 5 は、Opus シリーズにおいて 4.5 以来最大の飛躍です。フルスタックアプリの構築において、その進化はまずフロントエンドで顕著に現れます。Opus モデルがこれまで生み出した中で最も優れたアニメーション、ゲーム、3D 作品を体験できるのです。
Claude Opus 5 は、エージェントが扱うようなオープンEndedな分析作業において、Opus 4.8 よりも明確なアップグレードです。その効果は特に重要な領域で最大限に発揮され、難易度が高く曖昧なタスクほど顕著な改善が見られます。回答はより明快かつ簡潔になり、高負荷の処理においても効率性が向上しています。
Claude Opus 5 は、アナリストが毎日実行する金融調査ワークフローにおいて、Opus 4.8 よりも劇的な進化を遂げています。数値推論、表データ処理、そして精度が求められる場面での鋭い批判的思考において、その優位性が際立っています。
Claude Opus 5 は、専門的な企業コンテンツの分析に不可欠な業界レベルの知見と精度を提供します。Box の調査では、Opus 5 が Opus 4.8 を 8% 上回っており、特にテクノロジー、ヘルスケア、公共部門が日常的に依存するデータ分析(11% の改善)やデューデリジェンス(17% の改善)のワークフローにおいて、顕著なパフォーマンス向上を実現しています。
Claude Opus 5 は、Opus 4.8 と比較しても明確に世代を超えた進化を遂げています。ある週末、私はこのモデルに開発環境の統括官(チーフ・オブ・スタッフ)としての役割を与えてみました。自らモニターを構築し、各ボックスを駆使して作業を進め、判断が必要な局面のみで私を呼び出すという運用を行いました。

Claude Opus 5 は、Fundamental Research Assistant のコードベース全体に大規模な変更を加え、エージェント型ワークフローの中でフィードバックに対応しながら、これまで使ったどのモデルよりも明確に推論の根拠を説明しました。通常であれば細かく分割して処理する必要があるような作業も、このモデルなら一貫して扱えます。

金融モデリングの難易度の高いタスクにおいて、Claude Opus 5 は精度と効率性の両面で Opus 4.8 を明確に上回っています。特に深層の財務ドメインロジックを要するケースでは、性能の下限が大幅に引き上げられました。
あらゆる難易度レベルで平均して精度が 9 ポイント向上し、必要なターン数とツール呼び出しは約 3 分の 1 に減り、所要時間も 60% 短縮されています。
Claude Opus 5 は、実際のフロントエンド開発者が行うように、自身の作業を検証します。ベンチマークでは、デスクトップとスマホの両方の画面幅でページを開き、モバイル表示で隠れて見落としがちな商品や、画面外に位置するチェックアウトボタンを特定。それらを修正してから結果を返すことができました。
法務エージェント業務における性能も、以前の Opus モデルと比較して Claude Opus 5 は明確な進歩を示しています。特にコーポレートガバナンスや仲裁といった分野で大きな改善が見られました。また、推論レベルを下げていても品質を維持できる点にも感銘を受けました。最大推論設定の Opus 4.8 と同等のパフォーマンスを実現しながら、平均して必要なトークン数は 26% 削減しています。
Claude Opus 5 がもたらした最大の進歩は、より長い時間軸を要する作業におけるものです。例えば、資料の作成からその修正までを一貫して行うようなケースです。どのモデルを採用するかを決めるのはアートの質ですが、これはこれまでで最も明確なステップアップと言えます。視覚的な理解が向上し、フォーマットも洗練され、スライドに関する問題も減りました。
Claude Opus 5 の判断力が際立っています。プルリクエスト(PR)を任せる際、安易に公開しようとはしません。ブランチの確認やテンプレートのチェックを行い、テストへの影響まで考え抜くため、引き継ぎがスムーズになります。以前のモデルは先走りすぎて、私たちのチェックでつまずくことが多かったのです。
アーキテクチャの再設計セッションでは、Claude Opus 5 は私が提案した設計に対して異議を唱え、私の主張に屈しませんでした。その代わり、私のアイデアが持つ価値を具体的に説明し、異議の対象を単一の設計上の問いに絞り込みつつ、良い部分は残しつつ欠陥を修正する妥協案を提示しました。こうした判断力があれば、私たちはより少ない監視下で AI を信頼して任せることができます。
初回リダイレクトラインのテストにおいて、Claude Opus 5 は当社が検証したモデルの中で最高スコアを記録し、Opus 4.8 のほぼ倍に達しました。コメント機能も向上しており、機密保持契約(NDA)関連の処理では、より少ないパスで素早くリダイレクトラインに到達しながら、精度は維持またはそれ以上となっています。
Claude Opus 5 は、デッドコードを含まないクリーンでコンパクトな差分を生成します。また、コードベース固有の微妙な問題に対するリスク検出能力も優れています。このモデルはすでに本番環境での利用を開始しています。
当社の統合エージェントプラットフォーム「Cosmos」では、複数のユースケースを移行する予定です。コードレビューにおける Claude Opus 5 の活用をさらに拡大していくことを楽しみにしており、Opus 4.8 を使うよりも Opus 5 を利用すべきだと確信しています。
Claude Opus 5 が際立っているのは判断力です。1 行のコードを書く前に深く思考し、事後ではなく計画段階で自身の論理的な欠陥を捕捉します。また、「動作するかどうか」だけでなく「なぜその答えが正しいのか」という理由まで推論します。これは Claude モデル間で見られる問題解決能力における最も明確な飛躍であり、JetBrains IDE での採用を楽しみにしています。
Claude Opus 5 は、取引ベンチマークにおいて当社がテストした中で最も強力な Opus モデルです。Opus 4.8 と比較して、推論に要するトークン数は約 7 分の 1、レイテンシも半分以下で達成しています。計算リソースを大幅に削減しながら、より精度の高い回答を提供します。
Claude Opus 5 を採用することで、運用環境における監視エージェントが自身のメモリ管理の一部を自律的に行えるようになります。これにより、長時間の稼働においてもより自律的で信頼性の高い動作が可能になります。このモデルは文脈(コンテキスト)を生きたドキュメントとして扱います。例えば、あるサービスの異常を検知した際、その仮説が生産環境と矛盾しないか再検証し、問題がないことを確認すると、その訂正内容を自身のメモリに記録して、監視クエリを自動的に終了します。
Claude Opus 5 は、長時間実行され多段階の作業を要するタスクに特化した強力なエージェント型コーディングモデルです。コードベースを深く理解し、複雑なタスク全体を通じて文脈を一貫して維持します。機能開発やバグ修正における要件定義において、Opus 4.8 よりも効果的に要件を特定できます。開発者は Kiro で Opus 5 を利用可能となり、その高度な機能を駆使して大規模なプロジェクトに取り組むことができます。
01 /
24
アライメントと安全性
アライメント。 本番導入前のテストにおいて、自動行動監査を実施した結果、Opus 5 はこれまでのモデルの中で最もアライメント(目標整合性)が高いことが判明しました。以下に示すグラフのとおりです。
Opus 4.8、Sonnet 5、Fable 5 を上回り、Claude の憲章への準拠度が高く、欺瞞的な行動を示す割合が最も低く、悪用を誘発されるリスクも最小限に抑えられています。また、取り返しのつかない副作用をもたらす可能性のある無謀な行為を回避する点でも、これまでで最も安全なモデルです。

*自動行動監査の結果、Opus 5 の全体的なアライメントのズレスコアは 2.3 となり、直近のモデルの中で最低となりました。
安全性。 Opus 5 は、リスクの高い二重利用(デュアルユース)能力において新たなフロンティアを開拓したわけではありません。民間企業や政府機関とのパートナーと共に実施した厳格な評価では、生物学研究および攻撃的なサイバーセキュリティの分野において、Mythos 5 に劣っていることが確認されました。これらの評価に関する詳細は、システムカードをご覧ください。
前作の Opus 4.8 と同様、Opus 5 の学習にはサイバータスクは意図的に含まれていません。しかし、モデルがより汎用的な能力を獲得した結果、これらのタスクにおける性能は大幅に向上しました。特にセキュリティ脆弱性の「発見」においては Mythos 5 に迫るレベルです。ただし、脆弱性を具体的なサイバー脅威へと転換させる「悪用」の点では、依然として Mythos 5 を大きく下回っています。
これは、私たちが開発した評価指標 OSS-Fuzz の結果が示しています。この評価は、モデルが人間の広範な指導なしにどのようにして脆弱性を発見し、さらにそれを悪用できるかを測るものです。Opus 5 と Mythos 5 は脆弱性の発見において同様の成功率を収めていますが、悪用の開発に関するスコアでは Opus 5 が Mythos 5 に大きく遅れをとっています。
image*サイバーセキュリティ評価の OSS-Fuzz において、Opus 5 はソフトウェア脆弱性の特定(左)では Mythos 5 に近い性能を示しますが、それに対する悪用の開発(右)では considerably 成功率が低くなっています。
Opus 5 のセーフガード
Claude Opus 5 のセーフガードは、サイバーセキュリティおよび生物学の分野でモデルを有益に活用できるよう設計されています。これは Opus 4.8 に適用したものと概ね同様ですが、限定的な範囲のサイバータスクについてはより厳格な制限が設けられています。
サイバーセキュリティにおいて、Opus 5 の分類器は Fable 5 に比べて制限が緩やかになっています。これにより、ソースコード内の脆弱性を特定することは可能になりますが、「バイナリベース」の脆弱性スキャン(悪意のあるアクターに関連しやすい手法)やペネトレーションテスト、エクスプロイト生成はブロックされます。
私たちのテスト結果によると、Opus 5 の分類器が介入する頻度は Fable 5 に比べて約 85% 減少すると予想されます。Claude.ai、Claude Code、および Claude Cowork では、検知されたリクエストはデフォルトで Opus 4.8 にフォールバックされます。同様に、API でも Opus 4.8 へのフォールバックを有効にすることが可能です。
当社の サイバー検証プログラム(CVP)は、モデルのセキュリティ対策によって阻害されがちなサイバーセキュリティ作業を支援するものです。すでに CVP に参加している企業や研究者は、セキュリティ制限が少ない Opus 5 のバージョンに即座にアクセスできます。
生物学分野において、Opus 5 は Opus 4.8 と同様の安全対策を備えているため、現在では科学的研究に利用可能な最も高性能なモデルとなっています。ただし、長時間実行される自律的な研究タスクにおいては依然として重要な限界が見られ、AI モデルが生物関連リスクをもたらす可能性が高い領域でもあります。(この種の生物学作業については、Mythos 5 の方がより強力です。)今回のリリースに伴い、Fable 5 でブロックされていた生物学関連の問い合わせは、Opus 4.8 ではなく Opus 5 にルーティングされるようになります。
はじめに
Claude Opus 5 は本日、すべてのプラットフォームで利用可能になりました。料金は入力トークン 100 万あたり 5 ドル、出力トークン 100 万あたり 25 ドルで、Opus 4.8 と同じです。開発者は Claude API を通じて claude-opus-5 の利用を開始できます。
また、デフォルトの約 2.5 倍の速度で動作する「Fast モード」も提供されています。Opus 4.8 の場合と同様、Claude Platform では基本料金の 2 倍、Claude Code の利用クレジットでも同様の価格設定で Fast モードを利用可能です。
Opus 5 と併せて、ベータ版として以下の 2 つのアップデートをリリースします:
Claude プラットフォームでは、会話中にツールの設定を変更できるようになりました。これにより、開発者は会話の最中でも Claude が使用できるツールを切り替えることが可能になりながら、プロンプトキャッシュが無効化されることはありません。
API 側でも自動フォールバック機能が導入されました。Opus 5(または Fable 5)で安全性分類器によって検知されたリクエストについて、ユーザーは自動的に別のモデルへルーティングするオプションを選べるようになりました。この機能を有効にすると、API リクエストはブロックされるのではなく、デフォルトで利用可能な最良のモデルへ常に振り分けられます。
過去の Opus モデルと同様に、Opus 5 も一般アクセスにおいてはデータ保持要件を設けていません。
Opus 5 を最大限に活用するための詳細なガイドについては、プロンプティングガイドをご覧ください。
フットノート
Frontier-Bench v0.1 の結果(Effort プロット):**これらの数値は、mini-SWE-agent ハーネスと GKE バックエンド上で内部実行した Frontier-Bench v0.1 のデータです。各タスクで 5 回の試行を行った平均報酬を示しています。Opus 4.8 は、Opus 5 および Fable 5 における安全性分類器による拒否時のフォールバックモデルとして機能しました。
関連コンテンツ
経済未来研究基金のための研究アジェンダ
Anthropic Economic Futures Research Fund の研究アジェンダを公開します。
Claude で Anthropic 経済指数について質問する
Claude 用の「Anthropic 経済指数」コネクタをリリースしました。これにより、誰でも AI と労働に関する実データを探索できるようになります。
詳細は こちら
アンソロピック、Public First Action にさらに 2,000 万ドルを寄付
アンソロピックは Public First Action に対し、追加で 2,000 万ドルを拠出します。これにより、同団体への支援総額は 4,000 万ドルとなりました。
詳細は こちら
原文を表示
Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.
On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.
Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro.

Performance and cost-effectiveness
Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. The charts in this section show how performance changes according to the model’s effort setting, which customers can use to optimize for intelligence or conserve tokens for faster and cheaper results.
Opus 5 excels on valuable software engineering tasks. For example, on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2, at max effort, the model performs within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater performance at a given cost than all other models on high, xhigh, and max effort.
We see similar results on knowledge work and problem-solving tasks. For example:
- On ARC-AGI 3, an evaluation where the model has to solve novel problems, Opus 5’s score is three times as high as the next-best model.
- On Zapier AutomationBench, which measures whether models can complete business tasks from start to finish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its lowest effort setting, Opus 5 passes more tasks than any other model.
- On OSWorld 2.0, a computer use benchmark, Opus 5 outperforms every other model at any given cost, surpassing Fable 5’s best result at just over a third of the cost.
It’s also our best and most cost-efficient model on several related evaluations:
Opus 5 is a meaningful improvement over Opus 4.8 for scientific research. It shows better performance than Opus 4.8 on every one of our life sciences evaluations, which cover topics including structural biology, organic chemistry, and bioinformatics. Its improvements are most notable on organic chemistry tasks, like inferring molecular structures from spectroscopy data (it scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark), and on protein-related tasks like predicting how variations in a protein’s sequence affect how it functions (here, it scores 7.7 percentage points higher).
Finally, Opus 5 is capable of producing much stronger visual outputs:
Working with Claude Opus 5
Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. In evaluations and early-access testing, we and our users found many examples of Opus 5’s agency and thoroughness:
- On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded in doing so repeatedly; no competing model with the same setup could solve it after five attempts.
- Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case that the community’s patch had missed. A competing model fixed only the surface symptom (not the underlying cause), then reported the bug resolved.
- An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly.
Below are further reports from our early-access customers on their experience of working with Opus 5:
On FrontierCode 1.1, Claude Opus 5 approaches Fable-level performance at half the cost. Within Devin, it also shows particular strength on difficult debugging and root-cause analysis tasks.
Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it’s just under Fable 5 and has many of the same behaviors. We are excited to see how developers use it in Cursor.
Claude Opus 5 topped Zapier’s AutomationBench leaderboard without spending more tokens than prior Claude models. It took a raw account-health workbook and ran a full churn-prevention sequence end to end: flagging at-risk accounts, alerting the right owner, and summarizing for retention ops. Previous models didn’t pass; Opus 5 hit 100%.
On our genomics analysis work, Claude Opus 5 behaves more like a careful scientist than any model we’ve run. It reaches for the right statistical tests to rule out confounders, cross-checks its own results by independent methods, and stays on track through long multi-step analyses.
Claude Opus 5 came out ahead of every model in its family on our internal evals. It isn’t just better on our hardest agentic coding tasks, up 22% over Opus 4.7, it’s steadier, with far less variance run to run. For the millions of builders on Lovable, that consistency is the whole game. Reliable results, build after build.
Claude Opus 5 is the biggest leap in the Opus family since 4.5. On the same full-stack app builds, the front end shows it first: the best animations, games, and 3D work we have seen from an Opus model.
We’re loving Claude Opus 5. For the kind of open-ended analytical work our agent handles, it’s a strict upgrade over Opus 4.8, and the gains are biggest exactly where it matters: the harder, vaguer tasks. Responses are clearer and more concise, and we see improved efficiency at higher effort levels too.
Claude Opus 5 is a striking improvement over Opus 4.8 for the financial research workflows our analysts run every day. It stands out on numerical reasoning, table work, and sharper critical thinking where precision matters.
Claude Opus 5 delivers the industry intelligence and accuracy that is essential for the analysis of specialized enterprise content. Box found that Opus 5 outperforms Opus 4.8 by 8% and delivers notable performance gains in the data analysis (11% improvement) and due diligence (17% improvement) workflows that technology, healthcare, and public sector organizations rely on daily.
Claude Opus 5 is a clear generational step up from Opus 4.8. Over one weekend I gave it a chief-of-staff role over my dev environments: it built its own monitor, drove each box, and pulled me in only for the judgment calls.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant codebase, adapting to feedback throughout an agentic workflow and explaining its reasoning more clearly than any model we’ve used. It handled work we would normally have broken into much smaller pieces.

On some of our hardest financial-modeling tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both accuracy and efficiency. Its performance floor is materially higher, especially on deep finance domain logic. Across effort levels it averaged 9 percentage points higher accuracy with a third fewer turns and tool calls and 60% less time.
Claude Opus 5 checks its own work the way a real frontend developer would. On our benchmark it opened its pages in a browser at desktop and phone widths, caught a product hidden below the mobile fold and an off-screen checkout button, and fixed both before handing the work back.
Claude Opus 5 is a clear step up in performance on legal agent work compared to prior Opus models, and we saw the biggest gains in practice areas like corporate governance and arbitration. We were also impressed with Opus 5’s ability to maintain quality at lower reasoning levels, achieving similar performance while generating 26% fewer tokens on average compared to Opus 4.8 at max reasoning.
Claude Opus 5’s biggest gains for us are on longer-horizon work: building a full deck, then revising it. Artifact quality is what decides which model we ship, and this is the clearest step up we’ve seen — better visual understanding, cleaner formatting, fewer slide issues.
Claude Opus 5’s judgment is what stands out. Handing off a PR, it doesn’t rush to publish: it verifies the branches, checks the template, and thinks through test implications so the handoff is clean. The older models tended to jump ahead and get caught on our checks.
During a rearchitecting session, Claude Opus 5 pushed back on a design I proposed, and it didn’t fold when I insisted. Instead, it explained exactly what was valuable in my idea, narrowed its objection to a single design question, and proposed a compromise that kept the good part while fixing the flaw. That’s the kind of judgment that lets us trust it with less oversight.
On first-turn redlines, Claude Opus 5 scored the highest of any model we tested, nearly double Opus 4.8. Commenting is better too: on NDAs it gets to the redline in less time and with fewer passes, with accuracy maintained or better.
Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger hazard spotter on subtle, codebase-specific issues. We’re adopting it for production workloads.
We will definitely migrate a number of use cases in Cosmos, our unified agent platform. We’re looking forward to increasingly using Claude Opus 5 for code review, and I am confident in saying we would rather people be using Opus 5 than Opus 4.8.

What stands out about Claude Opus 5 is judgment. It thinks harder before it writes a single line, catches its own logical faults during planning rather than after the fact, and reasons about why an answer is right, not just whether it works. It’s the clearest jump in problem-solving we’ve seen from one Claude model to the next, and we’re looking forward to seeing it adopted in JetBrains IDEs.
Claude Opus 5 is the strongest Opus model we’ve tested on our trading benchmark, and it gets there using roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8. Better answers at a fraction of the compute.
Claude Opus 5 lets monitoring agents manage parts of their own memory in production, making them more autonomous and reliable over longer horizons. The agent treats its context as a living document: after flagging a potential anomaly in one of our services, it re-checked its own assumption against production, found the signal was benign, wrote the correction into its memory, and retired its monitoring queries on its own.
Claude Opus 5 is a strong agentic coding model built for long-running, multi-step work. It deeply understands your codebase, holds the thread across complex tasks, and pins down requirements for feature development and bug-fixing more effectively than Opus 4.8. Developers can now build with Opus 5 in Kiro, accessing its advanced capabilities to tackle ambitious projects.
01 /
24
Alignment and safety
*Alignment.* During pre-deployment testing, our automated behavioral audit found Opus 5 to be our most aligned model to date (as shown in the graph below). It adheres to Claude’s Constitution better than Opus 4.8, Sonnet 5, or Fable 5; exhibits the lowest rates of deceptive behavior; and is the least susceptible to being tricked into misuse. It’s also our safest model yet in terms of avoiding reckless actions that could have hard-to-reverse side effects.

*Safety.* Opus 5 does not advance the frontier in risky, dual-use capabilities. In rigorous evaluations conducted alongside private-sector and government partners, we found it remains behind Mythos 5 in both biology research and offensive cybersecurity. More information about these evaluations can be found in our System Card.
As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at *finding* cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the *exploitation *of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.
This is illustrated by Opus 5’s performance on OSS-Fuzz, an evaluation we’ve developed to assess how well models can find and then exploit vulnerabilities without extensive human guidance. Although Mythos 5 and Opus 5 identify vulnerabilities with similar success, Opus 5’s score on the development of exploits is far behind that of Mythos 5.

Safeguards for Opus 5
Claude Opus 5’s safeguards are designed to allow beneficial uses of the model in both cybersecurity and biology. They are similar to those we applied to Opus 4.8, with the exception of some stronger guardrails on a narrow range of cyber tasks.
*Cybersecurity. *Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation.
Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged requests will fall back to Opus 4.8 by default. Fallbacks to Opus 4.8 can also be enabled on the API.
Our Cyber Verification Program (CVP) facilitates cybersecurity work that would otherwise be impeded by the model’s safeguards. Enterprises and researchers who are already part of the CVP have immediate access to a version of Opus 5 with fewer security restrictions.
*Biology. *Since Opus 5 has a similar suite of safeguards to Opus 4.8, it is now our most capable generally available model for scientific research. Nevertheless, the model still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks. (Mythos 5 remains the stronger model for this type of biological work.) As part of this launch, biology-related requests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.
Getting started
Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.
It’s also offered in Fast mode, where it runs around 2.5 times the default speed. As with Opus 4.8, Fast mode is available at twice Opus 5’s base price on the Claude Platform and through usage credits in Claude Code.
Alongside Opus 5, we’re releasing two updates in beta:
- Mid-conversation tool changes on the Claude Platform. Within a conversation, developers can now change which tools Claude can use without invalidating the prompt cache.
- Automatic fallbacks on the API. Users can now choose to have requests that are flagged by our safety classifiers on Opus 5 (or Fable 5) automatically route to another model. With automatic fallbacks on, API requests always route to the best available model by default rather than being blocked.
Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.
For more guidance on how to get the best out of Opus 5, see our prompting guide.
Footnotes
Frontier-Bench v0.1, Effort plot: These results are from an internal run of Frontier-Bench v0.1, on the mini-SWE-agent harness and a GKE backend, mean reward over 5 attempts per task. Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
Related content
A research agenda for the Economic Futures Research Fund
We’re sharing the research agenda for the Anthropic Economic Futures Research Fund.
Ask Claude about the Anthropic Economic Index
We're launching the Anthropic Economic Index connector for Claude, which lets anyone explore real data about AI and work.
Anthropic is donating another $20 million to Public First Action
Anthropic is contributing an additional $20 million to Public First Action, bringing our total support to $40 million.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み