Anthropic、Claude Opus 5 を発表
アントロピックは最新大規模言語モデル「Claude Opus 5」を発表し、性能向上を謳っている。
AIニュース価値スコアβ
主要ニュースAI関連度、新規性、日本での有用性など6軸を公開検証中です。現在、掲載順には使用していません。
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 検索具体性
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
記事は AI モデルの最新バージョンである Claude Opus 5 の性能向上、コスト効率、および具体的なベンチマーク結果(Frontier-Bench, ARC-AGI 3 など)を詳細に報告しており、AI テクノロジーニュースとして極めて関連性が高い。新規性は公式発表であり、比較記事が存在しないため高い評価を得るが、日本固有の価格や規制情報がないため日本の関連性は低めとなる。
キーポイント
新モデルの発表
AI企業アントロピックが、次世代大規模言語モデルである「Claude Opus 5」を発表した。
性能向上の謳い文句
同社は今回のリリースにおいて、前作からの明確な性能向上を強調している。
重要な引用
AI企業アントロピックが、最新大規模言語モデル「Claude Opus 5」を発表した。
同社は性能向上を謳っている。
影響分析・編集コメントを表示
影響分析
このニュースはアントロピックの製品ラインナップが更新されたことを示すが、具体的な数値目標や実証データが欠けているため、業界全体への即時的なインパクトを評価するには情報が不足している。今後の詳細なベンチマーク発表や実際の導入事例に注目する必要がある。
編集コメント
発表された「Claude Opus 5」は、アントロピックの最新モデルとして注目されるが、現時点では詳細な性能比較や技術的革新性に関する具体的な情報が不足している。今後のベンチマーク結果や実利用におけるパフォーマンス次第で評価がさらに深まると考えられる。
Claude Opus 5 が本日、利用可能になりました。このモデルは思考深く、先回りして行動する能力を持ちながら、その知能の最前線である Claude Fable 5 に匹敵する性能を、半分の価格で提供します。
Frontier-Bench や GDPval-AA といったコーディングや知識作業の評価基準では、Opus 5 が新たな最高性能を記録しました。ただし、サイバーセキュリティのタスクに限れば、まだ Mythos 5 に劣ります。
Opus 5 は日常使いを想定して設計されています。他のモデルよりも効率的に動作するため、Claude Max のデフォルトモデルとなり、Claude Pro では最も強力な選択肢となっています。

パフォーマンスとコストパフォーマンス
Claude Opus 5 は、前作の Opus 4.8 と同じコストで、大幅に向上したパフォーマンスを実現しています。このセクションのグラフは、モデルの「努力設定(effort setting)」に応じた性能の変化を示しており、ユーザーはこの設定を調整して、知能の最大化と、より高速・低コストな結果を得るためのトークン節約のバランスを取ることができます。
Opus 5 は、価値の高いソフトウェアエンジニアリングタスクにおいて卓越した性能を発揮します。例えば「Frontier-Bench v0.1」では、他のすべてのモデルを上回り、タスクあたりのコストを下げながら Opus 4.8 のパフォーマンスを倍以上に引き上げました。「CursorBench 3.2」においても、最大限の努力を払った場合、Fable 5 の最高スコアと 0.5% 以内の差しか出さない結果を残しつつ、タスクあたりのコストは半分です。また、高負荷(high)、超高負荷(xhigh)、最大負荷(max)のいずれの試行設定においても、同じコストでより高いパフォーマンスを達成し、他モデルを上回っています。
知識労働や問題解決に関するタスクでも同様の結果が得られています。具体的には以下の通りです。
- 「ARC-AGI 3」では、モデルが新しい問題を解決する必要がある評価において、Opus 5 のスコアは次点のモデルの 3 倍に達しました。
- ビジネスタスクを最初から最後まで完了できるかを測定する「Zapier AutomationBench」では、同じコストで Opus 5 のパス率は次点のモデルの約 1.5 倍です。最低限の努力設定であっても、Opus 5 は他のどのモデルよりも多くのタスクをクリアしています。
- コンピュータ操作を評価するベンチマーク「OSWorld 2.0」では、あらゆるコスト水準において他モデルを上回り、Fable 5 の最良結果をコストの約 3 分の 1 で達成しました。
これらに関連する複数の評価項目においても、Opus 5 は最も高性能かつコスト効率に優れたモデルです。
Opus 5 は、科学研究の分野において Opus 4.8 よりも有意な進化を遂げています。構造生物学、有機化学、バイオインフォマティクスなどを含む生命科学に関するすべての評価項目で、Opus 4.8 を上回る性能を発揮しました。
特に目覚ましい改善が見られるのは、分光データから分子構造を推論するといった有機化学タスクです(内部ベンチマークでは Opus 4.8 よりも 10.2 ポイント高いスコアを記録)。また、タンパク質の配列変化がその機能にどう影響するかを予測するタンパク質関連タスクでも、7.7 ポイントの向上を示しています。
さらに、Opus 5 ははるかに強力な視覚出力も生成できるようになりました:
Claude Opus 5 の活用方法
Claude Opus 5 は、自身の作業を検証し、成功するまで慎重に反復を繰り返す能力が大幅に強化されています。評価テストや早期アクセス期間中のユーザー調査を通じて、Opus 5 が持つ自律性(エージェンシー)と徹底した取り組み姿勢の数多くの事例を確認しました:
Frontier-Bench のあるタスクでは、Opus 5 に機械部品の図面が提示され、それを 3D FreeCAD モデルとして再構築するコード作成を求められました。ただしこのタスクでは、モデルが図面を直接閲覧できないように意図的に制限されていました。それに対し Opus 5 は自らコンピュータビジョンパイプラインを記述して生ピクセルから幾何情報を抽出し、機械部品全体を再構築しました。これを何度試しても成功し、同じ環境下で競合モデルは 5 回挑戦しても解決できませんでした。
人気のオープンソースパッケージマネージャーに実際に存在するバグに対し、Opus 5 は根本原因を特定し、コミュニティの修正パッチでは見落とされていたエッジケースも正しく修正しました。一方、競合モデルは表面症状のみに対処して根本原因には手を付けず、「バグは解決した」と報告するにとどまりました。
ある取引所のエンジニアが、Opus 5 を使って単一のセッションで新取引所向けの市場データフィードを構築しました。以前のモデルではこのタスク自体完了できず、エンジニアから詳細な計画を提供しても無理でした。検証用のライブフィードが存在しない状況下でも、Opus 5 は自らのコードが取引所のデータを正しくパースしているか確認するためのテストハッチスを自ら構築しました。
以下は、早期アクセス顧客が Opus 5 との協業体験について寄せた追加報告です:
FrontierCode 1.1 では、Claude Opus 5 は Fable レベルの性能を半分のコストで達成します。 また Devin 内では、難易度の高いデバッグや根本原因分析タスクにおいて特に優れた能力を発揮しました。
Claude Opus 5 は、Opus の速度とコストで Fable 5 に匹敵する知能を実現しています。CursorBench では Fable 5 にわずかに及ばないものの、多くの点で同様の振る舞いを示します。開発者が Cursor でどのように活用するかを楽しみにしています。
Claude Opus 5 は、以前の Claude モデルよりも多くのトークンを消費することなく、Zapier の AutomationBench リーダーボードで首位に立ちました。生データのアカウント健康度ワークシートを受け取り、リスクのあるアカウントの特定、担当者の適切な通知、リテンションオペレーション向けの要約まで、チャーン防止シーケンスをエンドツーエンドで実行しました。以前のモデルはこれを突破できませんでしたが、Opus 5 は見事 100% の達成率を記録しました。
ゲノム解析の作業においては、Claude Opus 5 はこれまで試してきたどのモデルよりも慎重な科学者のように振る舞います。交絡因子を排除するために適切な統計検定を選択し、独立した手法で結果をクロスチェックし、長い多段階の分析プロセスでも軌道から外れることなく遂行します。
Claude Opus 5 は、同ファミリー内の全モデルを上回る成績を内部評価で記録しました。単に最も困難なエージェントコーディングタスクにおいて Opus 4.7 よりも 22% 向上しただけでなく、実行ごとのばらつきが大幅に減少し、はるかに安定しています。Lovable で活動する数百万人のビルダーにとって、この一貫性が何よりも重要です。結果の信頼性こそが、ビルドを積み重ねていく上で最も重要な要素だからです。
Claude Opus 5 は、Opus ファミリーにおいて 4.5 以来最大の飛躍です。フルスタックアプリの構築を比較した場合、その差はまずフロントエンドで顕著に現れます。Opus モデルがこれまで生み出した中で最も優れたアニメーション、ゲーム、3D デザインを実現しています。
Claude Opus 5 は、エージェントが扱うようなオープンエンドな分析作業において、Opus 4.8 を明確に上回る進化を遂げています。その効果は特に重要な領域で顕著で、難易度が高く曖昧なタスクほど大きな改善が見られます。回答はより明確かつ簡潔になり、高負荷の処理においても効率性が向上しています。
Claude Opus 5 は、アナリストが日常的に実行する金融調査ワークフローにおいて、Opus 4.8 よりも劇的な改善をもたらしました。数値推論、表処理、そして精度が求められる場面での鋭い批判的思考において、特にその能力を発揮しています。
Claude Opus 5 は、専門的な企業コンテンツの分析に不可欠な、業界最高水準の知性と精度を提供します。Box の調査によると、Opus 5 は Opus 4.8 を 8% 上回っており、特にテクノロジー、ヘルスケア、公共部門が日常的に依存するデータ分析(11% の改善)とデューデリジェンス(17% の改善)のワークフローにおいて、顕著な性能向上を実現しています。
Claude Opus 5 は、Opus 4.8 と比較しても明確に世代を超えた進化を遂げています。ある週末、私はこのモデルに開発環境の統括責任者(チーフ・オブ・スタッフ)のような役割を与えてみました。自らモニターを構築し、各ボックスを駆使して作業を進め、最終的な判断が必要な場面のみで私を呼び出すという運用を行いました。

Claude Opus 5 は、Fundamental Research Assistant のコードベース全体に大規模な変更を加え、エージェントワークフローを通じてフィードバックに対応しながら、これまで使用したどのモデルよりも明確に推論の根拠を説明しました。通常であれば細かく分割して処理する必要があるような作業も、単独で完遂できる能力を発揮しています。

Claude Opus 5 は、金融モデリングという最も難易度の高いタスクにおいて、Opus 4.8 と比べて精度と効率の両面で明確な進化を遂げています。特に深い金融ドメインの論理処理においては、その性能の下限が大幅に引き上げられました。
作業負荷レベルに関わらず、平均して精度は 9 ポイント向上し、必要なターン数やツール呼び出し回数は約 3 分の 1 に削減されました。さらに、所要時間は 60% 短縮されています。
Claude Opus 5 は、実際のフロントエンド開発者のように自ら作業を検証します。ベンチマークでは、デスクトップとスマホの両方の画面幅でページを開き、モバイル表示で隠れてしまう商品や、画面外にあるチェックアウトボタンを発見。それらを修正してから、成果物を提出しました。
法務エージェントのタスクにおいても、Claude Opus 5 は過去の Opus モデル群よりも明確に性能が向上しています。特に企業統治や仲裁といった分野で大きな進歩が見られました。
また、推論レベルを下げても品質を維持できる点にも感銘を受けました。Opus 4.8 の最大推論設定と比較して、平均して生成するトークン数を 26% 削減しながら、同等の性能を達成しています。
Claude Opus 5 が最も劇的に向上したのは、長期的な作業における成果です。例えば、資料作成からその修正まで一貫して行うようなタスクにおいて、その真価が発揮されます。どのモデルをリリースするかを決めるのはアーティファクトの品質ですが、今回の改善はこれまでで最も明確な進歩と言えるでしょう。視覚的な理解が深まり、フォーマットも洗練され、スライド作成時のトラブルも大幅に減っています。
Claude Opus 5 の際立った点は、その判断力にあります。プルリクエスト(PR)の引き渡しを行う際、安易に公開を急ぐのではなく、ブランチを確認し、テンプレートを見直し、テストへの影響まで十分に検討してから手渡します。これにより、引き渡しの品質が格段に向上しています。以前のモデルは、確認プロセスを飛び越えて先に進もうとする傾向があり、チェックでつまずくことが多かったです。
システム再設計のセッションでは、Claude Opus 5 は私が提案したデザインに対して異議を唱えることもありました。私が強く主張しても、すぐに折れることはありませんでした。むしろ、私のアイデアが持つ価値を明確に指摘し、異議の対象を単一の設計上の問いに絞り込みながら、良い部分は残しつつ欠点を修正する妥協案を提示しました。この種の判断力があれば、より少ない監視下で信頼して任せることができるようになります。
最初のターンにおける修正提案の精度では、Claude Opus 5 がテストしたモデルの中で最高を記録し、Opus 4.8 の約 2 倍に達しました。コメント機能も改善されており、機密情報を含む文書(NDA)への対応において、より短い時間で、かつ少ない試行回数で修正点を特定しながら、精度は維持または向上しています。
Claude Opus 5 は、デッドコードを含まないクリーンで簡潔な差分(diff)を生成します。また、コードベース固有の微妙な問題に対するリスク検知能力も優れており、本番環境での利用に採用する予定です。
統合エージェントプラットフォーム「Cosmos」では、多くのユースケースを移行していく方針です。コードレビューにおける Claude Opus 5 の活用をさらに拡大していく予定であり、Opus 4.8 よりも Opus 5 を使用していただくことを強く推奨します。
Claude Opus 5 が際立っているのは、その判断力です。1 行のコードを書く前に深く考え、計画段階で論理的な欠陥を自ら発見し、単に動作するかどうかだけでなく「なぜ正解なのか」を推論できます。これは Claude モデル間で見られる問題解決能力における最も明確な飛躍であり、JetBrains IDE での採用が期待されています。
Claude Opus 5 は、当社が実施した取引ベンチマークにおいて最も強力な Opus モデルです。Opus 4.8 と比較して、推論に使用するトークン数は約 7 分の 1 に抑えられ、レイテンシも半分未満に短縮されています。計算リソースを大幅に削減しながら、より高精度な回答を提供します。
Claude Opus 5 を採用することで、運用環境における監視エージェントが自身のメモリ管理を自律的に行えるようになります。これにより、長期にわたる運用でもより自立性が高く、信頼性の高い動作が可能になります。このモデルは文脈を「生きているドキュメント」として扱います。あるサービスの異常を検知した際、その仮説を実際の運用データで再検証し、問題がないことを確認すると、その修正内容をメモリに記録し、監視クエリを自ら停止します。
Claude Opus 5 は、長時間実行されるマルチステップの作業に特化した強力なエージェント型コーディングモデルです。コードベースを深く理解し、複雑なタスクを通じて一貫性を保ちながら、機能開発やバグ修正の要件を Opus 4.8 よりも効果的に把握します。開発者は Kiro で Opus 5 を利用可能となり、その高度な機能を駆使して大規模なプロジェクトに挑戦できます。
01 /
24
アライメントと安全性
*アライメント(整合性)*。本番導入前のテストにおいて、自動的な行動監査を実施した結果、Opus 5 はこれまでのモデルの中で最もアライメントが取れていることが確認されました(下のグラフ参照)。Claude の憲章 Claude's Constitution を遵守する点では Opus 4.8 や Sonnet 5、Fable 5 よりも優れており、欺瞞的な行動をとる割合が最も低く、悪用を誘発される可能性も最小限に抑えられています。また、取り返しのつかない副作用を引き起こす恐れのある無謀な行為を回避する点でも、これまでで最も安全なモデルです。
image*自動的な行動監査において、Opus 5 のアライメントのズレに関する総合スコアは 2.3 で、直近のモデルの中で最低値となりました。
*安全性*。Opus 5 は、リスクの高い二重利用(デュアルユース)能力の最前線に進出することはありません。民間企業や政府パートナーと共同で実施した厳格な評価では、生物学研究および攻撃的なサイバーセキュリティの分野において、Mythos 5 に劣っていることが判明しました。これらの評価に関する詳細は、システムカードでご確認ください。
前作の Opus 4.8 と同様、Opus 5 はサイバー攻撃タスクへの学習を意図的に避けています。しかし、モデル全体の能力が向上した結果、これらのタスクにおける性能は大幅に改善されました。特にセキュリティ脆弱性の「発見」においては Mythos 5 に迫る成果を上げていますが、「脆弱性を悪用して実際のサイバー脅威へと変換する」という点では、依然として Mythos 5 を大きく下回っています。
これは、私たちが開発した評価指標 OSS-Fuzz の結果がそれを示しています。この評価は、モデルがいかにして人間の広範な指導なしに脆弱性を発見し、さらに悪用コードを開発できるかを測るものです。Opus 5 と Mythos 5 は脆弱性の発見において同程度の成功を収めていますが、悪用コードの開発に関するスコアでは、Opus 5 は Mythos 5 に比べて大きく遅れをとっています。

*サイバーセキュリティ評価の 1 つである OSS-Fuzz において、Opus 5 はソフトウェア脆弱性の特定(左)では Mythos 5 に近い性能を示しますが、それに対する悪用コードの開発(右)では大幅に劣っています。
Opus 5 のセーフガード
Claude Opus 5 のセーフガードは、サイバーセキュリティおよび生物学の分野でモデルを有益に活用できるよう設計されています。これらは Opus 4.8 に適用されたものと同様ですが、限定的な範囲のサイバータスクについては、より強力なガードレールが追加されている点が異なります。
サイバーセキュリティ
Opus 5 のサイバー分類機能は、Fable 5 に比べて制限が緩やかになっています。これにより、ソースコード内の脆弱性を特定することは可能ですが、「バイナリベース」の脆弱性スキャン(悪意のあるアクターに関連しやすい手法)、ペネトレーションテスト、エクスプロイト生成といった行為はブロックされます。
私たちのテスト結果によると、Opus 5 の分類機能が介入するのは Fable 5 に比べて約 85% 少なくなると予想しています。Claude.ai、Claude Code、そして Claude Cowork では、検知されたリクエストはデフォルトで Opus 4.8 にフォールバックされます。同様に、API でも Opus 4.8 へのフォールバックを有効にすることが可能です。
サイバー検証プログラム(CVP)は、モデルのセキュリティ対策によって阻害されがちなサイバーセキュリティ関連の作業を支援する仕組みです。すでに CVP に参加している企業や研究者には、セキュリティ制限が少ない Opus 5 のバージョンへのアクセス権が即座に付与されます。
生物学研究について
Opus 5 は Opus 4.8 と同様の安全対策を備えているため、現在では科学的研究に利用できる最も高性能なモデルとなっています。ただし、長期間にわたる自律的な研究タスクにおいては依然として重要な限界が見られ、AI モデルが生物関連のリスクをもたらす可能性が高い領域でもあります(この種の生物学作業については、Mythos 5 の方がより強力です)。今回のリリースに伴い、Fable 5 ではブロックされていた生物学関連のリクエストは、Opus 4.8 ではなく Opus 5 にルーティングされるようになります。
はじめに
Claude Opus 5 は本日、すべてのプラットフォームで利用可能になりました。料金は入力トークン 100 万あたり 5 ドル、出力トークン 100 万あたり 25 ドルで、Opus 4.8 と同じです。開発者は Claude API を通じて claude-opus-5 の利用を開始できます。
また、デフォルトの約 2.5 倍の速度で動作する「Fast モード」も提供されています。Opus 4.8 の場合と同様、Claude Platform では基本料金の 2 倍、Claude Code の使用クレジットでも同様の価格設定で利用可能です。
Opus 5 と併せて、ベータ版として以下の 2 つのアップデートをリリースします:
Claude Platform では、会話中にツールの変更が可能になりました。これにより、開発者は会話を中断することなく、Claude が使用できるツールを柔軟に変更できます。
API には自動フォールバック機能も追加されました。Opus 5(または Fable 5)の安全分類器によって検知されたリクエストに対し、自動的に別のモデルへルーティングするオプションが選べるようになります。この機能を有効にすれば、リクエストはブロックされるのではなく、常に利用可能な最適なモデルへ優先的に振り分けられます。
過去の Opus モデルと同様、Opus 5 も一般アクセスにおいてデータ保持要件を設けていません。
Opus 5 の性能を最大限引き出すための詳細なガイドについては、プロンプトエンジニアリングのガイドをご覧ください。
フットノート
Frontier-Bench v0.1 の結果について:本データは、mini-SWE-agent ハーネスと GKE バックエンド上で内部実行した Frontier-Bench v0.1 の結果です。各タスクごとに 5 回の試行を行い、平均報酬を算出しています。なお、Opus 5 および Fable 5 で安全分類器による拒否が発生した場合のフォールバック先として、Opus 4.8 を使用しました。
関連コンテンツ
経済未来研究基金の研究アジェンダについて
Anthropic Economic Futures Research Fund の研究アジェンダを公開します。
Claude で Anthropic Economic Index を質問する
Claude 向けに「Anthropic Economic Index」コネクタをリリースしました。これにより、誰でも AI と労働に関する実データを検索・分析できるようになります。
アンソロピックが Public First Action にさらに 2,000 万ドルを寄付
アンソロピックは、Public First Action に対して追加で 2,000 万ドルを提供します。これにより、同団体への支援総額は 4,000 万ドルとなりました。
原文を表示
Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.
On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.
Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro.

Performance and cost-effectiveness
Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. The charts in this section show how performance changes according to the model’s effort setting, which customers can use to optimize for intelligence or conserve tokens for faster and cheaper results.
Opus 5 excels on valuable software engineering tasks. For example, on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2, at max effort, the model performs within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater performance at a given cost than all other models on high, xhigh, and max effort.
We see similar results on knowledge work and problem-solving tasks. For example:
- On ARC-AGI 3, an evaluation where the model has to solve novel problems, Opus 5’s score is three times as high as the next-best model.
- On Zapier AutomationBench, which measures whether models can complete business tasks from start to finish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its lowest effort setting, Opus 5 passes more tasks than any other model.
- On OSWorld 2.0, a computer use benchmark, Opus 5 outperforms every other model at any given cost, surpassing Fable 5’s best result at just over a third of the cost.
It’s also our best and most cost-efficient model on several related evaluations:
Opus 5 is a meaningful improvement over Opus 4.8 for scientific research. It shows better performance than Opus 4.8 on every one of our life sciences evaluations, which cover topics including structural biology, organic chemistry, and bioinformatics. Its improvements are most notable on organic chemistry tasks, like inferring molecular structures from spectroscopy data (it scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark), and on protein-related tasks like predicting how variations in a protein’s sequence affect how it functions (here, it scores 7.7 percentage points higher).
Finally, Opus 5 is capable of producing much stronger visual outputs:
Working with Claude Opus 5
Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. In evaluations and early-access testing, we and our users found many examples of Opus 5’s agency and thoroughness:
- On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded in doing so repeatedly; no competing model with the same setup could solve it after five attempts.
- Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case that the community’s patch had missed. A competing model fixed only the surface symptom (not the underlying cause), then reported the bug resolved.
- An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly.
Below are further reports from our early-access customers on their experience of working with Opus 5:
On FrontierCode 1.1, Claude Opus 5 approaches Fable-level performance at half the cost. Within Devin, it also shows particular strength on difficult debugging and root-cause analysis tasks.
Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it’s just under Fable 5 and has many of the same behaviors. We are excited to see how developers use it in Cursor.
Claude Opus 5 topped Zapier’s AutomationBench leaderboard without spending more tokens than prior Claude models. It took a raw account-health workbook and ran a full churn-prevention sequence end to end: flagging at-risk accounts, alerting the right owner, and summarizing for retention ops. Previous models didn’t pass; Opus 5 hit 100%.
On our genomics analysis work, Claude Opus 5 behaves more like a careful scientist than any model we’ve run. It reaches for the right statistical tests to rule out confounders, cross-checks its own results by independent methods, and stays on track through long multi-step analyses.
Claude Opus 5 came out ahead of every model in its family on our internal evals. It isn’t just better on our hardest agentic coding tasks, up 22% over Opus 4.7, it’s steadier, with far less variance run to run. For the millions of builders on Lovable, that consistency is the whole game. Reliable results, build after build.
Claude Opus 5 is the biggest leap in the Opus family since 4.5. On the same full-stack app builds, the front end shows it first: the best animations, games, and 3D work we have seen from an Opus model.
We’re loving Claude Opus 5. For the kind of open-ended analytical work our agent handles, it’s a strict upgrade over Opus 4.8, and the gains are biggest exactly where it matters: the harder, vaguer tasks. Responses are clearer and more concise, and we see improved efficiency at higher effort levels too.
Claude Opus 5 is a striking improvement over Opus 4.8 for the financial research workflows our analysts run every day. It stands out on numerical reasoning, table work, and sharper critical thinking where precision matters.
Claude Opus 5 delivers the industry intelligence and accuracy that is essential for the analysis of specialized enterprise content. Box found that Opus 5 outperforms Opus 4.8 by 8% and delivers notable performance gains in the data analysis (11% improvement) and due diligence (17% improvement) workflows that technology, healthcare, and public sector organizations rely on daily.
Claude Opus 5 is a clear generational step up from Opus 4.8. Over one weekend I gave it a chief-of-staff role over my dev environments: it built its own monitor, drove each box, and pulled me in only for the judgment calls.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant codebase, adapting to feedback throughout an agentic workflow and explaining its reasoning more clearly than any model we’ve used. It handled work we would normally have broken into much smaller pieces.

On some of our hardest financial-modeling tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both accuracy and efficiency. Its performance floor is materially higher, especially on deep finance domain logic. Across effort levels it averaged 9 percentage points higher accuracy with a third fewer turns and tool calls and 60% less time.
Claude Opus 5 checks its own work the way a real frontend developer would. On our benchmark it opened its pages in a browser at desktop and phone widths, caught a product hidden below the mobile fold and an off-screen checkout button, and fixed both before handing the work back.
Claude Opus 5 is a clear step up in performance on legal agent work compared to prior Opus models, and we saw the biggest gains in practice areas like corporate governance and arbitration. We were also impressed with Opus 5’s ability to maintain quality at lower reasoning levels, achieving similar performance while generating 26% fewer tokens on average compared to Opus 4.8 at max reasoning.
Claude Opus 5’s biggest gains for us are on longer-horizon work: building a full deck, then revising it. Artifact quality is what decides which model we ship, and this is the clearest step up we’ve seen — better visual understanding, cleaner formatting, fewer slide issues.
Claude Opus 5’s judgment is what stands out. Handing off a PR, it doesn’t rush to publish: it verifies the branches, checks the template, and thinks through test implications so the handoff is clean. The older models tended to jump ahead and get caught on our checks.
During a rearchitecting session, Claude Opus 5 pushed back on a design I proposed, and it didn’t fold when I insisted. Instead, it explained exactly what was valuable in my idea, narrowed its objection to a single design question, and proposed a compromise that kept the good part while fixing the flaw. That’s the kind of judgment that lets us trust it with less oversight.
On first-turn redlines, Claude Opus 5 scored the highest of any model we tested, nearly double Opus 4.8. Commenting is better too: on NDAs it gets to the redline in less time and with fewer passes, with accuracy maintained or better.
Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger hazard spotter on subtle, codebase-specific issues. We’re adopting it for production workloads.
We will definitely migrate a number of use cases in Cosmos, our unified agent platform. We’re looking forward to increasingly using Claude Opus 5 for code review, and I am confident in saying we would rather people be using Opus 5 than Opus 4.8.

What stands out about Claude Opus 5 is judgment. It thinks harder before it writes a single line, catches its own logical faults during planning rather than after the fact, and reasons about why an answer is right, not just whether it works. It’s the clearest jump in problem-solving we’ve seen from one Claude model to the next, and we’re looking forward to seeing it adopted in JetBrains IDEs.
Claude Opus 5 is the strongest Opus model we’ve tested on our trading benchmark, and it gets there using roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8. Better answers at a fraction of the compute.
Claude Opus 5 lets monitoring agents manage parts of their own memory in production, making them more autonomous and reliable over longer horizons. The agent treats its context as a living document: after flagging a potential anomaly in one of our services, it re-checked its own assumption against production, found the signal was benign, wrote the correction into its memory, and retired its monitoring queries on its own.
Claude Opus 5 is a strong agentic coding model built for long-running, multi-step work. It deeply understands your codebase, holds the thread across complex tasks, and pins down requirements for feature development and bug-fixing more effectively than Opus 4.8. Developers can now build with Opus 5 in Kiro, accessing its advanced capabilities to tackle ambitious projects.
01 /
24
Alignment and safety
*Alignment.* During pre-deployment testing, our automated behavioral audit found Opus 5 to be our most aligned model to date (as shown in the graph below). It adheres to Claude’s Constitution better than Opus 4.8, Sonnet 5, or Fable 5; exhibits the lowest rates of deceptive behavior; and is the least susceptible to being tricked into misuse. It’s also our safest model yet in terms of avoiding reckless actions that could have hard-to-reverse side effects.

*Safety.* Opus 5 does not advance the frontier in risky, dual-use capabilities. In rigorous evaluations conducted alongside private-sector and government partners, we found it remains behind Mythos 5 in both biology research and offensive cybersecurity. More information about these evaluations can be found in our System Card.
As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at *finding* cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the *exploitation *of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.
This is illustrated by Opus 5’s performance on OSS-Fuzz, an evaluation we’ve developed to assess how well models can find and then exploit vulnerabilities without extensive human guidance. Although Mythos 5 and Opus 5 identify vulnerabilities with similar success, Opus 5’s score on the development of exploits is far behind that of Mythos 5.

Safeguards for Opus 5
Claude Opus 5’s safeguards are designed to allow beneficial uses of the model in both cybersecurity and biology. They are similar to those we applied to Opus 4.8, with the exception of some stronger guardrails on a narrow range of cyber tasks.
*Cybersecurity. *Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation.
Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged requests will fall back to Opus 4.8 by default. Fallbacks to Opus 4.8 can also be enabled on the API.
Our Cyber Verification Program (CVP) facilitates cybersecurity work that would otherwise be impeded by the model’s safeguards. Enterprises and researchers who are already part of the CVP have immediate access to a version of Opus 5 with fewer security restrictions.
*Biology. *Since Opus 5 has a similar suite of safeguards to Opus 4.8, it is now our most capable generally available model for scientific research. Nevertheless, the model still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks. (Mythos 5 remains the stronger model for this type of biological work.) As part of this launch, biology-related requests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.
Getting started
Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.
It’s also offered in Fast mode, where it runs around 2.5 times the default speed. As with Opus 4.8, Fast mode is available at twice Opus 5’s base price on the Claude Platform and through usage credits in Claude Code.
Alongside Opus 5, we’re releasing two updates in beta:
- Mid-conversation tool changes on the Claude Platform. Within a conversation, developers can now change which tools Claude can use without invalidating the prompt cache.
- Automatic fallbacks on the API. Users can now choose to have requests that are flagged by our safety classifiers on Opus 5 (or Fable 5) automatically route to another model. With automatic fallbacks on, API requests always route to the best available model by default rather than being blocked.
Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.
For more guidance on how to get the best out of Opus 5, see our prompting guide.
Footnotes
Frontier-Bench v0.1, Effort plot: These results are from an internal run of Frontier-Bench v0.1, on the mini-SWE-agent harness and a GKE backend, mean reward over 5 attempts per task. Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
Related content
A research agenda for the Economic Futures Research Fund
We’re sharing the research agenda for the Anthropic Economic Futures Research Fund.
Ask Claude about the Anthropic Economic Index
We're launching the Anthropic Economic Index connector for Claude, which lets anyone explore real data about AI and work.
Anthropic is donating another $20 million to Public First Action
Anthropic is contributing an additional $20 million to Public First Action, bringing our total support to $40 million.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み