Anthropic、エージェント型知識作業で新リーダー「Claude Opus 5」を発表
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Artificial Analysis
Artificial Analysis が発表したベンチマーク結果によると、Anthropic の Claude Opus 5 は「AA-Briefcase」で新リーダーとなり、競合モデルを凌駕する性能とコスト効率を実現した。
AI深層分析を開く2026年8月1日 14:01
AI深層分析
キーポイント
Agentic Knowledge Work ベンチマークでの首位獲得
Anthropic が発表した Claude Opus 5 は、Artificial Analysis の独自ベンチマーク「AA-Briefcase」で最高スコア(Elo 1720)を記録し、Claude Fable 5 を約 146 ポイント引き離して新リーダーとなった。
コスト効率の劇的な向上
同社は Claude Opus 5 の Max エディションでタスクあたりのコストを 20% 削減し、Xhigh や High バリアントでは Fable 5 よりも大幅に安価ながら高い性能を発揮すると発表した。
多様なエディションによる柔軟な選択肢
Max、Xhigh、High の上位 3 つの設定がトップ 3 を独占し、Medium や Low エディションでも他社モデルと競合するスコアを記録することで、用途に応じた最適な設定の提供が可能となった。
複雑なタスク処理能力の評価
AA-Briefcase は数千ファイルに及ぶ非公開データを用いた研究報告書やスプレッドシートの作成など現実的なタスクを評価し、正確性、分析品質、プレゼンテーション品質の総合スコアとして Elo を算出している。
客観的基準と分析品質での優位性
Claude Opus 5 はルールの適合率と分析品質の向上により、最大努力時で分析品質 Elo が 2016 に達し、Claude Fable 5 より約 300 Elo 上回る。ただしプレゼンテーション品質では GPT-5.6 Sol にまだ劣っている。
重要な引用
Claude Opus 5 is the new leader on our agentic knowledge work benchmark, AA-Briefcase, outperforming Claude Fable 5 by nearly 150 Elo while reducing Cost per Task by 20%
Anthropic now holds a large majority of the top ten spots on the AA-Briefcase leaderboard.
AA-Briefcase is our proprietary benchmark for agentic knowledge work, evaluating models on realistic, private tasks spanning thousands of input files
Claude Opus 5's gains are driven primarily by rubric pass rate and analytical quality.
編集コメントを表示
編集コメント
今回の評価は、AI モデルが単なる情報検索や対話を超え、実際の業務を完遂する「エージェント」としての能力を問う段階に入ったことを示唆している。各エディションのパフォーマンスとコストのトレードオフを明確に提示した点は、実装担当者にとって極めて有用な判断材料となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
モデルページはこちら:https://artificialanalysis.ai/models/claude-opus-5
Claude Opus 5 は、エージェント型知識作業を評価する独自のベンチマーク「AA-Briefcase」において新たなリーダーとなりました。Claude Fable 5 を約 150 Elo 上回り、タスクあたりのコストも 20% 削減しています。
Anthropic が Claude Opus 5 をリリースしました。これは Artificial Analysis Intelligence Index の新トップモデルであり、私たちが独自に開発したエージェント型知識作業ベンチマーク「AA-Briefcase」でも第 1 位を獲得しました。最大努力設定(max effort)での Claude Opus 5 は AA-Briefcase Elo で 1720 を記録し、Claude Fable 5(1574)を 146 ポイント引き離しています。また、xhigh(1693)および high(1606)のバリアントも Fable 5 より性能が高く、かつコスト効率に優れています。
リリース前にすべての 5 つの努力設定でベンチマークを実施しました。Claude Opus 5(max)はタスクあたり 17.79 ドルで、Fable 5 の 22.30 ドルから 20% 削減されています。xhigh と high のバリアントはさらにコスト効率に優れており、どちらも Fable 5 より高性能を維持しながら、それぞれ 64%(14.26 ドル)と 47%(10.41 ドル)のコストで提供されます。
AA-Briefcase は、エージェント型知識作業を評価するための独自ベンチマークです。数千の入力ファイルにまたがる現実的で機密性の高いタスクを対象とし、調査レポートやプレゼンテーション、スプレッドシートなどの成果物作成能力を検証します。評価は「正しさ」「分析品質」「プレゼンテーション品質」の 3 つの観点から行われ、これらを統合して単一の指標である AA-Briefcase Elo として算出されます。

AA-Briefcase における Claude Opus 5 の全努力度設定での結果
エージェント型知識作業の明確なリーダー: Claude Opus 5 の上位 3 つの努力度設定(max、xhigh、high)が AA-Briefcase でトップ 3 を独占しました。Claude Fable 5、Sonnet 5(max)、Opus 4.8(max)と合わせると、Anthropic は現在 AA-Briefcase リーダーボードの上位 10 位中を大半占めています。Claude Opus 5 のスコアは、ミディアム努力度で 1470 Elo(GPT-5.6 Sol の max 設定である 1505 に次ぐ)、ロー努力度では 1223 Elo(GLM-5.2 の max 設定である 1254 に次ぐ)を記録しました。
客観的基準と分析の質では首位だが、プレゼンテーションの質ではやや劣る: AA-Briefcase は、客観的なルブリック基準、分析の質、プレゼンテーションの質という 3 つの観点で性能を測定しています。Claude Opus 5 のスコア向上は主にルブリックの合格率と分析の質によるものです。最大努力度設定では分析の質における Elo が 2016 に達し、Claude Fable 5 より約 300 Elo 上回っています。一方、Opus 4.8 と比較するとプレゼンテーションの質での改善幅は小さく、Presentation Elo は 1628 です。これは GPT-5.6 Sol(max 設定で 1666)にまだ約 40 Elo 及ばない状態です。

➤ ファブル級知能作業の半額コスト: Claude Opus 5 は、5 つの努力設定(effort settings)にわたって、性能とコストの幅広いトレードオフを提供します。最大努力設定では 1 タスクあたり 17.79 ドルで、Claude Fable 5 の 22.30 ドルより約 20% 安く、かつ AA-Briefcase Elo でより高いスコアを達成します。特筆すべきは、高努力設定の Claude Opus 5 が Fable 5 よりも 32 Elo 上回っており、1 タスクあたり 10.41 ドルという、Fable 5 の半額以下の価格で動作することです。トークン料金は Opus 4.8 と同様に、入力・出力とも 100 万トークンあたり 5 ドル/25 ドルを維持し、キャッシュ書き込みには 25% の割増、キャッシュヒットには 90% の割引が適用されます。

➤ 最前線性能のための 1 タスクあたり 25 分以上: Claude Opus 5 の上位 3 つの努力設定は、すべて AA-Briefcase タスクを平均して 25 分超かかります(最大・超高・高でそれぞれ 36.2 分、34.3 分、25.7 分)。最大努力時のこの時間は Opus 4.8 の 24.1 分より約 50% 長く、主にターン数の増加によるものです。Claude Opus 5 は上位 3 つの努力レベルで 1 タスクあたり平均 103 ターン、91 ターン、76 ターンを要しますが、Opus 4.8(最大)は 55 ターンです。

AA-Briefcase の完全な結果については、https://artificialanalysis.ai/evaluations/aa-briefcase をご覧ください。
原文を表示
Claude Opus 5 is the new leader on our agentic knowledge work benchmark, AA-Briefcase, outperforming Claude Fable 5 by nearly 150 Elo while reducing Cost per Task by 20%
Anthropic has released Claude Opus 5, the new leader on the Artificial Analysis Intelligence Index, and also the #1 model on our proprietary agentic knowledge work benchmark, AA-Briefcase. At max effort, Claude Opus 5 scores an AA-Briefcase Elo of 1720, 146 points ahead of Claude Fable 5 (1574). Its xhigh (1693) and high (1606) variants also outperform Fable 5 while being more cost efficient
We benchmarked all five effort settings ahead of release. Claude Opus 5 (max) costs $17.79 per task, a 20% reduction from Claude Fable 5 ($22.30). The xhigh and high variants offer stronger cost-efficiency tradeoffs, both scoring above Fable 5 while costing 64% ($14.26) and 47% ($10.41) as much respectively
AA-Briefcase is our proprietary benchmark for agentic knowledge work, evaluating models on realistic, private tasks spanning thousands of input files and requiring deliverables such as research reports, presentations, and spreadsheets. Performance is measured across correctness, analytical quality, and presentation quality, and combined into a single metric, the AA-Briefcase Elo

Key AA-Briefcase results across all 5 Claude Opus 5 effort settings:
➤ The clear leader in agentic knowledge work: Claude Opus 5’s top three effort settings (max, xhigh, and high) take the top three positions on AA-Briefcase. Combined with Claude Fable 5, Sonnet 5 (max), and Opus 4.8 (max), Anthropic now holds a large majority of the top ten spots on the AA-Briefcase leaderboard. Claude Opus 5 scores 1470 Elo at medium effort, just behind GPT-5.6 Sol (max, 1505), and 1223 at low effort, just below GLM-5.2 (max, 1254)
➤ Leads in objective criteria and analytical quality, but not presentation: AA-Briefcase measures performance across objective rubric criteria, analytical quality, and presentation quality. Claude Opus 5’s gains are driven primarily by rubric pass rate and analytical quality. At max effort it achieves an Analytical Quality Elo of 2016, nearly 300 Elo ahead of Claude Fable 5. Compared to Opus 4.8, presentation quality sees a smaller improvement, with a Presentation Elo of 1628, still ~40 Elo behind GPT-5.6 Sol (max, 1666)

➤ Half the cost for Fable-level knowledge work capabilities: Claude Opus 5 offers a wide range of intelligence-cost tradeoffs across its five effort settings. Max effort costs $17.79 per task, around 20% less than Claude Fable 5 ($22.30) while achieving a higher AA-Briefcase Elo. Notably, Claude Opus 5 with high effort outperforms Claude Fable 5 by 32 Elo while costing $10.41 per task, less than half the price. Claude Opus 5 retains Opus 4.8’s pricing of $5/$25 per million input/output tokens, with a 25% premium for cache writes and a 90% discount for cache hits

➤ Over 25 minutes per task for frontier performance: Claude Opus 5’s top three effort settings all average more than 25 minutes per AA-Briefcase task (36.2, 34.3, and 25.7 minutes for max, xhigh, and high). At max effort this is roughly 50% longer than Opus 4.8 (24.1 minutes), primarily due to an increase in the number of turns. Claude Opus 5 averages 103, 91, and 76 turns per task across its top three effort levels, compared to 55 for Opus 4.8 (max)

For full AA-Briefcase results, see https://artificialanalysis.ai/evaluations/aa-briefcase
AI算出
主要ニュースainew評価標準
記事は Claude Opus 5 の性能、コスト効率、および独自のベンチマーク結果という具体的な数値と事実を報じており、AI モデルの重大な更新として評価される。ただし、日本固有の情報や企業事例が含まれていないため、日本の関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み