Anthropic、Opus 5 を発表:Fable 5 並みの知能を低コストで提供
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Artificial Analysis
Claude Opus 5 は Artificial Analysis のインテリジェンス指数で Fable 5 と同程度のスコアを獲得しつつ、タスクあたりの平均コストを約 26% 削減することに成功した。
AI深層分析を開く2026年8月1日 14:01
AI深層分析
キーポイント
コスト効率と知能レベルの両立
Claude Opus 5 は Artificial Analysis のインテリジェンス指数で Fable 5 と同程度のスコアを獲得しつつ、タスクあたりの平均コストを約 26% 削減することに成功した。
エージェント作業における新リーダー
GDPval-AA v2 および AA-Briefcase ベンチマークにおいて Opus 5 は Fable 5 や GPT-5.6 Sol を上回り、専門的なアウトプット生成能力で新たなトップに立った。
コーディングと科学推論の性能
Opus 5 はコーディングエージェントインデックスで Fable 5 と並んで首位を獲得し、科学推論テストでも Fable 5 と同等の結果を示したが、事実知識の精度では Fable 5 に劣る。
コンテキストウィンドウと価格設定
100万トークンのコンテキストウィンドウを持ち、入力・出力は百万トークンあたり5ドル/25ドルでキャッシュ書き込みには25%のプレミアムが適用される。
ベンチマークでの新記録
GDPval-AA v2およびAA-Briefcaseにおいて、Claude Opus 5 (max) はそれぞれ1861 Eloと1720 Eloを記録し、Fable 5を大きく上回る性能を示した。
重要な引用
Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task
New leader in agentic knowledge work: Claude Opus 5 (max) scores 1861 Elo on GDPval-AA v2, >100 points ahead of Claude Fable 5 and GPT-5.6 Sol (max)
Joint first place on the Coding Agent Index: Claude Opus 5 (xhigh) with Claude Code leads the Artificial Analysis Coding Index
Claude Opus 5 (max) is the new leader on both GDPval-AA v2 (1861 Elo, +114 over Claude Fable 5) and AA-Briefcase (1720 Elo, +146 over Fable 5).
編集コメントを表示
編集コメント
今回の評価は、単なる性能の向上だけでなく、実運用におけるコスト対効果の劇的な改善を裏付けており、企業での大規模導入判断に重要な示唆を与える。ただし、事実知識やハルシネーション率の課題も明確になっているため、用途に応じたモデルの使い分けが求められる局面である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Claude Opus 5 は、Artificial Analysis の知能指数においてわずかにトップの座にあり、Fable 5 と同等の知能を持ちながら、タスクあたりのコストは 26% 低く抑えられています。
リリース前に Anthropic を支援して Claude Opus 5 を評価した結果、GDPval-AA v2 および AA-Briefcase のスコアでこれまでの最高記録を達成しました。Opus 5 (max) は Artificial Analysis Intelligence Index で 61 点を獲得し、Claude Fable 5 (max, 60 点) と実質的に同率首位となり、GPT-5.6 Sol (max, 59 点)、Kimi K3 (57 点)、そして Claude Opus 4.8 (max, 56 点) を上回っています。
主なポイント:
➤ エージェント型知識作業の新たなリーダー: Claude Opus 5 (max) は、GDPval-AA v2 で 1861 Elo のスコアを記録し、Claude Fable 5 や GPT-5.6 Sol (max) を 100 ポイント以上引き離しました。また、オープンソースの参照エージェントハッチ「Stirrup」を用いて、モデルが正確で高品質な専門的な成果物を生成する能力を測定するエージェント型知識作業ベンチマークである AA-Briefcase では 1720 Elo を獲得し、Fable 5 よりも 146 ポイント高いスコアを示しています。
➤ コーディングエージェント指数で首位タイ: Claude Code と組み合わせた Claude Opus 5 (xhigh) は、Artificial Analysis のコーディングインデックス全体をリードしており、SWE-Atlas-QnA においても最高スコアを記録しました。
コストを抑えた最先端知能:Claude Opus 5(max)の平均コストは、Intelligence Index タスクあたり 2.03 ドルで、Claude Fable 5(フォールバックあり)の 2.75 ドルを下回りますが、Claude Opus 4.8(max)の 1.80 ドルや Claude Sonnet 5(max)の 1.53 ドルよりは高いです。ただし、推論努力度を「high」または「xhigh」に設定すれば、Opus 5 は両モデルを上回る性能を発揮しながら、タスクあたりのコストは抑えられます。
最先端のエージェント端末利用:Terminal-Bench v2.1 で最大努力時 89% のスコアを記録し、リーダーである GPT-5.6 Sol(xhigh)とほぼ同等の水準です。
科学的推論での優位性:最先端のエージェント性能に加え、Claude Opus 5 は「Humanity's Last Exam」で Fable 5 と同率の 53% を達成しました。また、アルゴンヌ国立研究所とイリノイ大学シカゴ校(UIUC)の研究チームが開発した最先端物理評価基準「CritPt」でも Fable 5 に並びますが、GPT-5.6 Sol、GPT-5.5 Pro、GPT-5.6 Terra には及びません。
事実知識は依然として Fable 5 に劣る:モデルのサイズクラスから予想される通り、Opus 5 の AA-Omniscience における事実知識量は Fable 5 より低いです。AA-Omniscience の精度では Opus 4.8 より +7 ポイント向上しましたが、不確実な状況でも回答する傾向が強まり、ハルシネーション(幻覚)発生率は +14 ポイント上昇して 50% に達しています。
効率性の改善は限定的:Opus 5 は低コストで Fable 5 を上回りますが、推論努力度が低いレベルでは、「Intelligence vs. Cost per Task」のパレートフロンティアにおいて GPT-5.6 シリーズにわずかに劣ります。つまり、高い知能レベルにおけるコスト対効果の最適化においては改善が見られますが、それ以外の領域では依然として競合他社に譲る部分があります。
その他のモデル詳細:
コンテキストウィンドウは100万トークン(Opus 4.8と同等)です。
価格設定については、最近の Opus ラインアップと同様に、入力・出力ともに100万トークンあたり5ドル/25ドルです。キャッシュ価格は変更されず、キャッシュ書き込みには25%のプレミアム(100万トークンあたり6.25ドル、有効期限5分)が適用され、キャッシュヒット時には90%オフ(100万トークンあたり0.50ドル)となります。
「low」「medium」「high」「xhigh」「max」の5段階の努力設定に対応しており、Fable 5 と同様にサーバーサイドでのフォールバックもサポートされています。知能指数の評価は、Opus 4.8 のフォールバックを有効にした状態で実施されました。

Claude Opus 5(max)は、GDPval-AA v2(Elo 1861、Fable 5 より+114)および AA-Briefcase(Elo 1720、Fable 5 より+146)の両方で新たなリーダーとなりました。これらのベンチマークでは、オープンソースのリファレンスエージェントハーンである Stirrup を用いて、モデルが正確で質の高い専門的な成果物を生成できる能力を評価しています。

Claude Opus 5 の努力設定は、トークン使用量とパフォーマンスのトレードオフを広くカバーしています。GDPval-AA v2 では、努力レベルによって評価スコアが407ポイント変動し、出力に使用されるトークン数は低設定から最大設定まで約8倍の範囲になります。GPT-5.6 Sol と同様、Opus 5 は努力設定に応じて、他社のモデルよりもはるかに少ない、あるいははるかに多くのトークンを必要とします。
Claude Opus 5 の最大努力時の Artificial Analysis インテリジェンス指数における各評価の詳細分析


原文を表示
Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task
We supported Anthropic to evaluate Claude Opus 5 ahead of release: it sets the highest GDPval-AA v2 and AA-Briefcase scores so far. Opus 5 (max) scores 61 on the Artificial Analysis Intelligence Index, effectively tied with Claude Fable 5 (max, 60), and ahead of GPT-5.6 Sol (max, 59), Kimi K3 (57), and Claude Opus 4.8 (max, 56)
要点
➤ New leader in agentic knowledge work: Claude Opus 5 (max) scores 1861 Elo on GDPval-AA v2, >100 points ahead of Claude Fable 5 and GPT-5.6 Sol (max). On AA-Briefcase, our agentic knowledge work benchmark, it scores 1720 Elo, +146 ahead of Fable 5. These benchmarks test the ability of models to produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup
➤ Joint first place on the Coding Agent Index: Claude Opus 5 (xhigh) with Claude Code leads the Artificial Analysis Coding Index, including the highest score on SWE-Atlas-QnA
➤ Frontier intelligence with reduced cost: Claude Opus 5 (max) costs $2.03 on average per Intelligence Index task, below Claude Fable 5 (with fallback) at $2.75, but still above Claude Opus 4.8 (max) at $1.80 and Claude Sonnet 5 (max) at $1.53. However, at high and xhigh reasoning efforts Opus 5 can outperform both Opus 4.8 and Claude Sonnet 5 at a lower cost per task
➤ Frontier agentic terminal use: 89% on Terminal-Bench v2.1 at max effort, roughly in line with the leader, GPT-5.6 Sol (xhigh)
➤ Outperformance on scientific reasoning: Along with leading agentic performance, Claude Opus 5 scores 53% on Humanity’s Last Exam in line with Fable 5; on CritPt, a frontier physics evaluation developed by Argonne and UIUC researchers, it also matches Fable 5 but sits behind GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra
➤ Factual knowledge still lags Fable 5: As expected from the models’ size classes, Opus 5 still has lower factual knowledge on AA-Omniscience than Fable 5. It improves +7 points on AA-Omniscience Accuracy over Opus 4.8, but answers more often when uncertain - its hallucination rate rises +14 points to 50%
➤ Improving efficiency, but only on the Intelligence vs. Cost per Task Pareto frontier at high Intelligence levels: Opus 5 outperforms Fable 5 at lower cost, but at lower effort levels it sits just behind the GPT-5.6 family on the Intelligence vs. Cost per Task Frontier
Other model details:
➤ Context window: 1 million tokens (equivalent to Opus 4.8)
➤ Pricing: As with recent Opus launches tokens cost $5/$25 per million tokens of input/output; cache pricing remains at a 25% premium for cache writes ($6.25 per million tokens) with 5-minute time to live, and 90% discount for cache hits ($0.50 per million tokens)
➤ Five effort settings (low, medium, high, xhigh, max), and support for server-side fallback as with Fable 5. Intelligence Index evaluations were run with Opus 4.8 fallback enabled

Claude Opus 5 (max) is the new leader on both GDPval-AA v2 (1861 Elo, +114 over Claude Fable 5) and AA-Briefcase (1720 Elo, +146 over Fable 5). These benchmarks test the ability of models to produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup

Claude Opus 5's effort setting spans a wide range of token usage-performance tradeoffs. On GDPval-AA v2, effort levels span 407 Elo points, with output token usage ranging around 8x from low to max effort. Like with GPT-5.6 Sol, this means Opus 5 can use either far fewer or far more tokens to complete the evaluation than models from other labs, depending on effort settings

Full breakdown of the individual evaluations in the Artificial Analysis Intelligence Index for Claude Opus 5 with max effort

AI算出
主要ニュースainew評価標準
記事は Claude Opus 5 という具体的な新モデルの性能指数(Intelligence Index)、コスト削減率、および Fable 5 や競合他社との詳細なベンチマーク比較を報じており、AI モデル発表としての新規性と重要性が高い。ただし、日本企業や日本固有の価格・規制に関する情報は含まれていないため、日本の関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み