Claude モデル選定ガイド:用途別ベストモデル解説
Anthropic は、Claude モデルの各バージョンが持つ特性を比較し、開発者が利用ケースに応じて最適なモデルを選択するための具体的なガイドラインを提供した。
AIニュース価値スコアβ
技術分析AI関連度、新規性、日本での有用性など6軸を公開検証中です。現在、掲載順には使用していません。
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 検索具体性
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
Anthropic の公式ブログによる Claude モデルファミリーの特性比較と選定ガイドであり、AI モデルの実運用に関する深い分析を含んでいる。ただし、日本固有の情報や企業事例は含まれていないため、日本の文脈での関連性は低めとなる。
キーポイント
モデルの役割分担と特性比較
Claude の各バージョン(Sonnet, Haiku など)が持つコストパフォーマンス、推論速度、知的能力の違いを明確に定義し、それぞれの得意分野を解説している。
ユースケース別最適化戦略
複雑な推論が必要なタスクには高性能モデルを、大量のテキスト処理や低遅延が求められる場面には軽量モデルを選択するよう推奨しており、コストと性能のバランスを取る方法を提示している。
開発者への実用的アドバイス
単に機能比較を行うだけでなく、実際のプロジェクト設計においてどのようにモデルを切り替えるか、あるいは組み合わせるかの具体的な戦略を示している。
重要な引用
Anthropic is explaining the characteristics of each Claude model and how to choose the best one for your use case.
Choosing the right model can significantly impact cost, latency, and performance.
影響分析・編集コメントを表示
影響分析
この記事は、Anthropic の製品ポートフォリオが拡大する中で、ユーザーが自社のユースケースに合わせて最適なリソース配分を行うための重要な指針となる。技術的な詳細よりも「使い分け」に焦点を当てているため、実務レベルの開発者やアーキテクトにとってコスト削減とパフォーマンス最適化の即戦力となる情報である。
編集コメント
モデルの性能比較自体は技術ブログでよく見られますが、今回は「どのケースにどれを使うか」という実装レベルの判断基準を明確にした点が価値が高いです。特にコスト管理が重視される昨今の開発現場において、このガイドラインは導入時の意思決定を迅速化するのに役立つでしょう。
私たちのアドバイス:賢く始める
「このタスクにはどのモデルを選べばよいのか」という質問は、私たちが最も頻繁に受けるものの一つです。より多くのモデルクラスやバージョンをリリースするにつれて、その答えも次第に複雑なものとなっています。
この記事では、各モデルクラスの詳細説明に加え、モデル選定時に確認すべき主要な質問やその他のベストプラクティスについて解説します。
ただし、細かいニュアンスは一旦脇に置いておきましょう。私たちのデフォルトの推奨は、「最も賢い一般利用可能なモデル」から始めることです。その上で、必要に応じて「エフォートレベル(処理コスト)」を調整してパフォーマンスとコストのバランスを取ります。
より高度なモデルほど、トークンあたりの単価が高くても、タスクあたりのコストが低くなるケースが多いです。特に低エフォート設定ではその傾向が強まります。これは、能力の高いモデルほど、必要な対話回数や思考時間が短く済むためです。逆に、小さなモデルから始めると、モデル自体の性能不足による失敗と、セットアップの問題による失敗を区別しにくくなるというリスクがあります。
もちろん、レイテンシやコストに敏感なユースケースが発生した場合は、下位クラスのモデルも試して、最適な組み合わせを見つけることができます。
また、コスト効率の高いモデルから始め、品質基準を満たすまで上位クラスへ移行するアプローチを選ぶ組織もあります。私たちは、モデル選定に関するドキュメントで、これらの両方の方向性を紹介しています。
Claude モデルファミリー
Mythos / Fable
Mythos は Anthropic が提供する最も高性能なモデルクラスであり、あらゆる分野で最先端の能力を備えています。特にコーディング、長時間実行するエージェントタスク、そしてこれまで AI が安定して処理できなかった問題の解決において、その能力は際立っています。
この Mythos クラスには、同じ基盤モデルを用いた 2 つのパッケージが存在します。Claude Mythos は、デュアルユース(軍事転用可能性)を持つサイバーセキュリティや生物学の研究を扱う信頼できる組織向けに設計されています。一方、Claude Fable は、一般ユーザーが安全に利用できるよう追加の安全性対策を組み込んだパッケージです。どちらのモデルも、安全な利用を実現するために 限定的なデータ保持ポリシー を適用しています。
Opus
Opus は、推論を要するエンタープライズタスクに特化した強力なモデルクラスです。Opus モデルは、知識労働向けの GDPval-AA や、エージェントによるコーディング向けの Terminal-Bench 2.1 など、主要な業界ベンチマークにおいて常に上位にランクインしています。
表面上では Opus と Fable の使い分けが明確に見えないかもしれません。両者ともコーディング、長時間実行するエージェント、知識作業において卓越した能力を発揮します。しかし実際の運用環境では、Fable などの大規模モデルは、ベンチマークスコアが Opus と同等であっても、より高い知恵や創造性、そして優れた文章作成スキルを備えている傾向があります。
一般的な目安としては、評価や内部テストで Opus が特定のタスクに苦戦しているようであれば、Fable が最適解となります。一方、Opus がすでに品質基準を十分にクリアしている場合は、その速度と価格設定のバランスから、より良い選択肢となる可能性があります。
Sonnet
Sonnet は、日常的なタスクに対応する多目的なモデルクラスです。高頻度のサブエージェントが多数存在するマルチエージェント構成など、幅広い汎用ユースケースにおいて、性能・コスト・速度の最適なバランスを提供します。
Haiku
Haiku は、最も低コストかつ高速なモデルクラスです。レイテンシとコストが重要な、高頻度で動作するワークロード向けに設計されています。
Claude モデルの使い分け:どのモデルをワークロードに選ぶべきか
Claude の各モデルクラスは特定の作業タイプに特化しているわけではありません。金融分野には一つ、科学分野には別というように、用途に応じてモデルを分けることを推奨していません。すべての Claude モデルは、コーディング、エージェントタスク、知識処理といった領域で卓越するように訓練されています。
モデルクラス間の主な違いは、「どの程度の難易度の問題を確実に処理できるか」と、その能力を実現するために必要な価格と速度のバランスにあります。モデルを選ぶ際は、以下の点を自問してみてください。
このタスクはどれほど難しいですか? 多くの時間を要する、複数のステップを踏む必要がある、あるいはこれまで解決されたことのない問題である場合は、より高性能なモデルクラスが適しています。
レイテンシ要件はどうでしょうか。 モデルが高頻度で顧客と接するワークロードに組み込まれる場合、Sonnet が最適な選択肢となることが多いです。
アクセス制限はどうなっていますか?
Mythos は、Project Glasswing の対象組織にのみ提供されています。すべての組織が、すべての役割に対してあらゆるモデルクラスを公開しているわけではありません。
単位経済性はどのようになっていますか?
生産量が多い場合は、評価結果からタスクが十分に完了できていると判断される場合に限らず、より低クラスのモデルを使用する方が適しているケースがあります。トークンあたりの料金はモデルによって異なりますし、その能力や必要な処理量に応じて、1 タスクあたりのコストも変わります。
必要とされる処理量(Effort Level)は、品質・速度・コストのバランスにも影響します。高クラスのモデルを高処理量で利用すれば、最も優れたパフォーマンスが得られます。また、低処理量の設定でも、高クラスのモデルの方が小規模なモデルよりも効率的に動作する場合があります。

グラフは概略を示すものであり、ベンチマークデータからプロットされたものではありません。
image
グラフは概略を示すものであり、ベンチマークデータからプロットされたものではありません。詳しくは、Claude Code における Claude モデルと処理量の選択 をご覧ください。
アドバイザー戦略で各モデルの強みを組み合わせる
「アドバイザー戦略」advisor strategy を活用すれば、コスト効率に優れたワーカーモデルが、より高度なモデルを呼び出して計画の検証や成果物の評価を行わせることができます。これにより、パフォーマンスの向上が期待できます。
この手法では、実行役となるモデルは必要な時だけ指導を受けるため、性能が大幅に向上します。例えば、SWE-bench Pro の Sonnet 5 ベンチマークにおいて、Fable 5 をアドバイザーとして活用した場合、Fable 5 を単独で全タスクを実行する場合の約 63% のコストで、そのスコアの 90% に迫る結果を達成できます。
モデル選定における評価指標とベンチマークの役割
モデルの能力が自社のニーズを満たしているかを確認するには、主に「標準ベンチマーク」と「独自評価」の 2 つの方法があります。
ベンチマークとは、特定のドメインに特化した事前定義されたタスクやシナリオのセットで、正解が既知となっています。これらは、異なるモデルクラスやプロバイダー間での能力を方向付けるための有用な指標となります。ただし、Opus や Fable のような高性能モデルを対象とする場合、テスト問題のほとんどを正答してしまう「飽和(saturation)」現象が発生し、評価が難しくなるという課題があります。
こうしたケースでは、組織は実際の業務負荷でモデルを実行するか、自社の独自評価を用いてテストを行うことを推奨します。これにより、どのモデルが最適か判断できます。一般的に、評価とは生産現場から抽出した問題のキュレーションセットであり、既存のツールでは対応しきれない難易度の高いタスクや、チームが定義した成功基準が含まれるのが特徴です。

ここが、最先端モデルの能力や創造性が他社製品や他モデルと明確に差をつけるポイントです。私たちはこれまで、カスタムエージェントの評価を効果的に行うためのベストプラクティスについて詳しく解説してきました。
賢い選択をするために
AI モデルの選定に万能な正解はありません。だからこそ、私たちは複数のモデルクラスを用意しています。最適なモデルを選ぶためには、各モデルクラスの基礎知識を理解し、自社のユースケースを深く把握することが不可欠です。そのためには、堅牢な評価基準を構築・維持・運用する必要があります。
原文を表示
Our advice: start smart
One of the most frequent questions we hear is “what model should I choose for this workload?” As we have released more model classes and versions, the answer has become more nuanced.
This article covers those details including a description of each model class, the top questions to ask when selecting a model, and other best practices.
But to put aside the nuance for a moment, our default recommendation is to start with the most intelligent generally available model and use effort level to dial in performance and cost.
Cost-per-task is often lower for more intelligent models, especially at lower effort levels, even if the price-per-token is higher. This is because more capable models often take fewer turns and less thinking time to get most tasks right. Starting with a smaller model can also make it harder to distinguish between model failures and setup failures.
Of course, as use cases arise that are more latency or cost-sensitive, you can test lower tier models until you find your ideal fit.
Some organizations may also choose to start with the most cost effective model and move up classes until the quality bar is met. We include both directional approaches in our documentation on model selection.
The Claude model family
Mythos / Fable
Mythos is Anthropic’s most capable model class, with frontier capabilities across domains. This model class is especially capable at coding, long-running agent tasks, and solving problems AI has not reliably handled before.
The Mythos class ships in two packages of the same underlying model. Claude Mythos is for trusted organizations handling dual-use cybersecurity and biology work while Claude Fable is packaged with additional safeguards that make the model safe for use by the general public. Both require limited data retention so they can be used safely.
Opus
Opus is our powerful model class for reasoning-intensive enterprise tasks. Opus models consistently rank among leading models on key industry benchmarks such as GDPval-AA for knowledge work and Terminal-Bench 2.1 for agentic coding.
The choice between Opus and Fable may not seem clear on the surface, as both excel at coding, long-running agents, and knowledge work. In real-world situations, larger models such as Fable tend to have more wisdom, creativity, and writing skills despite having similar benchmark scores to models such as Opus.
The general rule of thumb is if your evals or internal testing show Opus struggling on some tasks, then Fable is the answer. If Opus already clears the quality bar, then its speed and price profile may make it the better choice.
Sonnet
Sonnet is our versatile model class for everyday tasks. Sonnet provides a balance of performance, cost, and speed for the widest set of general purpose use cases, including high-volume sub-agents in multi-agent orchestration setups.
Haiku
Haiku is our lowest cost and fastest model class. Haiku models are designed for high-frequency workloads where latency and cost matter.
How to choose which Claude model is best for your workload
Our model classes don’t specialize in one type of work. We don’t recommend one model class for finance and another for science. Every Claude model is trained to excel in areas like coding, agentic tasks, and knowledge work.
The main difference across model classes is in *how* *hard* *a problem* they can reliably carry, and what that capability costs in price and speed. When choosing a model, ask:
How hard is this task? If it typically takes a lot of time, involves multiple steps, or is previously unsolved then a more capable model class is appropriate.
What are the latency needs? If the model is involved in high-frequency customer facing workloads, then Sonnet is often the best choice.
What are the access constraints? Mythos is only available to organizations under Project Glasswing. Not all organizations make all model classes available to all roles.
What are the unit economics? Higher volumes of production may be more appropriate for lower classes of models, particularly if evaluations show those tasks are completed satisfactorily. Models are priced differently per token and will have different price-per-task costs based on their capabilities and effort level.
Effort level also impacts the balance of quality, speed, and cost. Higher-class models at higher efforts offer the best possible performance, and higher-class models at lower efforts can sometimes be more efficient than smaller models.


To learn more read Choosing a Claude model and effort level in Claude Code.
Combining models’ strengths with the advisor strategy
The advisor strategy allows faster, lower-cost worker models to call more intelligent models to check their plan and evaluate their work, leading to improved performance.
This method, where the executor model is coached only when needed, improves performance by a substantial amount. For example, on SWE-bench Pro Sonnet 5 with a Fable 5 advisor is within 10% of Fable 5’s score at 63% of the price of using Fable 5 for the whole task.
How evals and benchmarks help with model choice
Two common ways to see if model capabilities are sufficient for your needs are to use standard benchmarks and custom evaluations.
Benchmarks are a set of pre-determined tasks or scenarios, often for a specific domain, with known solutions. These can be helpful directional guides for evaluating capabilities across model classes and providers. The challenge arises when evaluating powerful models, such as Opus and Fable, which can solve almost all of the questions on the test (often referred to as saturation).
In these cases, we recommend organizations use the models on real workloads or test them with their own evaluations to make a decision on which model is the right choice. Typically, evaluations are a curated set of problems drawn from production — including difficult tasks where your current tools fall short, with success criteria your team defines.

This is where the capability and creativity of frontier models start to separate from the pack and from one another. We’ve written extensively on the best practices for developing custom agent evaluations.
Making the smart choice
There is no one-size-fits-all approach to AI model selection, which is why we make multiple model classes available. Ultimately, the best way to select a model is to understand the basics of each model class and understand your use case in-depth. That means building, maintaining, and deploying strong evaluations.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み