Together AI、Kimi K3とClaude Fable 5を比較
Together AI が、コーディング性能とコスト効率を評価する DeepSWE ベンチマークで、月之暗面の Kimi K3 と Anthropic の Claude Fable 5 を比較検証した。
主要ポイント
Kimi K3 は Claude Fable 5 と同等の品質を維持しながら、1 つのタスクを解決するあたりのコストは約 3 分の 1。さらにオープンモデルであるため、チームがデプロイを完全にコントロールできます。
- DeepSWE における Kimi K3 と Claude Fable 5 の比較では、pass@1(1 回の試行での成功率)は僅差です。Fable が 69.9% で K3 が 68.5% と、1.4 ポイントの差があります。
- しかし、試行回数を増やせば K3 が逆転します。pass@2(最大 2 回の試行)では 82.0% vs 80.2%、pass@4(最大 4 回の試行)では 89.4% vs 88.5% と、K3 が上回っています。
- コスト面でも K3 は圧倒的です。1 ロールアウトあたりの費用は K3 が 4.65 ドルに対し、Fable は 13.41 ドル。同じ予算で解決できるタスク数は K3 の方が約 2.8 倍です。
- ただし、Claude Fable 5 のほうが安定性は高いと言えます。4 回連続で成功するケース(4-for-4)では、Fable が 58 件に対し K3 は 45 件にとどまっています。
DeepSWE は多様なタスク種別やプログラミング言語にわたってモデルのソフトウェアエンジニアリング能力を評価するベンチマークです。今月の Kimi K3 と Claude Fable 5 の比較において、最も注目すべきはリーダーボードの首位ではなく、Fable にわずか 1.4 ポイント差で迫りながら、価格は約 3 分の 1 という新登場のオープンウェイトモデル「Kimi K3」です。
Kimi K3 は 2026 年 7 月 16 日に DeepSWE に参入しました。最大限の努力を払った 452 件の評価済みロールアウト(内訳:実稼働中のオープンソースリポジトリから抽出された 113 のリアルで長期にわたる機能リクエスト、各 4 回の試行)がテストされました。これらの結果は、非公開のテストスイートによって合格・不合格が判定されています。分析では、Anthropic が最も強力な設定(xhigh)としていた Claude Fable 5 と比較しました。これは当時のベンチマークリーダーでした。
以下の数値はいずれもこの一連の実行に基づいています。そのため、他の公開されている Kimi K3 と Claude Fable 5 のスコアカードとは異なる場合があります。
DeepSWE スコアボード:pass@1 と pass@k
DeepSWE の公式スコアリング基準では、Fable xhigh は 69.9% のタスクを初回で解決しています。一方、Kimi K3 は最大 68.5% です。オープンウェイトモデルと Anthropic のフラッグシップモデルの間にはわずか 1.4 ポイントの差しかありませんが、このデータセット全体におけるモデルリリース間での最大の飛躍は、K3 が 38 ポイントも向上した点にあります。
また、pass@k を大きく設定するとランキングは逆転します。pass@2 では Kimi が 82.0% で Fable の 80.2% を上回り、pass@4 では Kimi が 89.4% と Fable の 88.5% や Sol の 85.8% を大きく引き離しています。44 通りの設定構成全体を通じて、90.3% に達したのは、Luna max と GPT-5.5 high という 2 つの低コストな GPT 設定のみです。

カバレッジと信頼性:Kimi K3 と Fable 5 の違い
pass@1 を「カバレッジ(少なくとも一度は解決したタスクの割合)」と「信頼性(そのタスクでの達成率)」に分解すると、2 つのモデルは異なる特性を示します。Kimi はベンチマークの 89.4% に到達しており、これは過去最高値を上回る水準です。決して解けないタスクはわずか 12 件だけで、Fable の 13 件とほぼ同等です。しかし、4 回連続で成功させる信頼性では Kimi は 76.6% で、4 回とも成功したタスク数は 45 件に留まります。一方、Fable は信頼性が 79.0% と高く、4 回連続成功のタスク数も 58 件です。
Fable はより安定しており確定的な挙動を示しますが、Kimi はより広い網を張るタイプで、これが pass@2 や @4 での成績向上につながっています。

コスト比較:Kimi K3 と Claude Fable 5 の価格
- Kimi K3 はロールアウトあたり 4.65 ドル、Fable xhigh は 13.41 ドル。合計 452 ロールアウトのスイープでは、Kimi が 2,103 ドルに対し Fable は 6,010 ドルとなりました。
- 解決したタスクあたりのコスト効率を比較すると、Kimi は 100 ドルあたり 14.7 件の解決を実現するのに対し、Fable は 5.3 件。つまり Kimi の方が 2.8 倍の成果を出しています。
- Kimi K3 は処理に時間がかかりますが、モデルがオープンソース化され推論が最適化されれば、この点は確実に改善されるでしょう。

Kimi K3 と Claude Fable 5 の類似性はどれほどか?
- タスクごとの相関関係は 0.72 で、ベンチマーク全体で最も高いクロスベンダー間の類似性を示しています。実際、上位 4 つのクロスベンダー間での類似性すべてが、Kimi K3 と Anthropic(Fable)の組み合わせです。
- どちらかのモデルが 4 タスク中 4 件を解決し、もう一方が 0 件というケースは、どの方向でも存在しません。これはこれまで分析したすべてのペアで初めての現象です。
- 両モデルとも 96 タスクを解決しますが、Kimi だけが追加で 5 件、Fable だけが追加で 4 件のタスクを解決し、8 つのタスクは両方のモデルが失敗しています。
- 失敗のパターンも似ています。失敗の 65% が「あと一歩」の状態であり、リポジトリ内の既存テストスイートへの影響(ベースライン回帰)も、Kimi は 11%、Fable は 10% とほぼ同等です。

- 実務上、Kimi K3 と Claude Fable 5 はほぼ同じタスクで成功し、同じタスクで失敗します。つまり、両者を組み合わせても多様性はほとんど得られません。両者の併用でカバーできるのは 113 タスク中 105 件に過ぎず、Kimi 単独の 101 件と比べてもわずかな差です。

プログラミング言語別:Kimi K3 と Fable 5 の比較
Go 言語では Kimi が明確に勝利(スコアは 79 vs 71)しました。一方、Fable は残る 4 つの言語で優位に立ちました。Python では 74-68、JavaScript で 70-65、TypeScript で 64-60、そして Rust では 75-65 です。
特筆すべきは、Rust のタスクにおいて K3 が Fable に大きく肉薄している点です。他のどのモデルも、GPT-5.6 Sol でさえも、K3 ほど Rust に強みを持っていません。

この結果が示す意味
Kimi K3 は、今や「合理的なデフォルト」と言える存在になりました。フラッグシップクラスに匹敵する性能を持ちながら、Fable の 4 回試行時の成功率(pass@4)を凌駕し、コストは Fable の 1 ドルに対してわずか 35 セントです。
Together AI で Kimi K3 を実行する
Kimi K3 はオープンウェイトモデルとして公開されています。これにより、前述の pass@4 の性能やコスト構造を、自社の環境で自由に活用できます。重み(weights)が入手可能になった今、Together AI の推論スタックは Kimi K3 といったオープンモデルを生産規模で提供するために設計されています。これを使えば、最先端モデルの高額なトークン料金に頼らず、より広い試行回数(pass@k)の網羅性を追求できます。Together AI の推論価格ページ を確認し、自社のワークロードに見合ったコスト感を把握してください。
よくある質問 (FAQs)
Kimi K3 は Claude Fable 5 より優れていますか?
評価指標によります。Claude Fable 5 は、DeepSWE ベンチマークにおける単一試行の信頼性(pass@1:69.9% vs 68.5%)で勝利し、4 回連続でのタスク解決数も上回っています。
一方、Kimi K3 は pass@2 と pass@4 で勝っており、コストは圧倒的に低いです。そのため、大量処理や再試行を許容するエージェントワークにおいては、Kimi K3 の方が高い価値を持つ選択肢と言えます。
Kimi K3 は Claude Fable 5 よりどれくらい安いか?
今回の検証では、Kimi K3 のロールアウトあたりのコストは 4.65 ドルでした。一方、Claude Fable 5 を xhigh セットで実行した場合の費用は 13.41 ドルとなり、Kimi K3 はおおよそ 3 分の 1 の価格です。
解決したタスク数で換算すると、Kimi K3 は 100 ドルあたり 14.7 件の解決を達成しました。これに対し Fable 5 は 5.3 件でした。つまり、ドルあたりの生産性は約 2.8 倍、Kimi K3 の方が上回っています。
Kimi K3 はオープンウェイトモデルですか?
はい、Kimi K3 は Moonshot AI が提供するオープンウェイトモデルです。重み(weights)が公開されれば、自社でホストしたり、推論プロバイダーを通じて提供したりすることが可能です。一方、Claude Fable 5 はクローズドなモデルであり、Anthropic およびそのパートナー経由での利用のみが認められています。
コーディングにおいてどちらが優れているか?
コーディング能力は両者とも互角です。言語別の内訳を見ると、Kimi K3 は Go で明確にリードしています。一方、Claude Fable 5 は Python、JavaScript、TypeScript、Rust の各言語で先行しています。
Fable 5 は単一の試行において安定した結果を出しますが、Kimi K3 は複数回の試行を通じてより広い範囲をカバーする傾向があります。
DeepSWE における pass@k とは?
pass@k は、タスクに対して k 回行った試行のうち、少なくとも 1 回が非公開のテストスイートをパスしたかどうかを測定する指標です。pass@1 は「最初の試行で正解すること」を評価しますが、k の値が大きくなるほど、「複数回の試行を経て最終的に解決に到達できるか」という能力が重視されます。
Kimi K3 の優位性は、この k の値が大きくなるにつれて顕著になります。
原文を表示
Key Takeaways
Kimi K3 matches Claude Fable 5 on quality, costs a third as much per solved task, and as an open model gives teams full control over their deployment.
- Kimi K3 vs Claude Fable 5 is close on DeepSWE pass@1: Fable leads 69.9% to 68.5%, a 1.4 point gap.
- Give the models more attempts and Kimi K3 pulls ahead. It wins pass@2 (82.0 vs 80.2) and pass@4 (89.4% vs 88.5%).
- Kimi K3 is far cheaper: $4.65 per rollout vs $13.41, and 2.8x more solved tasks per dollar.
- Claude Fable 5 is the more reliable model: it solves more tasks four-for-four (58 vs 45).
In our Kimi K3 vs Claude Fable 5 comparison on DeepSWE, a benchmark that tests a model's software engineering capabilities across many task types and programming languages, the most interesting model this month is not the one at the top of the leaderboard. It is Kimi K3, the new open-weight model parked 1.4 points behind Claude Fable 5, at a third of the price.
Kimi K3 landed in DeepSWE on July 16, 2026, with 452 graded rollouts at max effort: 113 real, long-horizon feature requests from live open-source repos, four trials each, graded pass/fail by a hidden test suite. We analyzed all of them against Claude Fable 5 at its best setting (xhigh), Anthropic’s strongest configuration and the former benchmark leader. Every figure below comes from this run, so it can differ from other public Kimi K3 vs Claude Fable 5 scorecards.
The DeepSWE scoreboard: pass@1 and pass@k
Fable xhigh solves 69.9% of tasks on the first try under DeepSWE’s official scoring. Kimi K3 max solves 68.5%. One point four between an open-weight model and Anthropic’s flagship, the K3 +38 point leap is the largest between model releases in this entire dataset.
Also when you allow for larger pass@k's the ranking flips. pass@2 Kimi is ahead, 82.0 vs 80.2. pass@4 Kimi is 89.4% above Fable’s 88.5 and Sol’s 85.8; across the entire 44-config export, only two cheap GPT configs (Luna max and GPT-5.5 high, at 90.3) have ever reached more.

Coverage vs reliability: where Kimi K3 and Fable 5 differ
Decompose pass@1 into coverage (tasks solved at least once) and reliability (pass rate on those tasks) and the two models occupy different corners. Kimi reaches 89.4% of the benchmark - higher than any peak, with only 12 tasks it never cracks (Fable: 13). But it's less reliable on 4/4 tries: 76.6% reliability and only 45 tasks solved four-for-four, against Fable’s 79.0% and 58. Fable is steadier and more deterministic; Kimi is the wider net, which accounts for its gains at pass@2 and @4.

Cost comparison: Kimi K3 vs Claude Fable 5 pricing
- Kimi K3: $4.65 per rollout. Fable xhigh: $13.41. The full 452-rollout sweep: $2,103 vs $6,010.
- Per solved task, Kimi delivers 14.7 solves per $100 versus Fable’s 5.3 - 2.8x the work per dollar.
- Kimi takes a much longer time but this will no doubt improve when the model is open sourced and inference is optimized!

How similar are Kimi K3 and Claude Fable 5?
Per-task correlation between Kimi K3 and Fable is 0.72 - the highest cross-vendor similarity in the entire benchmark. In fact the top four cross-vendor similarities in the export are all Kimi-K3-versus-Anthropic pairs. There is not a single task where one goes four-for-four and the other zero-for-four, in either direction - a first across every pairing we have analyzed. Both solve 96 tasks, Kimi alone adds 5, Fable alone adds 4, and the same 8 resist both.
Their failure anatomies match too: 65% of failures are near misses for both, and both protect the repo’s existing test suite (11% vs 10% baseline regressions).

In practice, Kimi K3 and Claude Fable 5 succeed and fail on nearly the same tasks, so pairing them buys you almost no diversity: their union covers 105 of 113 tasks, barely above Kimi alone at 101.

Kimi K3 vs Fable 5 by programming language
Kimi takes Go decisively (79 vs 71). Fable holds the other four: Python 74-68, JavaScript 70-65, TypeScript 64-60, and Rust 75-65.
The really interesting detail here is how much K3 catches up with Fable on Rust tasks - no other model, not even GPT 5.6 Sol, is as good at Rust.

What it means
Kimi K3 is now the rational default: near-flagship reach, the best pass@4 of any flagship-tier config, at 35 cents on Fable’s dollar.
Run Kimi K3 on Together AI
Kimi K3 ships as an open-weight model, which is what makes the pass@4 reach and the cost profile above usable on your own terms. Once the weights are available, Together AI's inference stack is built to serve open models like Kimi K3 at production scale, so you can chase the wider pass@k net without frontier token prices. See Together AI inference pricing to size it against your workload.
FAQs
Is Kimi K3 better than Claude Fable 5?
It depends on the metric. Claude Fable 5 wins single-attempt reliability on DeepSWE (pass@1 69.9% vs 68.5%) and solves more tasks four-for-four. Kimi K3 wins pass@2 and pass@4 and costs far less, so it is the stronger value pick for high volume or retry-tolerant agent work.
How much cheaper is Kimi K3 than Claude Fable 5?
In our run, Kimi K3 cost $4.65 per rollout versus $13.41 for Claude Fable 5 at its xhigh setting, roughly a third of the price. Measured per solved task, Kimi K3 returned 14.7 solves per $100 against Fable's 5.3, about 2.8x the work per dollar.
Is Kimi K3 open weight?
Yes. Kimi K3 is an open-weight model from Moonshot AI, so once the weights are released it can be self-hosted or served through inference providers. Claude Fable 5 is a closed model available only through Anthropic and its partners.
Which is better for coding, Kimi K3 or Claude Fable 5?
For coding they are close. In our language breakdown Kimi K3 takes Go decisively, while Claude Fable 5 leads Python, JavaScript, TypeScript, and Rust. Fable is steadier on any single attempt; Kimi casts a wider net across multiple attempts.
What is pass@k on DeepSWE?
pass@k measures whether at least one of k attempts at a task passes the hidden test suite. pass@1 rewards getting it right first try; higher k rewards a model that can eventually reach a solution across several tries. Kimi K3's edge grows as k increases.
AI算出
技術分析ainew評価高い
Kimi K3 がオープンモデルとして Claude Fable 5 に匹敵する性能を持ちながらコストが約 1/3 という具体的な数値データを提示しており、技術的な実装や選定に有用な独自分析が含まれている。ただし、対象は中国・米国の主要モデルであり、日本固有の導入事例や規制情報は含まれていないため、日本の関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み