Together AI、Kimi K3とGPT-5.6 SolのDeepSWEベンチ比較結果を公開
Together AI の分析によると、GPT-5.6 Sol は単発性能で優位だが、Kimi K3 は複数試行とコスト効率で勝り、両モデルの特性差を活かしたルーティング戦略が最も効果的である。
AI深層分析を開く2026年7月27日 16:01
キーポイント
単発性能 vs 複数試行性能の逆転
DeepSWE ベンチマークにおいて GPT-5.6 Sol は pass@1 で 72.7% と Kimi K3 (68.5%) を上回るが、Kimi K3 は pass@2 や pass@4 で逆転し、複数試行時の成功率が高い。
コスト効率における圧倒的な差
Kimi K3 の 1 ロールアウトあたりの費用は 4.65 ドルで GPT-5.6 Sol (8.37 ドル) より 64% 安く、100 ドルあたりで解決できるタスク数は Kimi K3 が 14.7 と GPT-5.6 Sol の 5.3 を上回る。
信頼性と失敗パターンの相違
GPT-5.6 Sol は 4 回試行すべてで成功するタスク数 (61) が Kimi K3 (45) より多く信頼性が高いが、両モデルは解決・失敗するタスクの傾向が異なるため、相関が低い。
最適なルーティング戦略の提案
Kimi-first カスケード(まず Kimi K3 を試し、失敗時に GPT-5.6 Sol に昇格させる構成)により、113 タスク中 108 タスク (約 85.6%) をカバーできる戦略が示された。
重要な引用
GPT-5.6 Sol edges Kimi K3 on single-shot quality, but Kimi wins on pass@k with k > 1 and costs 64% less per completed task.
The models diverge (0.46 correlation across which tasks they solve correctly and incorrectly) and fail differently, so a Kimi-first cascade that escalates to Sol covers 108 of 113 tasks.
編集コメントを表示
編集コメント
単発性能のランキング争いだけでなく、実運用におけるコストと複数試行時の成功率を重視した比較は、実際のシステム設計において極めて示唆に富む。Kimi K3 のオープンウェイト属性と GPT-5.6 Sol の信頼性の違いを組み合わせる戦略は、リソース制約のある現場での有効なアプローチとなるだろう。
GPT-5.6 Sol は単一試行での品質で Kimi K3 をわずかに上回りますが、複数回の試行(pass@k, k > 1)では Kimi K3 が勝利し、完了したタスクあたりのコストは 64% も安くなります。両モデルの成功と失敗のパターンが異なるため、ベンチマーク上ではこの 2 つを状況に応じて使い分ける(ルーティングする)戦略が最も有効です。
- DeepSWE の pass@1 では Kimi K3 と GPT-5.6 Sol は互角ですが、Sol が 72.7% で K3 の 68.5% を上回り、4.2 ポイントの差がついています。
- しかし、試行回数を増やせば Kimi K3 が逆転します。pass@2 では 82.0%(K3)対 81.0%(Sol)、pass@4 では 89.4%(K3)対 85.8%(Sol)と、K3 が上回ります。
- コスト面では Kimi K3 の圧勝です。ロールアウトあたりの費用は K3 が 4.65 ドルに対し Sol は 8.37 ドルで、ドルあたりに解決できるタスク数は K3 の方が 2.8 倍多いです。
- GPT-5.6 Sol はより信頼性の高いモデルです。4 回すべて成功したタスクの数は Sol が 61 件に対し、K3 は 45 件にとどまります。
- 両モデルが正解・不正解とするタスクの傾向は大きく異なり(相関係数 0.46)、失敗する理由も異なります。そのため、まず Kimi K3 を試し、必要に応じて Sol にエスカレートさせるカスケード構成にすれば、113 タスク中 108 タスクを解決し、成功率は約 85.6% に達します。
DeepSWE · Head to Head
Kimi K3 vs GPT-5.6 Sol の比較サマリー
| Metric | Kimi K3 | GPT-5.6 Sol |
|---|---|---|
| DeepSWE pass@1 | 68.5% | 72.7% |
| DeepSWE pass@2 | 82.0% | 81.0% |
| DeepSWE pass@4 | 89.4% | 85.8% |
| Coverage (solved at least once) | 89.4% | 85.8% |
| Reliability (4/4 pass rate) | 76.6% | 84.5% |
| Tasks solved four-for-four | 45 | 61 |
コストとパフォーマンスの比較
ロールアウトあたりのコスト:Kimi K3 は 4.65 ドル、GPT-5.6 Sol は 8.37 ドル。
100 ドルあたりに解決できるタスク数:Kimi K3 は 14.7 件、Sol は 5.3 件。
ロールアウトの中央値所要時間:Kimi K3 は 66 分、Sol は 17 分。
オープンウェイト対応:Kimi K3 は Yes、Sol は No。
DeepSWE ベンチマークにおける Kimi K3 と GPT-5.6 Sol の比較では、誰がリーダーボードの頂点に立つかが焦点ではありません。このベンチマークは、多様なタスクタイプやプログラミング言語にわたってモデルのソフトウェアエンジニアリング能力を評価するものです。
GPT-5.6 Sol は 2 週間前に DeepSWE の公式スコアリングに基づき pass@1 で 72.7% を達成し、単一試行での最強かつ最高信頼性を記録しました。一方、オープンウェイトの newcomer である Kimi K3 は別の軸で応えています。Kimi K3 は pass@4 で 89.4% を達成し、Sol を含むすべてのフラッグシップモデルを上回っています。さらに、ロールアウトあたりのコストは Sol の 8.37 ドルに対し、Kimi K3 は 4.65 ドルと低く抑えられています。
私たちは、deepswe.datacurve.ai で公開された各トライアルの記録から、合計 904 の評価済みロールアウト(113 タスク×各 4 回トライアル、双方とも最大努力で実施)を分析しました。
DeepSWE スコアボード:pass@1 と pass@k
単一の試行では Sol が明確に勝利します。DeepSWE の公式スコアリングによると、Sol は 72.7%、Kimi K3 は 68.5% です。しかし、この差はすぐに縮まります。
2 回の試行では Kimi K3 がすでに上回ります(82.0% vs 81.0%)。4 回の試行では、Kimi の 89.4% が Sol の 85.8% を 3.6 ポイント引き離し、これは現在利用可能なすべてのフラッグシップ構成の中で最高の pass@4 パフォーマンスです。

コスト比較:Kimi K3 と GPT-5.6 Sol の価格
ロールアウト 1 回あたりのコストは、Kimi K3 が $4.65、GPT-5.6 Sol が $8.37 です。タスク解決数で換算すると、Kimi K3 は 100 ドルあたり 14.7 件の解決を達成し、Sol の 5.3 件と比較して約 2.8 倍の効率を示しています。
ただし、速度とステップ数のトレードオフがあります。Median(中央値)でのロールアウト所要時間は、Sol が 17 分であるのに対し、Kimi K3 は 66 分かかります(2024 年 7 月 25 日時点の Moonshot 社提供 API による)。Sol は必要なステップ数が約 40% 少ないのが特徴です。推論プロバイダーが Kimi K3 の最適化を進めるにつれて、エンドツーエンドのレイテンシは改善される見込みですが、ステップ数の多さはモデル固有の挙動によるものです。

プログラミング言語別:Kimi K3 と GPT-5.6 Sol の性能
Go 言語では両モデルが同率の 79 件で並んでいます。Python(74 対 68)、TypeScript(66 対 60)、JavaScript(75 対 65)では Sol が上回っています。一方、Rust では Kimi K3 が 65 件で Sol の 60 件を抜いています。

カバレッジと信頼性:対照的な特性
「pass@1」と「pass@4」を分解して分析すると、両モデルはそれぞれ異なる強みを持っています。カバレッジとは 4 回の試行のうち少なくとも 1 回でタスクが解決された割合、信頼性は 4 回すべてで成功した割合を指します。
Sol はベンチマークの 85.8% をカバーし、そのうち 61 タスクは 4 回中 4 回(100%)成功しました。この結果、信頼性スコアは 84.5% に達しています。一方、Kimi K3 はカバレッジが 89.4% と主要モデル中最も広く、45 タスクで完璧な成績を収めましたが、信頼性は 76.6% です。
Sol は安定性と決定論的な実行に優れ、Kimi K3 はより広い範囲のタスクに対応できる網羅性を備えています。

Kimi K3 と GPT-5.6 Sol、どこまで違うのか?
Kimi K3 と Fable の比較がほぼ同じ結果を返す「共鳴」状態(相関 0.72)だったのに対し、Kimi K3 と Sol は明確に別々の振る舞いを示します。タスクごとの相関は 0.46 に留まり、両モデルとも互いに完璧な成績を残したり、全く失敗したりするケースが対照的に見られます。つまり、成功と失敗のパターンが本質的に異なるのです。
この違いこそが、両モデルをラウティングやカスケード(順次処理)で組み合わせる際の強みになります。2 台を連携させることで、ベンチマークの全 113 タスクのうち 108 タスク(95.6%)をカバー可能となり、これは現在確認されている中で最も優れた 2 モデル構成です。

両モデル間のラウティングでどこまで精度を上げられるか?
両モデルが 113 タスク中 18 タスクで意見が分かれる場合、その不一致こそが「追加の正解」を生むチャンスになります。実際にどれほどの精度向上が見込めるかは、誤答を検知して上流へエスカレートできる検証機能(テストスイート)があるかどうかにかかっています。
現実的な答えは約 85.6% です。まず Kimi K3 を実行し、テストスイートで結果が却下された場合のみ Sol に切り替えるという構成です。これなら単独モデルや、完璧なラウター(83.4%)よりも高い精度を達成できます。その理由は、カスケード構成により難易度の高いタスクに対して 2 回の独立した試行が可能になるからです。
また、1 タスクあたりのコストも Sol 単独より安くなります。Kimi K3 が約 70% のタスクを処理し終えるまで Sol は起動しないためです。この仕組みの鍵は検証機能にあります。検証機能がなければ現実的なラウティングでも 70% 台に留まり、83.4% が理論上の上限となります。
DeepSWE · ルーティング戦略
ルーティング戦略:タスクごとの精度とコストのトレードオフ
| ルーティング戦略 | 精度 | タスクあたりのコスト |
|---|---|---|
| ベスト単一モデル(ルーティングなし、Sol) | 72.7% | $8.37 |
| 現実的な学習済みルーター(1 shot) | 72–83% | 約$6 |
| 完璧なオラクルルーター(1 shot、正解モデルを選択) | 83.4% | 約$6.07 |
| カスケード:Kimi → Sol(失敗時) | 85.6% | $7.30 |
| 両モデルを一度ずつ実行し、どちらかを採用 | 85.6% | $13.02 |
この組み合わせにおける理論上の上限は95.6%です。113件のタスクのうち、8回の試行すべてで両モデルが解決できなかったのは5件だけでした。95.6%を超える性能を出すには、より優れたルーターではなく、第三のモデルを追加する必要があります。

タスクタイプ別でどちらが有利か
効果的にルーティングを行うには、コーディングの要求内容に基づいてタスクを分類する必要があります。Sol が 8 ドメイン中 5 つで優位に立ち、Kimi K3.6 が残りを担います。
具体的には、シリアライゼーション(92対79)、並行処理(72対55)、プログラム解析(64対56)のタスクは Sol に任せます。一方、オペレーションツールリング(79対73)やランタイム内部の実装(77対75)は Kimi K3.6 が得意としています。
コンフォーマンス(適合性)テストは 61対59と僅差で、ほぼ五分の勝負です。なお、各タスクタイプはベンチマークのプロンプトを基に LLM によって分類されました。

失敗のモード
両モデルは全く異なる方法で失敗します。Kimi K3 はテストにほぼ合格しますが、すべての課題をクリアすることはできません。一方、Sol は既存のベースラインテストをより多く破綻させます。具体的には、Sol の失敗事例の 20% でリポジトリ内の既存テストが壊れており、これは他の GPT モデルでも同様の傾向が見られることです。
Kimi K3 では、11% の基本テストで失敗しましたが、65% のケースで「ニアミス」が発生しました。つまり、新規テストの 80% 以上は通過していたものの、一部が失敗したという結果です。

Together AI で Kimi K3 を実行する
Kimi K3 はオープンウェイトモデルとして提供されています。これにより、独自の条件で pass@4 の達成率を高め、低コストを実現し、上記のようなルーティング戦略を活用することが可能になります。
重みが公開されたことで、Together AI の推論スタックは Kimi K3 といったオープンモデルを生産規模で提供できるように設計されています。これにより、Kimi を優先するカスケード処理を実行しつつ、最先端モデルのトークン価格に頼らず、より広い pass@k の範囲を追求できます。
ご自身のワークロードに合わせてコスト感を把握するには、Together AI の推論料金ページをご覧ください。
Kimi K3 vs GPT-5.6 Sol FAQ
Kimi K3 は GPT-5.6 Sol より優れているか?
評価指標によります。GPT-5.6 Sol は DeepSWE における単一試行の品質(pass@1 で 72.7% vs 68.5%)で勝利し、4 つの課題すべてを一度に解決するケースも多いため、この点では優れています。
一方、Kimi K3 は pass@2 や pass@4 のスコアで上回り、コストが圧倒的に低いため、高ボリュームな処理や再試行を許容できるエージェントワークにおいては、より価値の高い選択肢となります。多くのチームにとって最適な答えは、両モデルの特性に応じて使い分けるルーティング戦略を採用することです。
Kimi K3 は GPT-5.6 Sol よりどれくらい安いか?
今回のテストでは、Kimi K3 の 1 ランあたりのコストは 4.65 ドルに対し、GPT-5.6 Sol は 8.37 ドルでした。つまり Kimi K3 は Sol の半分以下の価格です。解決したタスク単位で換算すると、Kimi K3 は 100 ドルあたり 14.7 件の解決を達成し、Sol の 5.3 件と比較して約 2.8 倍の成果を出しています。
Kimi K3 と GPT-5.6 Sol を使い分けるべきか?
結果を検証できる環境であれば、両モデルを組み合わせることをお勧めします。この 2 つのモデルは相関が低く(0.46)、失敗するパターンも異なります。そのため、まず Kimi K3 で試みて、テストスイートで出力が拒否された場合に GPT-5.6 Sol にエスカレートさせるアプローチをとると、解決率は約 85.6% に達します。これは単独で使用する場合や、完璧なルーティングモデルを使用した場合よりも優れた結果です。両者を組み合わせることで、113 タスク中 108 タスクをカバーできます。
コーディングタスクにおいてどちらが優れているか?
コーディング能力においては、両者の差は僅かです。Go では同率ですが、Python、TypeScript、JavaScript では GPT-5.6 Sol がリードし、Rust では Kimi K3 が上回ります。GPT-5.6 Sol は 1 回の試行で安定して高い性能を発揮する一方、Kimi K3 は複数回の試行を通じてより広い範囲の解決策を見つけ出す傾向があります。
DeepSWE における pass@k とは?
pass@k は、タスクに対して k 回試行したうち、少なくとも 1 回が非公開のテストスイートをパスするかどうかを測定する指標です。pass@1 は初回の試行で正解することを評価しますが、k の値が大きくなるほど、複数回の試行を通じて最終的に解決に到達できるモデルの評価が高まります。Kimi K3 の優位性は、k が増えるにつれて顕著になります。
原文を表示
Summary
GPT-5.6 Sol edges Kimi K3 on single-shot quality, but Kimi wins on pass@k with k > 1 and costs 64% less per completed task. The two models succeed and fail in different ways, which makes routing between them the strongest play on the benchmark.
- Kimi K3 vs GPT-5.6 Sol is close on DeepSWE pass@1: Sol leads 72.7% to 68.5%, a 4.2 point gap.
- Give the models more attempts and Kimi K3 pulls ahead. It wins pass@2 (82.0 vs 81.0) and pass@4 (89.4% vs 85.8%).
- Kimi K3 is far cheaper: $4.65 per rollout vs $8.37, and 2.8x more solved tasks per dollar.
- GPT-5.6 Sol is the more reliable model: it solves more tasks on all 4/4 tries (61 vs 45).
- The models diverge (0.46 correlation across which tasks they solve correctly and incorrectly) and fail differently, so a Kimi-first cascade that escalates to Sol covers 108 of 113 tasks and reaches about 85.6%.
DeepSWE · Head to Head
Kimi K3 vs GPT-5.6 Sol at a glance
Metric
Kimi K3
GPT-5.6 Sol
DeepSWE pass@1
68.5%
72.7%
DeepSWE pass@2
82.0%
81.0%
DeepSWE pass@4
89.4%
85.8%
Coverage (solved at least once)
89.4%
85.8%
Reliability (4/4 pass rate)
76.6%
84.5%
Tasks solved four-for-four
45
61
Cost per rollout
$4.65
$8.37
Solved tasks per $100
14.7
5.3
Median rollout time
66 min
17 min
Open weights
Yes
No
In our Kimi K3 vs GPT-5.6 Sol comparison on DeepSWE, a benchmark that tests a model's software engineering ability across many task types and programming languages, the story is not who sits at the top of the leaderboard. GPT-5.6 Sol took the pass@1 crown two weeks ago with 72.7% under DeepSWE's official scoring, the strongest single shot and the highest reliability ever recorded on the benchmark. Kimi K3, the open-weight newcomer, answers on the other axis: 89.4% pass@4, better than every flagship on the board, Sol included, while costing $4.65 a rollout to Sol's $8.37. We analyzed all 904 graded rollouts (113 tasks, four trials each, at max effort on both sides) from the published per-trial records at deepswe.datacurve.ai.
The DeepSWE scoreboard: pass@1 and pass@k
Given a single attempt, Sol wins clearly: 72.7% to 68.5% under DeepSWE's official scoring. But the gap closes fast. At two attempts Kimi K3 is already ahead, 82.0 to 81.0. At four attempts, Kimi's 89.4% beats Sol's 85.8% by 3.6 points, the best pass@4 of any flagship-tier config on the board.

Cost comparison: Kimi K3 vs GPT-5.6 Sol pricing
- Kimi K3: $4.65 per rollout. GPT-5.6 Sol: $8.37.
- Per solved task, Kimi delivers 14.7 solves per $100 versus Sol's 5.3, about 2.8x the work per dollar.
- The tradeoff is speed and steps: a median Sol rollout takes 17 minutes against Kimi's 66 (using the Kimi K3 API from Moonshot as of July 25th), with Sol using roughly 40% fewer steps. End-to-end latency should improve as inference providers optimize Kimi K3, though the extra steps are a model-behavior trait.

Kimi K3 vs GPT-5.6 Sol by programming language
The two are tied on Go (79 each). Sol leads Python (74-68), TypeScript (66-60), and JavaScript (75-65). Kimi K3 takes Rust (65-60).

Coverage vs reliability: opposite corners of the plane
Decompose pass@1 and pass@4 into coverage (tasks solved at least once across four tries) and reliability (tasks solved on all four), and the two models sit in opposite corners. Sol posts 84.5% reliability with 61 tasks solved four-for-four, across a modest 85.8% of the benchmark. Kimi K3 reaches 89.4% coverage, wider than any flagship, but only 76.6% reliability and 45 rock-solid tasks. Sol is steadier and more deterministic; Kimi K3 casts the wider net.

How different are Kimi K3 and GPT-5.6 Sol?
Where Kimi K3 versus Fable was essentially an echo chamber (0.72 correlation), Kimi K3 and Sol diverge: per-task correlation is 0.46, with four-for-four-versus-zero-for-four corners in both directions. The models succeed and fail in genuinely different ways. That is exactly what makes them a strong pair to route and cascade between: together they cover 108 of 113 tasks (95.6%), the best two-model portfolio on the benchmark.

How high can routing between them get you?
Because the two disagree on 18 of 113 tasks, routing turns that disagreement into free accuracy. How much you capture depends on whether you have a verifier (your test suite) to catch a bad answer and escalate.
The practical answer is about 85.6%: run Kimi K3 first and escalate to Sol only when the test suite rejects the result. That beats either model alone and even a perfect one-shot router (83.4%), because the cascade gives hard tasks two independent attempts instead of committing to one. It is also cheaper per solve than Sol alone, since Kimi K3 clears roughly 70% of the queue before Sol is ever invoked. The verifier does the heavy lifting: without one, a realistic router lands in the 70s and 83.4% is the ceiling.
DeepSWE · Routing strategies
Routing strategy: accuracy vs cost per task
Routing strategy
Accuracy
Cost / task
Best single model, no routing (Sol)
72.7%
$8.37
Realistic learned router (1 shot)
72–83%
~$6
Perfect oracle router (1 shot, pick right model)
83.4%
~$6.07
Cascade: Kimi → Sol on failure
85.6%
$7.30
Both models once, keep either
85.6%
$13.02
The hard ceiling for this pair is 95.6%. Five of the 113 tasks are solved by neither model across all eight combined attempts. Past 95.6% you need a third model, not a better router.

Where each wins, by task type
To route well, classify tasks by what the coding ask actually is. Sol leads 5 of 8 domains, Kimi K3. Send serialization (92-79), concurrency (72-55), and program analysis (64-56) to Sol. Send ops tooling (79-73) and runtime internals (77-75) to Kimi K3. Conformance is a 61-59 coin-flip. Task types here were classified by an LLM from each benchmark prompt.

Failure modes
They fail in very different ways - Kimi K3 gets close but doesn't pass all tests whereas Sol breaks more of the baseline tests. Sol breaks the repo's existing tests in 20% of failures - this is pretty consistent with other GPT models actually. Kimi K3: 11% base tests failed but 65% near misses where > 80% new tests were passing but some failed!

Run Kimi K3 on Together AI
Kimi K3 ships as an open-weight model, which is what makes the pass@4 reach, the low cost, and the routing story above usable on your own terms. Once the weights are available, Together AI's inference stack is built to serve open models like Kimi K3 at production scale, so you can run the Kimi-first cascade and chase the wider pass@k net without frontier token prices.
See Together AI inference pricing to size it against your workload.
Kimi K3 vs GPT-5.6 Sol FAQ
Is Kimi K3 better than GPT-5.6 Sol?
It depends on the metric. GPT-5.6 Sol wins single-attempt quality on DeepSWE (pass@1 72.7% vs 68.5%) and solves more tasks four-for-four. Kimi K3 wins pass@2 and pass@4 and costs far less, so it is the stronger value pick for high-volume or retry-tolerant agent work. For most teams the best answer is to route between them.
How much cheaper is Kimi K3 than GPT-5.6 Sol?
In our run, Kimi K3 cost $4.65 per rollout versus $8.37 for GPT-5.6 Sol, a little over half the price. Measured per solved task, Kimi K3 returned 14.7 solves per $100 against Sol's 5.3, about 2.8x the work per dollar.
Should I route between Kimi K3 and GPT-5.6 Sol?
Yes, if you can verify results. The two models diverge (0.46 correlation) and fail differently, so running Kimi first and escalating to Sol when your test suite rejects the output reaches about 85.6%, beating either model alone and even a perfect one-shot router. Together they cover 108 of 113 tasks.
Which is better for coding, Kimi K3 or GPT-5.6 Sol?
For coding they are close. They tie on Go; Sol leads Python, TypeScript, and JavaScript; Kimi takes Rust. Sol is steadier on any single attempt, while Kimi casts a wider net across multiple attempts.
What is pass@k on DeepSWE?
pass@k measures whether at least one of k attempts at a task passes the hidden test suite. pass@1 rewards getting it right first try; higher k rewards a model that can eventually reach a solution across several tries. Kimi K3's edge grows as k increases.
AI算出
技術分析ainew評価高い
Kimi K3 と GPT-5.6 Sol のコーディング能力、コスト、ルーティング性能を DeepSWE ベンチで比較した独自データが含まれており、技術的な深みがあるため technical_analysis に分類される。また、具体的なモデル名とバージョンがタイトルに含まれるため検索機会スコアは最高値となる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み