Together AI、DeepSWEベンチでGPT-5.6 LunaとDeepSeek-V4 Flashを比較
本文の状態
日本語全文を表示中
詳細モードで約15分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Together AI Blog
Together AIはDeepSWEベンチにおいて、品質ではGPT-5.6 Lunaが上回るものの、コスト効率を考慮したカスケード構成によりDeepSeek-V4 Flashの方が精度と費用面で優位であることを報告した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月7日 14:22
AI深層分析
キーポイント
性能とコストのトレードオフ分析
GPT-5.6 Luna は DeepSWE ベンチマークで Pass@1 が 67.2% と DeepSeek-V4 Flash の 53.3% を上回るが、そのコストは Luna が 0.61 ドルに対し DeepSeek は 0.10 ドルと約 6 倍の差がある。
カスケード構成による最適解
DeepSeek-V4 Flash をまず実行し、失敗した場合にのみ GPT-5.6 Luna に昇格させる戦略により、78.9% のタスクを 0.385 ドルで解決でき、単独使用よりも精度が高くコストも 37% 削減できる。
エラーの性質とリソース効率
DeepSeek-V4 Flash は失敗時に既存テストスイートを破損する割合が 9% と Luna の 15% より低く、また 100 ドルあたりの解決件数は DeepSeek が 532 件で Luna の 110 件を大きく上回る。
DeepSeek-V4 Flash のコスト効率性
Luna は DeepSeek より約14ポイント精度が高いが、コストは約6倍かかる。DeepSeek は2回の試行で Luna の単発試行を上回る性能を出し、価格は3分の1以下となる。
速度とトークン数のトレードオフ
DeepSeek は Luna よりも遅く(中位数23分対16分)、より多くのステップと出力トークンを消費するが、その利点は時間ではなく金銭的コストにある。
重要な引用
While GPT-5.6 Luna is the stronger engineer on every quality measure, DeepSeek-V4 Flash 0731 is cheap enough that a DeepSeek-first cascade beats Luna alone on both accuracy and cost.
DeepSeek-V4 Flash fails more cleanly, breaking the repo's existing test suite in 9% of failures vs Luna's 15%.
"Luna wins single shot clearly: 67.2% pass@1 to DeepSeek's 53.3% under DeepSWE's official scoring, a 14 point lead"
"DeepSeek is the slower of the two here... DeepSeek's advantage is money, not time."
編集コメントを表示
編集コメント
この分析は、高性能モデルが常に最良の選択肢ではないという現実を浮き彫りにしている。開発者はタスクの難易度や予算に応じて、複数のモデルを戦略的に組み合わせる「カスケード」アプローチの重要性を再認識する必要があるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
価格競争力において新基準を打ち立てたオープンウェイトモデル
主なポイント
GPT-5.6 Luna はあらゆる品質指標でより優れたエンジニアですが、DeepSeek-V4 Flash 0731 の低コストを活かした「DeepSeek 優先の連鎖(カスケード)」構成は、Luna を単独で使う場合よりも精度とコストの両面で優れています。
- GPT-5.6 Luna は DeepSWE ベンチマークの pass@1 で 67.2% と 53.3% を記録し、14 ポイントの差をつけて明確にリードしています。試行回数が増えるにつれてもこのリードは維持されます。
- DeepSeek-V4 Flash は DeepSWE ベンチボード上で最も安価なモデルです。ロールアウトあたり 0.10 ドル(Luna の 0.61 ドルの約 6 分の 1)で、100 ドルあたりの解決件数は Luna の 110 件に対し 532 件を達成しています。
- DeepSeek-V4 Flash は失敗する際にもクリーンに動作し、リポジトリの既存テストスイートを破綻させるケースは失敗時の 9% に留まります(Luna は 15%)。
- まず DeepSeek-V4 Flash を実行し、失敗した場合のみ Luna にエスカレーションする構成では、タスクの 78.9% を解決できます。コストは 0.385 ドル/件で、Luna 単独よりも精度が高く、かつ 37% 安価です。
DeepSeek-V4 Flash 0731 は DeepSWE ベンチボード全体の中で最も安価なモデルで、タスクあたり約 10 セントです。一方、堅実な上位フラッグシップである GPT-5.6 Luna はタスクあたり 0.61 ドルで、DeepSeek の価格の約 6 倍になります。したがって問われるべきは「どちらがリーダーボードで勝つか(Luna が圧倒的に勝利)」ではなく、「6 分の 1 の価格で何が得られ、何を得られないのか」、そして「両者を組み合わせることで単独使用よりも優位に立てるか」という点です。
113 の DeepSWE タスク(ライブのオープンソースリポジトリからの実在する長期的機能要件)において、DeepSeek-V4 Flash 0731 (max) と GPT-5.6 Luna (max) を比較しました。各モデルで 4 回の試行を行い、非公開のテストスイートによって合格・不合格を判定しています。公開された試行記録に基づくロールアウト総数は 900 回(DeepSeek が 452 回、Luna が 448 回)です。以下のすべての数値は本試験の結果に基づいているため、他のパブリックスコアカードとは異なる可能性があります。
一見してわかること
DeepSWE · ハード対ハード
DeepSeek-V4 Flash と GPT-5.6 Luna の比較
| モデル | Pass@1 | 平均コスト | 生成トークン数 | ステップ数 |
|---|---|---|---|---|
| gpt-5.6-luna [max] | 67% ± 4% | $0.61 | 73k | 102 |
| deepseek-v4-flash [max] | 53% ± 4% | $0.10 | 108k | 153 |
Luna は精度が約 14 ポイント高い代わりに、コストは DeepSeek の約 6 倍です。DeepSeek は出力トークン数とステップ数を増やして成果を出していますが、それでも 0.10 ドルという価格は、このセットの中で圧倒的に安価なランニングコストとなっています。

DeepSWE スコアボード:pass@1 と pass@k
Luna は単一試行(single shot)で明確に勝利しました。DeepSWE の公式スコアリングによると、Luna の pass@1 は 67.2% で、DeepSeek の 53.3% を上回り、14 ポイントの差(相対的には約 26% 高い精度)となっています。
試行回数を揃えた場合も Luna が常にリードしています。2 回の試行では Luna が 81.6%、DeepSeek が 70.1%、4 回の試行では Luna が 90.3%、DeepSeek が 80.5% です。純粋な解決能力だけで見れば、これは明確な差があります。
しかし、DeepSWE のようなタスクでは、複数の試行を並列実行して成功率を高めることが可能です。この視点に立つと、経済的な側面が状況を大きく変えます。DeepSeek の pass@2(70.1%)はすでに Luna の単一試行(67.2%)を上回っており、コストも DeepSeek が 2 回で約 0.20 ドルなのに対し、Luna は 0.61 ドールです。
もし検証器が正解となる実行結果を選別できるなら、安価なモデルでもフラッグシップ機種の初回試行品質に匹敵する結果を、3 分の 1 のコストで達成できます。これは後述のルーティング手法を導入する前段階での話です。

コスト比較:DeepSeek-V4 Flash と GPT-5.6 Luna の価格
1 件あたりのコストは $0.61 で、Luna は 100 ドルあたり 110 件の解決を提供するのに対し、DeepSeek は 532 件です。つまり、低価格モデルの方が 4.8 倍の価値があります。ロールアウトごとのコスト差は約 6 倍に達します。
ただし、安価なモデルが提供しないのは速度です。これが「フラッシュ」モデルの一般的な話とは異なる点です。実は DeepSeek の方が slower で、中位数では Luna の 16 分・92 ステップに対し、DeepSeek は 23 分・148 ステップを要します。処理に時間がかかり(より多くの出力も生成し、中位数で 104k トークンに対して Luna は 70k)、DeepSeek の優位性は時間ではなくコストにあります。
この実時間の差は、モデルが比較的新しいことによる影響の一部です。推論エンジンが最適化されるにつれて、この差は縮まると予想されます。

失敗のモード:DeepSeek は旗艦モデルより失敗を許容する
もう一つ、安価なモデルが静かに勝ち取っているのは「規律」です。DeepSeek が失敗した際、リポジトリ内の既存テストスイートを破綻させるのは 9% の場合だけです。一方、Luna ではそれが 15% に達します。これは GPT ファミリー特有の回帰シグネチャーであり、Sol や他の OpenAI 系モデルでも同様に 15〜20% で見られる現象です。
両モデルとも失敗の多くは「ほぼ成功」に近いケース(DeepSeek は 69%、Luna は 66%)ですが、高価なモデルほど既に動作していたコードを乱す可能性が高いのです。もし Luna をデプロイする場合は、完全な回帰テストのゲートを通す必要がありますが、DeepSeek ではそのガードレールはそれほど必要ありません。

タスクドメイン別 DeepSeek-V4 Flash と GPT-5.6 Luna の比較
コードの性質に基づいて 113 のタスクを分類すると、Luna は 8 ドメインのうち 7 で勝利しています。特に優位性が顕著なのは、推論が求められる分野です。
プログラム解析では 69 vs 33、並行処理と耐久性では 70 vs 38、言語やランタイムの内部構造では 86 vs 59 と、各ドメインで約 30 ポイントもの差がついています。まさにモデルの実力が問われる領域で、DeepSeek は大きく後れをとっています。
DeepSeek が勝利しているのはたった一つのドメインです。それは「クエリ言語と設定言語」の分野(78 vs 70)です。SQL ビルダー、ウィンドウ関数、キーセット分页、設定ファイルパーサーなど、構造化されており、スキーマに則り、決まり文句に従う作業においては、安価な DeepSeek も十分に競合可能であり、場合によっては Luna を上回ることもあります。
しかし、システム全体を通じて厳格な不変条件(インバリアント)の維持が求められるタスクでは、Luna の能力差が決定的なものとなります。

プログラミング言語別 DeepSeek-V4 Flash と GPT-5.6 Luna の比較
GPT-5.6 Luna は全 5 カテゴリで勝利しましたが、その差額から DeepSeek の弱点が浮き彫りになります。Rust(55 vs 60)や Go(62 vs 79)では健闘していますが、JavaScript では大惨敗です。スコアは 35 で、Luna の 60 と比べて 25 ポイントもの差があり、この対戦における最も低いスコアとなっています。Python も 49 と 65 で苦戦しています。もし開発スタックが JavaScript に依存しているなら、安価な DeepSeek を選ぶのは見栄えだけの節約になりかねません。一方、設定ファイルやクエリ処理、あるいは Rust が中心のプロジェクトであれば、DeepSeek はほぼ差を埋めることができます。

DeepSeek-V4 Flash と GPT-5.6 Luna の類似性はどの程度か?
安価なモデル同士のペアと比較すると、その類似性は低いです。タスクごとの相関は 0.50 に留まり、解決能力の偏りも顕著です。両モデルとも 87 タスクを共通で解決していますが、Luna だけが独自に解決したタスクが 15 件あるのに対し、DeepSeek が独自に解決したのはわずか 4 件です。重要なのは、DeepSeek が独占的に解決して Luna が全く失敗したケース(DeepSeek に有利な「4 勝 0 敗」の状況)が一つもないことです。一方、Luna は DeepSeek が一度も達成できなかった 4 つのタスクを独占しています(go-critic-doc-link-checker, langchain-request-coalescing, meriyah-explicit-resource-declarations, superjson-error-stack-serialization)。両モデルの解決タスクの総数は 113 中 106(93.8%)をカバーしていますが、そのほとんどは Luna の独自領域によるものです。純粋な多様性の観点から言えば、DeepSeek が追加する価値は限定的です。

ポートフォリオ戦略:なぜカスケードは機能するのか
この組み合わせは無意味であるべきなのに、実際にはそうではありません。その理由は価格にあります。
まず DeepSeek を実行し、テストスイートが回答を拒否した場合のみ Luna へエスカレーションします。このアプローチでは、タスクあたりの解決率は 78.9% で、コストは 0.385 ドルです。これをよく見てみましょう。Luna 単体(67.2%、11.7 ポイント向上)よりも精度が高く、Luna 単体(0.61 ドル)よりも安価で、タスクあたりのコストは Luna の約 63% に抑えられます。
最初の段階がほぼ無料であることが、この手法の核心です。DeepSeek は各タスクをわずかなコストで処理し、キューの約 53% をクリアします。そのため、Luna が課金するのは困難な残り部分のみとなり、Luna に到達したタスクには、2 つ目の独立した試行が追加されます。
このカスケード方式は、完璧な一発オラクルルーター(74.3%)さえも上回ります。なぜなら、2 回のチャンスがあれば、1 回の完璧な選択よりも勝つからです。どちらのモデルを先頭にするかで精度に差はありませんが、DeepSeek を先に使う方が圧倒的に安価です(0.385 ドル対 0.64 ドル)。したがって、常に安価な方から始めるべきです。
このペアの上限は 93.8% です。最終的にどちらのモデルも解決できない 7 つのタスク(gql-incremental-graphql-delivery や bandit-structured-nosec-directives など)には、より優れたルーターではなく、さらに強力な第 3 のモデルが必要です。

何を意味するか
単独で比較すれば、DeepSeek-V4 Flash 0731 は中堅モデルの域を出ません。クエリと設定に関するドメインでは唯一の実力を持ち、失敗パターンも明確ですが、JavaScript における弱点は決定的です。その代わり、価格競争力は圧倒的です。
一方、GPT-5.6 Luna はあらゆる品質指標において優れたエンジニアです。Pass@1 で 14 ポイント差をつけ、8 つあるドメインの 7 つ、そして全 5 つの言語で勝利し、処理速度も速いです。その対価として、DeepSeek の約 6 倍のコストがかかりますが、GPT ファミリー特有の回帰現象に対するガードレールが必要になる点も考慮すべきです。
タスクあたり 0.61 ドルという価格は、フラッグシップモデルとしては本格的に安価と言えます。品質を最優先する場合は、これが最も簡単なデフォルト選択肢となるでしょう。しかし、DeepSeek の真価は Luna の代替としてではなく、「Luna のフロントエンド」として発揮されます。
コストが数セントで済む第一段階で作業の半分を片付けておけば、フラッグシップモデル以上の精度を、その単価以下で購入できるのです。安価なモデルがこれほどまでに安ければ、カスケード(段階的処理)は妥協策ではなく、盤面において最良の手となります。
| 指標 | DeepSeek-V4 Flash 0731 (max) | GPT-5.6 Luna (max) |
|---|---|---|
| pass@1(公式スコアリング) | 53.3% | 67.2% |
| pass@1(エラーを失敗としてカウント) | 53.3% | 67.2% |
| pass@2 / pass@4 | 70.1 / 80.5% | 81.6 / 90.3% |
| カバレッジ / 信頼性 | 80.5 / 66.2% | 90.3 / 74.8% |
| 堅牢 (4/4) / 壁 (0/4) | 23 / 22 | 43 / 11 |
| ロールアウトごとのコスト / 合計 | 45 | 273 |
| $100 あたり解決数 | 532 | 110 |
| 中央値分 / ステップ数 | 23 / 148 | 16 / 92 |
| 中央値ピークコンテキスト / 出力トークン | 202k / 104k | 202k / 70k |
| 失敗の解剖分析 | 69% 近接、9% 回帰 | 66% 近接、15% 回帰 |
| 獲得ドメイン(8 件中) | 1 (クエリ & 設定) | 7 |
| 獲得言語 | なし | 全 5 |
| タスク別相関 / 和集合 | 0.50 / 113 中 106 (93.8%) | |
| Cascade DeepSeek > Luna(精度 / コスト) | 78.9% / $0.385 | Luna 単独 vs 67.2% / $0.61 |
| Oracle 1-shot ルーター | 74.3% | |
| リーダーボード順位(52 構成中) | 30 位 | 13 位 |
| インフラエラー | 0 | 0 |
よくある質問
GPT-5.6 Luna は DeepSeek-V4 Flash より優れているのでしょうか?
品質面では、GPT-5.6 Luna が勝っています。DeepSWE の pass@1 では 67.2% を達成し、同等の試行回数で常にリードしています。8 つあるタスクドメインのうち 7 つと、すべての 5 つ言語で勝利を収め、処理速度も速いです。
一方、DeepSeek-V4 Flash は価格面で優れています。1 ロールアウトあたりのコストは約 6 倍安く、失敗した際にもクリーンにエラーを返す傾向があります。
DeepSeek-V4 Flash は GPT-5.6 Luna よりどれくらい安価なのでしょうか?
今回のテストでは、DeepSeek-V4 Flash の 1 ロールアウトあたりのコストは約 0.10 ドルでした。対照的に、GPT-5.6 Luna は 0.61 ドルで、価格は約 6 分の 1 です。
解決したタスク数で換算すると、DeepSeek は 100 ドルあたり 532 タスクを解決しました。一方、Luna は 110 タスクです。つまり、ドルあたりの生産性は DeepSeek の方が約 4.8 倍高いことになります。
DeepSeek-V4 Flash と GPT-5.6 Luna を併用すべきでしょうか?
はい、カスケード(順次処理)形式での併用が推奨されます。まず DeepSeek で試行し、テストスイートで回答が拒否された場合にのみ GPT-5.6 Luna にエスカレートする構成です。
この方法では、78.9% のタスクを解決できました。1 タスクあたりのコストは 0.385 ドルです。これは Luna を単独で使う場合の精度(67.2%)よりも高く、かつコスト(0.61 ドル)も安くなります。
常に低コストモデルから始めるのが鉄則です。どちらを先に使っても最終的な精度は同じですが、DeepSeek から始める方が総コストを抑えられます。
コーディングにおいては、DeepSeek-V4 Flash と GPT-5.6 Luna のどちらが優れているのでしょうか?
言語別の分析では、Luna がすべての言語で勝利を収めています。ただし、その差は言語によって異なります。
DeepSeek は Rust において十分な性能を示し、クエリや設定関連の作業では競合するどころか、むしろリードしています。一方、JavaScript の処理能力が最も弱く(35 vs 60)、JS を多用するスタックでは安価なモデル単独での運用は避けるべきです。
DeepSWE における pass@k とは?
pass@k は、タスクに対して k 回の試行のうち少なくとも 1 回が非公開のテストスイートを通過したかどうかを測定する指標です。pass@1 は初回で正解した場合に評価が高くなりますが、より高い k の値は、複数回の試行を経て最終的に解決策に至るモデルを優遇します。
DeepSeek の pass@2(70.1%)はすでに Luna の pass@1(67.2%)を上回っており、これが並列実行やカスケード処理における経済性の根拠となっています。
原文を表示
The open-weight model setting the new price-intelligence bar
Key Takeaways
While GPT-5.6 Luna is the stronger engineer on every quality measure, DeepSeek-V4 Flash 0731 is cheap enough that a DeepSeek-first cascade beats Luna alone on both accuracy and cost.
- GPT-5.6 Luna leads DeepSWE pass@1 decisively at 67.2% vs 53.3%, a 14 point gap, and holds the lead at every equal attempt count.
- DeepSeek-V4 Flash is the cheapest model on the DeepSWE board: $0.10 per rollout vs $0.61, delivering 532 solves per $100 against Luna's 110.
- DeepSeek-V4 Flash fails more cleanly, breaking the repo's existing test suite in 9% of failures vs Luna's 15%.
- Running DeepSeek-V4 Flash first and escalating to Luna only on failure solves 78.9% of tasks at $0.385 each: more accurate than Luna alone and 37% cheaper.
DeepSeek-V4 Flash 0731 is the cheapest model on the entire DeepSWE board: about ten cents a task! GPT-5.6 Luna, a solid upper-tier flagship, runs $0.61 a task, roughly six times DeepSeek’s price. So the question is not which one wins the leaderboard (Luna, comfortably) but what six-times-cheaper buys you, what it costs you, and whether the two together beat either one alone.
We ran DeepSeek-V4 Flash 0731 (max) against GPT-5.6 Luna (max) on all 113 DeepSWE tasks: real, long-horizon feature requests from live open-source repos, four trials each, graded pass/fail by a hidden test suite. That is 900 rollouts in total from the published per-trial records (452 on DeepSeek's side, 448 on Luna's). Every figure below comes from this run, so it can differ from other public scorecards.
At a glance
DeepSWE · Head to Head
DeepSeek-V4 Flash vs GPT-5.6 Luna at a glance
| Model | Pass@1 | Avg cost | Out tok | Steps |
|---|---|---|---|---|
| gpt-5.6-luna [max] | 67% ± 4% | $0.61 | 73k | 102 |
| deepseek-v4-flash [max] | 53% ± 4% | $0.10 | 108k | 153 |
Luna is ~6x the cost for ~14 points more accuracy; DeepSeek uses more output tokens and steps to get less far, but at $0.10 it is the cheapest run in the set by a wide margin.

The DeepSWE scoreboard: pass@1 and pass@k
Luna wins single shot clearly: 67.2% pass@1 to DeepSeek's 53.3% under DeepSWE's official scoring, a 14 point lead, or about 26% more accurate in relative terms. At equal attempt counts Luna stays ahead at every k (81.6 vs 70.1 at two attempts, 90.3 vs 80.5 at four). On raw solving ability, this is not a close fight.
But DeepSWE is exactly the kind of workload where you can fan out several attempts in parallel, and there the economics rewrite the picture. DeepSeek's pass@2 (70.1%) already edges Luna's single shot (67.2%), and two DeepSeek attempts cost about $0.20 to Luna's $0.61. If a verifier can pick the winning run, the cheap model matches the flagship's first-try quality for a third of the price, before any of the routing tricks below even come into play.

Cost comparison: DeepSeek-V4 Flash vs GPT-5.6 Luna pricing
- At $0.61 a task, Luna returns 110 solves per $100 to DeepSeek's 532: a 4.8x value edge for the cheap model. The cost per rollout gap is about 6x.
- The one thing the cheap model does not buy you is speed, and this is the twist versus the usual "flash" model story: DeepSeek is the slower of the two here, taking a median 23 minutes and 148 steps to Luna's 16 minutes and 92. It grinds (and emits more output, 104k median tokens to Luna's 70k). DeepSeek's advantage is money, not time.
- That wall-clock gap is partly an artifact of how new the model is; expect it to shrink as inference engines get tuned for it.

Failure modes: DeepSeek fails more gracefully than the flagship
The other thing the cheap model quietly wins is discipline. When DeepSeek fails, it breaks the repository's existing test suite in only 9% of failures. Luna does so in 15%: the GPT-family regression signature, the same 15 to 20% we see across Sol and the other OpenAI-lineage models. Both fail mostly by near miss (DeepSeek 69%, Luna 66%), but the more expensive model is the one more likely to disturb code that already worked. If you deploy Luna, gate it behind a full regression run; DeepSeek needs that guardrail less.

DeepSeek-V4 Flash vs GPT-5.6 Luna by task domain
Classify the 113 tasks by what the code actually is, and Luna wins 7 of 8 domains. Its biggest edges are exactly the reasoning-heavy work: program analysis (69 vs 33), concurrency and durability (70 vs 38), language and runtime internals (86 vs 59), roughly a 30 point gap in each. This is where model capability actually shows up, and DeepSeek falls off hard.
DeepSeek holds exactly one domain, and it is a telling one: query and config languages, 78 vs 70. The SQL builders, window functions, keyset pagination, config parsers. Structured, schema-shaped, convention-following work is where the cheap model is genuinely competitive, even ahead. Everywhere the task demands holding a hard invariant across a whole system, Luna's capability separates.

DeepSeek-V4 Flash vs GPT-5.6 Luna by programming language
Luna wins all five, but the margins tell you where to be careful with DeepSeek. It is respectable on Rust (55 vs 60) and Go (62 vs 79), but its JavaScript is a collapse: 35 vs Luna's 60, the weakest single cell in the entire matchup, and a 25 point hole. DeepSeek's Python is also soft (49 vs 65). If your stack is JS-heavy, the cheap model is a false economy; if it is config, query, or Rust, DeepSeek closes most of the gap.

How similar are DeepSeek-V4 Flash and GPT-5.6 Luna?
Less than the cheap-tier pairs. Per-task correlation is 0.50, and the diversity is lopsided: they both solve 87 tasks, Luna alone gets 15, and DeepSeek alone gets only 4. Crucially, there is no task DeepSeek sweeps that Luna misses entirely (zero four-for-zero corners in DeepSeek's favor), while Luna sweeps four that DeepSeek never lands (go-critic-doc-link-checker, langchain-request-coalescing, meriyah-explicit-resource-declarations, superjson-error-stack-serialization). Their union covers 106 of 113 (93.8%), but almost all of that is Luna's own reach. As a raw diversity play, DeepSeek adds little.

The portfolio play: why the cascade works anyway
So the pairing should be pointless. It is not, and the reason is the price. Run DeepSeek first and escalate to Luna only when your test suite rejects the answer: 78.9% solved at $0.385 per task. Read that carefully. It is more accurate than Luna alone (67.2%, up 11.7 points) and cheaper than Luna alone ($0.61), landing at about 63% of Luna's per-task cost.
The near-free first stage is the entire trick: DeepSeek clears roughly 53% of the queue for a dime each, so Luna's charge only ever touches the hard remainder, and the tasks that reach Luna get a second independent attempt on top. The cascade even beats a perfect one-shot oracle router (74.3%), because two swings beat one perfect pick. Accuracy is identical whichever model leads, but DeepSeek-first is much cheaper ($0.385 vs $0.64), so always lead with the cheap one.
The pair's ceiling is 93.8%; the 7 tasks neither ever solves (including gql-incremental-graphql-delivery and bandit-structured-nosec-directives) need a stronger third model, not a better router.

What it means
On its own, DeepSeek-V4 Flash 0731 is a mid-pack model with one real domain (query and config), a clean failure profile, a catastrophic JavaScript weakness, and an unbeatable price. GPT-5.6 Luna is the better engineer by every quality measure: 14 points of pass@1, 7 of 8 domains, all 5 languages, faster too. You pay about 6x for it, plus a GPT-family regression habit to guardrail. At $0.61 a task the flagship is genuinely cheap, and it is the easy default when quality is what you care about. But the sharpest use of DeepSeek is not as a Luna replacement; it is as a Luna front-end. A first stage that costs a dime and clears half the work lets you buy flagship-beating accuracy for less than the flagship's own price. When the cheap model is this cheap, the cascade stops being a compromise and becomes the best row on the board.
Appendix: Data Table
DeepSWE · Appendix
DeepSeek-V4 Flash vs GPT-5.6 Luna: full data table
| Metric | DeepSeek-V4 Flash 0731 (max) | GPT-5.6 Luna (max) |
|---|---|---|
| pass@1 (official scoring) | 53.3% | 67.2% |
| pass@1 (errors as failures) | 53.3% | 67.2% |
| pass@2 / pass@4 | 70.1 / 80.5% | 81.6 / 90.3% |
| Coverage / reliability | 80.5 / 66.2% | 90.3 / 74.8% |
| Solid (4/4) / walls (0/4) | 23 / 22 | 43 / 11 |
| Cost per rollout / total | 45 | 273 |
| Solves per $100 | 532 | 110 |
| Median minutes / steps | 23 / 148 | 16 / 92 |
| Median peak context / output tokens | 202k / 104k | 202k / 70k |
| Failure anatomy | 69% near, 9% regression | 66% near, 15% regression |
| Domains won (of 8) | 1 (query & config) | 7 |
| Languages won | none | all 5 |
| Per-task correlation / union | 0.50 / 106 of 113 (93.8%) | |
| Cascade DeepSeek > Luna (accuracy / cost) | 78.9% / $0.385 | vs Luna alone 67.2% / $0.61 |
| Oracle 1-shot router | 74.3% | |
| Leaderboard rank (of 52 configs) | 30th | 13th |
| Infra errors | 0 | 0 |
FAQs
Is GPT-5.6 Luna better than DeepSeek-V4 Flash?
On quality, yes. GPT-5.6 Luna wins DeepSWE pass@1 67.2% to 53.3%, leads at every equal attempt count, wins 7 of 8 task domains and all 5 languages, and is faster. DeepSeek-V4 Flash wins on price by roughly 6x per rollout and fails more cleanly when it does fail.
How much cheaper is DeepSeek-V4 Flash than GPT-5.6 Luna?
In our run, DeepSeek-V4 Flash cost about $0.10 per rollout versus $0.61 for GPT-5.6 Luna, roughly a sixth of the price. Measured per solved task, DeepSeek returned 532 solves per $100 against Luna's 110, about 4.8x the work per dollar.
Should you use DeepSeek-V4 Flash and GPT-5.6 Luna together?
Yes, in a cascade. Running DeepSeek first and escalating to Luna only when the test suite rejects the answer solved 78.9% of tasks at $0.385 each in our run: more accurate than Luna alone (67.2%) and cheaper than Luna alone ($0.61). Always lead with the cheap model; accuracy is the same either way but DeepSeek-first costs less.
Which is better for coding, DeepSeek-V4 Flash or GPT-5.6 Luna?
Luna wins every language in our breakdown, but the margins vary. DeepSeek is respectable on Rust and competitive on query and config work, where it actually leads. Its JavaScript is the weakest cell in the matchup (35 vs 60), so JS-heavy stacks should not rely on the cheap model alone.
What is pass@k on DeepSWE?
pass@k measures whether at least one of k attempts at a task passes the hidden test suite. pass@1 rewards getting it right first try; higher k rewards a model that can eventually reach a solution across several tries. DeepSeek's pass@2 (70.1%) already edges Luna's pass@1 (67.2%), which is what makes the parallel-attempt and cascade economics work.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み