DeepSeek V4 Flash、人工知能指数で前モデルより10ポイント向上し50を記録
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Artificial Analysis
DeepSeek V4 Flash 0731 が Artificial Analysis のインテリジェンス指数で 50 を記録し、前世代より 10 ポイント向上してコスト効率とエージェント性能が大幅に改善された。
AI深層分析を開く2026年8月1日 13:57
AI深層分析
キーポイント
スコアと市場位置の向上
DeepSeek V4 Flash 0731 は Artificial Analysis Intelligence Index で 50 を獲得し、前世代(V4 Flash)より 10 ポイント、GPT-5.6 Luna よりもコスト効率で優位な位置を占める。
コスト効率とキャッシュ戦略
同社のファーストパーティ API ではキャッシュヒット率が約 98% と業界平均より高く設定されており、GPT-5.6 Luna よりもタスクあたりのコストが約 60% 低い。
エージェント性能の飛躍的向上
GDPval-AA v2 での Elo レーティングが 1559 に上昇し、Hallucination(幻覚)率が 84% と大幅に低下したことで、実務タスクにおける信頼性が強化された。
アーキテクチャと仕様の維持
1M トークンのコンテキストウィンドウと 284B の総パラメータ数(推論時 13B アクティブ)は前世代と同じであり、性能向上はアルゴリズムや評価指標の改善によるもの。
知能指数におけるすべての評価項目で前モデルを上回る
DeepSeek V4 Flash 0731 は、エージェント機能の向上に加え、CritPt、SciCode、Humanity's Last Exam、AA-LCR、GPQA Diamond の全評価項目でスコアを改善した。
重要な引用
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, a 10-point jump over DeepSeek V4 Flash
DeepSeek's ~98% cache hit discount on its first-party API comes in at ~60% lower than GPT-5.6 Luna (max)
Improvements in agentic performance: DeepSeek V4 Flash 0731 achieves an Elo rating of 1559 on GDPval-AA v2
AA-Omniscience improvements are driven by fewer hallucinations, rather than higher accuracy
編集コメントを表示
編集コメント
DeepSeek は前世代のアーキテクチャを維持したまま、キャッシュ戦略と評価指標の最適化により劇的なコストパフォーマンス向上を実現した。このモデルは、特にエージェントタスクや大規模な実務処理において、GPT-5.6 Luna を凌駕する選択肢となり得る。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
DeepSeek V4 Flash 0731、Artificial Analysis の知能指数で 50 点を記録。前作の DeepSeek V4 Flash(2026 年 4 月リリース)より 10 ポイント向上し、DeepSeek V4 Pro をも 6 ポイント上回る
DeepSeek V4 Flash 0731 は、知能指数で GPT-5.6 Luna(最大値:51)にわずか 1 ポイント差の 50 を記録しました。OpenAI が本日、GPT-5.6 Luna の価格を 80% 引き下げたにもかかわらず、DeepSeek V4 Flash 0731 の DeepSeek 公式 API におけるタスクあたりのコストは、知能レベルが同等である GPT-5.6 Luna(最大値)と比較して約 60% 低く抑えられています。この価格競争力の源泉は、同社が公式 API で提供する約 98% のキャッシュヒット割引です。これは業界全体で一般的に提供されている 90% という割引率を大きく上回る、極めて積極的な施策です。
本モデルは、既存の DeepSeek V4 Flash とアーキテクチャと価格体系を共有しており、「知能指数」と「タスクあたりのコスト」のパレート最適化フロンティアに位置しています。これは前世代である DeepSeek V4 Flash(スコア 40)からの大きな飛躍であり、GLM-5.2(最大値:51)にもわずか 1 ポイント差まで迫る性能です。
なお、オープンウェイトの最前線である Kimi K3(最大値:57)にはまだ 7 ポイント及ばないものの、直近で発表された Gemini 3.6 Flash(スコア 50)と同等のレベルにあり、Muse Spark 1.1(最高級:51)にも 1 ポイント差です。DeepSeek は今後数週間以内に、本モデルの全重みデータを公開する予定となっています。
DeepSeek V4 Flash 0731 は、1M トークンのコンテキストウィンドウを維持し、パラメータ数は前作の DeepSeek V4 Flash と同じく総数 284B、推論時にアクティブになるのは 13B のままです。
主要な結果:
➤ エージェント性能の向上: エージェントによる実世界のタスクに焦点を当てた評価「GDPval-AA v2」において、DeepSeek V4 Flash 0731 は Elo レーティング 1559 を達成しました。これは前作 DeepSeek V4 Flash の 1189 から大幅な向上です。重みの公開後は、このスコアは Kimi K3 (max, 1687) に次ぐオープンウェイトモデルとして第2位となり、GLM-5.2 (max, 1510) を上回ります。また、Terminal-Bench 2.1 では 17 ポイント上昇して 79%、τ³-Bench Banking でも 8 ポイント上昇し 31% に達しました。
➤ AA-Omniscience の改善は精度向上ではなく、ハルシネーションの減少によるもの: DeepSeek V4 Flash 0731 は AA-Omniscience Index で -16 を記録し、前作から 7 ポイント改善しました。この向上は正確性(正答率)の変化ではなく、ハルシネーション率の低下によってのみ実現されています。AA-Omniscience のハルシネーション率は 84% で、前作より 12 ポイント減少しており、GPT-5.6 Terra (max, 85%) や Mistral Medium 3.5 (82%) と同等の水準です。
➤ DeepSeek V4 Flash 0731 は知能指数の評価項目すべてで前作を上回る: エージェント性能の向上に加え、CritPt は 9 ポイント上昇して 17%、SciCode は 5 ポイント上昇して 50%、Humanity's Last Exam も 5 ポイント上昇して 37% に達しました。また、AA-LCR は 3 ポイント上昇して 66%、GPQA Diamond は 1 ポイント上昇して 91% を記録しています。
DeepSeek V4 Flash 0731 の知能指数スコアは 50 で、前作の DeepSeek V4 Flash よりも 10 ポイント向上しました。
総出力トークン使用量は前世代比で 12% 減少し、DeepSeek V4 Flash 0731 は知能指数の評価に約 2.06 億トーンを使用しましたが、前の DeepSeek V4 Flash では約 2.34 億トンが必要でした。
モデルの詳細:
- コンテキストウィンドウ: 1M トークン(DeepSeek V4 Flash と同等)
- サイズ: 総パラメータ数 284B(アクティブ数は 13B)
- 入力モダリティ: テキスト入出力のみ
- アクセシビリティ: DeepSeek の公式 API を通じて利用可能
- 価格設定: 入力/出力トークン 1M あたり $0.14/$0.28 で、DeepSeek V4 Flash と同じ。キャッシュヒット料金は 1M トークンあたり $0.0028 で、98% の割引となります。

DeepSeek V4 Flash 0731 は、エージェントによる実世界タスクを評価する GDPval-AA v2 で 1559 Elo を記録し、前作の DeepSeek V4 Flash の 1189 から大幅に向上しました。

DeepSeek V4 Flash 0731 は AA-Omniscience Index で -16 を記録し、DeepSeek V4 Flash の -23 よりも 7 ポイント改善しました。これは幻覚率の低下によるものであり、幻覚率は 11 ポイント減少して 84% に達しましたが、精度は 37% で変わっていません。これはモデルのサイズが総パラメータ数 284B のまま変更されていないことと一致しています。

Artificial Analysis Intelligence Index v4.1 における各評価項目の詳細

原文を表示
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, a 10-point jump over DeepSeek V4 Flash (released April 2026) that puts it 6 points ahead of DeepSeek V4 Pro
DeepSeek V4 Flash 0731 is one Intelligence Index point behind GPT-5.6 Luna (max, 51). Even after OpenAI’s 80% price cut on GPT-5.6 Luna today, DeepSeek V4 Flash 0731’s Cost per Task on DeepSeek’s first-party API comes in at ~60% lower than GPT-5.6 Luna (max), a model with comparable intelligence. A key driver of this is DeepSeek’s ~98% cache hit discount on its first-party API, a significantly more aggressive discount than the 90% cache hit discount offered by most of the industry. It shares identical architecture and pricing with the earlier DeepSeek V4 Flash, and lands on our Pareto frontier for Intelligence vs Cost per Task
The new model is a significant step up from the previous generation, DeepSeek V4 Flash (40), and places the model within 1 point of GLM-5.2 (max, 51). It remains 7 points behind the open weights frontier set by Kimi K3 (max, 57). For additional context, this places the model in line with recently released Gemini 3.6 Flash (50) and 1 point behind Muse Spark 1.1 (xhigh, 51). DeepSeek is expected to release the model’s full weights in the coming weeks
DeepSeek V4 Flash 0731 retains a 1M token context window, and its size remains unchanged from DeepSeek V4 Flash at 284B total parameters and 13B active at inference time
Key results:
➤ Improvements in agentic performance: DeepSeek V4 Flash 0731 achieves an Elo rating of 1559 on GDPval-AA v2, our evaluation focused on agentic real-world work tasks, up from 1189 for the previous DeepSeek V4 Flash. Once weights are released this will be the second highest open weights score, behind Kimi K3 (max, 1687) and ahead of GLM-5.2 (max, 1510). Terminal-Bench 2.1 rises 17 points to 79% and τ³-Bench Banking 8 points to 31%
➤ AA-Omniscience improvements are driven by fewer hallucinations, rather than higher accuracy: DeepSeek V4 Flash 0731 achieves an AA-Omniscience Index of -16, a +7 improvement from its predecessor. This improvement is purely driven by a reduced hallucination rate, with overall accuracy (percentage correct) unchanged. Its AA-Omniscience Hallucination Rate is 84%, a 12 point decrease from its predecessor, and comparable to models such as GPT-5.6 Terra (max, 85%) and Mistral Medium 3.5 (82%)
➤ DeepSeek V4 Flash 0731 improves over its predecessor on every evaluation in the Intelligence Index: Alongside the agentic gains, CritPt gains 9 points to 17%, SciCode 5 points to 50%, Humanity's Last Exam 5 points to 37%, AA-LCR 3 points to 66% and GPQA Diamond 1 point to 91%
➤ Total output token usage falls 12% against the predecessor: DeepSeek V4 Flash 0731 used ~206M output tokens to run the Intelligence Index, against ~234M for the previous DeepSeek V4 Flash
Additional model details:
➤ Context window: 1M tokens (equivalent to DeepSeek V4 Flash)
➤ Size: 284B total parameters (13B active)
➤ Input modalities: Text input and output only
➤ Accessibility: Available through DeepSeek’s first-party API
➤ Pricing: $0.14/$0.28 per 1M input/output tokens, unchanged from DeepSeek V4 Flash. Cache hit price of $0.0028 per 1M tokens, a 98% discount

DeepSeek V4 Flash 0731 scores 1559 Elo on GDPval-AA v2, our evaluation for agentic real-world work tasks, up from 1189 for the previous DeepSeek V4 Flash

DeepSeek V4 Flash 0731 scores -16 on the AA-Omniscience Index, a 7 point improvement over DeepSeek V4 Flash (-23), driven entirely by a lower hallucination rate. The hallucination rate falls 11 points to 84% while accuracy is unchanged at 37%, consistent with the model being unchanged in size at 284B total parameters

Breakdown of the individual evaluations in the Artificial Analysis Intelligence Index v4.1

AI算出
主要ニュースainew評価標準
記事は特定のモデル(DeepSeek V4 Flash 0731)の性能向上、コスト構造の変化、および詳細なベンチマークデータ(GDPval-AA, AA-Omniscience など)を報じており、AI モデルの最新動向として極めて関連性が高い。また、バージョン番号と具体的な数値が明記されているため検索機会も最大となる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み