アリババ「Qwen3.8 Max」、Claude Opus 4.8 に並ぶも Kimi K3 が低価格で上回る
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Decoder
アリババのQwen3.8 Maxは人工知能分析指数でClaude Opus 4.8と同等の評価を得たが、コスト効率の悪化や推論精度の低下により、Kimi K3には及ばない結果となった。
AI深層分析を開く2026年8月6日 23:07
AI深層分析
キーポイント
ベンチマークでの位置づけ変化
Qwen3.8 MaxはArtificial Analysis Intelligence Indexで56点を記録し、Claude Opus 4.8と同等となりGLM-5.2を上回ったが、Kimi K3には1点及ばない。
コストパフォーマンスの悪化
単一タスクのコストがQwen3.7 Maxの約2倍となる1.14ドルに上昇し、Kimi K3やGLM-5.2と比較して価格競争力が低下した。
推論プロセスと性能のトレードオフ
タスクあたりのステップ数が64に増加し入力トークンが15倍になった結果、処理は丁寧になったものの速度が遅くなり、コストが増大した。
精度とハルシネーションの悪化
長文情報統合テストや知識回答の正確性で前バージョンより低下し、ハルシネーション率が23%から40%へと急増した。
Kimi K3の価格対性能比
Kimi K3はQwen3.8 MaxやClaude Opus 4.8を上回るスコアを記録しながら、コストは25%低い。
重要な引用
Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max (46).
A single task in the Intelligence Index now costs $1.14, more than double Qwen3.7 Max ($0.53).
The accuracy rate stays around 31 percent, but the hallucination rate jumped from 23 to 40 percent.
編集コメントを表示
編集コメント
Qwen3.8 Maxの発表は、性能向上が必ずしもコスト効率や信頼性の向上に直結しないことを示す事例となっている。開発者はベンチマークスコアだけでなく、実際の運用環境におけるハルシネーション率とコスト構造を重視する必要がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
アリババの「Qwen3.8 Max」は、Artificial Analysis の知能指数(Intelligence Index)で 56 を記録し、前世代の Qwen3.7 Max(46)から 10 ポイント上昇しました。同サイトによると、このスコアは Claude Opus 4.8 と同等であり、GLM-5.2(51)を上回りますが、Kimi K3(57)には及びません。なお、Kimi K3 はコストが約 25% 安いという特徴があります。
業務関連タスクに特化したベンチマーク「GDPval-AA」では、Qwen のスコアは 468 ポイント上昇して 1,739 に達し、Kimi K3(1,685)を抜きました。ただし、このスコアを上回るのは Claude Opus 5(1,852)のみです。
注目すべきは、このスコア獲得に至るプロセスの違いです。Qwen3.8 Max は 1 つのタスク処理に 64 ステップを要し、従来の 14 ステップと比較して大幅に増加しています。また、各ステップで会話履歴全体がモデルへ再送信されるテスト仕様のため、入力トークン数は 15 倍に膨張しました。

このモデルはより徹底的に動作しますが、実行速度は低下し、コストも上昇します。Alibaba の価格対性能比は悪化しました。トークン単価自体は低下しているにもかかわらずです(入力:100 万トークンあたり 2.50 ドルから 2.00 ドルへ、出力:7.50 ドルから 6.00 ドルへ、キャッシュヒット:0.50 ドルから 0.25 ドルへ)。Intelligence Index の単一タスクにかかるコストは現在 1.14 ドルで、Qwen3.7 Max(0.53 ドル)の倍以上です。一方、Kimi K3 はタスクあたりわずか 0.86 ドルで 1 ポイント高いスコアを記録し、GLM-5.2 は 0.57 ドルです。
前バージョンと比較すると性能が後退した項目もあります。AA-LCR は 2 ポイント低下しました。これは非常に長いテキストから情報を正しく統合できるかをテストする項目です。また、AA-Omniscience(知識への回答や「知らない」と正直に答える能力を測定)は 10 ポイント下落しました。正確率は約 31% で横ばいですが、ハルシネーション(幻覚・誤答)の発生率が 23% から 40% に急増しています。Qwen3.8 Max は「知らない」と答える代わりに、より頻繁に推測して回答する傾向が強まっています。
AI News Without the Hype – Curated by Humans
広告なしで読める THE DECODER の購読、週刊 AI ニュースレター、年 6 回の独占「AI Radar」フロンティアレポート、アーカイブへの完全アクセス、コメント欄の利用権などをご提供しています。
原文を表示
Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max (46). According to Artificial Analysis, that puts it on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but behind Kimi K3 (57), which also runs 25 percent cheaper.
On GDPval-AA, a benchmark for work-related tasks, Qwen jumps 468 Elo points to 1,739, passing Kimi K3 (1,685). Only Claude Opus 5 (1,852) scores higher. The catch is how it gets there. Qwen3.8 Max needs 64 steps per task instead of 14, and input tokens grew 15x because the test resends the full conversation history to the model at each step.

The model works more thoroughly but runs slower and costs more. Alibaba's price-to-performance ratio takes a hit despite lower token prices (input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25). A single task in the Intelligence Index now costs $1.14, more than double Qwen3.7 Max ($0.53). Kimi K3 scores one point higher at just $0.86 per task, and GLM-5.2 comes in at $0.57.
There are also regressions compared to the previous version. AA-LCR dropped 2 points, a test that checks whether a model can correctly pull together information from very long texts. AA-Omniscience fell 10 points, measuring whether a model answers knowledge questions correctly or honestly admits it doesn't know. The accuracy rate stays around 31 percent, but the hallucination rate jumped from 23 to 40 percent. Qwen3.8 Max guesses far more often instead of saying it doesn't know.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み