アリババ、Qwen3.8-Maxを発表しGPT-5.6 Sol Maxなどを上回ると主張
本文の状態
日本語全文を表示中
詳細モードで約15分の本文を読めます。
同じ出来事の情報源
6媒体で確認
TLDR AI · Qwen Blog · Alibaba Engineering · MarkTechPost · The Decoder · VentureBeat AI
各社の報じ方を比較 ↓アリババのQwenチームが、自律的なソフトウェア開発や長期間の企業業務に特化した新モデル「Qwen3.8-Max」を発表し、同社によると主要ベンチマークで競合他社を上回る性能を示した。
AI深層分析を開く2026年8月4日 08:56
AI深層分析
キーポイント
新モデルの発表と性能主張
アリババは2.4兆パラメータのMoE構造を持つマルチモーダルLLM「Qwen3.8-Max」を発表し、OSWorld-Verifiedベンチマークで86.1点を記録して競合を上回ったと発表した。
自律型エージェントとしての能力
同社によると、このモデルは数日単位のプロジェクト実行や数千行のコードを含む研究再現を自律的に完了できる「自律的な同僚」として設計されている。
オープンウェイト公開の可能性
アリババは来週にQwen3.8-Maxおよび27Bモデルのオープンウェイト版を公開すると発表したが、ライセンス条件はまだ明言されていないため、企業による自己ホストの実現性は不透明である。
業界における戦略的ポジション
汎用性やコーディング特化から、長期間にわたる自動化とエンタープライズワークフローに焦点を当てたことで、アリババは市場の競合環境において独自の立ち位置を確立しようとしている。
自律的な実行能力を重視するベンチマークへの移行
従来の推論やコーディング問題だけでなく、OSWorld-Verifiedなどの評価で長期的なワークフローの完了能力が測られるようになっている。
重要な引用
Qwen3.8-Max scores 86.1 on the OSWorld-Verified benchmark measuring how well ahead of GPT-5.6 Sol Max (83.2) and Fable 5 (85.0)
Alibaba is positioning the model as an autonomous coworker capable of executing projects that span days rather than minutes
The release also signals a potentially significant strategic shift for Alibaba: the company says open weights for Qwen3.8-Max will be released next week
On OSWorld-Verified, which evaluates computer-use agents interacting with desktop environments, Qwen3.8-Max posts 86.1, ahead of GPT-5.6 Sol Max's 83.2, Fable 5's 85.0, and Gemini 3.1 Pro's 76.2.
編集コメントを表示
編集コメント
アリババが「自律的な同僚」という概念を前面に出したモデルリリースは、エンタープライズ市場におけるAIの役割定義を再構築する可能性がある。今後のライセンス条件と独立検証の結果次第で、業界全体のオープンウェイト戦略に大きな影響を与えるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
中国のECおよびクラウド大手アリババが、昨晩、AI研究チーム「Qwen」から新フラッグシップモデル「Qwen3.8-Max」を発表しました。これは2.4兆パラメータを備えた混合専門家(MoE)方式のマルチモーダル大規模言語モデル(LLM)で、自律的なソフトウェアエンジニアリングや長期にわたるエンタープライズ業務といった、最先端AI市場の中でも特に過激な競争領域を狙っています。
同社が公開したベンチマーク結果が、より広範な独立系のテストでも裏付けられるのであれば、Qwen3.8-Maxは現在の主要なプロプライエタリモデルと単に競合するだけでなく、エージェントコンピューティングにおけるいくつかの重要指標でそれらを上回る可能性があります。
特筆すべきは、OSWorld-Verifiedベンチマークでのスコアです。これは自律的なソフトウェア操作能力を測定するもので、Qwen3.8-Maxは86.1点を記録し、GPT-5.6 Sol Max(83.2点)やFable 5(85.0点)を抜いています。また、PaperBenchでも最高スコアを達成しており、ソフトウェアエンジニアリング、研究の再現性、マルチモーダル推論、視覚的なウェブ開発などの分野でも首位を維持するか、極めて高い競争力を示しています。
今回の発表は、アリババにとって戦略的な転換点となる可能性も示唆しています。同社は、Qwen3.8-MaxとQwen3.8-27Bのオープンウェイト版を来週公開する予定だと述べています。
もしこれが寛容なライセンスの下で実現されれば、MaxクラスのQwenモデルがセルフホスト環境でのデプロイに利用可能になるのは初めてとなります。これは企業の導入状況を大きく変える動きになり得ます。
ただし、重要な注意点として、アリババはまだライセンス条件を公開していません。そのため、最近の中国の競合他社である Moonshot がオープンな Kimi K3 フロントティアモデルで採用したように、Apache 2.0 のような広く許可されたライセンスではなく、より制限の厳しい独自ライセンスが適用される可能性も残されています。
「フロンティア」の異なる定義
過去1年間、基盤モデルの競争環境はますます専門化してきました。
OpenAI は GPT シリーズを主に一般推論、マルチモーダルな対話、そして企業向け生産性の向上に注力しています。
Anthropic の Claude シリーズは、コーディングと信頼性の高い長文脈推論を強調しています。Google は Gemini をマルチモーダルな生産性と Web ネイティブのワークフローへと押し進めています。
Moonshot AI の Kimi K3 は、フロンティアクラスの性能とオープンウェイトでの公開を組み合わせて、この議論に参入しました。
Qwen3.8-Max は、これらの多くの強みを単一のモデルに統合し、企業自動化を明確なターゲットとしたアプローチを試みています。
対話型知能を強調するのではなく、アリババはこのモデルを数分ではなく数日にわたるプロジェクトを実行できる自律的な同僚として位置づけています。
同社によると、Qwen3.8-Max は10日以上続くソフトウェアプロジェクトの完了、数千行に及ぶコードを含む研究論文の再現、反復的なチップ設計の最適化、そしてマルチモーダルフィードバックループを用いた計画の継続的な見直しを自律的に行うことができます。
これらのデモは企業が生産したものであり、まだ独立した評価者によって広く再現されていません。しかし、これらは業界の明確なトレンドを示しています。最先端モデルが競うのはもはや個別のプロンプトへの回答ではなく、一連のワークフローを完遂する能力です。
ベンチマークは自律的な実行をより重視する方向へシフト
Qwen3.8-Max と同時に発表されたベンチマークスイートは、この変化を反映しています。
従来の推論試験やコーディングパズルに焦点を当てるだけでなく、注目すべき評価の多くが長期にわたる実行能力を測定しています。
デスクトップ環境と対話するコンピューター使用エージェントを評価する「OSWorld-Verified」では、Qwen3.8-Max は 86.1 のスコアを記録し、GPT-5.6 Sol Max(83.2)、Fable 5(85.0)、Gemini 3.1 Pro(76.2)を上回りました。
同モデルは以下の項目でも首位に立ちます:
PaperBench: 93.0
TerminalBench 2.1: 86.6
Vision2Web: 69.0
LVBench: 81.8
ERQA: 77.8
一方で、一部の分野では依然として競合他社と拮抗しており、いくつかのカテゴリーではやや劣る結果となっています。
例えば、専門的なソフトウェアエンジニアリング向けのベンチマーク「SWE-Pro」では、OpenAI のモデルが最高スコアを記録しています。また、Opus 4.8 は特定のソフトウェアエンジニアリング評価や「Agents' Last Exam」において引き続きリードしています。
あらゆるベンチマークで圧倒するわけではありませんが、Qwen は現在利用可能な最もバランスの取れたパフォーマンスプロファイルの一つを提供していると言えます。このバランスこそが、孤立したベンチマークでの勝利以上に、企業購入層にとって重要になる可能性があります。
多くの組織では、モデルの評価基準が特定の狭い機能の最適化から、コード作成やドキュメント閲覧、インターフェース操作、レポート生成、画像解析、複数の下位タスクの調整など、多様なワークフローをいかに確実に完了できるかという点へとシフトしています。
Qwen3.8-Max の強み
アリババが発表している結果が生産環境でも通用すると仮定すれば、Qwen3.8-Max が特に適したエンタープライズ・ワークロードはいくつか存在します。
- 長時間稼働するソフトウェアエンジニアリング
アリババの主なデモは、10 日以上続く自律的なソフトウェア開発です。
これらのデモは独立して再現されるまでベンダーの主張として扱うべきですが、対話型ではなく継続的に動作する永続型のコーディング・エージェントへの関心が高まっている現状と合致しています。
自律的なエンジニアリングチーム、CI/CD 自動化、リポジトリ管理、回帰テスト、機能実装などを試行している組織にとって、そのエージェントとしての性能が実験室環境以外でも一貫して証明されれば、Qwen は非常に魅力的な選択肢となるでしょう。
- コンピュータ操作型エージェント
最大の差別化要因は「コンピュータ操作」にあるかもしれません。
OSWorld は、モデルがテキストを生成するだけでなくオペレーティングシステムと対話できる能力を測定するため、業界で最も注目されるベンチマークの一つに急速に成長しています。
デスクトップソフトウェアを確実に操作できるモデルは、API が存在しない文書処理、企業システム統合、内部業務、レガシーワークフローなど、無数の反復的なビジネスプロセスの自動化を実現できます。
ベンチマークでのパフォーマンスが生産環境でも一般化されれば、OSWorld のような成果は実際の運用上の優位性へと直結する可能性があります。
- 研究自動化
Qwen が PaperBench で首位を維持していることは、科学計算や文献レビュー、実験の再現、技術分析を行う組織にとって大きな可能性を示しています。
研究機関、製薬企業、産業用 R&D チームでは、LLM を単なる要約ツールとしてだけでなく、再現可能な計算ワークフローを実行する手段としても活用するケースが増えています。こうした環境では、長時間にわたるセッション全体で文脈を維持できるモデルの価値がさらに高まっています。
- 多モーダルな産業ワークフロー
アップロードされた画像を主に分析する従来の多モーダルシステムとは異なり、Qwen は視覚情報を計画と実行に統合された継続的なフィードバックメカニズムとして位置づけています。
このアーキテクチャは、製造業、物流、エンジニアリング検査、設計レビューなどにおいて特に有用となるでしょう。これらの分野では、視覚入力が単なる孤立したプロンプトではなく、運用上の意思決定を絶えず支える情報源となります。
経済性もまた極めて重要な要素となり得ます
最大の競争圧力はベンチマークスコアではなく、中国に拠点を置く QwenCloud 上の Qwen API を通じた価格設定から生じています。
Qwen3.8-Max は、入力・出力ともに百万トークンあたり 2 ドル/6 ドルで登場しました。これは中価格帯のモデルですが、比較対象となる上位の米国製プロプライエタリ製品を意味ある割合で下回る価格設定です。具体的には、Claude Opus 5 の入出力合計料金の 3 分の 1 未満、GPT-5.6 Sol Max の料金の 4 分の 1 未満となっています。
- モデル:入力 ($/1M) / 出力 ($/1M) / 合計 ($/1M) / ソース
- MiMo-V2.5 Flash:$0.10 / $0.30 / $0.40 / Xiaomi
- deepseek-v4-flash:$0.14 / $0.28 / $0.42 / DeepSeek
- deepseek-v4-pro:$0.435 / $0.87 / $1.305 / DeepSeek
- GPT-5.6 Luna:$0.20 / $1.20 / $1.40 / OpenAI
- MiniMax-M3:$0.30 / $1.20 / $1.50 / MiniMax
- LongCat-2.0 — 期間限定プロモ:$0.30 / $1.20 / $1.50 / LongCat
- Gemini 3.1 Flash-Lite:$0.25 / $1.50 / $1.75 / Google
- Qwen3.7-Plus:$0.40 / $1.60 / $2.00 / Alibaba Cloud
- MiMo-V2.5:$0.40 / $2.00 / $2.40 / Xiaomi
- Gemini 3.5 Flash-Lite:$0.30 / $2.50 / $2.80 / Google
- LongCat-2.0 — 通常価格:$0.75 / $2.95 / $3.70 / LongCat
- MiMo-V2.5 Pro (≤256K):$1.00 / $3.00 / $4.00 / Xiaomi
- GLM-5.2:$1.40 / $4.40 / $5.80 / Z.ai
- Grok 4.5:$2.00 / $6.00 / $8.00 / xAI
- MiMo-V2.5 Pro (>256K):$2.00 / $6.00 / $8.00 / Xiaomi
- Qwen3.8-Max:$2.00 / $6.00 / $8.00 / QwenCloud
- Gemini 3.6 Flash:$1.50 / $7.50 / $9.00 / Google
- Qwen3.7-Max:$2.50 / $7.50 / $10.00 / Alibaba Cloud
- Gemini 3.5 Flash:$1.50 / $9.00 / $10.50 / Google
- Gemini 3.1 Pro Preview (≤200K):$2.00 / $12.00 / $14.00
GPT-5.6 Terra
2 ドル
12 ドル
14 ドル
OpenAI
GPT-5.4
2.50 ドル
15 ドル
17.50 ドル
OpenAI
Kimi K3
3 ドル
15 ドル
18 ドル
Moonshot AI
Gemini 3.1 Pro Preview (>200K)
4 ドル
18 ドル
22 ドル
Claude Opus 5
5 ドル
25 ドル
30 ドル
Anthropic
GPT-5.5
5 ドル
30 ドル
35 ドル
OpenAI
GPT-5.5 Instant (chat-latest)
5 ドル
30 ドル
35 ドル
OpenAI
Sakana Fugu Ultra (≤272K)
5 ドル
30 ドル
35 ドル
Sakana AI
GPT-5.6 Sol — Standard mode
5 ドル
30 ドル
35 ドル
OpenAI
Claude Fable 5 / Claude Mythos 5
10 ドル
50 ドル
60 ドル
Anthropic
GPT-5.6 Sol — Fast mode
10 ドル
60 ドル
70 ドル
OpenAI
エージェントシステムは従来のチャットボットと比較して桁違いに多くのトークンを消費するため、推論コストの低減がますます重要になっています。この現実を踏まえ、OpenAI は先週後半、ミドルおよびローエンドに位置する GPT-5.6 ラインナップ(Terra および Luna)の API 価格をそれぞれ 20% と 80% 引き下げると決定しました。
実際にこれらのシステムを運用している人々が証言するように、数時間にわたる自律的なワークフローや反復的な計画策定、継続的な自己修正を行うと、単一のタスクで数百万トークンが発生することさえあります。
数百あるいは数千のエージェントを同時に展開する企業にとって、推論コストは最大の運営費の一つになり得ます。そのため、1 トークンあたりの価格がわずかに下がっただけでも、累積効果によって大きな節約につながります。
米国の最先端モデルとの比較
ヘッドラインのベンチマーク比較結果だけで、Qwen3.8-Max が米国の主要モデルを全面的に置き換える存在だと考える必要はありません。
むしろ、このモデルの強みは異なる導入戦略を示唆しています。
OpenAI の GPT ファミリーは、成熟したツール類やエコシステムとの統合、広範な商用展開を背景に、汎用的なエンタープライズ推論プラットフォームとして引き続き高い評価を得ています。すでに Microsoft エコシステムや OpenAI の企業向けサービスに投資している組織にとっては、特定のエージェントベンチマークで Qwen が上回っていても、これらの運用上の優位性を重視し続ける可能性があります。
Anthropic の Claude Opus は、慎重なソフトウェアエンジニアリングや長文脈の推論において、依然として最強のコーディングアシスタントの一つと見なされています。信頼性や予測可能な挙動が、生来の自律性よりも優先される「人間をループに組み込んだ」開発プロセスにおいては、一部の企業は引き続き Claude を選ぶでしょう。
Google Gemini は、Workspace との深い統合、マルチモーダル機能、そして Google Cloud サービスを通じて独自性を保ち続けており、すでに Google のエンタープライズスタックを採用している組織にとって魅力的な選択肢です。
Qwen が最も注目されるのは、最先端レベルのパフォーマンスを維持しつつ、自律的な実行、長期にわたる計画策定、そして有利な推論コストを最優先する企業においてです。
オープンウェイト版の是非は未だ結論が出ていません
Qwen3.8-Max に関する最大の不透明さは、ベンチマークの結果とはほとんど関係ありません。
アリババは、来週にオープンウェイトの公開を予定している。しかし、発表内容や提供されたドキュメントには、そのウェイトを管理するライセンスの詳細は明記されていない。
この違いが極めて重要になる可能性がある。
Apache 2.0 のような寛容なライセンスであれば、組織がモデルをセルフホストしたり、ファインチューニングを行ったり、制限の少ない条件で独自製品に統合したりできるため、企業の採用範囲が大幅に広がるだろう。
一方、いくつかの最近のフロンティアモデルで採用されているようなカスタムライセンスの場合、商業的な展開や再配布、利用分野、あるいはモデルの改変に対して制限を課す可能性がある。こうした制約は、モデルの技術的性能に関わらず、長期的なインフラ投資を求める企業にとって魅力を削ぐことになる。
Moonshot AI が最近発表した Kimi K3 の事例が、なぜこの違いが重要なのかを示している。Kimi K3 はウェイトを誰でも利用可能にしたものの、「Model as a Service」として提供する場合には開示と商用ライセンスの取得を義務付けるなど、特定のライセンス条項を含んでいた。
アリババが Qwen3.8-Max のライセンスを公表するまで、セルフホストを検討している組織は、オープンウェイトの発表を有望ではあるが不十分なものとして捉えるべきだ。
競争が激化するフロンティア領域
Qwen3.8-Max は、ファウンデーションモデルの歴史の中で最も急速に変化している時期の一つに登場した。
数週間以内に、Moonshot AI、OpenAI、Anthropic などの主要開発者から相次ぐ大規模リリースが発表され、それぞれが推論能力、コーディング支援、マルチモーダル処理、自律型エージェント、コスト効率など異なる強みを前面に押し出しています。
アリババの貢献は特筆すべきものです。競合他社と対等に戦えるベンチマーク性能に加え、圧倒的な低価格設定、100 万トークンという広大なコンテキストウィンドウ、そしてフラッグシップモデルの重み(ウェイト)公開への明確なコミットメントをすべて兼ね備えているからです。
このモデルが企業の自律型エージェント向けプラットフォームとして選ばれるかどうかは、単なるリーダーボード上の順位よりも、独立した第三者による検証結果、実運用での信頼性、そして今後公開される重みのライセンス条件といった要素に大きく依存します。
ベンチマークの数字だけではないこれらの要因こそが、Qwen3.8-Max が米国の大手プロプライエタリモデルに対する真の代替案となるのか、それとも過熱する最先端 AI 競争における単なる目立つ新規参入者に留まるのかを決定づける鍵となります。
原文を表示
Chinese e-commerce and cloud giant Alibaba's famed Qwen team of AI researchers last night unveiled Qwen3.8-Max, a new flagship 2.4-trillion-parameter mixture-of-experts (MoE) multimodal large language model (LLM) that targets one of the most competitive corners of the frontier AI market: autonomous software engineering and long-horizon enterprise work.
If the company's published benchmarks hold up under broader independent testing, Qwen3.8-Max doesn't merely compete with today's leading proprietary models — it surpasses several of them on some key benchmarks in agentic computing.
Most notably, Qwen reports that Qwen3.8-Max scores 86.1 on the OSWorld-Verified benchmark measuring how well ahead of GPT-5.6 Sol Max (83.2) and Fable 5 (85.0), while also posting the highest reported score on PaperBench and leading or remaining highly competitive across software engineering, research reproduction, multimodal reasoning, and visual web development benchmarks.
The release also signals a potentially significant strategic shift for Alibaba: the company says open weights for Qwen3.8-Max will be released next week, alongside Qwen3.8-27B.
If that happens under a permissive license, it would represent the first time a Max-class Qwen model becomes available for self-hosted deployment—a move that could substantially reshape enterprise adoption.
One important caveat remains, however: Alibaba has not yet disclosed the licensing terms, leaving open the possibility that the release could use a more restrictive custom license, as we saw recently with Chinese rival Moonshot's open Kimi K3 frontier model, rather than a broadly permissive one such as Apache 2.0.
A different definition of 'frontier'
Over the past year, the competitive landscape for foundation models has become increasingly specialized.
OpenAI has largely focused its GPT series on general reasoning, multimodal interaction and enterprise productivity.
Anthropic's Claude series has emphasized coding and dependable long-context reasoning. Google continues to push Gemini toward multimodal productivity and web-native workflows.
Moonshot AI's Kimi K3 recently entered the conversation by pairing frontier-class performance with an open-weight release.
Qwen3.8-Max attempts to combine many of these strengths into a single model aimed squarely at enterprise automation.
Rather than emphasizing conversational intelligence, Alibaba is positioning the model as an autonomous coworker capable of executing projects that span days rather than minutes.
According to the company, Qwen3.8-Max can autonomously complete software projects lasting more than 10 days, reproduce research papers involving thousands of lines of code, perform iterative chip-design optimization, and continuously revise plans using multimodal feedback loops.
Those demonstrations remain company-produced and have not yet been broadly replicated by independent evaluators. Nevertheless, they illustrate a growing industry trend: frontier models are increasingly competing on their ability to finish entire workflows rather than answer individual prompts.
Benchmarks increasingly reward autonomous execution
The benchmark suite released alongside Qwen3.8-Max reflects this shift.
Instead of focusing solely on traditional reasoning exams or coding puzzles, many of the highlighted evaluations measure long-horizon execution.
On OSWorld-Verified, which evaluates computer-use agents interacting with desktop environments, Qwen3.8-Max posts 86.1, ahead of GPT-5.6 Sol Max's 83.2, Fable 5's 85.0, and Gemini 3.1 Pro's 76.2.
The model also leads:
PaperBench: 93.0
TerminalBench 2.1: 86.6
Vision2Web: 69.0
LVBench: 81.8
ERQA: 77.8
Elsewhere, it remains competitive with proprietary leaders while trailing in several categories.
On the professional software engineering benchmark SWE-Pro, for example, OpenAI's model posts the highest reported score, while Opus 4.8 continues to lead on certain software engineering evaluations and Agents' Last Exam.
Rather than dominating every benchmark, Qwen appears to offer one of the broadest balanced performance profiles currently available.
That balance may ultimately matter more for enterprise buyers than isolated benchmark wins.
Many organizations increasingly evaluate models based on how reliably they complete heterogeneous workflows—writing code, reading documents, navigating interfaces, generating reports, inspecting images and coordinating multiple subtasks—rather than optimizing for one narrow capability.
Where Qwen3.8-Max appears strongest
Assuming Alibaba's published results translate into production deployments, several enterprise workloads stand out as particularly well suited for Qwen3.8-Max.
- Long-running software engineering
Alibaba's primary demonstration involves autonomous software development extending beyond ten days.
While enterprises should treat these demonstrations as vendor claims until independently reproduced, they align with a growing interest in persistent coding agents that operate continuously rather than interactively.
Organizations experimenting with autonomous engineering teams, CI/CD automation, repository maintenance, regression testing or feature implementation may find Qwen particularly attractive if its agentic performance proves consistent outside laboratory settings.
- Computer-use agents
The strongest differentiator may be computer use.
OSWorld has rapidly become one of the industry's most closely watched benchmarks because it measures a model's ability to interact with operating systems instead of simply generating text.
Models capable of reliably navigating desktop software can automate countless repetitive business processes, including document processing, enterprise software integration, internal operations and legacy workflows where APIs may not exist.
Leading OSWorld could therefore translate into real operational advantages if benchmark performance generalizes to production environments.
- Research automation
Qwen's PaperBench leadership suggests strong potential for organizations performing scientific computing, literature review, experiment reproduction and technical analysis.
Research institutions, pharmaceutical companies and industrial R&D teams increasingly use LLMs not only for summarization but also for executing reproducible computational workflows. Models capable of maintaining context across extended sessions become increasingly valuable in these environments.
- Multimodal industrial workflows
Unlike earlier multimodal systems that primarily analyze uploaded images, Qwen describes vision as an ongoing feedback mechanism integrated into planning and execution.
That architecture could prove particularly useful in manufacturing, logistics, engineering inspection and design review, where visual inputs continuously inform operational decisions rather than serving as isolated prompts.
The economics may prove just as important
Perhaps the biggest competitive pressure comes not from benchmark scores but from pricing through Qwen's application programming interface (API) on QwenCloud (based in China):
Qwen3.8-Max launches at $2/$6 per million input/output tokens, a mid-priced model but undercutting the top U.S. proprietary offerings to which it is benchmarked against by meaningful percentages, less than 1/3 the combined in/out price of Claude Opus 5 and less than 1/4 the price of GPT-5.6 Sol Max.
Model
Input ($/1M)
Output ($/1M)
Total ($/1M)
Source
MiMo-V2.5 Flash
$0.10
$0.30
$0.40
Xiaomi
deepseek-v4-flash
$0.14
$0.28
$0.42
DeepSeek
deepseek-v4-pro
$0.435
$0.87
$1.305
DeepSeek
GPT-5.6 Luna
$0.20
$1.20
$1.40
OpenAI
MiniMax-M3
$0.30
$1.20
$1.50
MiniMax
LongCat-2.0 — limited-time promo
$0.30
$1.20
$1.50
LongCat
Gemini 3.1 Flash-Lite
$0.25
$1.50
$1.75
Qwen3.7-Plus
$0.40
$1.60
$2.00
Alibaba Cloud
MiMo-V2.5
$0.40
$2.00
$2.40
Xiaomi
Gemini 3.5 Flash-Lite
$0.30
$2.50
$2.80
LongCat-2.0 — standard
$0.75
$2.95
$3.70
LongCat
MiMo-V2.5 Pro (≤256K)
$1.00
$3.00
$4.00
Xiaomi
GLM-5.2
$1.40
$4.40
$5.80
Z.ai
Grok 4.5
$2.00
$6.00
$8.00
xAI
MiMo-V2.5 Pro (>256K)
$2.00
$6.00
$8.00
Xiaomi
Qwen3.8-Max
$2.00
$6.00
$8.00
QwenCloud
Gemini 3.6 Flash
$1.50
$7.50
$9.00
Qwen3.7-Max
$2.50
$7.50
$10.00
Alibaba Cloud
Gemini 3.5 Flash
$1.50
$9.00
$10.50
Gemini 3.1 Pro Preview (≤200K)
$2.00
$12.00
$14.00
GPT-5.6 Terra
$2.00
$12.00
$14.00
OpenAI
GPT-5.4
$2.50
$15.00
$17.50
OpenAI
Kimi K3
$3.00
$15.00
$18.00
Moonshot AI
Gemini 3.1 Pro Preview (>200K)
$4.00
$18.00
$22.00
Claude Opus 5
$5.00
$25.00
$30.00
Anthropic
GPT-5.5
$5.00
$30.00
$35.00
OpenAI
GPT-5.5 Instant (chat-latest)
$5.00
$30.00
$35.00
OpenAI
Sakana Fugu Ultra (≤272K)
$5.00
$30.00
$35.00
Sakana AI
GPT-5.6 Sol — Standard mode
$5.00
$30.00
$35.00
OpenAI
Claude Fable 5 / Claude Mythos 5
$10.00
$50.00
$60.00
Anthropic
GPT-5.6 Sol — Fast mode
$10.00
$60.00
$70.00
OpenAI
Lower inference costs increasingly matter because agentic systems consume dramatically more tokens than conventional chatbots — a reality that likely factored into OpenAI's decision late last week to cut the API prices of its mid- and lower-end GPT-5.6 lineup of models (Terra and Luna) by 20% and 80%, respectively.
Indeed, as those running these systems can attest, multi-hour autonomous workflows, iterative planning and continuous self-correction can generate millions of tokens during a single task.
For enterprises deploying hundreds or thousands of agents simultaneously, inference costs often become one of the largest operational expenses. Small reductions in per-token pricing therefore compound rapidly.
How it compares with American frontier models
Despite headline benchmark comparisons, Qwen3.8-Max should not necessarily be viewed as a wholesale replacement for leading American models.
Instead, its strengths suggest different deployment strategies.
OpenAI's GPT family continues to excel as a broadly capable enterprise reasoning platform with mature tooling, ecosystem integration and extensive commercial deployment. Organizations already invested in Microsoft ecosystems or OpenAI's enterprise offerings may continue to value those operational advantages even if Qwen leads on selected agent benchmarks.
Anthropic's Claude Opus remains widely regarded as one of the strongest coding assistants, particularly for careful software engineering and long-context reasoning. Some enterprises may still prefer Claude for human-in-the-loop development where reliability and predictable behavior outweigh raw autonomy.
Google Gemini continues to differentiate itself through deep Workspace integration, multimodal capabilities and Google Cloud services, making it attractive for organizations already standardized on Google's enterprise stack.
Where Qwen appears most compelling is for enterprises prioritizing autonomous execution, extended planning horizons and favorable inference economics without sacrificing frontier-level performance.
The open-weight question remains unanswered
The largest unknown surrounding Qwen3.8-Max has little to do with benchmarks.
Alibaba says open weights are coming next week. However, neither the announcement nor the provided documentation specifies the license that will govern those weights.
That distinction could prove critical.
A permissive license such as Apache 2.0 would significantly broaden enterprise adoption by allowing organizations to self-host, fine-tune and integrate the model into proprietary products with relatively few restrictions.
A custom license—similar to approaches used by several recent frontier releases—could impose limitations on commercial deployment, redistribution, field of use or model modification. Such restrictions would narrow the appeal for enterprises seeking long-term infrastructure investments, regardless of the model's technical performance.
Moonshot AI's recent Kimi K3 release illustrates why this distinction matters. While Kimi K3 made its weights openly available to all, its licensing terms included specific terms including a disclosure and a commercial license requirement for those offering it as a "Model as a Service."
Until Alibaba publishes Qwen3.8-Max's license, organizations considering self-hosting should treat the open-weight announcement as promising but incomplete.
An increasingly crowded frontier
Qwen3.8-Max arrives during one of the fastest-moving periods in the history of foundation models.
Within weeks, developers have seen major releases from Moonshot AI, OpenAI, Anthropic and others, each emphasizing different strengths: reasoning, coding, multimodality, autonomous agents or economics.
Alibaba's contribution is notable because it combines competitive benchmark performance, aggressive pricing, a million-token context window and a stated commitment to releasing weights for its flagship model.
Whether it becomes the preferred platform for enterprise autonomous agents will ultimately depend less on leaderboard positions than on broader independent validation, production reliability and the licensing terms accompanying the forthcoming weight release.
Those factors—not benchmark charts alone—will determine whether Qwen3.8-Max becomes a genuine alternative to the leading American proprietary models or simply another impressive entrant in an increasingly crowded frontier AI race.
AI算出
主要ニュースainew評価高い
記事は Qwen3.8-Max の詳細な仕様(2.4 兆パラメータ、MoE)と、OSWorld-Verified や PaperBench などの最新ベンチマークにおける具体的な数値スコアを提供しており、単なる再報ではなく独立した情報増分がある。ただし、日本企業への直接的な影響や日本語一次情報の欠如により、日本の関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 75
- 新規性
- 75
- 調べる価値
- 100
- 重複の少なさ
- 43
- 日本での有用性
- 25
同じ出来事を6媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み