アリババ、最上位モデル「Qwen 3.8 Max」を発表しオープンウェイト公開を予告
本文の状態
原文本文あり
日本語全文が未生成の場合も、詳細モードで原文を確認できます。
同じ出来事の情報源
6媒体で確認
Alibaba Engineering · Smol AI News · MarkTechPost · The Decoder · VentureBeat AI · Latent Space
各社の報じ方を比較 ↓アリババは新フラッグシップモデル「Qwen3.8-Max」を発表し、2.4兆パラメータを備えコーディングや長期間の自律作業に特化すると同時に、来週のオープンウェイト公開とAPI価格を明示した。
AI深層分析を開く2026年8月5日 11:29
AI深層分析
キーポイント
新モデル「Qwen3.8-Max」の発表と仕様
アリババは「史上最高性能」とする2.4兆パラメータのフラッグシップモデルを発表し、コーディングや長期間の自律作業、マルチモーダル推論に特化していると主張した。
オープンウェイト公開と価格設定
同社は来週に「Qwen3.8-Max」および「Qwen3.8-27B」のオープンウェイトを公開すると発表し、API価格は入力100万トークンあたり2ドル、出力6ドル、キャッシュ済みトークン0.25ドルと定めた。
具体的な機能強化とエコシステム展開
同社は10日以上続く自律コーディングや365日続くeコマース戦略などの特徴を強調し、Qwen StudioやAPIに加え、BasetenやHermes Agentなどのパートナー企業によるサポート体制も確認された。
中国のオープンウェイトモデルが西側のクローズドモデルと直接競合
複数の観察者により、コーディングやアジェンシーワークフロー、マルチモーダルタスクにおいて中国のオープンウェイトフロンティアがトップクラスの西側クローズドモデルと直接競争している証拠と見なされている。
公式発表された極めて攻撃的な性能主張
ベンダー報告によると、総パラメータ数は2.4T、1Mトークンのコンテキストウィンドウ、および10日以上続く自律的なコーディング実行などの驚異的な数値が提示されている。
重要な引用
"most capable model to date"
"open weights will be released next week"
"Chinese open-weight frontier is now competing directly with top Western closed models"
Chinese open-weight frontier is now competing directly with top Western closed models
編集コメントを表示
編集コメント
アリババが2.4兆パラメータという巨大規模モデルをオープンウェイトとして公開する方針を示したことは、業界の競争構造に大きな影響を与える可能性がある。特に長期間の自律作業やマルチモーダル処理における性能向上は、実務現場でのAI活用範囲を広げる可能性を秘めている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
a quiet day.
AI News for 8/3/2026-8/1/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews' website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
Top Story: Qwen 3.8 Max open model launch
What happened
Alibaba Qwen announced Qwen3.8-Max as its new flagship and said open weights are coming next week.
- Alibaba introduced Qwen3.8-Max as its “most capable model to date,” describing it as a 2.4T-parameter model focused on coding, long-horizon agentic work, and multimodal reasoning, with the explicit claim that open weights will be released next week, alongside Qwen3.8-27B also going open-weight @Alibaba_Qwen
- The launch tweet also included API pricing: $2.00 / M input tokens, $6.00 / M output tokens, and $0.25 / M cached tokens @Alibaba_Qwen
- Alibaba framed the model around several headline capabilities: 10+ days of autonomous coding, 500+ turns of chip design optimization, 365 days of e-commerce strategy, and native multimodal intelligence where vision is part of the execution loop rather than just an input channel @Alibaba_Qwen
- The company simultaneously pushed availability across its own surfaces and partners: Qwen Studio, API, Command Code, and later Venice; infra and app builders quickly confirmed support plans or integrations including Baseten, Hermes Agent, and Command Code @Alibaba_Qwen @Alibaba_Qwen @baseten @Teknium
- The announcement landed as part of a broader pattern: multiple observers described it as evidence that the Chinese open-weight frontier is now competing directly with top Western closed models, especially in coding, agentic workflows, and multimodal tasks @kimmonismus @matvelloso
Official claims and reported specs
Vendor-reported model details and performance claims were unusually aggressive for an open-weight release.
- Alibaba’s own framing:
2.4T total parameters @Alibaba_Qwen
- Long-horizon agentic/cowork focus @Alibaba_Qwen
- Autonomous coding over 10+ days with a public GitHub trace @Alibaba_Qwen
- 500+ turns for chip design optimization @Alibaba_Qwen
- 365 days of e-commerce strategy execution @Alibaba_Qwen
- Native multimodal planning loop rather than vision-only input @Alibaba_Qwen
- Third-party summary tweet from ZhihuFrontier added more claimed or reported technical details:
95B active parameters per token, implying an MoE activation ratio of roughly 4%
- 1M-token context window
- API exposes low / medium / xhigh reasoning-effort modes
- Compatibility with OpenAI and Anthropic protocols
- Benchmark claims: PaperBench 93.0, CoWorkBench 74.8, WideSearch 81.9 @ZhihuFrontier
- Vals AI independently posted concrete eval/runtime settings:
1M token context
- 128k max output tokens
- Tested at temperature 0.7 with default top-p / top-k @ValsAI
These numbers matter because they place Qwen3.8-Max in the same deployment class as other giant sparse open models like Kimi K3 and GLM-5.2, not the more practical 30B–70B local tier.
Independent evaluations and leaderboard placements
The model immediately posted strong third-party results, especially in coding-adjacent, vision, and design-heavy arenas.
- Frontend Code Arena: Qwen3.8-Max debuted at #4 overall with 1,668 Elo, trailing only Claude Opus 5 [Max] at 1,705 and Kimi K3 [Max] at 1,676, and roughly tied with Claude Opus 5 [High] at 1,669 @arena
- In Frontend Code Arena subslices, it ranked:
#2 Consumer Product
- #3 Brand & Marketing, Reference-based design, Gaming, Content Creation Tools
- #4 Data & Analytics
- #5 Simulations @arena
- Vision Arena: Qwen3.8-Max ranked #2 with 1,305, only 13 points behind Claude Fable 5 [High] @arena
- Vals Index: Qwen3.8-Max ranked #2 among open-weight models, #10 overall out of 43, with a score of 66.1 @ValsAI
- Vals also reported:
It matched Claude Opus 4.7 on the Index, 66.1 vs 66.1
- At about 2.3x lower cost per test: $2.68 vs $6.17 @ValsAI
- Vals’ benchmark-specific numbers:
SWE-bench: 87.3%, ahead of GPT-5.5 (82.6%) and GLM-5.2 (83.3%), but behind Claude Opus 4.8 (89.2%)
- Terminal-Bench 2.1: 67.4, up from 61.0 for Qwen 3.7 Max @ValsAI
- Vals also highlighted the pace of progress:
Qwen 3.7 Max = 57.5
- Qwen 3.8 Max = 66.1
- Gain of 8.6 points in ~2.5 months
- Price cut from $2.50/$7.50 to $2.00/$6.00 input/output @ValsAI
There were also more anecdotal but technically relevant claims:
- One user visualized benchmark deltas and argued “Opus 4.8 is mostly subsumed by 3.8-Max” on the chart they reconstructed @deliprao
- Another claimed Qwen 3.8 surpassed Fable 5 on Terminal Bench and said Anthropic was now under visible pressure @kimmonismus
- A separate tweet called Qwen 3.8 Max the “best object detection VLM” across satellite, infrared, documents, technical drawings, sketches, crowded scenes, and small objects, though this was based on examples rather than a cited benchmark paper @skalskip92
Facts vs. opinions
Facts / directly attributable claims
- Alibaba announced Qwen3.8-Max and said open weights arrive next week; Qwen3.8-27B will also go open-weight @Alibaba_Qwen
- Alibaba disclosed API pricing of $2 input / $6 output / $0.25 cached per million tokens @Alibaba_Qwen
- Arena reported #4 in Frontend Code Arena at 1,668 and #2 in Vision Arena at 1,305 @arena @arena
- Vals reported 66.1 on Vals Index, #2 among open-weight models, 87.3% SWE-bench, 67.4 Terminal-Bench 2.1, 1M context, 128k output, and lower cost-per-test than Opus 4.7 @ValsAI @ValsAI @ValsAI
- ZhihuFrontier stated 95B active parameters and protocol compatibility; this appears to be a secondary summary rather than an original Alibaba spec sheet @ZhihuFrontier
Opinions / extrapolations / rhetoric
- “China is no longer lagging behind but competing on equal footing” @kimmonismus
- “Open models are winning now” @JonathanRoss321
- “Looks like Opus 4.8 is mostly subsumed” @deliprao
- “Anthropic is under pressure” and “mood shifted drastically” are ecosystem readings, not measurements @kimmonismus
- “Best object detection VLM” is an informed product judgment, but not one tied in-thread to a standard benchmark table @skalskip92
- Claims that Qwen3.8-Max plus open agents prove open models have “caught up” are user-level interpretations rather than consensus eval conclusions @omarsar0
The central factual story is strong even after stripping out the hype: a very large sparse model, open-weight promise, lower pricing than prior Qwen Max, and high placements on multiple third-party leaderboards.
The infrastructure reality: “open-weight” does not mean easy to run
A major counterpoint in the discussion was that frontier open models are operationally open, but not broadly accessible in the local-inference sense.
- Jamin Ball argued that pricing comparisons were overstated because “vanilla” token prices ignore token efficiency and because these models are enormous:
Qwen 3.8 Max >2T params
- Kimi K3 ~104B active per token
- GLM 5.2 = 744B total, 40B active
- For K3, loading weights alone is >1TB memory
- Requires at least 8 H100/B200 GPUs to run
- Moonshot recommends 64+ accelerators in supernode-style setups @jaminball
- This same critique implicitly applies to Qwen3.8-Max, even if its active-parameter count is somewhat lower than K3’s: a 2.4T-class MoE is not a commodity local model @jaminball
- StableQuan made the practical version of the same point more bluntly: long, RAM-heavy prompts and slow tool calls make giant models painful on consumer hardware, recommending API use instead @stablequan
- At the same time, the excitement around Qwen3.8-27B shows where many developers think the real adoption wave may come from: a smaller open-weight descendant in the same family, possibly inheriting some of the flagship’s post-training or distilled capabilities @kimmonismus @TheZachMueller
This is the key split in the open-model story: ecosystem influence and benchmark legitimacy come from releasing the 2.4T flagship; practical deployment at scale may come from the 27B release.
Licensing controversy and geographic restrictions
The most concrete skeptical reaction was not about performance, but about the license.
- OstrisAI flagged what they read as a license prohibition covering the USA, EU, UK, and Korea, saying the terms appeared to forbid even downloading the model from the US @ostrisai
- That concern echoed a broader discussion happening simultaneously around another open-weight release, MiniMax H3, where users argued that geographic restrictions undercut claims of openness @kimmonismus
- No clarifying Qwen license tweet appears in this dataset from Alibaba itself, so the restrictive-license reading remained unresolved within these tweets
For engineers, this matters more than the marketing label. “Open weights” can still mean:
- no OSI-style open-source rights,
- use-case restrictions,
- export/jurisdiction limits,
- or no legal permission for commercial deployment in key regions.
That licensing ambiguity is one of the main reasons some of the reaction was more cautious than celebratory.
Why the launch matters strategically
This was widely read as a strategic shift by Alibaba, not just a routine product update.
- ZhihuFrontier explicitly framed the move as Alibaba choosing ecosystem influence over exclusivity, arguing that earlier Max models stayed closed while the open line had previously topped out around Qwen3-235B @ZhihuFrontier
- In that reading, DeepSeek, Kimi, and other Chinese open models weakened the premium of keeping top-tier systems API-only, pushing Alibaba to compete on ecosystem adoption as well as model quality @ZhihuFrontier
- Multiple observers connected Qwen3.8-Max to a broader Chinese-model surge:
“Top three spots in front-end design are now shared between two Chinese and one Western model” @kimmonismus
- “Remember when China was 2 years behind?” @matvelloso
- “The open weights frontier has been consistently dominated by labs from China for the last two years” @_micah_h
- Some posters escalated this into a geopolitical concern that US labs cannot rely on closed-model leads forever, especially if Chinese labs keep pushing frontier-ish systems into open-weight channels @kimmonismus
A subtext here is that the moat may be shifting:
- not just raw pretraining,
- but post-training, agent harnesses, inference infra, distillation pipelines, and developer lock-in.
That is exactly why an open-weight flagship at 2.4T is strategically valuable even if relatively few teams ever self-host it.
Model architecture and sparsity implications
The technical profile suggests Alibaba is leaning harder into sparse MoE than some rivals.
- If the 95B active / 2.4T total number quoted by ZhihuFrontier is accurate, Qwen3.8-Max activates only about 4% of total parameters per token @ZhihuFrontier
- ZhihuFrontier contrasted this to Qwen3-235B-A22B, which they say activates closer to 10% @ZhihuFrontier
- Elie Bakouch’s broader comment—“the two biggest OSS models in the world use linear attention?”—captures another architectural thread in the ecosystem conversation, though it was not directly tied to Qwen3.8-Max with a cited source in-thread @eliebakouch
- The wider thread around sparse MoE and Switch Transformers reflects why people care about these parameter numbers: frontier open models can look “huge to store yet still cheap to run” by only activating a narrow expert slice per token @ProfTomYeh
This is likely part of how Alibaba can cut API pricing while scaling total parameter count upward: bigger expert pool, lower active footprint, lower effective inference cost, assuming routing and systems optimizations hold up in production.
Long-horizon agents, cowork, and benchmark fit
Qwen3.8-Max was pitched less as a chatbot and more as a model-harness substrate for long-running work.
- Alibaba’s own language emphasized “coding and cowork” rather than generic assistant use @Alibaba_Qwen
- The launch claims map unusually well to the current “long-horizon agents” discourse:
10+ day autonomous coding
- 500+ turns in chip optimization
- 365-day business strategy @Alibaba_Qwen
- ZhihuFrontier’s benchmark picks—PaperBench, CoWorkBench, WideSearch—all emphasize persistent objective maintenance, tool use, and trajectory coherence rather than one-shot Q&A @ZhihuFrontier
Omar Sa
同じ出来事を6媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み