Qwen、コーディング・協働向け新オープンウェイトモデル「3.8 Max」などを公開へ
本文の状態
日本語全文を表示中
詳細モードで約26分の本文を読めます。
同じ出来事の情報源
6媒体で確認
Qwen Blog · Alibaba Engineering · MarkTechPost · The Decoder · VentureBeat AI · Latent Space
各社の報じ方を比較 ↓アリババのQwenチームは、コーディングや長期的な自律作業に特化した2.4Tパラメータ規模の「Qwen 3.8 Max」を発表し、来週にはオープンウェイト版を公開すると明言した。
AI深層分析を開く2026年8月4日 14:43
AI深層分析
キーポイント
大規模モデルとオープン化の発表
アリババは新フラッグシップモデル「Qwen3.8-Max」を発表し、API利用料金の詳細を提示した上で、来週にオープンウェイト版をリリースすると発表した。
自律的なコーディングと研究能力
同モデルは10日以上無人で動作するコード生成や、125時間にわたる自律的な研究ループによる新手法の発明など、高度な自律作業を実証している。
ハードウェア設計と実世界業務
RTL編集から物理レイアウトまでの完全なシリコン設計フローを完了し、同時に法的レビューやETFクオンツ研究などの数百の専門ワークフローでも生産 grade の出力を示した。
マルチモーダルエージェント機能
計画、コーディング、GUI操作にネイティブな視覚フィードバックを統合し、既存のエージェントフレームワークへ拡張するための「Qwen-MM-Plugins」も公開された。
新モデルの概要と機能
Qwen3.8-Max は 2.4T パラメータを持つオープンウェイトモデルで、コーディングや長期自律作業に特化している。
重要な引用
Qwen 3.8 Max is a MONSTER 2.4T model
10+ Days Unattended Coding: Built a self-evolving coding harness from scratch over a multi-week autonomous run.
Alibaba introduced Qwen3.8-Max as its 'most capable model to date,' describing it as a 2.4T-parameter model focused on coding, long-horizon agentic work, and multimodal reasoning
"most capable model to date"
編集コメントを表示
編集コメント
昨年の「Qwen Exodus」や管理体制の変化による懸念を払拭し、同社が再びオープンモデルの先駆者としての地位を確立しようとする姿勢が明確に表れている。特に自律的なハードウェア設計や長期間の無人作業といった実証事例は、次世代AIエージェントの実用化に向けた重要なマイルストーンとなるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
昨年の Qwen Exodus と新しい経営陣によるクローズドモデル API の展開以降、この主要なオープンモデル研究ラボが引き続き関連性の高いモデルをリリースし続けるのかどうかについて、懐疑的な声も聞かれました。
しかし、その不安はもうありません。Qwen 3.8 Max は、すでに紹介済みの Kimi K3 の登場を除けば世界最高のオープンモデルとなる、2.4T パラメータを持つモンスター級のモデルです。
API では 100 万トークンあたり入力$2、出力$6 で提供されていますが、両モデルの重み(weights)をオープンソース化することを約束しています。
主要な機能と画期的な成果
自律的な長期間コーディング
- 10 日以上無人のコーディング: 数週間にわたる自律的な実行を通じて、ゼロから自己進化型のコード生成ハッチ(harness)を構築しました。
- 自律的な AI 研究: 「LLM の推論のための統一データ選択」という論文の全体パイプラインをゼロから再構築し、125 時間にわたって自律的に反復研究ループを実行。その結果、オリジナル論文のベンチマークを +2.71 ポイント上回る新しいデータ選択手法を発明しました。
- 競争力のあるデータサイエンス: WWW2025 のマルチモーダル対話意図認識チャレンジに 526 チームが参加する中、24 時間以内に上位 13%(全チームの 87% を上回る成績)を達成しました。
自律的なハードウェア・チップ設計
- 完全なシリコン設計フローの実行: RTL エディティングからシミュレーション、合成、物理レイアウトに至るまで、GCD/RSA 暗号化アクセラレータの設計を完遂しました。
- ゲート数の劇的削減: ゲート数を 8,298 から 678 に減らしながら、チップ面積を 81% 削減。さらに 500 MHz で物理タイミングクロージャ(timing closure)も達成しました。
実世界での業務とオペレーション:
数百の専門的なワークフローにわたって、本番環境で通用する出力を実証しました。具体的には、企業の法務レビュー、UI/UX デザイン、構造力学モデル、ETF 定量的研究の自動化などが含まれます。
E-Commerce Bench(365 日間の店舗運営シミュレーション)では競合モデルを上回る結果を達成し、継続的なゲーム理論に基づく交渉と在庫計画を通じて、4.16 倍のリターン(残高¥416,252)を生み出しました。
マルチモーダルエージェントと視覚フィードバック:
プランニング、コーディング、GUI 操作の各段階にネイティブな視覚フィードバックを統合し、デスクトップ、モバイル、Web といったプラットフォーム間でのアプリケーションの直接再現を可能にしています。
既存のエージェントフレームワークに対してマルチモーダル機能を拡張する「Qwen-MM-Plugins」も公開されました。
オープンウェイトモデルにとって非常に素晴らしい成果です。本日 Baseten と共に放送したポッドキャストでは、これらの大規模なモデルリリースをどのようにサポートするかについて議論しました。
7/25/2026〜7/27/2026 の AI ニュース。12 のサブレッドと 544 のツイートを確認し、Discord での追加情報は確認できませんでした。AINews のウェブサイトでは過去のニュースをすべて検索可能です。なお、AINews は現在 Latent Space の一部となっています。メール購読の頻度設定はご自身で選択できます。
AI Twitter レビュー
トップストーリー:Qwen 3.8 Max オープンモデルの発表
出来事
Alibaba Qwen が新フラッグシップ「Qwen3.8-Max」を発表し、オープンウェイト版は来週公開されると明言しました。
アリババは「Qwen3.8-Max」を「これまでに最も能力が高いモデル」として紹介しました。このモデルは2.4兆パラメータを持ち、コーディングや長期にわたるエージェント作業、多モーダル推論に特化しています。また、来週にはオープンウェイト版がリリースされ、同様にオープンウェイトとなる「Qwen3.8-27B」も同時に公開されると明言しました。
発表時のツイートでは、APIの利用料金も併せて発表されました。入力トークン100万あたり2ドル、出力トークン100万あたり6ドル、キャッシュされたトークン100万あたり0.25ドルです。
アリババはこのモデルの主要な能力として、いくつかの点を強調しました。具体的には、10日以上続く自律的なコーディング作業や、500ターンを超える半導体設計の最適化、365日分のEC戦略策定などです。また、視覚情報を単なる入力チャネルとして扱うのではなく、実行ループの一部として組み込むネイティブな多モーダル知能も特徴の一つとしています。
同社は同時に、自社プラットフォームやパートナー企業を通じて利用可能にする動きを加速させました。Qwen Studio、API、Command Code、そして将来的にVeniceでの提供です。インフラ構築者やアプリ開発者はすぐにサポート計画や統合の発表を行いました。Baseten、Hermes Agent、Command Codeなどがその例です。
今回の発表は、より広範なトレンドの一部として捉えられています。複数の観察者が、中国発のオープンウェイトモデルが、特にコーディング、エージェントワークフロー、多モーダルタスクにおいて、トップクラスの西側クローズドモデルと直接競合する段階に入ったことを示す証拠だと指摘しています。
公式発表と報告された仕様
ベンダーが報告したモデルの詳細や性能に関する主張は、オープンウェイトのリリースとしては異例ほどに積極的なものでした。
アリババ自身の見解:
総パラメータ数 2.4T @Alibaba_Qwen
長期にわたる自律型エージェントや共同作業(コワーク)への焦点を当てたモデル @Alibaba_Qwen
公開された GitHub のトレースに基づき、10 日以上継続して自律的にコードを生成・実行する能力 @Alibaba_Qwen
半導体設計の最適化において 500 回以上の対話ターンを実現 @Alibaba_Qwen
EC(電子商取引)戦略の実行を 365 日間継続 @Alibaba_Qwen
視覚情報のみの入力に依存するのではなく、ネイティブなマルチモーダルな計画ループを採用 @Alibaba_Qwen
ZhihuFrontier の第三者による要約ツイートでは、さらに多くの技術的詳細が報告されています:
トークンあたり 95B のアクティブパラメータ。これは MoE(混合专家モデル)の活性化比率が約 4% であることを示唆しています。
1M トクンのコンテキストウィンドウ。
API は、推論効率を調整するための低(low)、中(medium)、超高(xhigh)のモードを提供します。
OpenAI および Anthropic のプロトコルとの互換性を備えています。
ベンチマーク結果:PaperBench 93.0、CoWorkBench 74.8、WideSearch 81.9 @ZhihuFrontier
Vals AI による独立した評価では、具体的な評価・実行設定が公開されました:
1M トクンのコンテキスト。
最大出力トークン数は 128k。
温度(temperature)0.7 でテストされ、デフォルトの top-p および top-k が使用されています @ValsAI
これらの数値は重要であり、Qwen3.8-Max を Kimi K3 や GLM-5.2 のような他の巨大なスパースオープンモデルと同じ「展開クラス」に位置づけるものです。これは、より実用的とされる 30B〜70B レベルのローカルモデルとは異なるカテゴリです。
独立した評価とリーダーボードでの順位
このモデルはすぐに強力な第三者による結果を発表しました。特にコーディング関連、視覚処理、デザイン重視の分野で顕著な成果を収めています。
Frontend Code Arena において、Qwen3.8-Max は総合 4 位(Elo 1,668)でデビューを果たしました。これは Claude Opus 5 [Max] の 1,705 と Kimi K3 [Max] の 1,676 に次ぐ順位です。また、Claude Opus 5 [High] の 1,669 とほぼ同率となっています @arena
フロントエンドコードアリーナのサブスライスでは、以下の順位を記録しました。
#2 消費者向け製品
#3 ブランド・マーケティング、リファレンスベースのデザイン、ゲーム、コンテンツ作成ツール
#4 データ&アナリティクス
#5 シミュレーション @arena
ビジョンアリーナでは、Qwen3.8-Max が総合 1,305 点で #2 にランクイン。Claude Fable 5 [High] @arena との差はわずか 13 ポイントです。
Vals Index では、オープンウェイトモデルの中で #2、全体(43 モデル中)では #10 の 66.1 点を獲得しました @ValsAI。
Vals が報告したその他の主要データ:
- インデックス値は Claude Opus 4.7 と同等の 66.1 点(Claude Opus 4.7 も 66.1)
- テストあたりのコストは約 2.3 倍低く、$2.68 vs $6.17 @ValsAI
Vals のベンチマーク別数値:
- SWE-bench: 87.3%。GPT-5.5(82.6%)や GLM-5.2(83.3%)を上回る一方、Claude Opus 4.8(89.2%)には及びません。
- Terminal-Bench 2.1: 67.4。Qwen 3.7 Max の 61.0 から向上しました @ValsAI
Vals はまた、進歩のスピードにも注目しています:
- Qwen 3.7 Max = 57.5
- Qwen 3.8 Max = 66.1
- 約 2.5 ヶ月で 8.6 ポイントの向上
- 価格も入力/出力ともに $2.50/$7.50 から $2.00/$6.00 に引き下げられました @ValsAI
技術的な裏付けのある事例報告もいくつかあります:
- あるユーザーはベンチマークの差分を可視化し、「Opus 4.8 は主に 3.8-Max に吸収されている」と、自身が再構築したチャート上で主張しました @deliprao。
- 他の投稿では「Qwen 3.8 が Terminal Bench で Fable 5 を上回り、Anthropic に明確な圧力がかかっている」と指摘されました @kimmonismus。
- さらに別のツイートでは、Qwen 3.8 Max が衛星画像、赤外線、文書、技術図面、スケッチ、混雑したシーン、小物体などあらゆる対象検出において「最良のオブジェクト検出 VLM」であると称賛されました。ただしこれは引用されたベンチマーク論文に基づくものではなく、あくまで具体例に基づいた評価です @skalskip92
事実と意見
事実/直接帰属可能な主張
アリババは Qwen3.8-Max の発表を行い、オープンウェイト版は来週公開されると明言しました。また、Qwen3.8-27B も同様にオープンウェイト化されます(@Alibaba_Qwen)。
API 利用料金の詳細も公表されました。入力 100 万トークンあたり 2 ドル、出力 100 万トークンあたり 6 ドル、キャッシュされたトークンは 100 万あたり 0.25 ドルです(@Alibaba_Qwen)。
Arena の結果では、Frontend Code Arena で 1,668 ポイントを獲得し 4 位、Vision Arena では 1,305 ポイントで 2 位となりました(@arena @arena)。
Vals の評価では、総合スコアが 66.1 とオープンウェイトモデルの中で 2 位にランクイン。SWE-bench は 87.3%、Terminal-Bench 2.1 は 67.4% を達成しました。コンテキスト長は 1M、最大出力トークンは 128k で、Opus 4.7 よりもテストあたりのコストが低いことが確認されています(@ValsAI @ValsAI @ValsAI)。
ZhihuFrontier は「アクティブパラメータ数が 95B」と「プロトコル互換性がある」と報じましたが、これはアリババ公式の仕様書ではなく、二次的な要約情報の可能性が高いです(@ZhihuFrontier)
意見/推測/修辞的表現
「中国はもはや後れを取っているわけではなく、対等な立場で競い合っている」(@kimmonismus)
「オープンモデルが今や勝利を収めている」(@JonathanRoss321)
「Opus 4.8 はほぼ吸収されたようだ」(@deliprao)
「Anthropic に圧力がかかっている」「ムードが劇的に変わった」といった見解は、生態系全体の傾向を示すものであり、数値測定に基づくものではありません(@kimmonismus)
「最良の物体検出 VLM」という評価は専門的な製品判断ですが、スレッド内で標準ベンチマーク表と結びついたものではないため注意が必要です(@skalskip92)
Qwen3.8-Max とオープンエージェントの組み合わせが「オープンモデルがついに追いついた」ことを証明するという主張は、ユーザーレベルの解釈であり、コンセンサスに基づく評価結論ではありません(@omarsar0)
hype を取り除いても、核心となる事実は依然として強力です。非常に大規模なスパースモデルであり、オープンウェイトの提供が約束され、以前の Qwen Max よりも低価格で、複数のサードパーティリーダーボードで上位にランクされています。
インフラの実情:「オープンウェイト」は実行しやすいことを意味しない
議論における主要な反論として、最先端のオープンモデルは運用上はオープンであっても、ローカル推論という観点では広く利用可能ではないという点があります。
Jamin Ball は、価格比較が過大評価されていると指摘しました。その理由は、「バニラ」トークン価格がトークンの効率性を無視していること、そしてこれらのモデルが巨大であることです。
Qwen 3.8 Max >2T パラメータ
Kimi K3 トークンあたり約 104B アクティブ
GLM 5.2 総計 744B、アクティブ 40B
K3 の場合、ウェイトの読み込みだけでメモリが 1TB を超えます。実行には少なくとも 8 枚の H100 または B200 GPU が必要です。
Moonshot は、supernode スタイルのセットアップで 64 基以上のアクセラレータの使用を推奨しています @jaminball
この批判は、アクティブパラメータ数が K3 よりもやや低いとしても、Qwen3.8-Max に暗黙的に適用されます。2.4T クラスの MoE は、汎用的なローカルモデルではありません @jaminball
StableQuan は、同じ点をより率直に実践的な観点から説明しました。長く RAM を大量に消費するプロンプトや、遅いツール呼び出しにより、巨大モデルは消費者向けハードウェアでは扱いにくいものです。そのため、API の利用を推奨しています @stablequan
同時に、Qwen3.8-27B への注目度の高さは、多くの開発者が実用的な採用の波がどこから来ると考えているかを示しています。それは同じファミリーに属するより小型のオープンウェイトモデルであり、フラッグシップ版のポストトレーニングやディストillation(知識蒸留)による能力の一部を継承している可能性があります @kimmonismus @TheZachMueller。
これはオープンモデル界における重要な分岐点です。エコシステムへの影響力とベンチマークでの正当性は 2.4T のフラッグシップ版の公開から得られますが、大規模な実用展開は 27B モデルのリリースによって実現するかもしれません。
ライセンス論争と地理的制限
最も具体的な懐疑的な反応は性能に関するものではなく、ライセンスに関するものでした。
OstrisAI は、米国、EU、英国、韓国をカバーするライセンス禁止条項があると指摘しました。その条件は米国内からのモデルダウンロードさえも禁じているように見えると述べています @ostrisai。
この懸念は、同時に別のオープンウェイトリリースである MiniMax H3 を巡って行われていた議論とも響き合いました。ここではユーザーが地理的制限が「オープン性」を主張する根拠を損なうと論じていました @kimmonismus。
アリババ自身によるライセンスに関する明確なツイートはこのデータセットには含まれていないため、制限的なライセンス解釈はこれらのツイート内では解決されませんでした。
エンジニアにとってこれはマーケティング上のラベルよりも重要です。「オープンウェイト」という言葉が意味するものは以下の通りです:
- OSI 基準のようなオープンソース権限の欠如
- 利用ケースに関する制限
- エクスポートや管轄区域に関する制限
- 主要地域における商用展開のための法的許可の欠如
ライセンスの曖昧さが、一部で祝賀よりも慎重な反応が示された主な理由の一つです。
なぜこの発表が戦略的に重要なのか
これはアリババによる戦略的転換と広く受け止められました。単なる通常の製品アップデートではありません。
ZhihuFrontier はこの動きを、アリババが独占性ではなくエコシステムへの影響力を選んだものとして明確に位置づけました。同氏は、以前の Max モデルはクローズドなままであり、オープン系では以前から Qwen3-235B が最高峰だったと指摘しています @ZhihuFrontier。
この見方によれば、DeepSeek や Kimi、その他の中国製オープンモデルが、最上位システムを API のみで提供し続けることのプレミアム価値を低下させました。その結果、アリババはモデルの品質だけでなく、エコシステムの採用率でも競争せざるを得なくなったのです @ZhihuFrontier。
複数の観察者が、Qwen3.8-Max を中国製モデルの更なる台頭というより大きな文脈に結びつけています:
「フロントエンドデザインのトップ 3 は現在、2 つの中国製と 1 つの西洋製モデルで分け合われています」 @kimmonismus
「かつて中国は 2 年遅れだったことを覚えていますか?」 @matvelloso
「オープンウェイトの最前線は、過去 2 年間一貫して中国の研究所によって支配されてきました」 @_micah_h
一部の投稿者はこれを地政学的な懸念へとエスカレートさせました。米国の研究所がクローズドモデルでのリードを永遠に頼りにできるわけではなく、特に中国の研究所が最前線に近いシステムをオープンウェイトのチャネルへ送り出し続ける限り、その限界は明白です @kimmonismus。
ここにはもう一つの背景があります。競争優位性の源泉(モート)が変化しているという点です:
単なる生データの前学習だけでなく、
ポストトレーニング、エージェントのハルネス、推論インフラ、蒸留パイプライン、そして開発者のロックインです。
2.4T という巨大な重みを持つオープンモデルが、実際にそれを自己ホストするチームは限られるとしても、戦略的に極めて価値が高い理由がここにあります。
モデルアーキテクチャとスパース性の影響
技術的なプロファイルから、アリババは競合他社よりもさらにスパースな MoE(Mixture of Experts)に注力していることが伺えます。
ZhihuFrontier が引用した「アクティブパラメータ 95B / 総パラメータ 2.4T」という数字が正確であれば、Qwen3.8-Max はトークンごとに全パラメータの約 4% しか活性化しません @ZhihuFrontier。
これに対し、ZhihuFrontier は Qwen3-235B-A22B を比較対象として挙げ、同モデルはより高い 10% に近い割合で活性化すると述べています @ZhihuFrontier。
Elie Bakouch の「世界最大のオープンソースモデル 2 つが線形アテンションを使用している」という広範なコメントは、エコシステムにおける議論の別の側面を捉えていますが、これはスレッド内で引用された出典に基づいて Qwen3.8-Max に直接結びつけられたものではありません @eliebakouch。
スパース MoE やスイッチトランスフォーマーに関するより広い議論がなぜ重要視されるのか。それは、フロンティア級のオープンモデルが「保存には巨大だが、実行コストは低い」という特性を持つからです。これは、トークンごとに狭いエクスパートの断片のみを活性化させることで実現されます @ProfTomYeh。
おそらくこれが、アリババが総パラメータ数を増大させつつ API 価格を引き下げられる理由の一部でしょう。より大きなエクスパートプールを持ち、アクティブなフットプリントを抑え、推論コストを実質的に下げる。ただし、これはルーティングやシステム最適化が生産環境でも機能することを前提とした話です。
長期ホライゾンのエージェント、コワーク、ベンチマーク適合性
Qwen3.8-Max は単なるチャットボットとしてではなく、長時間実行される作業のためのモデルハネス(基盤)として提案されています。
アリババが公式に強調したのは、汎用的なアシスタントとしての利用ではなく、「コーディングと共同作業」です。
今回の発表内容は、現在の「長期ホライズン・エージェント」という議論と驚くほど一致しています。具体的には、10 日以上続く自律的なコーディングや、チップ最適化における 500 回以上の対話ターン、そして 365 日規模のビジネス戦略策定などが挙げられています。
ZhihuFrontier が選定したベンチマーク(PaperBench、CoWorkBench、WideSearch)もまた、単発的な Q&A の正解率ではなく、目標の持続的維持、ツールの活用、そして行動経路の一貫性を重視しています。
Omar Sar0 は今回のリリースがエージェント用フレームワークと明確に結びついていると指摘し、「Hermes Agent で Qwen3.8-Max を使用すると、オープンな最先端モデルがクローズドなシステムとの差をいかに縮めたかが否定できなくなる」と述べています。
Cline によるオープンウェイトモデルに関する別スレッドも文脈として重要です。彼らは多くのオープンモデルが RL(強化学習)によってトレーニングされており、検証に多くのトークンを費やすように設計されているため、その特性を活かせるフレームワークと組み合わせることで最大の効果を発揮すると主張しています。実際、フレームワークの変更だけで約 20% の性能向上が見られるケースもあるといいます。
これは Qwen3.8-Max の発表内容とも非常に良く合致しています。単に「モデルが賢くなった」というだけでなく、「長期にわたる検証を重視する作業において、それに特化したフレームワークと組み合わせれば特に競争力が高まる」という示唆が含まれているのです。
反応における異なる視点
支持派
オープンモデル開発者やインフラプロバイダーからは強い熱意が寄せられています。
「Qwen 3.8 Max と、新しいローカル向けの 27B モデルである Qwen 3.8 が登場する」
「はい、Qwen3.8-Max を提供します」
「Hermes Agent で Qwen3.8-Max を試してみてください」
「素晴らしい!オープンソースの最大規模モデルだ」
@NerdyRodent
いくつかのコメントでは、このリリースがオープンモデルの実用的な負荷において最前線と同等かそれに近づいたことの証明だと捉えられていました。
@JonathanRoss321 @kimmonismus
中立・分析的な反応
Jamin Ball のスレッドが「しかし」という反論の中心となりました。彼は、価格差は過大評価されている可能性や、トークン効率の重要性、そして 2T パラメータを超えるオープンモデルにおけるインフラ負荷が依然として極めて大きい点を指摘しました。
@jaminball
Nrehiew は、性能向上が新しい事前学習よりも微調整(ポストトレーニング)に偏って寄与しているのではないかという疑問を呈しました。つまり、結果の差は「レシピ」によるものか、それとも「スケール」によるものかを問うたのです。
@nrehiew_
Vals は重要な方法論的な指摘を加えました。Alibaba が報告した Terminal Bench の結果ではベンチマークのタイムアウト時間が変更されていますが、Vals は元のタイムアウト設定を維持して評価を行ったと述べています。
@ValsAI
懐疑的・反対意見
ライセンスに関する懸念は、最も明確な実質的な批判でした。主要市場での利用に制限がある場合、「オープン」という主張は限定的なものになると指摘されました。
ostrisai
いくつかの強い懐疑論は間接的なものでした。これらの巨大なオープンウェイトモデルがスーパーノードを必要とし、慎重なハーン(環境構築)エンジニアリングを要するのであれば、その実用的な競争効果はリーダーボードの見出しが示唆するほど劇的ではないという見方です。
@jaminball
また、ベンチマークでの飛躍的な向上だけで最強のクローズドモデルとの完全な同等性が証明されることへの、より広範なエコシステムレベルの懐疑論もありました。例えば、一部のユーザーはオープンソースが「非常に近い」状態にあるとしても、トップクラスのエージェント型コーディングにおいてはまだ到達していないと主張しました。
@scaling01
背景:2026 年のオープンモデルサイクルにおける Qwen3.8-Max
今回の発表は、中国のラボから相次いで行われている大規模なオープンまたは準オープンのリリース群の中に位置づけられます。
議論で繰り返し言及されている比較対象モデルは以下の通りです。
- Kimi K3 (2.8T)
- GLM-5.2
- DeepSeek V4 Flash / Pro
- 多モーダル・動画分野では MiniMax H3 (@jaminball, @kimmonismus)
スレッド内で引用された Artificial Analysis の見解によれば、中国の最先端モデルは米国トップモデルに比べて概ね 3〜9 月遅れですが、オープンウェイトの最先端領域自体は約 2 年間、中国ラボが主導してきたとされています (@_micah_h)。
これが今回の発表がこれほど大きな注目を集めた理由を説明しています。単なるモデルの発売ではなく、以下のような明確な再編成の一部だからです。
- オープンウェイトの最先端規模において中国が最も強い
- 米国ラボは依然としてクローズドモデルの最高性能で先行している
- コーディング、デザイン、一部の多モーダルタスクなど特定領域ではその差が縮まっている (@_micah_h, @kimmonismus)
エンジニアにとっての実践的な影響
エンジニアにとって最も重要なのはマーケティング上の主張ではなく、実際の導入形態です。
最先端に近いオープンウェイトの品質を求める場合、Qwen3.8-Max は以下のようなトレードオフの空間を示唆しています。
- 非常に高い評価性能
- 積極的なトークン価格設定
- 巨大なサービング規模
- ライセンスや管轄区域に関する制約の可能性
1M のコンテキスト長と 128k の出力数は、トランスクリプトの再利用やキャッシュ価格が重要なリポジトリスケールおよびワークフロースケールのタスクにおいて実現可能であることを意味します (@ValsAI, @Alibaba_Qwen)。
キャッシュトークンあたりの価格が 0.25 ドル/M という点は、コードベースやツールトレース、大規模な指示プレフィックスを繰り返し再生するエージェントにとって特に重要です。@Alibaba_Qwen
Qwen3.8-27B の発表はフラッグシップモデルと同様に重要となる可能性があります。これは、より広範なオープンソーススタックやローカルサービングエコシステムで実際に利用可能になる層として最も期待されているからです。@Alibaba_Qwen @kimmonismus
すでに複数の開発者が、このリリースをチャット UX だけでなく、ダウンストリームハルネスやエージェントの文脈で捉えています。Hermes Agent、Command Code、Baseten など、OpenAI や Anthropic と互換性のあるプロトコルをサポートするあらゆるプロバイダーが、既存のワークフローにすばやく組み込むことができるでしょう。@Alibaba_Qwen @Alibaba_Qwen @baseten
TeortaxesTex による注目すべき解釈の一つは、Qwen 3.8 Max が以下の特性を持つ可能性があるというものです:
- 画像認識やラベリングにおいて極めて強力であること
- サンプル効率に優れている可能性
- タスク特化型の同等性能を維持しつつ、Qwen 3.8 27B に蒸留(OPD)可能であること。これは、フラッグシップモデルの能力からノートパソコンで展開可能な専門化への道筋を示唆しています。
@teortaxesTex
その他のトピック
エージェントインフラストラクチャ、ハルネス、長期ホライズンシステム
詳細な調査サマリーでは、長期ホライズンの能力はモデル単体の性質ではなく、「モデル×ハルネス」の組み合わせによる性質であると論じられています。この能力は、ゴールのドリフト、コンテキストの破損、報酬が希薄または不可逆的なアクションを伴う問題といった失敗要因に分解され、制御プレーンがプロンプトエンジニアリングからランタイムハルネスへとシフトするべきだと示唆しています。@ZhihuFrontier
Cloudflare は、各エージェントに「専用のコンピューター」を提供するために、アイソレートと Linux コンテナの間を動的にルーティングするエージェントランタイム「@cloudflare/computer」を発表しました。
一方、Cursor は 20〜30% のベータ版を報告しています。
原文を表示
After the Qwen Exodus last year and new management took over launching more closed model APIs, there was some real doubt as to whether or not this leading open models lab would continue to release relevant models.
That doubt is now gone. Qwen 3.8 Max is a MONSTER 2.4T model that would have been the top open model in the world but for the Kimi K3 release we already covered.
Qwen offers them on API for $2 input/$6 output per million tokens, but they have promised to open-weight both models.
Key Capabilities & Breakthrough Highlights
Autonomous Long-Horizon Coding:
10+ Days Unattended Coding: Built a self-evolving coding harness from scratch over a multi-week autonomous run.
Autonomous AI Research: Rebuilt a complete paper’s pipeline (Unified Data Selection for LLM Reasoning) from scratch, then autonomously ran an iterative research loop over 125 hours to invent a new data selection method beating the original paper’s benchmark by +2.71 points.
Competitive Data Science: Competed against 526 human teams in the WWW2025 Multimodal Dialogue Intent Recognition Challenge, placing in the top 13% (outperforming 87% of human teams) within 24 hours.
Autonomous Hardware & Chip Design:
Executed a complete silicon design flow (GCD/RSA cryptographic accelerator) from RTL editing to simulation, synthesis, and physical layout.
Reduced gate count from 8,298 to 678 gates while achieving an 81% die area reduction and meeting physical timing closure at 500 MHz.
Deep Real-World Work & Operations:
Demonstrated production-grade outputs across hundreds of professional workflows (e.g., corporate legal reviews, UI/UX design, structural engineering models, and automated ETF quant research).
Outperformed competing models in the E-Commerce Bench (a 365-day store operation simulation), generating a 4.16x return (¥416,252 balance) through continuous game-theoretic negotiation and inventory planning.
Multimodal Agents & Visual Feedback:
Integrates native visual feedback across planning, coding, and GUI interaction, enabling direct application recreation across platforms (desktop, mobile, web).
Released Qwen-MM-Plugins to extend multimodal capabilities to existing agent frameworks.
A very nice win for open weights! On today’s pod with Baseten we talked about what it’s like to support these massive model drops on release.
AI News for 7/25/2026-7/27/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
Top Story: Qwen 3.8 Max open model launch
What happened
Alibaba Qwen announced Qwen3.8-Max as its new flagship and said open weights are coming next week.
Alibaba introduced Qwen3.8-Max as its “most capable model to date,” describing it as a 2.4T-parameter model focused on coding, long-horizon agentic work, and multimodal reasoning, with the explicit claim that open weights will be released next week, alongside Qwen3.8-27B also going open-weight @Alibaba_Qwen
The launch tweet also included API pricing: $2.00 / M input tokens, $6.00 / M output tokens, and $0.25 / M cached tokens @Alibaba_Qwen
Alibaba framed the model around several headline capabilities: 10+ days of autonomous coding, 500+ turns of chip design optimization, 365 days of e-commerce strategy, and native multimodal intelligence where vision is part of the execution loop rather than just an input channel @Alibaba_Qwen
The company simultaneously pushed availability across its own surfaces and partners: Qwen Studio, API, Command Code, and later Venice; infra and app builders quickly confirmed support plans or integrations including Baseten, Hermes Agent, and Command Code @Alibaba_Qwen @Alibaba_Qwen @baseten @Teknium
The announcement landed as part of a broader pattern: multiple observers described it as evidence that the Chinese open-weight frontier is now competing directly with top Western closed models, especially in coding, agentic workflows, and multimodal tasks @kimmonismus @matvelloso
Official claims and reported specs
Vendor-reported model details and performance claims were unusually aggressive for an open-weight release.
Alibaba’s own framing:
2.4T total parameters @Alibaba_Qwen
Long-horizon agentic/cowork focus @Alibaba_Qwen
Autonomous coding over 10+ days with a public GitHub trace @Alibaba_Qwen
500+ turns for chip design optimization @Alibaba_Qwen
365 days of e-commerce strategy execution @Alibaba_Qwen
Native multimodal planning loop rather than vision-only input @Alibaba_Qwen
Third-party summary tweet from ZhihuFrontier added more claimed or reported technical details:
95B active parameters per token, implying an MoE activation ratio of roughly 4%
1M-token context window
API exposes low / medium / xhigh reasoning-effort modes
Compatibility with OpenAI and Anthropic protocols
Benchmark claims: PaperBench 93.0, CoWorkBench 74.8, WideSearch 81.9 @ZhihuFrontier
Vals AI independently posted concrete eval/runtime settings:
1M token context
128k max output tokens
Tested at temperature 0.7 with default top-p / top-k @ValsAI
These numbers matter because they place Qwen3.8-Max in the same deployment class as other giant sparse open models like Kimi K3 and GLM-5.2, not the more practical 30B–70B local tier.
Independent evaluations and leaderboard placements
The model immediately posted strong third-party results, especially in coding-adjacent, vision, and design-heavy arenas.
Frontend Code Arena: Qwen3.8-Max debuted at #4 overall with 1,668 Elo, trailing only Claude Opus 5 [Max] at 1,705 and Kimi K3 [Max] at 1,676, and roughly tied with Claude Opus 5 [High] at 1,669 @arena
In Frontend Code Arena subslices, it ranked:
#2 Consumer Product
#3 Brand & Marketing, Reference-based design, Gaming, Content Creation Tools
#4 Data & Analytics
#5 Simulations @arena
Vision Arena: Qwen3.8-Max ranked #2 with 1,305, only 13 points behind Claude Fable 5 [High] @arena
Vals Index: Qwen3.8-Max ranked #2 among open-weight models, #10 overall out of 43, with a score of 66.1 @ValsAI
Vals also reported:
It matched Claude Opus 4.7 on the Index, 66.1 vs 66.1
At about 2.3x lower cost per test: $2.68 vs $6.17 @ValsAI
Vals’ benchmark-specific numbers:
SWE-bench: 87.3%, ahead of GPT-5.5 (82.6%) and GLM-5.2 (83.3%), but behind Claude Opus 4.8 (89.2%)
Terminal-Bench 2.1: 67.4, up from 61.0 for Qwen 3.7 Max @ValsAI
Vals also highlighted the pace of progress:
Qwen 3.7 Max = 57.5
Qwen 3.8 Max = 66.1
Gain of 8.6 points in ~2.5 months
Price cut from $2.50/$7.50 to $2.00/$6.00 input/output @ValsAI
There were also more anecdotal but technically relevant claims:
One user visualized benchmark deltas and argued “Opus 4.8 is mostly subsumed by 3.8-Max” on the chart they reconstructed @deliprao
Another claimed Qwen 3.8 surpassed Fable 5 on Terminal Bench and said Anthropic was now under visible pressure @kimmonismus
A separate tweet called Qwen 3.8 Max the “best object detection VLM” across satellite, infrared, documents, technical drawings, sketches, crowded scenes, and small objects, though this was based on examples rather than a cited benchmark paper @skalskip92
Facts vs. opinions
Facts / directly attributable claims
Alibaba announced Qwen3.8-Max and said open weights arrive next week; Qwen3.8-27B will also go open-weight @Alibaba_Qwen
Alibaba disclosed API pricing of $2 input / $6 output / $0.25 cached per million tokens @Alibaba_Qwen
Arena reported #4 in Frontend Code Arena at 1,668 and #2 in Vision Arena at 1,305 @arena @arena
Vals reported 66.1 on Vals Index, #2 among open-weight models, 87.3% SWE-bench, 67.4 Terminal-Bench 2.1, 1M context, 128k output, and lower cost-per-test than Opus 4.7 @ValsAI @ValsAI @ValsAI
ZhihuFrontier stated 95B active parameters and protocol compatibility; this appears to be a secondary summary rather than an original Alibaba spec sheet @ZhihuFrontier
Opinions / extrapolations / rhetoric
“China is no longer lagging behind but competing on equal footing” @kimmonismus
“Open models are winning now” @JonathanRoss321
“Looks like Opus 4.8 is mostly subsumed” @deliprao
“Anthropic is under pressure” and “mood shifted drastically” are ecosystem readings, not measurements @kimmonismus
“Best object detection VLM” is an informed product judgment, but not one tied in-thread to a standard benchmark table @skalskip92
Claims that Qwen3.8-Max plus open agents prove open models have “caught up” are user-level interpretations rather than consensus eval conclusions @omarsar0
The central factual story is strong even after stripping out the hype: a very large sparse model, open-weight promise, lower pricing than prior Qwen Max, and high placements on multiple third-party leaderboards.
The infrastructure reality: “open-weight” does not mean easy to run
A major counterpoint in the discussion was that frontier open models are operationally open, but not broadly accessible in the local-inference sense.
Jamin Ball argued that pricing comparisons were overstated because “vanilla” token prices ignore token efficiency and because these models are enormous:
Qwen 3.8 Max >2T params
Kimi K3 ~104B active per token
GLM 5.2 = 744B total, 40B active
For K3, loading weights alone is >1TB memory
Requires at least 8 H100/B200 GPUs to run
Moonshot recommends 64+ accelerators in supernode-style setups @jaminball
This same critique implicitly applies to Qwen3.8-Max, even if its active-parameter count is somewhat lower than K3’s: a 2.4T-class MoE is not a commodity local model @jaminball
StableQuan made the practical version of the same point more bluntly: long, RAM-heavy prompts and slow tool calls make giant models painful on consumer hardware, recommending API use instead @stablequan
At the same time, the excitement around Qwen3.8-27B shows where many developers think the real adoption wave may come from: a smaller open-weight descendant in the same family, possibly inheriting some of the flagship’s post-training or distilled capabilities @kimmonismus @TheZachMueller
This is the key split in the open-model story: ecosystem influence and benchmark legitimacy come from releasing the 2.4T flagship; practical deployment at scale may come from the 27B release.
Licensing controversy and geographic restrictions
The most concrete skeptical reaction was not about performance, but about the license.
OstrisAI flagged what they read as a license prohibition covering the USA, EU, UK, and Korea, saying the terms appeared to forbid even downloading the model from the US @ostrisai
That concern echoed a broader discussion happening simultaneously around another open-weight release, MiniMax H3, where users argued that geographic restrictions undercut claims of openness @kimmonismus
No clarifying Qwen license tweet appears in this dataset from Alibaba itself, so the restrictive-license reading remained unresolved within these tweets
For engineers, this matters more than the marketing label. “Open weights” can still mean:
no OSI-style open-source rights,
use-case restrictions,
export/jurisdiction limits,
or no legal permission for commercial deployment in key regions.
That licensing ambiguity is one of the main reasons some of the reaction was more cautious than celebratory.
Why the launch matters strategically
This was widely read as a strategic shift by Alibaba, not just a routine product update.
ZhihuFrontier explicitly framed the move as Alibaba choosing ecosystem influence over exclusivity, arguing that earlier Max models stayed closed while the open line had previously topped out around Qwen3-235B @ZhihuFrontier
In that reading, DeepSeek, Kimi, and other Chinese open models weakened the premium of keeping top-tier systems API-only, pushing Alibaba to compete on ecosystem adoption as well as model quality @ZhihuFrontier
Multiple observers connected Qwen3.8-Max to a broader Chinese-model surge:
“Top three spots in front-end design are now shared between two Chinese and one Western model” @kimmonismus
“Remember when China was 2 years behind?” @matvelloso
“The open weights frontier has been consistently dominated by labs from China for the last two years” @_micah_h
Some posters escalated this into a geopolitical concern that US labs cannot rely on closed-model leads forever, especially if Chinese labs keep pushing frontier-ish systems into open-weight channels @kimmonismus
A subtext here is that the moat may be shifting:
not just raw pretraining,
but post-training, agent harnesses, inference infra, distillation pipelines, and developer lock-in.
That is exactly why an open-weight flagship at 2.4T is strategically valuable even if relatively few teams ever self-host it.
Model architecture and sparsity implications
The technical profile suggests Alibaba is leaning harder into sparse MoE than some rivals.
If the 95B active / 2.4T total number quoted by ZhihuFrontier is accurate, Qwen3.8-Max activates only about 4% of total parameters per token @ZhihuFrontier
ZhihuFrontier contrasted this to Qwen3-235B-A22B, which they say activates closer to 10% @ZhihuFrontier
Elie Bakouch’s broader comment—“the two biggest OSS models in the world use linear attention?”—captures another architectural thread in the ecosystem conversation, though it was not directly tied to Qwen3.8-Max with a cited source in-thread @eliebakouch
The wider thread around sparse MoE and Switch Transformers reflects why people care about these parameter numbers: frontier open models can look “huge to store yet still cheap to run” by only activating a narrow expert slice per token @ProfTomYeh
This is likely part of how Alibaba can cut API pricing while scaling total parameter count upward: bigger expert pool, lower active footprint, lower effective inference cost, assuming routing and systems optimizations hold up in production.
Long-horizon agents, cowork, and benchmark fit
Qwen3.8-Max was pitched less as a chatbot and more as a model-harness substrate for long-running work.
Alibaba’s own language emphasized “coding and cowork” rather than generic assistant use @Alibaba_Qwen
The launch claims map unusually well to the current “long-horizon agents” discourse:
10+ day autonomous coding
500+ turns in chip optimization
365-day business strategy @Alibaba_Qwen
ZhihuFrontier’s benchmark picks—PaperBench, CoWorkBench, WideSearch—all emphasize persistent objective maintenance, tool use, and trajectory coherence rather than one-shot Q&A @ZhihuFrontier
Omar Sar0 explicitly linked the release to agent harnesses, saying using Qwen3.8-Max in Hermes Agent makes it hard to deny how much open frontier models have closed the gap with closed frontier systems @omarsar0
Cline’s separate thread about open-weight models is relevant context: they argue many open models are RL-trained to spend more tokens on verification and work best when the harness lets them lean into that behavior, producing ~20% gains from harness changes alone @cline
That fits Qwen3.8-Max’s launch narrative unusually well. The implication is not simply “model is smarter,” but “model may be especially competitive when paired with a harness designed for long-running verification-heavy work.”
Different perspectives in the reaction
Supportive
Strong enthusiasm from open-model developers and infra providers:
“Qwen 3.8 Max and a new local 27B Qwen 3.8 is coming” @Teknium
“Yes, we will have Qwen3.8-Max” @baseten
“Try Qwen3.8-Max on Hermes Agent…” @omarsar0
“Nice! An open source max model” @NerdyRodent
Several commenters treated the release as proof that open models are at or near frontier parity on meaningful workloads @JonathanRoss321 @kimmonismus
Neutral / analytical
Jamin Ball’s thread was the main “yes, but” reaction:
pricing gap may be overstated,
token efficiency matters,
infra burden remains extreme for >2T open models @jaminball
Nrehiew questioned whether performance gains might come disproportionately from post-training rather than novel pretraining, essentially asking how much of the delta is recipe vs scale @nrehiew_
Vals added an important methodological note: Alibaba’s reported Terminal Bench results modify benchmark timeouts, whereas Vals preserved original timeouts @ValsAI
Skeptical / opposing
License concern was the clearest substantive criticism: if usage is restricted in major markets, “open” becomes a narrower claim @ostrisai
Some of the strongest skepticism was indirect: if these giant open-weight models require supernodes and careful harness engineering, then their practical competitive effect may be less dramatic than leaderboard headlines suggest @jaminball
There was also broader ecosystem skepticism that benchmark jumps alone prove full parity with the strongest closed models; e.g. some users argued open source is “very close” but not actually there yet on top-end agentic coding @scaling01
Context: Qwen3.8-Max inside the 2026 open-model cycle
The launch sits in a dense cluster of giant open or quasi-open releases from Chinese labs.
The comparison set repeatedly mentioned in the discussion:
Kimi K3 at 2.8T
GLM-5.2
DeepSeek V4 Flash / Pro
MiniMax H3 on the multimodal/video side @jaminball @kimmonismus
Artificial Analysis commentary cited in-thread said Chinese frontier models have generally trailed top US models by about 3–9 months, while the open-weight frontier itself has been dominated by Chinese labs for roughly two years @_micah_h
This helps explain why the release drew such outsized attention: it is not just another model launch, but part of a visible realignment where:
China is strongest in open-weight frontier scale
US labs still often lead in top closed-model performance
the gap is narrowing on select domains like coding, design, and some multimodal tasks @_micah_h @kimmonismus
Practical implications for engineers
For engineers, the most important questions are less about marketing claims and more about deployment shape.
If you want frontier-ish open-weight quality, Qwen3.8-Max suggests the tradeoff space is now:
very strong eval performance
aggressive token pricing
huge serving footprint
possible license/jurisdiction constraints
The 1M context and 128k output numbers make it viable for repository-scale and workflow-scale tasks where transcript reuse and cache pricing matter @ValsAI @Alibaba_Qwen
The cached-token price of $0.25/M is especially relevant for agents repeatedly replaying codebases, tool traces, and large instruction prefixes @Alibaba_Qwen
The announcement of Qwen3.8-27B may be just as consequential as the flagship, because it is the tier likeliest to become actually usable across broader open-source stacks and local-serving ecosystems @Alibaba_Qwen @kimmonismus
Several developers already framed the release in terms of downstream harnesses and agents, not just chat UX: Hermes Agent, Command Code, Baseten, and likely any provider supporting OpenAI/Anthropic-compatible protocols can slot it into existing workflows quickly @Alibaba_Qwen @Alibaba_Qwen @baseten
One notable interpretation from TeortaxesTex was that Qwen 3.8 Max may be:
exceptionally strong on image recognition/labeling
potentially sample efficient
and distillable/OPD-able into Qwen 3.8 27B for task-specific parity, implying a route from flagship capability to laptop-deployable specializations @teortaxesTex
Other Topics
Agent infrastructure, harnesses, and long-horizon systems
A detailed survey summary argued that long-horizon capability is a model × harness property, not just a model property; it breaks failures into goal drift, context corruption, and sparse-reward/irreversible-action issues, and frames the control plane as shifting from prompt engineering to runtime harnesses @ZhihuFrontier
Cloudflare launched @cloudflare/computer, an agent runtime that dynamically routes between isolates and Linux containers so each agent gets “a computer of its own” @Cloudflare
Cursor reported 20–30% bet
AI算出
主要ニュースainew評価高い
既存の同クラスター記事と比較して、自律的な長期タスク実行(10 日以上)や RTL 設計フローの完全自動化など、具体的な技術的証拠と数値データが追加されており、単なる発表の繰り返しではない。また、2.4T パラメータという明確なバージョン情報が検索意図に合致する。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 75
- 新規性
- 75
- 調べる価値
- 100
- 重複の少なさ
- 52
- 日本での有用性
- 25
同じ出来事を6媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み