アリババ、2.4兆パラメータのオープンウェイト「Qwen3.8-Max」を公開へ
本文の状態
日本語全文を表示中
詳細モードで約9分の本文を読めます。
アリババは長期間の自律的な複雑タスク処理を目的とした新モデル「Qwen3.8-Max」を発表し、2.4兆パラメータ規模でコード生成や研究再現を実証した上で、来週に重み公開を予定している。
AI深層分析を開く2026年8月3日 23:32
AI深層分析
キーポイント
自律的な長期タスク処理能力の確立
本モデルは単発のプロンプト応答ではなく、数日間にわたる複雑なタスクを人間の手を介さずに完遂する設計となっている。
2.4兆パラメータ規模のアーキテクチャ
Qwen3.5のアーキテクチャを基盤とし、クエリごとに950億のパラメータが活性化される構造を持つ。
人間介入なしでの実証事例
16日間のコード開発や研究論文の再現・改良など、3つのケーススタディで完全自律的な動作を実証した。
オープンウェイト版の公開予定
Qwen-Maxクラス初のモデルとして、来週に重み(weights)が一般公開される予定である。
長期間にわたる複雑なタスクでの計画能力
暗号回路の設計において論理ゲートを約8,300個から678個へ削減し、チップ面積を81%縮小した。また、eコマースシミュレーションでは詐欺師を見抜き、初期資金を4倍に増やして競合を上回る成果を出した。
重要な引用
Alibaba's new flagship model Qwen3.8-Max is built to handle complex tasks on its own over days at a time
Qwen3.8-Max spent 16 days building the command-line tool oh-my-cli... without a single human touch.
It first reproduced all six of the paper's main results. Then it tested 18 of its own ideas across four rounds and beat the paper's method on the AIME24 math benchmark by 2.7 points.
The model kept making deep structural changes even after hundreds of iterations instead of settling for surface-level tweaks.
編集コメントを表示
編集コメント
アリババはQwen-Maxシリーズにおいて初めて重み公開を決定し、自律的なタスク実行能力を実証した。これは長期間にわたる複雑な作業をAIが担うという新しいパラダイムを示唆しており、開発現場の自動化可能性を大きく広げる内容である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
アリババの最新フラッグシップモデル「Qwen3.8-Max」は、研究論文の再現やチップ設計など、数日単位で複雑なタスクを自律的に処理できるように設計されています。同チームは今週中にこのモデルの重み(ウェイト)を公開する予定です。
アリババの Qwen チームは、これまでで最も能力の高い言語モデル「Qwen3.8-Max」を発表しました。このモデルは総パラメータ数が 2.4 兆に達し、1 クエリあたり 950 億のパラメータがアクティブになります。Qwen3.8-Max は Qwen3.5 のアーキテクチャを基盤としており、チームは単発のプロンプトへの回答だけでなく、長期にわたる複雑なタスクの自律的な完了に重点を置いていると述べています。
アリババはこのモデルを 7 月中旬にプレビュー版として発表し、アリババのトークンプラン「Qoder」や「QoderWork」を通じて標準価格の 10% で利用可能にしました。当時もチームは 2.4 兆パラメータを有し、Fable 5 に次ぐ性能を持つと評価 していましたが、ベンチマークデータは公開されていませんでした。Qwen3.8-Max は、重みが一般に公開される Qwen-Max クラス初のモデルとなります。
自律的なコーディング実行 3 回で性能を検証
Qwen3.8-Max のコーディング能力を披露するため、チームは人間の手を一切介さずにモデルが作業を行った 3 つの事例を紹介しました。
まず、Qwen3.8-Max は 16 日間にわたり、コマンドラインツール「oh-my-cli」の開発を行いました。このモデルはユーザーからのリクエストを受け取り、GitHub のイシューとして登録し、自らを担当者として割り当ててコードを書き、テストを実行して結果を反復的に改善しました。2026 年 7 月 30 日時点では、人間の介入なしにコミット数 265 件、プルリクエスト 127 件、イシュー 151 件を達成しています。
二つ目のケースでは、モデルは研究論文「Unified Data Selection for LLM Reasoning」を受け取りましたが、初期コードはありませんでした。その役割は論文の結果を再現し、さらにそれを改善することです。チームによると、約 5 日間で計算時間 125 時間を費やし、Qwen3.8-Max は 7,600 行のコードを作成して GPU 学習ジョブ 33 回を実行しました。まず論文の主要な 6 つの結果をすべて再現し、その後 4 ラウンドにわたって独自のアイデア 18 個を検証。その結果、AIME24 の数学ベンチマークでは論文手法よりも 2.7 ポイント高いスコアを記録しました。
三つ目のケースは、アリババの Tianchi プラットフォームで開催された「WWW2025 Multimodal Dialogue Intent Recognition Challenge」です。ここには 526 チームが参加する中、モデルは 24 時間以内に複数の中国語言語モデルと Qwen2.5-VL-7B を微調整し、製品スクリーンショット用の統合投票システムを構築しました。45 回の提出を通じて精度は 0.60 から 0.853 に向上し、参加した 526 チームのうち 458 チームを上回る結果となりました。
チップ設計とシミュレーションされた会計年度テストにおける長期計画の検証
2 つのケーススタディでは、数百回の対話ラウンドにわたるタスクへの挑戦が試されました。最初の事例は、暗号化スキームのための暗号学的構成要素を設計するものです。このような回路における主要な効率指標は、必要な論理ゲートの数です。論理ゲートはチップ上の基本要素であり、ゲート数が少ないほど、より小型で効率的なチップを実現できます。
Qwen3.8-Max は当初、8,298 個のゲートを使用する動作はするものの冗長な設計からスタートしました。約 500 回の反復を経て、最終的には 678 個のゲートにまで削減することに成功しました。
オープンソースツールである OpenROAD を用いた自動レイアウト処理の後、物理的なチップ面積は 106x106 マイクロメートルから 46x46 マイクロメートルへと縮小し、81% の削減を達成しました。Qwen チームによると、このモデルは数百回の反復を経ても表面的な調整で満足するのではなく、構造的な深い変更を継続的に行っていました。
2 つ目のケーススタディ「E-Commerce-Bench」は、淘宝(タオバオ)と天猫(ティーマル)からの匿名データを基に構築された、オンライン小売の 1 年間の財政年度をシミュレーションしたものです。モデルは初期資本として 10 万元からスタートし、1 年間を通じて複数のオンラインストアを並行して運営する必要があります。具体的には、商品の仕入れ、自然言語によるサプライヤーとの交渉、価格調整、返品管理、そして台風やサプライチェーンの混乱といった危機への対応などを行います。
サプライヤープールには 152 人の詐欺師が潜んでおり、モデルはこれらを見抜かなければなりません。Qwen3.8-Max は最終的に 416,252 元の残高を達成し、初期資金の 4 倍に増やしました。これは上位陣の GLM 5.2 よりも 38% 多く、前世代の Qwen3.7-Max が成し遂げた成果の 2.5 倍以上です。モデルは年初から積極的な投資を行い、休暇期間中には 10 万元を超える純利益を上げました。
ソースコードなしでのマルチモーダルスキルとアプリ再構築
チームによると、Qwen3.8-Max は 200 ページを超える文書や 100 時間以上の動画も処理可能です。また、ソースコードにアクセスできなくても実行中のアプリケーションを再構築することを要求する新しいベンチマーク「RecreationBench」も発表しました。
このモデルは、クリックやキーボード入力といった操作を通じてのみ対象のアプリを観察できます。テスト環境には Ubuntu、macOS、Windows、Android、および Web が含まれます。これに併せて、既存のエージェントシステムに画像・動画処理機能、視覚ツールの利用能力、マルチモーダルメモリを追加する拡張ライブラリ「Qwen-MM-Plugins」も公開されます。

Qwen チームが公開したベンチマーク表では、このモデルは多くのカテゴリーで Claude Opus 4.8、Claude Fable 5、GPT-5.6 Sol に匹敵するか、それ以上の成績を収めています。PaperBench では Qwen3.8-Max が 93 を記録し、比較対象の中で最高スコアとなりました。一方、TerminalBench 2.1 では 86.6 のスコアで、GPT-5.6 Sol の 88.8 にわずかに及びません。モデル開発者が発表する数値には一般的に言えることですが、これらの結果は内部での実行に基づくものです。独立した第三者による検証はまだ行われていません。
Qwen チームによると、このモデルが長時間のタスクを維持できる能力は、強化学習におけるトレーニング環境の大幅な拡大によるものとしています。トレーニングではもはや単一のタスクに焦点を当てるだけでなく、数日間にわたるワークフローや、個別ファイルではなくネストされたディレクトリ構造、そして多様なエージェント・ハーン(harness)などにも対応するようになりました。
チームが評価した 10 以上のベンチマークにおける総合スコアインデックスは、0.474 から 0.725 に向上しました。モデルの性能は約 4,000 の環境で最も高くなり、それ以降はわずかに低下する傾向が見られました。

中国におけるオープンモデル競争が激化
Qwen3.8-Max の最も直接的なライバルもまた中国から登場しました。Moonshot AI は 7 月 27 日、Hugging Face でオープンウェイトの「Kimi K3」をリリースしています。これはマルチモーダルな混合専門家(MoE)モデルで、パラメータ数は 2.8 兆、コンテキストウィンドウは 100 万トークンに達します。Moonshot は重量データとともに、アテンションカーネルや MoE 用通信ライブラリ、大規模エージェント実行のためのツールなど、自社インフラの一部も公開しました。
ただし、独立したテスト結果は企業の主張を少し冷たく受け止めました。K3 はサイバーセキュリティ能力と複雑な数学の両方で、トップクラスの西側モデルに大きく劣ることが判明しています。
Qwen3.8-Max は現在、QwenCloud を通じて利用可能です。重量データは来週、Hugging Face と ModelScope で公開される予定です。このモデルは OpenAI の Chat Completions 形式と Anthropic の API プロトコルの両方をサポートしており、Claude Code、Codex、Qoder CLI、Qwen Code、OpenClaw にそのまま接続して使用できます。また、reasoning_effort というパラメータにより、ユーザーは速度と精度のバランスが取れる 3 つのレベルから選択可能です。

Qwen は、このモデルが異なるエージェント環境でも同様の性能を発揮し、独自の環境に特化して調整されていないと述べています。| 画像:Alibaba
Qwen3.8-Max が Alibaba が最近追加した唯一の成果ではありません。数週間前には、情報密度の高いレイアウト向けの画像生成モデル「Qwen-Image-3.0」も発表されました。このモデルは最大 4,500 トークンの入力を処理でき、10 ピクセルという極めて小さな文字でも読みやすいテキストをレンダリングできます。一方、スケールの小さい方では、Alibaba は小規模なオープンモデルの推進を続けています。例えば「Qwen3.6-35B-A3B」は総パラメータ数が 350 億ですが、同時に活性化するのは 30 億のみです。
原文を表示
Alibaba's new flagship model Qwen3.8-Max is built to handle complex tasks on its own over days at a time, from reproducing research papers to designing chips autonomously. The team plans to release the weights next week.
Alibaba's Qwen team has unveiled Qwen3.8-Max, its most capable language model to date. The model scales to 2.4 trillion total parameters, with 95 billion active per query. Qwen3.8-Max builds on the Qwen3.5 architecture, and the team says the focus is on completing complex tasks independently over extended periods rather than just answering one-off prompts.
Alibaba first announced the model in mid-July as a preview version available through Alibaba's Token Plan, Qoder, and QoderWork at ten percent of the standard price. Even then, the team cited 2.4 trillion parameters and ranked the model just behind Fable 5, but didn't share benchmarks. Qwen3.8-Max is the first model in the Qwen-Max class whose weights will be made publicly available.
Three autonomous coding runs put the model through its paces
To show off Qwen3.8-Max's coding chops, the team presented three case studies in which the model worked without any human help.
In the first, Qwen3.8-Max spent 16 days building the command-line tool oh-my-cli. The model took incoming user requests, turned them into GitHub issues, assigned them to itself, wrote the code, ran tests, and improved the results iteratively. By July 30, 2026, it had racked up 265 commits, 127 pull requests, and 151 issues, all without a single human touch.
In the second case, the model received the research paper "Unified Data Selection for LLM Reasoning" but no starter code. Its job was to reproduce the paper's results and then improve on them. Over roughly five days and about 125 hours of compute time, Qwen3.8-Max wrote 7,600 lines of code and ran 33 GPU training jobs, according to the team. It first reproduced all six of the paper's main results. Then it tested 18 of its own ideas across four rounds and beat the paper's method on the AIME24 math benchmark by 2.7 points.
The third case involved the WWW2025 Multimodal Dialogue Intent Recognition Challenge on Alibaba's Tianchi platform, where 526 human teams competed. Within 24 hours, the model fine-tuned several Chinese language models along with Qwen2.5-VL-7B for product screenshots and combined them into a voting system. Across 45 submissions, accuracy climbed from 0.60 to 0.853. That put Qwen3.8-Max ahead of 458 of the 526 human teams.
Chip design and a simulated fiscal year test long-horizon planning
Two more case studies target tasks that stretch across hundreds of interaction rounds. In the first, Qwen3.8-Max had to design a cryptographic building block for encryption schemes. The key efficiency metric for such a circuit is the number of logic gates it needs, the basic elements on a chip. Fewer gates mean a smaller, more efficient chip. The model started with a working but bloated design using 8,298 gates and whittled it down to 678 gates over roughly 500 iterations.
After an automated layout pass with the open-source tool OpenROAD, the physical chip area shrank from 106x106 to 46x46 micrometers, an 81 percent reduction. The Qwen team says the model kept making deep structural changes even after hundreds of iterations instead of settling for surface-level tweaks.
The second case study is E-Commerce-Bench, a simulation of an entire fiscal year in online retail based on anonymized data from Taobao and Tmall. The model starts with 100,000 yuan in capital and has to run multiple online stores in parallel for a full year. That means buying products, negotiating with suppliers in natural language, adjusting prices, managing returns, and dealing with crises like typhoons or supply chain disruptions.
Hidden in the supplier pool are 152 scammers that the model has to spot. Qwen3.8-Max ended up with a balance of 416,252 yuan, quadrupling its starting capital. That's 38 percent more than the runner-up GLM 5.2 and more than 2.5 times what its predecessor Qwen3.7-Max managed. The model invested aggressively early in the year and pulled in a net profit of over 100,000 yuan during the holiday season.
Multimodal skills and app reconstruction without source code
For multimodal tasks, Qwen3.8-Max can process documents with over 200 pages and videos longer than 100 hours, the team says. They're also introducing RecreationBench, a new benchmark that requires the model to rebuild running applications without access to source code.
The model can only observe the target app through interaction, meaning clicks and keyboard input. Testing covers Ubuntu, macOS, Windows, Android, and the web. Alongside this, the team is releasing Qwen-MM-Plugins, an extension library that adds image and video processing, visual tool use, and multimodal memory to existing agent systems.

In the benchmark tables the Qwen team published, the model lands near or above Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol across many categories. On PaperBench, Qwen3.8-Max hits 93, the highest score in the comparison. On TerminalBench 2.1, it scores 86.6, trailing GPT-5.6 Sol's 88.8. As is typical with self-reported numbers from model makers, these results come from internal runs. Independent verification is still pending.
The Qwen team attributes the model's ability to sustain such long-running tasks to a major expansion of training environments during reinforcement learning. Training no longer focused on single tasks alone but also covered multi-day workflows, nested directory structures instead of individual files, and a variety of agent harnesses.
The team's internal score index across more than ten benchmarks rose from 0.474 to 0.725. The model performed best at around 4,000 environments, after which scores dipped slightly.

China's open-model race heats up
Qwen3.8-Max's most direct rival also comes from China. Moonshot AI released Kimi K3 with open weights on Hugging Face on July 27, a multimodal mixture-of-experts model with 2.8 trillion parameters and a one-million-token context window. Along with the weights, Moonshot also published parts of its own infrastructure, including attention kernels, an MoE communication library, and tools for running agents at scale. Independent testing tempered the company's claims, though. K3 fell well short of top Western models in both cyber capabilities and complex math.
Qwen3.8-Max is available now through QwenCloud. The weights are set to go live on Hugging Face and ModelScope next week. The model supports both OpenAI's Chat Completions format and Anthropic's API protocol, so it plugs directly into Claude Code, Codex, Qoder CLI, Qwen Code, and OpenClaw. A parameter called reasoning_effort lets users choose between three levels that trade speed for thoroughness.

Qwen3.8-Max isn't the only piece Alibaba has added recently. A few weeks ago, the team introduced Qwen-Image-3.0, an image generator for information-dense layouts that handles inputs up to 4,500 tokens and renders readable text as small as ten pixels. At the other end of the scale, Alibaba continues to push small open models like Qwen3.6-35B-A3B, which has 35 billion total parameters but only activates three billion at a time.
同じ出来事を4媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み