アリババ、270億パラメータのQwen 3.8をApache 2.0で公開
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The New Stack AI
アリババは270億パラメータのQwen3.8モデルをApache 2.0ライセンスで公開し、同社ベンチマークではAnthropicのOpus 4.6と同等以上の性能を示すと発表している。
AI深層分析を開く2026年8月15日 03:21
AI深層分析
キーポイント
高性能なローカル実行可能モデルの公開
アリババは270億パラメータのQwen3.8をApache 2.0ライセンスで公開し、高性能なMacBook ProやMac Studioでのローカル実行が可能であることを示した。
競合モデルとの性能比較
アリババのベンチマークによると、このモデルはAnthropicのOpus 4.6(Max設定)と同等かそれ以上の性能を持ち、特にコーディングや知識作業で優れた結果を記録している。
マルチモーダル機能の統合
同モデルは画像だけでなく動画を含むビジョン機能を備えており、ローカル環境での多様なタスク処理に適した選択肢となっている。
ベンチマークと実使用の乖離に関する注意
アリババが示す数値はオリジナルチェックポイントに基づくものであり、実用化される量子化版では性能低下が生じる可能性や、エージェント機能での過剰思考といった課題も指摘されている。
小規模モデルの未公開と推論速度
Alibaba は現在、Qwen3.8-27B のより軽量な混合専門家(MoE)版をまだリリースしていない。この種のモデルはトークンあたりの活性化パラメータ数を減らし、密度の高い 27B リリースよりも高速に動作する可能性がある。
重要な引用
according to Alibaba's benchmarks, the model's performance is in the same league as Anthropic's Opus 4.6 running at its Max setting
On quite a few benchmarks, it even outperforms Anthropic's former flagship model... especially when it comes to some computer use, coding and knowledge work tests.
Alibaba notes that the model significantly outperforms the previous-generation Qwen 3.7-Plus, especially in coding and knowledge work tasks
That's not what most people will run.
編集コメントを表示
編集コメント
アリババが公開したQwen3.8-27Bは、ローカル環境での高品質なAI利用を現実的なものにする重要な一歩である。ただしベンチマーク値と実運用時の性能差に注意し、ハードウェア要件や量子化の影響を慎重に評価する必要がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

アリババは先月、パラメータ数 2.4 兆の「Qwen3.8」モデルのオープンウェイト版を公開しました。これは非常に大規模なモデルであり、ベンチマーク結果からは米国のクローズド型最前線モデルと直接競合するレベルであることがわかります。しかし、それ以上に注目すべきは、アリババが金曜日に Qwen 3.8 の密集型(Dense)バージョン、パラメータ数 270 億のモデルを Apache 2.0 ライセンスの下で公開した点です。
このモデルは、スペックの高い MacBook Pro や Mac Studio などのローカル環境でも動作可能です。実際に実行する価値があるかもしれません。アリババが発表したベンチマークによると、このモデルのパフォーマンスは、Anthropic の「Opus 4.6」を最大設定で稼働させた場合と同等のレベルにあるからです。
多くのベンチマークにおいて、このモデルは Anthropic のかつてのフラッグシップモデルをも上回っています。同モデルは昨年 2 月に発表され、半年前まで最先端技術でしたが、特にコンピューター操作、コーディング、知識処理に関するテストではその差が顕著です。
さらに、このモデルには動画を含むビジョン機能も備わっており、ハードウェアの要件を満たすのであれば、ローカル利用における非常に魅力的な選択肢となります。
注意点
いつものことですが、ベンチマーク結果は必ずしも実世界でのパフォーマンスを反映するわけではありません。特にエージェントとしての活用においては、モデルそのものよりも、それを実行する環境(ハネス)の方が重要になるケースもあります。初期の報告では、このモデルが過度に思考してしまう傾向があるという指摘もなされています。
また、Alibaba のベンチマーク結果は、ローカルユーザーが実際に使用する量子化バージョンではなく、元のチェックポイントに基づいている可能性もあります。量子化はコンシューマー向けハードウェアでのモデル運用を現実的なものにする一方で、性能と品質の間には必ずトレードオフが生じます。
それでも全体として、ローカル環境で実行可能なモデルとしては非常に印象的な結果です。
imageCredit: Alibaba.
Alibaba によると、このモデルは前世代の Qwen 3.7-Plus を大きく上回り、特にコーディングや知識処理タスクでの性能が顕著です。DeepSWE エージェント型コーディングベンチマークでは 14.2 ポイントから 42.2 ポイントへと大幅な向上を達成しました。
これは依然として最前線のモデル(および newly released なオープンソースモデルである GLM-5.3 など)には及びませんが、Google のミドルレンジモデルである Gemini 3.6 Flash でさえこのテストで 49% に留まっており、そのモデルをローカルノートパソコンで実行するのは現実的ではありません。
imageCredit: Alibaba.
Alibaba はさらに、Meta が最近リリースした同規模のローカルモデルである Muse Glimmer-30B とも比較しています。Qwen3.8-27B は、両モデルのスコアが報告されているすべてのテストで Glimmer を上回っています。
どうやら「Glimmer」という名前は、このモデルにふさわしいものでしたね。
Qwen 3.8 27B をローカル環境で実行するために必要なハードウェア
速度もまた重要な課題です。現時点では、Alibaba は Qwen3.8 27B の軽量版である MoE(Mixture of Experts)モデルをまだ公開していません。もしそのようなモデルがあれば、トークンあたりの活性化パラメータ数を減らすことで、密度の高い 27B モデルよりも高速に動作させることが可能になります。実際、同社は以前 Qwen 3.6-35B モデルにおいて、この手法で高速化を実現しています。
量子化されていない Qwen3.8-27B のリポジトリは、推論ランタイムやコンテキストキャッシュを考慮する前にすでに 55.6GB に達します。これは多くの人が実際に運用できるサイズではありません。
Mac ユーザーにとっては朗報です。Apple シリコン向けのコミュニティ作成による MLX 変換版が既に利用可能です。4 ビット量子化版は約 16.1GB、8 ビット版は 29.5GB です。これにより、統一メモリ 32GB を搭載した Mac で、中程度のコンテキスト長であれば 4 ビット版を動作させるのに十分な環境が整います。もし 48GB や 64GB のメモリを搭載していれば、より高い精度や長いプロンプトに対応する余地も十分にあります。
デフォルトでは、このモデルは 262,000 トークンのコンテキスト長をサポートしています。さらに、YaRN メソッドを使用することで、100 万トークンまで拡張可能です(Alibaba がホストする本番環境版でも同様に拡張される見込みです)。ただし、これはデスクトップ環境でフルのコンテキストウィンドウを利用できることを意味しません。プロンプトが長くなるにつれて、キーバリューキャッシュのメモリ消費量が増大するためです。
この記事は元々 The New Stack に掲載された「Alibaba の新モデルは、あなたのラップトップ上で Opus 4.6 レベルのパフォーマンスを発揮する」というタイトルのものです。
原文を表示

Alibaba recently made the open weights of its 2.4 trillion parameter Qwen3.8 model available. That’s a massive model, and its benchmarks put it in direct competition with closed frontier models from American labs. But what’s maybe even more interesting is that on Friday, Alibaba also made a dense 27 billion parameter version of Qwen 3.8 available under the Apache 2.0 license.
That’s a model that you could run locally on a well-specced Macbook Pro or Mac Studio, for example — and it might be worth doing so, because according to Alibaba’s benchmarks, the model’s performance is in the same league as Anthropic’s Opus 4.6 running at its Max setting.
On quite a few benchmarks, it even outperforms Anthropic’s former flagship model, which was state-of-the-art six months ago when it launched in February, especially when it comes to some computer use, coding and knowledge work tests.
Add to that the fact that the model also has vision capabilities — including videos — and this becomes a compelling option for local usage (assuming you have the hardware to support it).
Caveats
As usual, the caveat here is that benchmarks don’t always measure how well a model performs in the real world, and for agentic use cases, the harness they run in can be as important as the model itself. Early reports say the model tends to overthink, for example.
It also looks like Alibaba’s benchmarks describe the original checkpoint, not the quantized versions most local users will run. Quantization may make a model practical on consumer hardware, but there is always a quality tradeoff.
Still, overall, these are impressive results for a model that you could run locally.
imageCredit: Alibaba.
Alibaba notes that the model significantly outperforms the previous-generation Qwen 3.7-Plus, especially in coding and knowledge work tasks, and some of the performance jumps are quite impressive, including a step up on the DeepSWE agentic coding benchmarks from 14.2 points to 42.2.
That’s still behind the current frontier (and other open models like the newly released GLM-5.3), but even Google’s mid-tier Gemini 3.6 Flash only got to 49% on this test, and you’re not running that model on your laptop anytime soon.
imageCredit: Alibaba.
Alibaba also compares the model with Meta’s Muse Glimmer-30B, another recently released local model in roughly the same size class. Qwen3.8-27B leads it on every test for which Alibaba reports scores for both models.
Glimmer, it seems, was the right name for that model.
The hardware you need to run Qwen 3.8 27B locally
Speed is another issue, of course. As of now, Alibaba hasn’t released a smaller Qwen3.8 27B mixture-of-experts sibling. Such a model could activate fewer parameters per token and run faster than the dense 27B release. It previously did so for the Qwen 3.6-35B model, for example.
The unquantized Qwen3.8-27B repository is 55.6GB, before accounting for the inference runtime and context cache. That’s not what most people will run.
The good news for Mac users is that community-created MLX conversions for Apple silicon are already available. The 4-bit version is about 16.1GB, while the 8-bit version is 29.5GB. That makes a Mac with 32GB of unified memory a resonable platform for running a 4-bit version at a moderate context length. If you have it, 48GB or 64GB would obviously provides considerably more room for higher precision or longer prompts.
By default, the model supports 262,000 tokens of context, which can be extended (and will be by Alibaba for its hosted production version) to 1 million tokens by using the YaRN method. That doesn’t mean you can use that full context window on your desktop (the key-value cache consumes more memory as the prompt grows).
The post Alibaba’s new model promises Opus 4.6-level performance on your laptop appeared first on The New Stack.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み