Modal、Qwen3.8-2.4T-A95B の提供を開始し高速化を強化
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
Modal は Qwen3.8-2.4T-A95B のオープンウェイトモデルを公開し、SGLang と DFlash 仕様に最適化したスペキュレーターを搭載した Auto Endpoints で即日サポートを開始した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月13日 12:46
AI深層分析
キーポイント
新モデルの公開とプラットフォーム対応
Qwen3.8-2.4T-A95B がオープンウェイトとしてリリースされ、Modal の Shared Endpoint で利用可能になった。
性能向上の具体領域
同モデルは Qwen 3.7 を上回る性能を示し、コーディング、業務タスク、研究、および長期ホライズンタスクにおいて大幅な改善が見られる。
技術スタックの最適化
Modal は Qwen と事前協力し、SGLang を基盤としつつ、Qwen3.8 の形状にチューニングされた DFlash 仕様のスペキュレーターを採用して即日サポートを実現した。
カスタム DFlash スペキュレーターによる推論速度の向上
モデルを実行可能にするだけでなく、DFlash スペキュレーターを活用して推論を高速化している。
スペキュレーションの精度向上とトレーニングデータの調整
Max のコーディングやリサーチにおける改善によりツール呼び出しが増えたため、トレーニングデータにその傾向を組み込んでトークンの採択率を高めた。
重要な引用
just launched as an open weights model, and it's now available on Modal
substantial improvement across coding, work, research, and long-horizon tasks
backed by SGLang and a custom DFlash speculator tuned to Qwen3.8's shape
Speculation is all you need
編集コメントを表示
編集コメント
Qwen3.8 の性能向上は具体的なタスク領域に焦点を当てており、実務での利用価値が高い。Modal が DFlash 仕様を独自にチューニングして即日サポートを提供した点は、インフラ面での迅速な対応力を示している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
「Qwen3.8-2.4T-A95B」がオープンウェイトモデルとして公開され、Modal での利用が可能になりました。
Qwen 3.7 を上回る性能を誇り、コーディング、業務支援、研究、長期タスクの各領域で大幅な改善が見られます。
リリースに先立ち、SGLang と Qwen3.8 の形状に合わせて最適化された独自仕様の DFlash 推測器を基盤とした Modal Auto Endpoints への Day Zero サポートを実現するため、Qwen チームと連携して開発を行いました。
Shared Endpoint として今すぐお試しください。

カスタム DFlash 推測による推論速度の向上
モデルを動かせるようになることと、それを高速化することは別問題です。その解決策として再び DFlash 推測器に注目しました。私たちが常に強調している通り、Speculation is all you need です。
推論器が価値を発揮するのは、ターゲットモデルがドラフトされたトークンを受け入れる場合に限られます。その可否は、推測側が予測するシーンスimilarなものを学習データで見たことがあるかどうかにかかっています。Max のバージョンではコーディング、研究、業務支援の能力が向上しており、これらはツール呼び出しを多用する傾向があるため、トレーニングデータにおいてこれらの要素を強化し、トークンの受け入れ率を引き上げました。
今すぐお試しください
Qwen3.8-2.4T-A95B(テキスト専用)は、今後1ヶ月間、OpenAI 互換の Shared Endpoint として提供されます。利用料金はトークンベースで課金されます。
原文を表示
Qwen3.8-2.4T-A95B just launched as an open weights model, and it’s now available on Modal.
Over Qwen 3.7, the latest model sees substantial improvement across coding, work, research, and long-horizon tasks.
We worked with Qwen ahead of the drop to bring day zero support to Modal Auto Endpoints, backed by SGLang and a custom DFlash speculator tuned to Qwen3.8’s shape.
Try it out now as a Shared Endpoint.

Speeding up inference with custom DFlash speculation
Just getting the model running is one thing, making it fast is another. For this, we once again turn to a DFlash speculator model because—say it with us now—Speculation is all you need.
A speculator only earns its keep when the target accepts the tokens it drafts, and acceptance comes down to whether the drafter has seen sequences like the ones it's predicting. Because of Max’s improvements in coding, research, and work (things that tend to use more tool calls) we leaned into that in our training data to increase accepts.
Try it now
Qwen3.8-2.4T-A95B (text only) is available for the next month as an OpenAI compatible Shared Endpoint with token-based pricing.
同じ出来事を3媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み