Moonshot の Kimi K3 が Modal で利用可能に
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
Moonshot が公開した 2.8 兆パラメータの多機能モデル「Kimi K3」を、Modal が同日サポートを開始し、高速推論と専用リソース提供を実現した。
AI深層分析を開く2026年8月4日 01:57
AI深層分析
キーポイント
Kimi K3 モデルの公開
Moonshot は 2.8 兆パラメータ、100 万トークンのコンテキストウィンドウを備え、ネイティブビジョン機能を持つ多機能モデル「Kimi K3」を発表した。
Modal による即日サポート
Modal は Moonshot と vLLM とのパートナーシップにより、リリース当日から Kimi K3 の提供を開始し、トークンベース課金と専用キャパシティの両方を可能にした。
DFlash 推論技術の統合
Modal は Kimi K3 のアーキテクチャに合わせて調整されたカスタムトレーニング済み DFlash スペキュレーターを併用し、1 秒あたり 460 トークンの推論速度を達成した。
Kimi K3の性能とアーキテクチャ
Kimi K3は公開知能指数で最強のオープンモデルであり、2.8Tパラメータと1Mトークンのコンテキストウィンドウを備えた混合専門家トランスフォーマーである。
実行効率のための最適化
MXFP4重みとMXFP8活性化による量子化学習や、vLLMへの新しい実装の貢献により、大規模なハードウェアでもスループットを維持して動作可能となっている。
重要な引用
Moonshot released Kimi K3, a 2.8 trillion parameter multimodal model with a 1M token context window and native vision.
Modal runs it at 460 tokens per second, on release day.
Kimi K3 is the strongest open model on public intelligence indexes, fourth overall in a leaderboard dominated by closed source models.
It's a lot of unglamorous work in service of people who aren't Moonshot, and it's the reason a 3T-class open model is servable at all.
編集コメントを表示
編集コメント
Moonshot の Kimi K3 は、パラメータ数とコンテキストウィンドウの規模において業界最高水準を維持しており、Modal がそれを即日サポートしたことは開発者にとって大きな利点となる。特に DFlash 技術との組み合わせによる高速推論は、実運用におけるコスト対効果を高める重要な要素である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Moonshot が「Kimi K3」をリリースしました。これは 2.8 トリオンパラメータを持つマルチモーダルモデルで、100 万トークンのコンテキストウィンドウとネイティブビジョン機能を備えています。Modal ではリリース当日から、このモデルを秒間 460 トークンの速度で実行可能です。
私たちは Moonshot と vLLM と連携し、リリース初日からサポートを開始しました。これにより、Kimi K3 はトークン課金制の Shared API で利用可能になったほか、専用リソースが必要な場合は Auto Endpoint でも提供しています。さらに、K3 のアーキテクチャに最適化された独自トレーニング済みの DFlash スペキュレーターも用意しました。
今すぐ試すか、なぜこのモデルとそのアーキテクチャが重要なのか、以下でお読みください。
フロンティア、オープン、高速。そのうち 3 つを選べ。
Kimi K3 は、パブリックな知能指数において最強のオープンモデルです。クローズドソースモデルに支配されたリーダーボードでは総合 4 位を記録しています。
K3 は Mixture-of-Experts (MoE) トランスフォーマーで、総パラメータ数は 2.8T です。トークンごとに 896 個のエキスパートのうち 16 個が活性化し、100 万トークンのコンテキストウィンドウとネイティブビジョン機能を搭載しています。このモデルは長期にわたるエージェント作業を目的として設計されており、Moonshot は開発中にチームのカーネル最適化作業の大半を早期バージョンに任せることで実証しました。
Kimi Delta Attention はシーケンスが長くなるにつれてアテンションのコストを抑え、Attention Residuals は深い層が直下の層だけでなく、より前のアテンション出力にもアクセスできるようにします。この 2 つの技術により、K2 と比較して約 2.5 倍のスケーリング効率を実現しています。
提供されるのは、実行が容易なモデルではありません。Moonshot 社もこの課題に多くのアーキテクチャを割いており、SFT(Supervised Fine-Tuning)段階から MXFP4 重みと MXFP8 活性化を用いた量子化対応トレーニングを実施しました。その結果、幅広いハードウェアで動作可能となり、大規模スケールでもスループットが低下しないよう専門家の並列処理を再調整しています。
KDA が従来のプレフィックスキャッシュの仕組みを破綻させることが判明した際、Moonshot 社は新しい実装を開発し、リリース前に vLLM に貢献しました。これは Moonshot 以外のユーザーのために地道に行われた仕事であり、3T クラスのオープンモデルが実際に提供されるに至った理由です。
Modal のエンドポイントでは、K3 の形状に最適化されたカスタム DFlash スペキュレーターに対して Day 0 サポートを提供し、さらに一歩踏み込みました。K3 はタスクあたり大量のトークンを生成するため、ユーザーが待機している時間の大部分はデコード時間に占められますが、スペキュレーションはこのデコードを高速化します。
エージェントワークロードにおいては、これは決定的な差となります:
- 対話速度が 360% 向上(1 秒あたり 100 トークンから 460 トークンへ)
- スループットが 88% 増加(GPU あたり 80 万 TPM から 150 万 TPM へ)

フロンティアモデルにはフロンティアな推測が必要
K3 は、私たちがこれまでトレーニングした中で最大の対象となるスペキュレーターです。
DFlash などのカスタムドラフトモデルアーキテクチャは、慎重にトレーニングすることで、本番環境での推論を劇的に高速化できます。教師モデルが対象モデルそのものとなるため、DFlash のトレーニングは主にデータ生成の問題となります。
最終的な実行では、32 個の B300 ノードを使用しました。そのうち 28 個は K3 を TP8 で稼働させて隠れ状態を生成し、残りの 4 個でそれらに対してドラフトモデルのトレーニングを行いました。
このドラフターには多くの設計上の判断が背後にあり、今後の投稿でさらに詳しく共有できることを楽しみにしています。その間、なぜ私たちがスペキュレイティブ・ディコーディング(推測的デコード)に全幅を信頼しているのか については、すでに詳細な記事を書いています。
Modal で Kimi K3 を実行する
Kimi K3 は本日、トークンベースの課金体系を持つ OpenAI 互換の共有 API または専用 Auto Endpoint として利用可能です。
Modal の標準オファーの対象となっており、毎月 $30 の無料コンピューティングリソースが付与されます。そのため、Shared API を通じて K3 を毎月無料で使い続けることができます。
今日から試して、ぜひ感想をお聞かせください!
原文を表示
Today Moonshot released Kimi K3, a 2.8 trillion parameter multimodal model with a 1M token context window and native vision. And Modal runs it at 460 tokens per second, on release day.
We partnered with Moonshot and vLLM on day zero support, making K3 available with token-based pricing on our Shared API, and as an Auto Endpoint for dedicated capacity, alongside a custom-trained DFlash speculator tuned to K3's architecture.
Try it now or read on for why we think this model, and its architecture, matter.
Frontier, open, fast. Pick three.
Kimi K3 is the strongest open model on public intelligence indexes, fourth overall in a leaderboard dominated by closed source models.
K3 is a mixture-of-experts transformer: 2.8T total parameters, 16 of 896 experts active per token, a 1M token context window, and native vision. It's built for long-horizon agentic work, and Moonshot put that to the test internally, handing an early version most of the team's kernel optimization work during development. Kimi Delta Attention holds down the cost of attention as sequences grow, and Attention Residuals let deeper layers reach back to earlier attention outputs rather than only the layer below, which together give roughly 2.5x the scaling efficiency of K2.
What it doesn't give you is a model that's easy to run, and Moonshot spent a lot of the architecture on that problem too. They did quantization-aware training from the SFT stage onward with MXFP4 weights and MXFP8 activations, so the model runs on a wide range of hardware, and they rebalanced expert parallelism to keep throughput up at large scales. When KDA turned out to break conventional prefix caching, they wrote a new implementation and contributed it to vLLM ahead of the release. It's a lot of unglamorous work in service of people who aren't Moonshot, and it's the reason a 3T-class open model is servable at all.
For endpoints on Modal, we took this even further with Day 0 support for a custom DFlash speculator tuned to K3's shape. K3 generates a large number of tokens per task, which means most of the time a user spends waiting is decode time, and decode is what speculation speeds up.
On agentic workloads, that's a huge difference:
- 360% faster interactivity (from 100 to 460 tokens per second)
- 88% higher throughput (from 800k to 1.5 million TPM per GPU)

Frontier models need frontier speculation
K3 is the largest target we've trained a speculator against.
Custom draft model architectures like DFlash can drastically accelerate production inference, especially when carefully trained. DFlash training is mainly a data generation problem, since the teacher is the target model. The final run used 32 B300 nodes: 28 running K3 at TP8 to produce hidden states, 4 training the draft model against them.
There were a lot of design decisions behind this drafter, which we look forward to sharing more about in a forthcoming post. In the meantime, we've written at length about why we're all-in on speculative decoding.
Run Kimi K3 on Modal
Kimi K3 is available today as an OpenAI compatible Shared API with token-based pricing, or as a dedicated Auto Endpoint.
It's covered by Modal's standard offer: $30 of free compute every single month, so you can keep using K3 on the Shared API for free, month over month.
Try it today and let us know what you think!
AI算出
主要ニュースainew評価高い
記事は特定の AI モデル(Kimi K3)のプラットフォーム展開と技術的特徴(MoE アーキテクチャ、DFlash スペキュレーターなど)を具体的に報じており、新規性の高い一次情報として評価できる。ただし、対象が中国企業およびグローバルクラウドプラットフォームであり、日本固有の導入事例や規制情報は含まれていないため、日本の関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
同じ出来事を3媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み