Moonshot が超大規模モデル「Kimi K3」を発表
月之暗面(Kimi)が世界初のオープン 3T クラスモデル「Kimi K3」を発表し、2.8兆パラメータと100万トークンのコンテキストウィンドウを特徴とする新アーキテクチャを公開した。
キーポイント
世界初のオープン 3T クラスモデルの登場
Kimi K3 は2.8兆パラメータ規模を持ち、現在まで存在したオープンソースモデルの中で最大規模となる。
革新的なアーキテクチャと効率化
「Kimi Delta Attention」と「Attention Residuals」を採用し、MoE 構造の最適化により計算資源から知能への変換効率が K2 比で約 2.5 倍向上した。
長文コンテキストとネイティブビジョン機能
100万トークンのコンテキストウィンドウをサポートし、コード生成、知識作業、推論において先端的なパフォーマンスを発揮する。
段階的な公開ロードマップ
即時利用可能だが、完全なモデル重みは 2026 年 7 月 27 日にリリースされ、技術レポートも併せて公開される予定である。
アーキテクチャの革新と効率化
KDA(Kimi Delta Attention)とAttnRes(Attention Residuals)を採用し、MoEのスパーシティを拡大することで、K2比で計算効率を約2.5倍向上させました。
長期的なコーディング能力
最小限の人間の手助けで長時間にわたるエンジニアリングセッションを維持し、大規模リポジトリのナビゲーションやターミナルツールの操作が可能です。
視覚推論とコード生成の融合
スクリーンショットやビジュアル情報を活用してゲーム開発、フロントエンド、CADなどのタスクを最適化する能力に優れています。
重要な引用
Kimi K3 is the world's first open 3T-class model
It marks the latest step in Kimi's sustained push at the scaling frontier: for nine of the past twelve months, Kimi models have set the upper bound of open-model sizes.
The full model weights will be released by July 27, 2026.
Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth.
Operating with minimal human oversight, it can sustain long engineering sessions, navigate massive repositories, and orchestrate terminal tools.
Kimi K3 can build a coherent end-to-end compiler — from DSL frontend and IR passes to PTX codegen and runtime — rather than isolated kernels
影響分析・編集コメントを表示
影響分析
この発表は、オープンソース AI モデルのスケール競争において決定的な転換点となり、従来はクローズドモデルに限定されていた「3T クラス」の性能を一般開発者が利用可能にする可能性を開いた。特に計算効率の向上と長文コンテキストの実装は、複雑な推論タスクや大規模コード生成におけるオープンソースの競争力を劇的に高めるものであり、業界全体の技術基準を再定義する影響を持つ。
編集コメント
月之暗面(Kimi)がオープンソースモデルのスケール限界を再び塗り替え、2.8T パラメータという驚異的な規模で登場しました。ただし、完全な重み公開まで 2026 年と長期化している点は、大規模モデルの品質維持やインフラ整備における現実的な課題を浮き彫りにしています。
本日、当社で最も能力の高いモデル「Kimi K3」をご紹介します。Kimi K3 は、当社独自の「Kimi Delta Attention」と「Attention Residuals」を基盤に構築され、ネイティブな視覚機能と 100 万トークンという超長文コンテキストウィンドウを備えた、2.8T パラメータのモデルです。これは世界初のオープン化された 3T クラスのモデルであり、長期にわたるコーディング、知識作業、推論といった最先端の知能タスクに対応するために設計されています。
全体的な性能においては、依然として Claude Fable 5 や GPT 5.6 Sol といった最強のクローズドモデルには及ばないものの、Kimi K3 は当社の評価スイートにおいて最先端レベルのパフォーマンスを発揮し、テストされた他のモデルを常に上回る結果を示しました。
Kimi K3 は本日より、Kimi.com、Kimi Work、Kimi Code、および Kimi API で利用可能です。リリース当初はデフォルトで最大限の思考努力(max thinking effort)が適用され、低・高効率モードについては今後のアップデートで順次導入される予定です。現在、推論パートナーやオープンソースのメンテナーと緊密に連携し、技術仕様の整合性を図りながらエコシステム全体での安定した展開を進めています。モデルの重み(weights)は 2026 年 7 月 27 日までに公開されます。アーキテクチャ、トレーニング手法、評価結果の詳細については、Kimi K3 の公式技術レポートと併せて発表いたします。
オープン化された 3T クラスのモデル
Kimi K3 は、2.8 トリオンパラメータを達成した初のオープンモデルです。これは Kimi がスケーリングの最前線に注力し続けてきた最新の一歩であり、過去 12 ヶ月のうち 9 ヶ月で Kimi モデルがオープンモデルのサイズ上限を更新し続けてきました。
Kimi K3 は、情報フローをシーケンス長とモデル深さ全体で最適化するために設計された 2 つのアーキテクチャ改良である「Kimi Delta Attention (KDA)」と「Attention Residuals (AttnRes)」を基盤としています。さらに、Mixture of Experts (MoE) のスパース性を強化し、Stable LatentMoE フレームワークと組み合わせることで、896 個のエキスパートのうち 16 個を有効に活用します。これら構造上の改良に加え、トレーニング手法やデータレシピも洗練された結果、Kimi K2 と比較してスケーリング効率が約 2.5 倍向上しました。そのおかげで、計算リソースをより効果的に知能へと変換できるようになっています。
αwKDAαwStable LatentMoEαwGated MLAαwStable LatentMoEwα3×1×Block n−1Block n−2Block n−3EmbeddingRouterLinear12123NNormLinearShared ExpertRouted ExpertLinearConvL2LinearConvL2LinearConvσσLinearσKimi Delta AttentionNormLinearOutputKimi K3 architecture: the Stable LatentMoE and KDA modules (left), the AttnRes operation α (top right), and the Block Attention Residuals backbone (right).
コーディング
Kimi K3 は、長期にわたるコーディングタスクにおいても強力なパフォーマンスを発揮します。人間の監視を最小限に抑えながら、長時間のエンジニアリングセッションを継続し、巨大なリポジトリを自在にナビゲートし、ターミナルツールを統括して操作できます。
Kimi K3 は、ソフトウェアエンジニアリングと視覚推論を融合したタスクにおいても卓越した能力を発揮します。スクリーンショットや画像情報を活用して、ゲーム開発、フロントエンド実装、CAD 設計の最適化を支援します。
以下の事例では、Kimi K3 のコーディング能力がどのようにオープンエンドなソフトウェア作成や科学研究に活かされているかを示しています。
カーネル最適化
各モデルを独立した同一のサンドボックス環境でテストし、GPU カーネルの最適化能力を検証しました。AttnRes、KDA、および 512 ヘッド次元の MLA カーネルなど 4 つのタスクについて、最大 24 時間かけてプロファイリング、書き換え、ベンチマークを行いました。対象ハードウェアは NVIDIA H200 と、他社製の GPGPU です。
その結果、Kimi K3 は Fable 5(フォールバック機能あり)と互角の性能を示し、Opus 4.8、GPT 5.6 Sol、GPT 5.5 を大きく上回りました。
Claude の Fable 5 は第三者機関が評価したものであり、その結果にはフォールバック動作が含まれている可能性があります。対象モデルの多くで、数値的な許容範囲内に収まる小さな精度を犠牲にする短絡的な処理が見られるケースもありました。なお、GPGPU とはグラフィック描画以外の計算処理に用いられる汎用 GPU を指します。
Kimi K3 の開発後期には、その初期バージョンがチームのカーネル最適化作業の大部分を担っていました。
GPU コンパイラ開発
さらに、Kimi K3 がゼロから GPU プログラミングシステムを構築できるか検証しました。その結果、Kimi K3 は「MiniTriton」という独自の実装を開発しました。これは MLIR 上にタイルレベルの IR レイヤーを持ち、最適化パスと PTX コード生成パイプラインを備えた、コンパクトな Triton 互換コンパイラです。
サポートされているルーフラインベンチマークでは、MiniTriton は Triton や torch.compile と同等かそれ以上の性能を発揮し、特定のワークロードでは Triton を上回る結果となりました。マイクロベンチマークを超えて、NanoGPT のトレーニング全体でも安定した収束を示しました。損失曲線は参照値とほぼ同じ軌跡をたどり、わずかな乖離のみが見られる程度です。これは、現実的なワークロードにおいてフルパイプラインが正しく機能していることを裏付ける結果です。
これらの成果は、Kimi K3 が個別のカーネルではなく、DSL フロントエンドから IR パス、PTX コード生成、そしてランタイムに至るまで一貫したエンドツーエンドのコンパイラを構築できる能力を持っていることを示しています。特に、ゼロから実装された Tensor Core 用パスは、すでに Triton の広範に最適化されたスタックと互角の性能を誇ります。
ゲーム開発とデジタルクリエイション
Kimi K3 は、強力な 3D 推論能力、コーディングスキル、そしてビジョン機能を組み合わせることで、コンセプトや画像、動画から完全にプレイ可能なインタラクティブな体験へと変換します。コードとライブスクリーンショットの間をシームレスに反復処理し、出力を即座に確認して改善できる「Vision in the Loop」を実現しています。
チップ設計
Kimi K3 の初期概念実証の一例として、同社の独自アーキテクチャを基盤とするナノモデル向けにチップ設計を行いました。単なる 48 時間の自律実行において、K3 はオープンソースの EDA ツールと Nangate 45nm ライブラリを活用し、チップの構築から最適化、検証までを一貫して完了させました。
面積は 4mm²未満で、タイミングは 100MHz でクロージングされ、シミュレーション上では毎秒 8,700 トークンを超えるデコードスループットを維持します。標準セル数は 146 万個、SRAM は 0.277MB、そして量子化解除機能を統合した INT4 MAC アレイを搭載しています。モデルがモデルのために設計したこのチップは、Kimi K3 の長期にわたる自律的なエージェント能力を如実に示すものです。
コーディングによる研究支援
Kimi K3 は科学文献と実行可能なコードをつなぐ役割を果たし、複雑な計算研究ワークフローの自動実装、検証、分析を実現します。
ある事例では、熟練した研究者であれば通常 1〜2 週間を要する作業を、Kimi K3 は約 2 時間で完了させました。計算天体物理学における「I–Love–Q」の普遍関係の再現において、同モデルは 20 編以上の論文を検証・相互検証し、数値パイプライン全体を実装。さらに 300 種類以上の状態方程式を評価し、公開された式の不整合を特定しました。その結果、3,000 行を超える Python コードを生成し、結果を検証するためのインタラクティブな HTML ダッシュボードも作成しています。
知識労働
Kimi K3 は、エンドツーエンドの知識作業を推進します。公開ベンチマークを超えて、Kimi K3 (max) は、実際のユーザーエージェントワークフローで観察される反復パターンや課題に基づいた内部評価においても、一貫した性能向上を示しています。このように異なる生産指向のワークフロー全体で一貫して優位性を示すことは、Kimi K3 のエージェント型知識作業能力が全体的に向上していることを反映しています。
対話型可視化を活用した研究
以下は、Kimi Work の Kimi K3 が金融コンサルティングや科学研究の分野で生成できる事例の一部です。
ケース 1: AI ASIC 産業に関する 42 年分のインタラクティブな研究ウェブサイト
掘り下げて閲覧可能なインタラクティブな調査レポートです。AI ASIC 業界の 42 年間の歴史を、再帰的な自己改善を 120 回以上繰り返すことで構築しました。Kimi K3 は証拠に基づき、カスタムチャートやアニメーション図解、インタラクティブな視覚的ナラティブ(物語)へと変換します。このレポート作成には、ウェブ検索・データ取得が 2,800 件超、ターミナルからのデータ抽出が 1,100 件超に及び、87 四半期報告書と 99 のオリジナル PDF を含む 11,000 ページ以上の情報を跨いで行われました。
ケース 2: 融合産業の研究
タイムライン、ファネルチャート、レンジバーチャート、ガントチャート、出版品質のスライドなど、インタラクティブな可視化を備えたコンサルティングスタイルの産業レポートです。
ケース 3: GWTC-5 重力波分析
20 以上の並列サブエージェントを活用して 391 の重力波イベントを分析し、7 つの科学的な可視化図、2 つの表、そして 10 篇以上の論文を統合した文献レビューを生成しました。
Kimi K3 は、以下に示すように、完全編集可能なヒートマップや年次報告書など、インフォグラフィックスタイルのプレゼンテーション作成においても特に優れた能力を発揮します。
ウィジェットとダッシュボード
Kimi Work では、Kimi K3 との対話をより視覚的かつ持続的なものにするため、「ウィジェット」と「ダッシュボード」の 2 つの新機能を導入しました。ウィジェットを使えば、チャット内で直接インタラクティブなコンポーネントを生成でき、ローカルデータや外部プラグインと連携して継続的な更新も可能です。一方、ダッシュボードは、特定のトピック、プロジェクト、あるいは目標を中心に、ユーザーが最も重視するウィジェットを 1 つの持続的なパーソナライズされたビューにまとめます。
動画編集
Kimi K3 は、テキスト、画像、動画を同一モデル内で理解できるネイティブなマルチモーダルアーキテクチャを持つため、モーションデザインやアニメーション、動画編集において卓越した能力を発揮します。
例えば、K3 は自らのアーキテクチャを解説する 3Blue1Brown 風のモーショングラフィックス explainer(解説動画)を作成しました。これは技術的な概念をアニメーションされた図形とトランジションに変換したものです。
別の事例では、Kimi K3 が 56 のソースクリップから自身のティザービデオを編集しました。クリップの選定、モーションに合わせたカット、フレーム単位のビート同期、音声処理、そして複数回の修正対応までを一手に引き受けたのです。このような高密度なショート動画の制作は、熟練のエディターでも 1〜2 営業日、初心者であれば 3〜5 日を要するのが一般的です。
アーキテクチャとインフラ
Kimi K3 は、Kimi Delta Attention(KDA)と Attention Residuals(AttnRes)を基盤に構築されています。KDA はアテンションの拡張に向けた効率的な基盤を提供し、AttnRes は深さ方向で表現を均一に蓄積するのではなく、必要な表現を選択的に取り出す仕組みです。これら 2 つの技術が組み合わさり、1 兆パラメータを超える規模でも十分にスケール可能なモデルのアーキテクチャの中核を形成しています。
Kimi K3 では、Stable LatentMoE を採用し、896 個あるエクスパートのうち 16 個のみを動的に活性化します。このレベルの高いスパース性において、ルーティングと最適化が最大の課題となります。そこで導入された Quantile Balancing は、ルーターのスコア分布から直接エクスパートの割り当てを決定し、経験則に基づく更新や調整が難しいハイパーパラメータを不要にしました。また、Per-Head Muon は従来の Muon オプティマイザーを拡張し、アテンションヘッドごとに独立して最適化を行うことで、大規模学習における適応的な学習を実現します。さらに、Sigmoid Tanh Unit(SiTU)と Gated MLA を採用することで、それぞれ活性化の制御とアテンションの選択性を向上させています。これらの技術的進歩により、2.8 兆パラメータという巨大なスケールでも安定かつ効率的なトレーニングが可能になっています。
Kimi K3 は SFT(Supervised Fine-Tuning)段階から量子化対応トレーニングを適用し、MXFP4 形式の重みと MXFP8 形式の活性化値を採用することで、幅広いハードウェア環境での互換性を確保しています。
大規模なエキスパート並列処理においてスループットが低下するのを防ぐため、静的形状を使用し、クリティカルパス上のホスト同期を不要とする完全バランス型のエキスパート並列トレーニング手法を導入しました。また、推論効率も高帯域幅の通信ドメインの拡大によって向上するため、64 基以上のアクセラレータを搭載したスーパーノード構成でのデプロイを推奨します。
さらに、KDA(Key-Value Cache Dynamic Allocation)は従来のプレフィックスキャッシュに新たな課題をもたらすため、vLLM コミュニティへ対応実装を提供し、モデル公開と同時にリリースする予定です。プリフィルキャッシュ付きの KDA を活用することで、Kimi K3 の大規模さと長いコンテキスト処理能力を維持しつつ、極めて競争力のあるトークン価格での提供を実現しています。
より詳細な技術情報は、近日公開予定のレポートにて発表いたします。
利用方法
- Kimi K3 エージェント: iOS、Android、HarmonyOS の各モバイルアプリストアから最新の Kimi アプリをダウンロードまたは更新するか、kimi.com にアクセスしてください。
- Kimi K3 ワーク: Windows および Apple シリコン搭載 Mac 向けに、バージョン 3.1.0 以降の最新デスクトップアプリ「Kimi Work」をダウンロードしてご利用ください。
- Kimi K3 でコーディング: ターミナルで Kimi Code を実行し、/model コマンドを使用して Kimi K3 を選択してください。
- Kimi API を利用して構築する:Kimi API プラットフォームにアクセスし、kimi-k3 を選択してください。キャッシュヒット時の入力料金は MTok あたり 0.30 ドル、キャッシュミス時は 3.00 ドル、出力は 15.00 ドルです。Mooncake の非集約推論アーキテクチャを採用しているため、公式 Kimi API はコーディングタスクにおいてキャッシュヒット率が 90% を超えています。
- 組織に Kimi を導入する:Kimi Enterprise では、エンタープライズレベルのデータプライバシーとメンバー管理を提供し、個人アカウントと組織アカウントを完全に分離しています。チーム向けに契約するには、料金ページから「Get Kimi Enterprise」を選択してください。
Full Benchmark Table
Footnotes
以下に報告するすべての Kimi K3 の結果は、推論エフォートを「max」に設定し、温度パラメータ(temperature)を 1.0、top-p を 1.0 に固定した条件下で取得されたものです。ベンチマークによっては、以下の注記にある通り、KimiCode、Claude Code、Codex のいずれかのエージェントハネス下で各モデルが評価されます。
Coding benchmarks
- DeepSWE:Kimi K3 は KimiCode ハネスで評価されました。GLM-5.2 のスコアは GLM-5.2 公式リリースブログ(https://z.ai/blog/glm-5.2)から引用しており、それ以外のすべてのスコアは公式の DeepSWE リーダーボード(https://deepswe.datacurve.ai/)からのものです。同リーダーボードにおいて、Kimi K3 は mini-SWE-agent ハネスを使用した場合に 67.3 のスコアを達成しています。
- Terminal-Bench 2.1。Kimi K3 は「KimiCode」ハーンネスを用いて評価されました。他のモデルについては、各ハーンネスで得られた最高スコアを報告しています:GLM-5.2 は Claude Code(https://z.ai/blog/glm-5.2)、Claude Opus 4.8 と Claude Fable 5 は Terminus 2(https://artificialanalysis.ai/evaluations/terminalbench-v2-1)、GPT 5.5 と GPT 5.6 Sol は Codex(https://openai.com/index/previewing-gpt-5-6-sol/)それぞれで測定した結果です。
- Program Bench。Kimi K3 の評価には「KimiCode」ハーンネスを使用しました。GLM-5.2 のスコアは https://z.ai/blog/glm-5.2 から引用し、それ以外のスコアは https://www.vals.ai/benchmarks/programbench に基づいています。
- SWE Marathon。Kimi K3、Claude Opus 4.8、Claude Fable 5 は「Claude Code」ハーンネスで評価され、GPT 5.6 Sol は Codex ハーンネスで評価されました。GLM-5.2 のスコアは https://z.ai/blog/glm-5.2 から引用しています。
- FrontierSWE。Kimi K3 は「KimiCode」ハーンネス、GPT 5.6 Sol は Codex ハーンネスで評価され、それ以外の結果は https://www.frontierswe.com/ から取得しました。ドミナンススコアは公式の評価スクリプトを用いて生データから再計算したもので、2026 年 7 月 16 日時点の最新値です。
- PostTrain Bench。GLM-5.2、GPT 5.5、Claude Opus 4.8 のスコアは公式の PostTrainBench 結果から引用しています。Kimi K3、Claude Fable 5、GPT 5.6 Sol は、最大推論能力を発揮する状態で公式の Harbor 実装を用いて評価しました。これは 3 回の試行の平均値です。Kimi K3 と Claude Fale 5 は Claude Code ハーネスを、GPT 5.6 Sol は Codex ハーネスを使用しています。Claude Code ハーネスでは、利用ポリシーにより Claude Fable 5 がリクエストを拒否した場合、自動的に Claude Opus 4.8 にフォールバックします。
- MLS Bench Lite。Kimi K3 は KimiCode ハーネスで、GLM-5.2 と Claude モデルは Claude Code ハーネスで、GPT 5.5 と GPT 5.6 Sol は Codex ハーネスでそれぞれ評価されました。
- KCB 2.0。Kimi K3 は KimiCode と Claude Code の両方のハーネスで評価され、GLM-5.2、Claude Opus 4.8、Claude Fable 5 は Claude Code ハーネスを、GPT 5.5 と GPT 5.6 Sol は Codex ハーネスを使用しました。すべてのモデルは最大推論能力で評価されていますが、GPT 5.5 のみ「xhigh」設定を使用しています。
生産性とエージェントベンチマーク
- OfficeQA Pro では、各テストケースでエージェントに PDF コーパス全体が提供されます。PDF はすべて画像としてレンダリングされており、機械可読なテキストは存在しません。
- OfficeQA Pro と SpreadsheetBench 2。Kimi K3、GLM-5.2、Claude Opus 4.8、Claude Fable 5 は Claude Code ハーネスで評価され、GPT 5.5 と GPT 5.6 Sol は Codex ハーネスで評価されました。
- MCP Atlas。すべてのモデルは、Gemini 3.1 Pro をジャッジ(評価者)に用い、500 タスクからなる公開サブセットで 100 ターン制限付きで評価されました。
- AutomationBench。すべてのモデルは、600 タスクの公開サブセット上で評価され、それ以外の点は公式 GitHub の設定に従っています。
- BrowseComp。Claude モデルカードで採用されているコンテキスト圧縮戦略を 30 万トークンでトリガーして適用しました。100 万トークンのコンテキストウィンドウを使用し、コンテキスト管理を行わない場合、Kimi K3 は 90.4 のスコアを記録します。Claude Fable 5、Claude Opus 4.8、GPT 5.6 Sol、GPT 5.5 の結果は、それぞれ https://www.anthropic.com/news/claude-fable-5-mythos-5 および https://openai.com/index/gpt-5-6/ から引用しています。
- GDPval-AA v2 と AA-Briefcase のスコアは、https://artificialanalysis.ai/ から引用しました。
多モーダルベンチマーク
- ZeroBench は公式設定に従い 5 回実行したものを除き、すべての多モーダルスコアは 3 回の平均値です。MMMU-Pro は公式プロトコルに従って評価され、入力順序が保持された上で画像がテキスト入力の先頭に付加されています。
- PerceptionBench。これは原子レベルの視覚知覚能力に焦点を当てた社内ベンチマークです。
限界
- 思考履歴への感応性。K3 は「保存された思考履歴モード」でトレーニングされています。もしエージェントのハーンが、必要なすべての歴史的な思考内容を正しく返却できない場合や、他のモデルとのセッション中に K3 に切り替えられた場合は、生成品質が著しく不安定になる恐れがあります。Kimi Code のように互換性が検証されたハーンを使用し、セッション途中で K3 に切り替えないことを強く推奨します。
- 過度な主体性。K3 のトレーニングでは、長期にわたる課題や難易度の高いタスクへの対応が特に重視されています。その結果、タスク実行中に些細な問題や曖昧なユーザーの意図に出会った際、ユーザーに代わって予期せぬ判断を下す可能性があります。アプリケーションでエージェントが明確に定義された範囲内で動作し、過度な即興を控えることが求められる場合は、システムプロンプトまたは AGENTS.md において K3 に対してより明示的な行動制約を設けてください。
- 全体として非常に競争力のあるモデルであるにもかかわらず、K3 はユーザーエクスペリエンスの面で Claude Fable 5 や GPT 5.6 Sol と比較して、明確な差が見て取れます。
原文を表示
Today, we are introducing Kimi K3 — our most capable model. Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.
While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models.
Kimi K3 is available today on Kimi.com, Kimi Work, Kimi Code, and the Kimi API. At launch, Kimi K3 will use max thinking effort by default, with low- and high-effort modes to be introduced in subsequent updates. We are currently working closely with inference partners and open-source maintainers to align technical details and ensure a reliable rollout across the ecosystem. The full model weights will be released by July 27, 2026. Further details on the architecture, training, and evaluations will be released alongside the Kimi K3 technical report.
An Open 3T-Class Model
Kimi K3 is the first open model to reach 2.8 trillion parameters. It marks the latest step in Kimi's sustained push at the scaling frontier: for nine of the past twelve months, Kimi models have set the upper bound of open-model sizes.
Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth. We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework. Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to Kimi K2, allowing the model to convert compute into intelligence more effectively.
Coding
Kimi K3 has strong long-horizon coding performance. Operating with minimal human oversight, it can sustain long engineering sessions, navigate massive repositories, and orchestrate terminal tools.
Kimi K3 also excels in tasks blending software engineering with visual reasoning — it leverages screenshots and visuals to optimize game dev, frontend, and CAD.
The case studies below show how Kimi K3's coding capability translates into open-ended software creation and scientific research.
Kernel Optimization
We tested the models' capability to optimize GPU kernels. Each model works independently in an identical sandbox, with up to 24 hours to profile, rewrite, and benchmark four tasks spanning AttnRes, KDA, and a 512-head-dimension MLA kernel across NVIDIA H200 and GPGPU from an alternative vendor. Kimi K3 performed competitively with Fable 5 (with fallback) and substantially outperformed Opus 4.8, GPT 5.6 Sol, and GPT 5.5.
Claude Fable 5 was evaluated by a third party, and its results may include fallback behavior. Across most models, some trajectories include small, acceptable precision shortcuts that remain within our numerical tolerance. GPGPU denotes general-purpose GPUs used for computation beyond graphics rendering.
In the late stages of Kimi K3 development, an early version of Kimi K3 handled the majority of the team's kernel optimization works.
GPU Compiler Development
We further tested whether Kimi K3 could build a GPU programming system from scratch. Kimi K3 developed MiniTriton, a compact Triton-like compiler with its own tile-level IR layer over MLIR, optimization passes, and a PTX code-generation pipeline. Across supported roofline benchmarks, MiniTriton delivers performance on par with or better than Triton and torch.compile — beating Triton on certain workloads. Beyond microbenchmarks, MiniTriton sustains end-to-end nanoGPT training with stable convergence, the loss curve closely tracking the reference with only minor divergence — validating the full pipeline on a realistic workload. These results demonstrate that Kimi K3 can build a coherent end-to-end compiler — from DSL frontend and IR passes to PTX codegen and runtime — rather than isolated kernels; its from-scratch Tensor Core path already rivals Triton’s extensively optimized stack.
Game Dev and Digital Creation
Kimi K3 combines strong 3D reasoning, coding, and vision capabilities to turn concepts, images, and videos into fully playable interactive experiences. Kimi K3 achieves true "vision in the loop" by seamlessly iterating between code and live screenshots—instantly seeing and refining outputs.
Chip Design
As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM, and an INT4 MAC array with fused dequantization. A chip built by a model, for a model, reflects K3's long-horizon agentic capabilities.
Coding for Research
Kimi K3 bridges scientific literature and executable code, autonomously implementing, validating, and analyzing complex computational research workflows.
In one case, Kimi K3 completed in about two hours what would typically require one to two weeks of work by an experienced researcher. To reproduce the I–Love–Q universal relations in computational astrophysics, it reviewed and cross-validated 20+ papers, implemented the full numerical pipeline, evaluated 300+ equations of state, identified inconsistencies in published formulas, generated 3,000+ lines of Python code, and produced an interactive HTML dashboard for exploring the results.
Knowledge Work
Kimi K3 advances end-to-end knowledge work. Beyond public benchmarks, Kimi K3 (max) demonstrates consistent gains across our internal evaluations, which are derived from recurring patterns and challenges observed in real-world user-agent workflows. These consistent advantages across distinct production-oriented workflows reflect a broad improvement in Kimi K3's agentic knowledge work capabilities.
Research with Interactive Visualization
Below are a few examples of what Kimi K3 in Kimi Work can produce across financial consulting and scientific research:
Case 1: Interactive 42 years of AI ASIC industry research website
An interactive research report you can drill into: 42 years of the ASIC industry, created through 120+ rounds of recursive self-improvement. Kimi K3 transforms evidence into bespoke charts, animated diagrams, and interactive visual narratives. It pulled data via 2.8k+ web searches/fetches and 1.1k+ terminal data pulls, across 11k+ pages spanning 87 quarterly reports and 99 original PDFs.
Case 2: Fusion Industry Research
A consulting-style industry report with interactive visualizations—including timelines, Funnel Chart, Range Bar Chart, Gantt Charts, and publication-quality slides.
Case 3: GWTC-5 Gravitational-wave Analysis
An analysis of 391 gravitational-wave events using 20+ concurrent subagents, producing 7 scientific visualizations, 2 tables, and a literature synthesis from 10+ papers.
Kimi K3 is also particularly effective at producing infographic-style presentations, such as the fully editable heatmap and annual report shown below:
Widgets and Dashboard
In Kimi Work, we introduce two new features - Widgets and Dashboard - which make interactions with Kimi K3 more visual and persistent. Widgets let you generate interactive components directly within a chat, with connections to local data or external plugins for continuous updates. Dashboard brings the widgets you care about most into one persistent, personalized view organized around a topic, project, or goal.
Video Editing
Kimi K3 excels at motion design, animation, and video editing because its native multimodal architecture understands text, images, and video within the same model.
In one example, K3 created a 3Blue1Brown-style motion-graphics explainer of its own architecture, translating technical ideas into animated diagrams and transitions.
In another, Kimi K3 edited its own teaser video from 56 source clips, handling clip selection, motion-matched cuts, frame-accurate beat synchronization, audio processing, and multiple rounds of revision. A high-density short video like this would typically take an experienced editor one to two working days, or a beginner three to five.
Architecture and Infrastructure
Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). KDA provides an efficient foundation for scaling attention, while AttnRes selectively retrieves representations across depth rather than accumulating them uniformly. Together, they form the architectural backbone of a model designed to scale well beyond the trillion-parameter regime.
Kimi K3 uses Stable LatentMoE, effectively activating 16 of 896 experts. At this level of sparsity, routing and optimization become first-order challenges. Quantile Balancing derives expert allocation directly from router-score quantiles, eliminating heuristic updates and a sensitive balancing hyperparameter, while Per-Head Muon extends Muon by optimizing attention heads independently for more adaptive learning at scale. Sigmoid Tanh Unit (SiTU) and Gated MLA improve activation control and attention selectivity respectively. Together, these advances enable stable and efficient training at the 2.8-trillion-parameter scale.
Kimi K3 applies quantization-aware training from the SFT stage onward, using MXFP4 weights with MXFP8 activations for broad hardware compatibility. To prevent expert imbalance from degrading throughput at large expert-parallel scales, we introduce a fully balanced expert-parallel training method with static shapes and no host synchronization on the critical path. Since inference efficiency likewise benefits from larger high-bandwidth communication domains, we recommend deploying Kimi K3 on supernode configurations with 64 or more accelerators. Finally, as KDA poses new challenges for conventional prefix caching, we have contributed a corresponding implementation to the vLLM community, to be released alongside the model. KDA with prefill cache allows us to serve Kimi K3 at a highly competitive token price despite its scale and long context.
More technical details will be available in our coming report.
Availability
- Kimi K3 Agents: Download or update to the latest Kimi app from your mobile app store, available on iOS, Android, and HarmonyOS, or visit kimi.com.
- Work with Kimi K3: Download the latest Kimi Work desktop app, version 3.1.0 or later, available for Windows and Apple silicon Macs.
- Code with Kimi K3: Run Kimi Code in your terminal and select Kimi K3 using the /model command.
- Build with the Kimi API: Visit the Kimi API Platform and select kimi-k3. Pricing is $0.30/MTok for cache-hit input, $3.00/MTok for cache-miss input, and $15.00/MTok for output. Powered by Mooncake's disaggregated inference architecture, the official Kimi API achieves a cache hit rate above 90% in coding workloads.
- Bring Kimi to your organization: Kimi Enterprise provides enterprise-grade data privacy and member management, with complete separation between personal and organization accounts. Visit the pricing page and select “Get Kimi Enterprise” to subscribe for your team.
Full Benchmark Table
Footnotes
All Kimi K3 results reported below are obtained with the reasoning effort set to 'max', setting temperature = 1.0 and top-p = 1.0. Depending on the benchmark, each model is evaluated under one of three agentic harnesses — KimiCode, Claude Code, or Codex — as specified in the notes below.
Coding benchmarks
- DeepSWE. Kimi K3 is evaluated with the KimiCode harness. The GLM-5.2 score is taken from the GLM-5.2 release blog (https://z.ai/blog/glm-5.2); all remaining scores are from the official DeepSWE leaderboard (https://deepswe.datacurve.ai/), under which Kimi K3 attains 67.3 with the mini-SWE-agent harness.
- Terminal-Bench 2.1. Kimi K3 is evaluated with the KimiCode harness. For all other models, we report the best score across harnesses: GLM-5.2 with Claude Code (https://z.ai/blog/glm-5.2); Claude Opus 4.8 and Claude Fable 5 with Terminus 2 (https://artificialanalysis.ai/evaluations/terminalbench-v2-1); GPT 5.5 and GPT 5.6 Sol with Codex (https://openai.com/index/previewing-gpt-5-6-sol/).
- Program Bench. Kimi K3 is evaluated with the KimiCode harness. The GLM-5.2 score is from https://z.ai/blog/glm-5.2; all other scores are from https://www.vals.ai/benchmarks/programbench.
- SWE Marathon. Kimi K3, Claude Opus 4.8, and Claude Fable 5 are evaluated with the Claude Code harness; GPT 5.6 Sol is evaluated with the Codex harness. The GLM-5.2 score is from https://z.ai/blog/glm-5.2.
- FrontierSWE. Kimi K3 is evaluated with the KimiCode harness and GPT 5.6 Sol with the Codex harness; all other results are from https://www.frontierswe.com/. Dominance scores are recomputed from the raw scores using the official evaluation script and are current as of July 16, 2026.
- PostTrain Bench. Scores for GLM-5.2, GPT 5.5, and Claude Opus 4.8 are adopted from the official PostTrainBench results. Kimi K3, Claude Fable 5, and GPT 5.6 Sol are evaluated with the official Harbor implementation at maximum reasoning effort, averaged over three runs — Kimi K3 and Claude Fable 5 with the Claude Code harness, and GPT 5.6 Sol with the Codex harness. Under the Claude Code harness, requests refused by Claude Fable 5 due to its usage policy automatically fall back to Claude Opus 4.8.
- MLS Bench Lite. Kimi K3 is evaluated with the KimiCode harness; GLM-5.2 and the Claude models with the Claude Code harness; GPT 5.5 and GPT 5.6 Sol with the Codex harness.
- KCB 2.0. Kimi K3 is evaluated with both the KimiCode and Claude Code harnesses; GLM-5.2, Claude Opus 4.8, and Claude Fable 5 with the Claude Code harness; GPT 5.5 and GPT 5.6 Sol with the Codex harness. All models are evaluated at maximum reasoning effort, except GPT 5.5, which uses the "xhigh" setting.
Productivity and agentic benchmarks
- For OfficeQA Pro, each test case provides the agent with the entire PDF corpus, with all PDFs rendered as images and no machine-readable text available.
- OfficeQA Pro and SpreadsheetBench 2. Kimi K3, GLM-5.2, Claude Opus 4.8, and Claude Fable 5 are evaluated with the Claude Code harness; GPT 5.5 and GPT 5.6 Sol are evaluated with the Codex harness.
- MCP Atlas. All models are evaluated on the 500-task public subset with a 100-turn limit, using Gemini 3.1 Pro as the judge.
- AutomationBench. All models are evaluated on the 600-task public subset, following the official GitHub setup in all other respects.
- BrowseComp. We adopt the context-compaction strategy used in the Claude model cards, triggered at 300K tokens. When evaluated with a 1M-token context window and no context management, Kimi K3 achieves a score of 90.4. The results of Claude Fable 5, Claude Opus 4.8, GPT 5.6 Sol, and GPT 5.5 are cited from https://www.anthropic.com/news/claude-fable-5-mythos-5 and https://openai.com/index/gpt-5-6/.
- GDPval-AA v2 and AA-Briefcase scores are cited from https://artificialanalysis.ai/.
Multimodal benchmarks
- Except for ZeroBench, which follows the official setting and is run five times, all multimodal scores are averaged over three runs. MMMU-Pro is evaluated following the official protocol, preserving the original input order and prepending images to the text input.
- PerceptionBench. PerceptionBench is an in-house benchmark that focuses on atomic visual perception capabilities.
Limitations
- Sensitivity to thinking history. K3 was trained in the preserved thinking history mode. If the agent harness fails to pass back all the historical thinking content as required, or if an ongoing session with another model is switched over to K3, generation quality may become highly unstable. We recommend using a harness with verified compatibility, such as Kimi Code, and avoiding switching to K3 in the middle of a session.
- Excessive proactiveness. K3's training places particular emphasis on long-horizon, challenging tasks. As a result, when it encounters minor issues or ambiguous user intent during task execution, it may make unexpected decisions on the user's behalf. If your application requires the agent to operate within well-defined boundaries and refrain from excessive improvisation, please impose more explicit behavioral constraints on K3 in the system prompt or in AGENTS.md.
- Despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み