Kimi K3 の能力と課題を分析
The Zvi は、Kimi K3 が強力なオープンモデルとなる一方で、ベンチマークと実性能の乖離や遅さといった課題を指摘し、Moonshot AI の IPO 計画についても言及している。
キーポイント
能力と立ち位置の評価
Kimi K3 は重み公開予定により最強のオープンモデルとなる可能性が高いが、クローズドモデルの最前線とは数ヶ月(4〜6ヶ月)の差があり、特に事前学習段階での遅れが目立つ。
ベンチマークと実性能の乖離
最大限の努力でスコア付けされたベンチマークでは卓越するが、実際の使用状況(トークン数など)では性能が低下する傾向があり、モデルの特性に偏りがある。
アーキテクチャとコスト
2.8T パラメータという巨大な規模を持つが、推論速度が遅くトークン消費量が多いため、コストパフォーマンスは小型モデルやトップクローズドモデルに劣る可能性がある。
Moonshot AI の将来展望
Kimi K3 の成功を受け、Moonshot AI は香港での IPO を今後 6 ヶ月以内に計画しており、市場の熱気に乗った戦略的動きと見られる。
中国モデル発表への過度な反応の危険性
中国のモデル発表をきっかけに「米国のAIリーダーシップは終わった」と即断する傾向があり、これが政策決定者に誤ったパニックや規制緩和を招くリスクがある。
DeepSeek 時の類似した市場動向とパニック
Google や Nvidia の株価下落など、過去の DeepSeek モーメントと同様の不合理な市場の混乱が再び起きつつあり、同様の過剰反応を警戒する必要がある。
単純なベンチマークによる結論の誤り
Arena などの単一のベンチマーク結果だけで米国の優位性が消滅したと断じる論理は、Washington D.C. の政策決定者を何度も惑わし、実害をもたらしてきた。
重要な引用
In aggregate it is several months behind the closed model frontier, at least four and my median guess is six
It likely outperforms on benchmarks relative to practical performance.
Kimi K3 will be excellent at some things, less so at other things.
Moonshot AI has informed investors that it plans an IPO in Hong Kong within the next six months
It is actively suicidal to respond to 'the Chinese have better models now' with 'then we had better sell them the compute so they can run them and also build even better ones.'
The pattern is: Chinese model releases. There is some impressive benchmark cited. Therefore, America's lead is gone, QED.
影響分析・編集コメントを表示
影響分析
この記事は、Kimi K3 の技術的ポテンシャルと現実的な制約を冷静に分析することで、業界が過度な期待を抱くのを防ぎつつ、実用的な導入判断の基準を提供しています。特に、ベンチマークスコアの過大評価に対する警告と、巨大モデル特有のコスト・パフォーマンス課題は、開発者や意思決定者にとって重要な示唆となります。また、Moonshot AI の IPO 計画発表は、中国発 AI モデルが資本市場でも注目される段階に至ったことを象徴する出来事として捉えられます。
編集コメント
Kimi K3 の評価において、ベンチマークの数値に踊らされず、実運用でのコストと速度をどうバランスさせるかが問われています。Moonshot AI の IPO 計画は、中国の AI スタートアップが技術力だけでなく資本市場でも存在感を示し始めた重要な転換点と言えます。
Kimi K3 はベンチマーク結果に優れる非常に優れたモデルです。計画通りに重みが公開されれば、純粋な能力だけで見れば最強のオープンモデルとなるでしょう。
しかし、油断は禁物です。Kimi K3 の相対的な強さだけを評価してはいけません。総合的に見ると、クローズドモデルの最前線には数ヶ月遅れています。少なくとも 4 ヶ月、私の中央値の推測では 6 ヶ月ほど遅れていると見ています。後処理(ポストトレーニング)の差は小さく、事前学習(プレトレーニング)の差の方が大きいです。以前に比べて遅れ幅は縮まりましたが、その 1 ヶ月の密度は以前とは比べ物になりません。
このモデルはいくらか知識蒸留(ディストillation)されています。そのため、ベンチマーク上のスコアは実際の運用性能よりも高く出る傾向があります。すべてのベンチマークテストで最大限の努力を払っており、Fable や Sol が類似テストで使用するトークン数とは比較にならないほど多くのトークンを消費しています。結果を見ると、性能にムラがあるように見えます。Kimi は特定のタスクでは卓越した能力を発揮しますが、他の分野ではやや劣るでしょう。
今後数週間でさらに詳細が明らかになるはずです。現時点ではアクセス環境が不安定で、実際に Kimi K3 を試せる人も限られているため、その能力に関する私の推測には通常よりも大きな誤差が含まれています。しかし時間は待ってくれません。私たちは引き続き進んでいきます。
現時点で最大のオープンモデルは、2.8T パラメータを誇る Kimi K3 です。このサイズ感は Claude Opus の上限付近にあり、Mythos の下限に近い位置にあります。これが多くの性能向上の要因となっています。
ただし、処理速度は遅く、トークン消費量も多めです。好意的な評価の多くは「大規模モデルである」という点自体への反応であり、確かにそこには「大規模モデル特有の匂い」が感じられます。一般的に、速度とコストを犠牲にして性能向上を図るトレードオフですが、これは素晴らしい戦略ではあります。ただ、同様の規模で再度このアプローチが取れるようになるまでには、まだ時間がかかるでしょう。
Claude からの知識蒸留(ディストillation)は物語の一部であり、おそらく Fable が大きな役割を果たしています。しかし、これが全てを説明するわけではありません。Moonshot 社がこれを実行した背景には、わずかな効果であっても実行する意義があったはずです。より大きなモデルをリリースするタイミング自体にも、何らかの意図が感じられます。
Kimi K3 は確かに優れたモデルですが、ベンチマークでの過剰なパフォーマンスを補正し、実務での期待される性能を考慮すると、2.8T パラメータを持つ Kimi K3 から予想されるものとそれほど変わらないように見えます。
Kimi K3 の(未公式の暫定版)Epoch 能力指数は、中国のトレンドライン上に正確に位置しています。
Kimi K3 は、自社のワークフローに適しているか確認する価値は十分にあります。API とサブスクリプションの両方でこの価格帯であれば、より小さく安価なオープンモデルの役割を担うことはなく、トップクラスのクローズドモデルとの比較では通常劣る結果になるでしょう。それでも、Kimi K3 が最適な選択肢となるケースは確実に存在します。
アンディ・カラン:Kimi K3 の成功を受け、ムーンショット AI はブルームバーグによると、今後 6 ヶ月以内に香港での IPO(株式公開)を計画していることを投資家に通知した。
これは良い考えだ。鉄が熱いうちに打つべきだろう。
目次
DeepSeek モーメント:またか、再び始まる。
あの瞬間の再訪(2025 年 6 月の記事より)
その後の物語
Kimi K3 の発表、主張、そして基本的事実
現代におけるベンチマーク競争について
他社のベンチマーク
ベンチマークは現実世界ではない
技術的なセーフガードとは何か?
Kimi ができること
Kimi ができないこと
Kimi に実行させるのが難しいこと
オープンウェイトモデルは安全ではなく、それを解決するものは存在しない
ディーン・ボールの建設的試み
OpenAI の社員はこの件に対して比較的楽観的だ
Kimi K3 は、典型的なエージェント型コーディングと 3D 処理において相対的に最も強力である
反応
あなたは誰か?
彼らはどうやってこれを成し遂げたのか?
結論
DeepSeek モーメント:またか、再び始まる
中国製モデルに関するすべての議論は、DeepSeek の衝撃という影の中にあります。
もう一度 DeepSeek のような瞬間が訪れることを強く望む人々が大勢います。
彼らは、中国のオープンモデルが米国のクローズドモデルに追いつきつつあり、AI や推論がコモディティ化されるだろうという物語を語りたいと切実に願っています。
動機は様々です。オープンモデルや「技術スタック」の重要性を主張したい人もいれば、OpenAI や Anthropic の凋落を望む人、あるいはそうした主張が売れることを知っている人もいます。最終的には、すべての AI 規制や、「進歩を遅らせる」「中国に負ける」といった懸念に対する反対論を展開することが目的となるケースも少なくありません。
「中国には今、より優れたモデルがある」という指摘に対して、「ならば計算リソースを売却して、彼らが実行し、さらに良いものを作れるようにすべきだ」と答えるのは、自滅行為です。しかし、毎回のように人々はそう主張します。ため息が出ます。
また、多くの人はアメリカの AI 企業に対し、予防措置を中止し、煩わしい規制を取り払い、分類器(クラシファイア)を撤廃するよう求めています。「瓶から出た genie は戻せないのだから、願いが叶う無限の genie を解放すべきだ。それしかない」という論調です。彼らは分類器と向き合うよりも、重大な壊滅的なリスクを冒すことを選び、その点に対して激しく憤慨しています。
Google の株価は当日 4.4% 下落し、SpaceX は 3.1%、Nvidia は 2% を超えて下落。金曜日のテクノロジー株も再び低迷しました。つまり、また同じことが起こりつつある可能性が高いのです。
そして、確かに私たちは再びそのリスクに直面しています:
Axios(誤報):中国が米国の AI リードを完全に消し去った
Seán Ó hÉigeartaigh:いいえ、そんなことはありません。(Kimi K3 が圧倒的に印象的であることは疑いようがありませんが)
このパターンは以下の通りです。
中国のモデル発表。
いくつかの印象的なベンチマーク結果が引用される。
したがって、米国のリードは終わった。Q.E.D.(証明済み)、これで終わりだ。本当に、これで終わりだ。
Axios の事例では、対象となったベンチマークは Arena です。それだけを根拠に、彼らは「米国の優位性はもうない」と断言しています。
私はこの論理を無視したいところですが、こうした思考様式がワシントン D.C. で何度も人々を説得し、実際に大きな政策影響を与えてきたことを考えると、軽視できません。
今や彼らが Mythos についても知ることになったことで、いったいどのような行動に出るのかと思うとゾッとします。Fable の脱獄問題に関する混乱が、より広範な根拠のないパニックへと拡大する可能性は十分にあります。
元々の DeepSeek モーメントは、複数の出来事が重なった結果として発生しました。
振り返ってみましょう。
私たちが経験した瞬間(2025 年 6 月の再掲)
誰もが「DeepSeek モーメント」を覚えているでしょう。これは App Store でのパニックや、根本的な合理性に欠ける株式市場の混乱を引き起こしましたが、結果としてその騒ぎはあまりにも馬鹿げたものであり、結局のところパニックする必要はなかったという結論に至りました。
数ヶ月を経て、何が起きたのか(少なくとも大部分が)明確な姿を現しました。物語的な要素がいくつか重なり合うことで、DeepSeek の r1 は単に更新すべき素晴らしいモデルから、直接的な派手さがないにもかかわらず世界に衝撃を与えた出来事へと変貌したのです。
特に以下の要因がすべて組み合わさり、この現象を引き起こしました。
「600 万ドルのモデル」という物語です。人々は v3 の限定的な計算コストを、OpenAI や Anthropic といった米国の研究機関全体の予算と同一視しました。これはまるで、「DeepSeek はリンゴにかけた費用が OpenAI が食料にかけた費用よりずっと少ない」と言っているようなものです。条件を揃えて比較すれば DeepSeek の支出は確かに少なかったものの、その差はそれほど劇的なものではありませんでした。
DeepSeek は、無料でありながら極めて洗練されたデザインと可視化された思考連鎖(CoT)を備えたアプリを同時にリリースしました。同社は迅速な対応で CoT を隠す理由もなかったため、比較検討では DeepSeek の主要ユースケースのみが他社と同じ基準で評価され、DeepSeek が欠いている、あるいは苦手とする機能やユースケースは考慮されませんでした。
そのため、「初日の無料クエリ」を利用したいユーザーにとっては、当時としてはユニークで話題を呼ぶ体験を提供できました。この動きが他の研究機関にも CoT の表示を促し、各種モデルや機能のリリース加速へとつながりました。
モデルの実力を正確に把握するには時間がかかります。異なるスタイルと可視化された CoT、そして周囲の興奮が相まって、人々は「r1」が実際以上に優れていると感じていました。
タイミングは完璧でした。DeepSeek は他社の連続したモデルリリースの直前に参入しました。2 週間も経たないうちに、米国の研究機関が依然として先行していることが明確になりました。これが DeepSeek のサイクルにおける最高潮であり、他社にとっては底辺に当たる時期でした。
技術面においてもタイミングは絶妙でした。これは強化学習(RL)のスケーリングがまだ初期段階にあったため、トレーニングプロセスを低コストで実施できた時代です。DeepSeek は自社のチップから最大限の性能を引き出すことに成功しましたが、今後は計算資源における不利さが次第に顕在化し、対応が難しくなっていく可能性が高いでしょう。
DeepSeek は「そもそも安全性テストとは何か」という議論や、他社に先駆けて追随する戦略を巧みに利用しました。可能な限り迅速に新モデルをリリースすることで、実際よりもはるかに進んでいるように見せかけ、遅れを取り戻したかのような印象を与えたのです。
Teortaxes 氏は、R1 の論文で指摘された多くの課題について述べています。当時は時間的余裕がなかったため修正できなかった点ですが、R1-0528 ではそれらが解決されました。これらの改善点は、パニック状態にあった当時の評価では「カウントされ」ていなかったのです。
DeepSeek は「モメンタム(勢い)」という議論をうまく展開しました。中国はこれまでリリースされたモデルの数において米国に大きく遅れをとっていましたが、現在はその差が縮まり、一部からは「すでに先行している」との声さえ上がりました。「これで近い将来、中国が米国を抜くだろう」という楽観論が広がったのです。
しかし、それは誤りです。そのような推定は成り立ちませんし、追随者からリーダーへと立場を変えるには、想像以上に大きな飛躍が必要なのです。
これと深く関連して、「中国が米国に追いついた」という物語を望む声が、中国擁護派だけでなく、あらゆる立場の中国警戒論者からも広がっています。その結果、私たちは今や「ミサイル・ギャップ」のようなストーリーに直面しているのです。
また、「オープンモデルこそが勝つ」と主張し、クローズドなモデルは破滅的な運命をたどり、評価に値しないと考える人々も多数います。彼らは非常に声高で、感情論(バイブス)を武器として使いこなしています。中にはトランプ政権と密接なつながりを持つ者さえいるのです。
株式市場は状況認識に著しい欠陥があり、今回の発表を実際以上に大きなニュースと捉えてしまいました。その結果、既に知られていた事柄について人々が「目覚めた」と錯覚したり、他の人も同様に目覚めるだろうと予測し始めたりしました。さらに、ジェボンズの逆説や、「r1 を実行するにはより多くのチップ(Nvidia 製を含む)を購入する必要がある」といった根本的なメカニズムの理解が広く欠如していました。また、DeepSeek に対する株式市場の反応の一部は、トランプ政権の政策発表に関するインサイダー取引に起因した可能性も十分にあります。つまり、効率的市場仮説は誤りだったのです。
その後の展開
以来、「中国が追いついた、あるいは追いつきつつある」という考えが繰り返し持ち上がり、ワシントンを中心に多くの議論を巻き起こしました。これはまさに、あの瞬間の反響が繰り返されたようなものです。
新しい強力な、あるいは注目すべき中国製モデルが登場するたびにこの話題は再燃します。中国が一日でもモデルを発表しない間は、彼らはさらに遅れをとっているように見えます。一方、発表が行われると「追いついた」とみなされ、遅れが縮まったように評価されます。
r1 以降、K3 以前に起きたこうした潜在的な転換点の上位 10 はおそらく以下の通りです:
Manus
DeepSeek r1-0528
Kimi K2
GPT-5(逆方向から来たもの)— これは多くの人を驚かせ、非常に愚かな反応を引き起こしました。
DeepSeek v3.1
Kimi K2 Thinking
DeepSeek v3.2
Kimi K2.5
DeepSeek v4
GLM-5.2
全体として、中心となったのは Kimi と DeepSeek です。Manus は一時的なブームを巻き起こし人々を驚かせましたが、GLM-5.2 は彼らにとって最も強力な製品であり、GLM シリーズの存在を確立する役割を果たしました。
これら多くのモデルは確かに優秀でしたが、業界のゲームチェンジャーとなる根本的な転換点をもたらしたものはありませんでした。
DeepSeek の登場以降、約 8 ヶ月遅れながらも迅速な追従戦略で一部の主要機能を早期に達成したあの瞬間から、私たちは状況が流動的に推移しています。当初は中国勢がかなり後れを取っているように見えた時期もありましたが、現在は様相が変わりつつあります。
GLM-5.2 や Kimi K3 は非常に印象的な成果を収めました。現在の推測では、両者との技術格差は過去最低水準まで縮まっています。AI 業界全体で出来事が加速しているため、以前よりも製品サイクル数で数えるほどの遅れがあるとは限りません。むしろ来週には Qwen の話題でも同じような分析が必要になるかもしれません。Kimi K3 はパラメータ数が 2.8T と推定されていますが、4 月 7 日に発表された Mythos Preview に比べるとまだ明確に後れを取っているように見えます。この事実から、少なくとも 3 ヶ月の格差があるという下限を想定できます。
イーサン・モリック氏はこう述べています。「Kimi は私がこれまで言ってきた通り非常に優れたモデルです。しかし、これは DeepSeek r1 のような劇的な転換点(DeepSeek moment)ではありません。予想される曲線上の位置にあり、驚くべき飛躍があったわけではありません。ただし、認知度が広がるにつれて、さまざまな理由から『DeepSeek moment』として扱われるようになるでしょう」。
ライアン・グリーンブラット氏は Kimi K3 に意外なほど良い印象を抱きつつも、フォワードパスの計算に基づけば、事前学習モデルの品質は Opus 4 と Opus 4.5 の中間程度と推定しています。他の利点はあるものの、やはり予想通り約 8 ヶ月の遅れがあると考えられます。
ポストトレーニングの進展は、少なくとも一部では知識蒸留(ディストillation)によるものです。Moonshot 社は明確にイノベーションを起こしていますが、同時に直接的な手法や、他社の出力を分析して技術を模倣する迅速な追従も実施しています。
UK AISI は Kimi K3 の発表前に包括的なレポートを発表しており、狭義のサイバータスクにおける時間的格差が徐々に縮まっていることを示しました。詳細は以下のリンクから確認できます。

私の理解では、オープンウェイトモデルが相対的に最も強いのは、狭義で比較的簡単なコーディングタスクです。しかし、この分野のベンチマークはすでに飽和状態に近づいています。
実際、ポスト全体を詳しく見ると、より長いタスクに対する答えは異なります。

UK AI Security Institute の調査によると、TLO(Task Level Optimization)における GLM-5.2 は、リリースから 7 ヶ月後に登場した Opus 4.5 と互角の性能を示しました。一方、DeepSeek の V4-Pro は、同じく 7 ヶ月前にリリースされたサブサイバーフロンティアモデルである Sonnet 4.5 を下回っています。これらの結果は、他のサイバーレンジでも概ね一貫して確認されています。
特筆すべきは、GLM-5.2 が平均して他モデルよりもわずかに少ないトークン数でステップ 7 に到達した点です。このモデルは Opus 4.6 のようにステップ 11 まで軌道に乗るものの、その先で頭打ちとなりました。
AI 研究開発の自動化やサイバー攻撃といった、最も懸念すべき 2 つの観点から考えると、より長いタスクの実行能力が重要になります。
バイオリスクへの懸念も無視できません。流行りの話題ではないかもしれませんが、オープンモデルにおいて標準的なテストが行われていない現状は少し気になります。この課題には早急に対処が必要です。
OpenAI の GeneBench-Pro に関するスコアについては、Andrew Ho 氏から入手しています。これは計算生物学における長期的な不確実性下での判断力を測定するものです。Kimi K3 は期待を上回る結果を出しました。一方、Mythos はまだテストされておらず、Fable はベンチマークの要求をほとんど拒否しました。


Opus や GPT-5.5 を凌駕したという事実は、確かに印象的です。これまでに公開されたモデルの中で、これに匹敵するものは存在しませんでした。私たちは今まさに、「やってみて初めてわかる」という局面に立っています。普段は静かな場所でも、ある瞬間に劇的な変化が起きるような場所です。
「オープンソースモデルは、現在の独自性を除く一般能力において Mythos に追いつく」という結論は疑う余地がありません。それは必ずや到来しますが、問題はいつかという点と、その時点で実効性のある差がどうなるかです。Kimi K3 の登場を踏まえると、この変化は数ヶ月以内に起きると予想されます。
UK AISI が提示した推計の両端を考慮すると、Kimi K3 以前のギャップは 4〜7 ヶ月でした。これは昨年の 6〜10 ヶ月から縮小しており、同社が相対的に強い分野での話です。絶対的な進捗の差としてはほぼ同等ですが、全体的な加速は止まりません。
Kimi K3 の発表概要と基本事実
パラメータ数は 2.8 兆個。同時に活性化する領域は 896 中 16 で、これは約 500 億のパラメータが動作していることを意味します。ローカル環境での実行は容易ではなく、コストも決して安くはありません。
API 利用料金は $3.00/$15.00 で、Opus や Sol よりもやや安価です。
サブスクリプションプランは月額 $19、$39、$99、$199 の 4 つ。高額なプランほど若干の優遇特典があり、クォータは価格に比例して拡大します。
コンテキスト長は 100 万トークンです。
API リンクと技術ブログへのリンクが用意されています。
学習データの cutoff は 2026 年初頭 reportedly と報じられています。
オープンウェイト(重み)の公開は 7 月 27 日までに約束されています。
すべてのベンチマークは最大限の努力設定で実施されました。
Kimi:本日、最も能力の高いモデル「Kimi K3」をご紹介します。Kimi K3 は、2.8 兆パラメータを備え、独自の「Kimi Delta Attention」と「Attention Residuals」アーキテクチャを採用しています。ネイティブな視覚機能と 100 万トークンのコンテキストウィンドウを特徴とし、長期的なコーディング、知識作業、推論における最先端の知能を実現するために設計された、世界初のオープン 3T クラスモデルです。
全体的な性能はまだ Claude Fable 5 や GPT 5.6 Sol といった最強のプロプライエタリモデルには及びませんが、Kimi K3 は評価スイート全体で最先端レベルの性能を発揮し、他社がテストした他のモデルを常に上回りました。
Kimi.ai:Kimi K3 は現在、http://Kimi.com、Kimi Work、Kimi Code、および Kimi API で利用可能です。オープンウェイトは 2026 年 7 月 27 日に公開予定です。
K3 は、情報フローをシーケンス長とモデル深さの両面で改善するために設計された 2 つのアーキテクチャ改良、「Kimi Delta Attention (KDA)」および「Attention Residuals (AttnRes)」を基盤としています。
また、Mixture of Experts (MoE) のスパース性を拡大し、Stable LatentMoE フレームワークと組み合わせることで、896 個のエキスパートのうち 16 個が有効に動作するようにしています。
これらの構造的変更に加え、訓練手法やデータレシピを洗練させることで、K2 と比較して全体のスケーリング効率がおおむね 2.5 倍向上しました。これにより、計算リソースをより効果的に知能へと変換できるようになっています。
前述の通り、公式ベンチマークでは Sol や Fable に劣るものの、強力な結果を示しています。
タイラー・カウエン氏が「非常にポジティブで素晴らしい」と評したこの広告だが、私には全く響かず、実用的な情報も含まれていない。
現代のベンチマーク最適化について
かつては、より露骨なベンチマーク最適化が見られたものだ。各ラボはテストデータそのものや、テストがカバーする極めて限定的な領域でモデルを訓練していた。なぜなら、対象となるターゲットは限られたセットしかなかったからだ。数値を見る際、どのラボがどこまで、いつそのような行為を行ったかを把握しておく必要があった。
しかし現在、ベンチマーク技術は進化を遂げている。今や各テストは現実的な課題を網羅的に測定しており、特定の分野に偏りすぎた対策に対しても、複数のバックアップ手段が用意されているため、単一の指標だけで判断することは難しくなっている。
OpenAI の roon は、ベンチマーク最適化(ベンチマックス)について触れています。最近、ベンチマーク技術自体が他の技術と同様に向上し、人々はより懐疑的になり、こうした評価を慎重に構築するようになっています。また、下流のコーディング顧客の間でも、社内で保有したデータを用いた評価(内部ホールドアウト評価)を構築する動きが急増しています。
さまざまなベンチマークの全体像(ゲシュタルト)を見渡すことも価値があります。すべてのベンチマークは、背後にある実態を反映する共通のパターンに収まる地図の一部として捉えるべきです。
以前ほど露骨でなくても、ベンチマーク最適化を行うことは依然として可能です。
ベンチマークは特定の能力のみを測定し、表面的なタスクには深く対応しますが、多くの貴重な特性や潜在的なリスク要因については除外してしまいます。また、研究機関によって注力する側面や、それらでの成功度合いに差があります。
すべてのテストに対して最大限の努力を傾けることもできますが、それが Moonshot の行ったことです。
ベンチマークは下限値と考えるべきです。Kimi のベンチマーク結果は、同モデルが本物であることを証明しており、その性能が実際の能力から外れる幅には限りがあります。私は引き続き、Kimi K3 の相対的な能力を、同ベンチマークは若干過大評価していると考えています。
他人のベンチマークについて
我々が「唯一絶対のベンチマーク」と呼べるものに最も近いものにおいて、Kimi K3 は良好な結果を残しており、このモデルが全体としてベンチマークで第 3 位であるという主張を裏付けています。
 Epoch Capabilities Index is exactly on the Chinese trend line.
Kimi K3 is absolutely worth checking to see if it fits into your workflows. At this price point, for both the API and the subscription, it is not going to fill the role of the smaller cheaper open models, and I expect it to usually lose out in a fight with the top closed models, but there are going to be some places where Kimi K3 is a good choice.
Andrew Curran: Following the success of Kimi K3, Moonshot AI has informed investors that it plans an IPO in Hong Kong within the next six months, according to Bloomberg.
Good idea. Strike while the iron is hot.
Table of Contents
DeepSeek Moments: Here We Go Again.
We Had a Moment (Reprise from June 2025).
The Story Since Then.
The Kimi K3 Announcement, Pitch and Basic Facts.
On Modern Benchmaxxing.
Other People’s Benchmarks.
Benchmarks Are Not The Real World.
Technical Safeguards? What Are Those?
Things Kimi Can Do.
Things Kimi Cannot Do.
Things It Is Not Easy To Get Kimi To Do.
Open Weight Models Are Unsafe And Nothing Can Fix This.
Dean Ball Attempts To Be Constructive.
OpenAI Employees Are Relatively Bullish On This One.
Kimi K3 Is Relatively Strongest At Typical Agentic Coding and 3D.
Reactions.
Who Are You?
How Did They Do It?
Conclusion.
DeepSeek Moments: Here We Go Again
All discourse about Chinese models lives in the shadow of the DeepSeek moment.
There are a lot of people who really, really want another DeepSeek moment to happen.
These people really, really want to tell the story that Chinese open models are catching up to American closed models, that AI and inference will become commoditized.
Their motivations vary. They often want to affirm open models, or the importance of the ‘tech stack.’ Others simply want to see OpenAI and Anthropic go down, or know that such claims sell. Often the ultimate objective is to argue against all AI regulations, or anything that might ‘slow us down’ or cause us to ‘lose to China.’
It is actively suicidal to respond to ‘the Chinese have better models now’ with ‘then we had better sell them the compute so they can run them and also build even better ones.’ Yet every time, yes, people will argue that. Sigh.
Often they simply want to tell American AI to stop taking precautions, to stop being annoying and take down the classifiers, as in ‘genie is out of the bottle, so release the bigger genie with unlimited wishes, it’s the only way.’ People really would take major catastrophic risks rather than deal with classifiers, and are Big Mad about this.
Google was down 4.4% on the day, SpaceX was down 3.1% and Nvidia down over 2%, and tech stocks were down again on Friday, so plausibly we’re doing this again.
And yep, we are at risk of doing this again:
Axios (being wrong): China just erased America's AI lead
Seán Ó hÉigeartaigh: No it didn't. (although Kimi K3 is undoubtedly impressive)
The pattern is:
Chinese model releases.
There is some impressive benchmark cited.
Therefore, America’s lead is gone, QED, that’s it, no, really, that’s it.
In the Axios case the benchmark in question is Arena. Based on that alone, they state as fact that America’s lead is gone.
I would ignore, but this style of logic has convinced a lot of Washington D.C. multiple times, and that has had substantial policy impact.
I shudder to think what such folks might do now that they also know about Mythos. The confusion over Fable jailbreaks could easily extend to a broader dumb panic.
The original DeepSeek moment happened because of a confluence of events.
Let’s review.
We Had a Moment (Reprise from June 2025)
We all remember The DeepSeek Moment, which led to Panic at the App Store, lots of stock market turmoil that made remarkably little fundamental sense and that has been borne out as rather silly, a very intense week and a conclusion to not panic after all.
Over several months, a clear picture emerged of (most of) what happened: A confluence of narrative factors transformed DeepSeek’s r1 from an impressive but not terribly surprising model worth updating on into a shot heard round the world, despite the lack of direct ‘fanfare.’
In particular, these all worked together to cause this effect:
The ‘six million dollar model’ narrative. People equated v3’s marginal compute costs with the overall budget of American labs like OpenAI and Anthropic. This is like saying DeepSeek spent a lot less on apples than OpenAI spent on food. When making an apples-to-apples comparison, DeepSeek spent less, but the difference was far less stark.
DeepSeek simultaneously released an app that was free with a remarkably clean design and visible chain-of-thought (CoT). DeepSeek was fast following, so they had no reason to hide the CoT. Comparisons only compared DeepSeek’s top use cases to the same use cases elsewhere, ignoring the features and use cases DeepSeek lacked or did poorly on. So if you wanted to do first-day free querying, you got what was at the time a unique and viral experience. This forced other labs to also show CoT and accelerate release of various models and features.
It takes a while to know how good a model really is, and the different style and visible CoT and excitement made people think r1 was better than it was.
The timing was impeccable. DeepSeek got in right before a series of other model releases. Within two weeks it was very clear that American labs remained ahead. This was the peak of a DeepSeek cycle and the low point in others cycles.
The timing was also impeccable in terms of the technology. This was very early days of RL scaling, such that the training process could still be done cheaply. DeepSeek did a great job extracting the most from its chips, but they are likely going to have increasing trouble with its compute disadvantage going forwards.
DeepSeek leveraged the whole ‘what even is safety testing’ and fast following angles, shipping as quickly as possible to irrevocably release its new model the moment it was at all viable to do so, making it look relatively farther along and less behind than they were. Teortaxes notes that the R1 paper pointed out a bunch of things that needed fixing but that DeepSeek did not have time to fix back then, and that R1-0528 fixes them, and which weren’t ‘counted’ during the panic.
DeepSeek got the whole ‘momentum’ argument going. China had previously been much farther behind in terms of released models, DeepSeek was now less behind (and some even said was ahead), and people thought ‘oh that means soon they’ll be ahead.’ Whereas no, you can’t assume that, and also moving from a follower to a leader is a big leap.
There was highly related to a widespread demand for a ‘China caught up to the USA’ narrative, from China fans and also from China hawks of all sorts. Going forward, we are left with a ‘missile gap’ style story.
There are also a lot of people always pushing the ‘open models win’ argument, and who think that non-open models are some combination of doomed and don’t count. These people are very vocal, and vibes are a weapon of choice, and some have close ties to the Trump administration.
The stock market was highly lacking in situational awareness, so they considered this release much bigger news than it was, and it caused various people to ‘wake up’ to things that were already known and anticipate others waking up, and there was widespread misunderstanding of how any of the underlying dynamics worked, including Jevon’s Paradox and also that if you want to run r1 you go out and buy more chips, including Nvidia chips. It is also possible that a lot of the DeepSeek stock market reaction was actually about insider trading of Trump policy announcements. Essentially: The Efficient Market Hypothesis Is False.
The Story Since Then
Since then, the idea that China had caught up, or was catching up, kept coming up, drove much discussion around Washington, as echoes of this one moment.
This is then renewed every time a new strongest or exciting Chinese model comes out. Every day that China does not release a model, they look one day farther behind. When they do release, they ‘catch up’ and look less behind.
The top 10 such potential moments since r1 and before K3 were likely these:
Manus.
DeepSeek r1-0528.
Kimi K2.
GPT-5 (in reverse) which spooked a lot of people in highly stupid ways.
DeepSeek v3.1.
Kimi K2 Thinking.
DeepSeek v3.2.
Kimi K2.5.
DeepSeek v4.
GLM-5.2.
Mostly it’s been Kimi and DeepSeek. Manus got a hype train going and spooked people, and GLM-5.2 was by far their strongest offering, putting GLMs on the map.
Many of these were good models, but none fundamentally changed the game.
Roughly, since the DeepSeek moment, when DeepSeek was roughly eight months behind but had matched some key aspects faster via fast following, we have bounced around. For a while it looked like China was quite a lot behind.
GLM-5.2 and Kimi K3 have been impressive. The current best estimate of the time gap is at its lowest point. Events in AI have accelerated all around, so it is not clear that the gap is fewer product cycles than before, and I half expect to be doing this again next week for Qwen. Kimi K3 is 2.8T and seems to still be solidly behind Mythos Preview, which was announced on April 7, so that provides a starting point lower bound of a three month gap.
Ethan Mollick: Kimi is, as I have been saying, a very good model. But it is not a DeepSeek r1 moment, in that it is roughly where I would expect on the curve rather than an unexpected leap. It will be treated as a DeepSeek moment for a variety of reasons especially as more people hear about it.
Ryan Greenblatt, despite being pleasantly surprised by Kimi K3, estimates that the pretrain quality is about halfway between Opus 4 and Opus 4.5 based on forward pass math, but with some other advantages, so ~8 months behind, as one might expect.
The post-training is closer, at least in part because of distillation. Moonshot is clearly innovating, but it is also clearly distilling, both directly and also fast following via looking at outputs and copying techniques.
UK AISI issued a report on everything prior to Kimi K3, showing the time gap for narrow cyber tasks narrowing somewhat over time. Their full report is here.

My understanding is that narrow and relatively easy coding tasks are where open weights model are at their relative strongest, and the benchmark here is approaching saturation.
Indeed, when you look at the full post, you get a different answer for longer tasks.

UK AI Security Institute: On TLO, GLM-5.2 reaches as far as Opus 4.5, a model released less than 7 months before it, while DeepSeek’s V4-Pro falls below Sonnet 4.5 (a sub-cyber-frontier model released 7 months before it). These results are broadly consistent across our other cyber ranges. Notably, GLM-5.2 reached step 7 with marginally fewer tokens than any other model on average, tracking Opus 4.6’s trajectory to step 11 before stalling.
Longer tasks are more relevant in terms of both of the most important things to worry about: Automation of AI R&D and cyber attacks.
One might also worry about bio risks, even if that is not as in fashion, and it is a little concerning we don’t see standard testing on that at all for the open models. That needs to be addressed. We do have the score on OpenAI’s GeneBench-Pro via Andrew Ho. This measures judgment under long-horizon ambiguity in computational biology. Kimi K3 exceeded expectations. Mythos have not been tested. Fable refused most requests in the benchmark.


Beating Opus and GPT-5.5 is impressive stuff. No previous open model came close. We are on the verge of doing some f***ing around and thus finding out. This is a place where plausibly not much happens until suddenly quite a lot happens.
The conclusion of ‘open models will catch up to Mythos in general capability including the thing that currently makes it unique’ is indisputable. That is coming, the question is when, and that will establish the effective gap. Given Kimi K3 we should expect this to happen a few months from now.
If we take both ends of UK AISI’s estimates, the pre-Kimi gap was 4-7 months, down from 6-10 months last year, in an area of relative strength. That’s roughly a similar amount of progress gap in absolute terms, and everything is accelerating.
The Kimi K3 Announcement, Pitch and Basic Facts
2.8 Trillion parameters, 16 of 896 areas active at once which implies ~50B active. This is not easy to run locally, and won’t be that cheap.
$3.00/$15.00, modestly cheaper than Opus and Sol.
Subscription plans are $19/$39/$99/$199 per month. The larger buys have some modest advantages and quota scales linearly with price.
1M token context.
API link, Tech blog link.
Training cutoff is reportedly early 2026.
Open weights promised by July 27th.
All benchmarks run under maximum effort settings.
Kimi: Today, we are introducing Kimi K3 — our most capable model. Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.
While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models.
Kimi.ai: Kimi K3 is now live on on http://Kimi.com, Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026.
K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth.
We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework.
Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to K2, allowing the model to convert compute into intelligence more effectively.
As stated above, they claim strong official benchmarks, although short of Sol or Fable.



The ad, which Tyler Cowen called very positive and very good, falls flat to me, and contains zero useful information.
On Modern Benchmaxxing
We used to see rather explicit benchmaxxing. Labs would train on the test, or on the very narrow thing the test would cover, because we had a limited set of known targets. You had to know which labs did this, to what extent, when looking at numbers.
Our benchmarking tech has improved, and now they collectively measure real things, and there are a variety of backups in case you aim too narrowly.
roon (OpenAI): on the subject of benchmaxxing - it seems benchmarking technology has gotten better recently, like most other technology. people are more skeptical and build these things more carefully. there’s also a explosion of downstream coding customers building internal heldout evals
Looking at the gestalt of different benchmarks is also valuable. Everything should be part of a map that fits into a common pattern that reflects the underlying territory.
You can still absolutely benchmaxx without being as explicit as you used to be.
Benchmarks measure some types of abilities rather than others, and measure shallow rather than deep tasks, and exclude many valuable properties or potential liabilities. And some labs focus more on those aspects, or have more success on them, than others.
You can also set effort to maximum for all the tests, which Moonshot did.
Think of the benchmarks as a lower bound. Kimi’s benchmarks prove it is for real, and it could only underperform (or outperform) them by so much. I still expected, and continue to believe, that they modestly overstate Kimi K3’s relative capabilities.
Other People’s Benchmarks
On the closest thing we have to the One True Benchmark, Kimi K3 does well, confirming claims that overall this model has the third highest benchmarks:
![image](https://substackcdn.com/image/fetch/$s_!XsAE!,w_1456,c_limit,f_auto,q_auto:good,fl_pro
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み