Open models 動向:Kimi K3・Qwen 3.8・WAIC 演説を解説
Interconnects の解説において、Kimi K3 や Qwen の最新動向を踏まえ、中国モデルの台頭が米中地政学やオープンソース戦略に与える影響、およびディストillationに関する議論が深掘りされています。
キーポイント
中国モデルの急激な進歩と要因
Kimi K3 のリリースを皮切りに、GLM や Qwen など中国発のオープンモデルが急速に性能を向上させており、その背景にあるデータ環境や研究機関の戦略について分析されています。
地政学とオープンソース戦略の転換
習近平国家主席の WAIC での演説で中国がオープンソースを明確な戦略として位置づけたこと、および Qwen が次期モデルをオープンウェイト化すると発表したことが、米中対立における AI の役割に大きな影響を与えています。
ディストillation とベンチマークの議論
「オープンモデルはクローズドより何ヶ月遅れているか」という単純な比較への批判や、Ben Thompson 氏との対立を含むディストillation(知識蒸留)に関する技術的・経済的な議論が展開されています。
サイバーセキュリティと規制の行方
フロンティアモデルにおけるセキュリティリスクを踏まえ、オープンソースモデルへの輸出規制や禁止措置が逆にセキュリティを低下させる可能性について言及されています。
オープンモデルとクローズドフロンティアのギャップ評価の難しさ
ベンチマークの数値化は提供者や指標によって大きく異なり、オープンモデルがクローズドモデルより何ヶ月遅れているかという単純な数値での議論は誤解を招きやすい。
特定タスク向けポストトレーニングの重要性と可能性
コーディングやコンピューター操作などのニッチな領域では、クローズドモデルとの差が数ヶ月であっても大きな影響を与えるため、Qwen や GLM 5.2 のような基盤モデルを特定の業務に微調整する動きが活発である。
Kimi K3 の大規模化による微調整の技術的課題
Kimi K3 はスケールが大きすぎて単一の B300 ノードで重みを読み込むことさえ困難であり、実用的な微調整状態にするには高度なエンジニアリングと時間が必要となる。
重要な引用
Xi gave his speech where he directly committed to openness and open source as a strategy.
Qwen announced their next big model is going to be open weight, which is a big change of things.
People love to put a definitive number onto this... which is really mudding because we have so many different benchmark providers these days.
people love to put a definite uh definitive number onto this uh which is really really mudding because we have so many different benchmark providers these days
if it's like what is the market for Claude Code and Codex right now and if it is software engineering then like the fact that the models are say a couple months behind on that can be a very, very big deal
Kimi K3 did some more interesting things... found the interesting models one or two months before the download numbers usually take off like those kind of analysis is groundbreaking
影響分析・編集コメントを表示
影響分析
この解説は、単なるモデル性能の比較を超え、米中対立という地政学的文脈の中で AI のオープンソース化が戦略的兵器として再定義されつつある点を浮き彫りにしています。特に中国の国家戦略と主要企業の方向性が一致し始めている点は、今後のグローバルな AI エコシステムにおける「オープン vs クローズド」の勢力図を大きく変える可能性を秘めています。
編集コメント
今回の解説は、モデルの性能比較という表面的な話題から、国家戦略やセキュリティ規制といったマクロな視点へと議論を深める重要な転換点を示しています。特に中国がオープンソースを明確な国策として掲げた点は、今後の開発者コミュニティや企業戦略に大きな影響を与えるでしょう。
朗報です!ポストトレーニングの知識を世界と共有しようとした私の本がついに完成し、近日出荷されます。Manning または Amazon でご注文ください。皆様のおかげで、現在は Amazon の AI 書籍ランキング第 1 位を獲得しています(笑)。
ネイサンとフローリアンが、オープンモデルを取り巻く最新動向について語り合います。先週 Kimi K3 がリリースされたこともあり、地政学(米中対立)、経済性(オープン vs クローズドモデル)、AI の最前線におけるセキュリティなど、あらゆる分野で加速度的な変化が起きていると感じられます。
チャプター構成:
00:00 ウェルカムと背景説明
04:38 Kimi K3 との共存・活用方法
08:53 GLM 5.2 の継続的な役割
12:47 なぜ中国製モデルはこれほど高性能なのか?
17:41 データ、環境、そして中国のラボ見学ツアー
19:47 中国プロバイダー総括:Qwen、DeepSeek、MiniMax など…
24:08 米国のオープンモデルエコシステム
30:25 フロンティアとニアフロンティアの違い、および禁止措置に対するサイバーセキュリティの観点
34:58 ディスティルレーション(知識蒸留)とベン・トンプソン氏の議論
44:12 予測とフロンティア層級リスト
48:36 まとめ
Apple Podcasts、Spotify、またはお好みのポッドキャストアプリで聴取可能です。他の Interconnects インタビューはこちらをご覧ください。
シェアする
ポストトレーニングに関する教育動画をもっと見たい方は、現在準備中のコースをご覧ください。
文字起こし
00:00:06 Nathan Lambert: こんにちは、Interconnects へようこそ。今回は四半期ごとのオープンモデル総括です。今回は主に、なぜ多くの「知識蒸留(distillation)」に関する議論が間違っているのかを解説し、現状の立ち位置を理解することに焦点を当てます。
先週木曜日に Kimi K3 がリリースされました。今後さらに多くの情報が公開されるのはほぼ確実でしょう。週末には習主席が演説を行い、オープン性とオープンソースを戦略として直接約束しました。ただし、詳細なロードマップを示したわけではありませんでした。
Qwen は次期大型モデルを「オープンウェイト(重み付きで公開)」する方針を発表し、これは大きな転換点です。取り上げるべき話題は山ほどありますね。
Flo さん、すでにパフォーマンスの格差や蒸留に関する議論について触れておられましたか?まずはそこから始めましょう。私の方でもいくつかトピックをリストアップしていますので、私が書いたブログ記事の内容も踏まえながら、一つずつ掘り下げていきましょう。記事は非常にニュアンスに富んでおり、尽きることなく話せる内容です。
では、引き続き本題に入りましょう。
00:01:17 Florian Brand:はい、モデルリリース、特にオープンソースモデルのリリースにおいて最も議論されるのは、「クローズドな最前線のモデルと比べてどれくらい遅れているのか」、あるいは「何ヶ月分の差があるのか」という点です。多くの人がこのギャップに明確な数字を当てはめようとしますが、それはかえって状況を混乱させています。なぜなら現在ではベンチマークを提供する機関が多数あり、利用可能なテスト項目も多岐にわたるからです。各サイト(私自身も例外ではありません)は、自社の好きなベンチマークを採用して「現在のモデルや新リリースのモデルが最前線にある」と主張し、それに対して他側は別のベンチマークを持ち出して「実は1年遅れている」などと反論します。
この議論の多くは、結局のところ「オープンソースモデルはクローズドな最前線から何ヶ月遅れているのか」という問いに依存しているように見えます。
00:02:26 ナタン・ランバート:はい。私の主張は、いくつかのベンチマークが実際に人々が取り組んでいる内容とある程度相関しているという点です。具体的には、エージェントによるコーディングやコンピューター操作タスクなどが該当します。一方で、他のベンチマークはロングテール領域との相関を示しており、こここそが Claude や GPT が真価を発揮する場所だと考えています。
しかし、現在の Claude Code や Codex の市場規模を考えると、もしそれがソフトウェアエンジニアリング分野であるなら、モデルの性能が数ヶ月遅れているという事実は非常に重大な問題になり得ます。そこで私は、このモデルも概ね大丈夫だろうと予想しています。ただし、モデルの重みは 7 月 27 日に公開される予定ですが、まだ未公開です。議論の多くはこの公開を前提とした仮定に依存することになる点にご注意ください。
しかし、多くの人がこのモデルをポストトレーニング(追加学習)することで、Opus や GPT と同様に、特定のニッチな分野で高い性能を発揮できるようになる可能性は十分にあります。私は、ポストトレーニングされたオープンモデル業界の現状について、あなたとは異なる見解を持っているかもしれませんが、こうしたモデルを特定の高価値タスク向けに微調整する取り組みには、非常に大きな期待と進展があります。
歴史的に見れば、Qwen や GLM を組み合わせたアプローチが主流でしたが、GLM 5.2 の登場がこの動きをさらに加速させました。私は、誰かが「私たちのタスク向けに Kimi K3 を微調整した」というブログ記事を初めて公開する瞬間を楽しみにしています。きっと大きな成果が得られると確信しています。
私も Kimi K3 を使っていますが、あなたの使用頻度には及びません。ただの直感ですが、このモデルはスケールアップの規模が大きすぎるため、ポストトレーニングを施すとまだ荒削りな状態になる可能性があります。しかし、その分、まだ引き出せる性能が大量に残っているはずです。
00:04:03 Florian Brand: はい。実行、特にポストトレーニングの実行は非常に困難です。B300 を 1 ノード用意して重みを読み込むだけで十分すぎるほどで、スケールの大きさに驚かされます。そのため、実際に微調整可能な状態にするには、時間と多くのエンジニアリングの労力が必要になるでしょう。
ところで、このモデルを実際に使いたいという話に戻りましょう。あなたはコーディングプログラムに登録し、実際に使用したとのことですね。その体験談は、このモデルを広く普及させる上で非常に有益な文脈となります。
00:04:38 ナタン・ランバート:はい。リリースの翌日かその辺りに、200 ドルプランに登録しました。これが同社の最大プランで、他の大手サービスと同様に、他にも 40 ドルや 100 ドルのプランも用意されていますが、この最大プランではコンテキスト長が 100 万トークンに達します。また、API リクエストに対して優先度が付与されているように感じられますね。ネット上では「常に API エラーが発生する」という声をよく聞きますが、私自身はこれまで特に問題なく利用できています。
モデルの能力については、フロントエンド周りは非常に優れていますが、それ以外にもいくつかの点で際立って優秀です。私の期待値を超えた部分もありますね。例えば研究タスクでは、インターコネクト(Interconnects)で公開しているオープンモデルに関する 1 年以上にわたるデータ分析を基に、「これまでやったことのない新しい分析結果を出してほしい」と Frontier モデルたちに依頼しました。私たちはすでに独自の分析を行って発表済みなので、彼らには「何か新しいことをして驚かせてくれ」という趣旨の指示を出したのです。
しかし、多くのモデル、特に Frontier モデルたちは、私たちが既に行っているデータ分析部分をそのまま再現し、その後に奇妙で難解な部分を追加する傾向がありました。
Kimi K3 はさらに興味深い動きを見せてくれました。私は Reddit のスクレイピングを明確に指示したところ、これまで私が考慮していなかったサブレディットまで発見しました。例えば、Reddit 上の議論が、ダウンロード数が急増する通常よりも 1〜2 ヶ月先行して行われていることが分かりました。つまり、Kimi は Qwen のような注目モデルを、ダウンロード数の増加が一般的に始まるよりも 1〜2 ヶ月早く見抜いているのです。こうした分析は画期的ですが、他の最先端モデルと比較しても Kimi が私を驚かせた点です。
「このモデルを使って、あなたが普段行うコアな作業の大部分をカバーできますか?」という単純な質問に対して、私の回答は「どれくらい裁量を与えるかによる」です。Kimi K3 について私が特に感じたのは、Prime Intellect で取り組んでいるフレームワークでのコード生成において、その出力が非常にシンプルで読みやすいということです。ただし、Codex(特にバージョン 54 や 55)が持つような一部の機能やレベルにはまだ届いていません。したがって、こうしたタスクにおいては、Kimi K3 は Codex のバージョン 54〜55 レベルの性能を持つと評価できます。
00:07:24 Florian Brand: うーん、それは私がどれくらい裁量を与えるかによりますね。私が Kimi K3 で特に感じたのは、Prime Intellect で取り組んでいるフレームワークでのコード生成において、その出力が非常にシンプルで読みやすいということです。ただし、Codex(特にバージョン 54 や 55)が持つような一部の機能やレベルにはまだ届いていません。したがって、こうしたタスクにおいては、Kimi K3 は Codex のバージョン 54〜55 レベルの性能を持つと評価できます。
コードを読んで「これは素晴らしい」と評価し、Codex でチェックすると、得意ではないニッチなケースが見つかることもあります。しかし、監督的な実行や実験の実行としては十分に実用性があります。また、他の一部の特殊な用途でも、そのまま実行させることができます。
唯一の欠点は、API がユーザーで溢れかえっており、サーバーが中国にあるためです。その結果、実際の経過時間は GPT に比べて大幅に長くなります。それでも、もし私がこれを日常ワークフローに組み込んで使おうとした場合、処理速度は遅くなるかもしれませんが、「これは使えない」と言うほどには支障をきたさないでしょう。
00:08:53 ネイサン・ランバート:では、GLM 5.2 と比較するとどうでしょうか?私はまだ GLM 5.2 の状況が変化し続けていると考えています。例えば、私がサンフランシスコを歩き回って人々に「このアジェンシーコーディングやワークフローの一部で実際に使っています」と尋ねると、彼らはそう答えます。あなたは GLM を使っている立場の方ですか?
00:09:20 Florian Brand: はい、どこで使うかという点ですが、私は主に GLM を使っています。理由は社内のエンドポイントが非常に高速だからです。以前は 1 秒間に 200〜300 トークンを処理できる API も利用していましたが、それよりも GLM の方が圧倒的に速いですね。
大量のタスクを十分な精度で高速に処理できるなら、そのモデルを使うのが一番です。Codex に切り替えてから、さらに劣ったモデルを選んだり、推論レベルを調整したり、速度優先の設定に変更したりといった手間を考えれば、GLM を使えば同じ結果が得られますし、十分実用できます。
能力面では Sonnet レベルに匹敵すると言え、特に単純な下処理やクリーンアップ作業などには非常に効果的です。Kimi K3 をメインのエージェントとして、GLM をサブエージェントとして使い分ければ、多くの業務で十分に通用するでしょう。
00:10:24 ナタン・ランバート:今回のキミ(Kimi)の発表と、そのスケール感には大きな違いがあります。これらのオープンモデルが推論プロバイダー全体で実際に最適化され、利用可能になるまでには、もう少し時間がかかるでしょう。
例えば GLM 5.2 は非常に高速ですが、まず重み(weights)がまだ公開されていません。さらに、500B や 700B の MoE モデルのようなケースほど、採用と普及のスピードは速くないと考えられます。そこには新たな課題が生じるでしょう。
これは過去とは全く異なる状況です。以前は中国のモデル開発チームが RL(強化学習)のトレーニングを完了すると、数時間から数日、遅くとも一週間以内にオープンウェイトでモデルを公開していました。そうすると、すぐにエコシステム側がその扱い方を理解し、対応が始まるという流れでした。
次期オープンウェイトモデルのスケールアップには、より大きなインフラストラクチャの向上が必要だと考えています。これは、クローズドな研究機関がモデルを発表する前に裏側で実施しているような取り組みを、私たちも考慮に入れるべきです。
つまり、この時間差を利用した戦略によって、Kimi をワークフローに大規模に適用し、ポストトレーニングを行うまでに実際にはさらに 1 ヶ月ほどかかる可能性があります。オープンウェイトの支持者たちは、「クローズドモデルが利用可能になってからでないと時間差は生じない」とよく言いますが、現在ではオープンモデル側にも同様のダイナミクスが生じています。具体的には、Kimi の API が完全に機能不全に陥っている状態です。
供給過多と需要過多のバランスが崩れ、実際には十分な供給がありません。そのため、このモデルは即座に拡散されるわけではありません。私は、これがパフォーマンスにおける時間差の問題としてどう関連するかを常に考えています。
00:11:54 Florian Brand: その通りですが、一方でここ数ヶ月でオープンエコシステムは驚くほどプロフェッショナル化しました。初期リリース時には、パートナーが事前に重み(weights)を入手し、vLLM のパッチも公開から数日、あるいは数週間前に提供されていました。これは 1 年前の状況とは全く異なります。当時は単に重みが公開されただけで、「モデル作成者は自分で解決してください」という状態でした。
そのため、一般提供初日の品質はまずまずのものになるでしょう。その後は、各プロバイダーが「より高速化」を競い合うレースが始まります。なぜなら、そこには莫大な名誉がかかるからです。
00:12:47 Nathan Lambert: なるほど。では 2 つの方向性について考えましょう。なぜ中国のモデルはこれほどまでに高性能なのでしょうか?
私の Discord サーバーで JSD と Epoch の間で行われた議論を記事にしましたが、私は今、中国のラボが「資本効率」において優れているという見解に傾いています。彼らは資本を計算資源(compute)、データ、そして人材へと変換する能力に長けており、それがモデルの性能向上につながっているのです。
この構造上の優位性が事実であれば、その原因は教育システムによって訓練された人材が、LLM の性能を高める問題解決に適しているからかもしれません。これは非常に重要な視点です。
どうやら中国では、計算資源や人材、そしてあらゆるコストが何らかの形で大幅に安いのでしょう。補助金があるのか、あるいは平均給与が低いからなのかは定かではありません。しかし、モデルの迭代が進む中でこれは非常に大きな意味を持ちます。もし次世代モデルが Anthropic では 100 億ドルかかるのに、Kimi のような中国のモデルでは 40 億ドルで済むとしたら、それは極めて大きな差です。ただ、なぜそうなるのかは明確ではありません。
例えば、Kimi のエンジニアである Big Eagle が私のツイートに返信し、「私たちはフロンティアを押し広げようとしているのではなく、追いつこうとしているだけだ」と述べていました。これはまさにマインドセットの問題であり、中国のラボが設定する目標のスコープの違いが、モデル構築にかかるコストを劇的に下げる要因になっているのかもしれません。
しかしここ一年、私は「中国製モデルは後退するのか?」という問いを何度もしてきました。このように訓練に巨額の資本が必要となるため、クローズドとオープンなモデルの格差は広がるだろうと考えていました。ところが実際には逆の方向へ進んでいるようです。これは簡単に説明できることではありませんが、中国のラボが我々が予想していた以上に追いつきつつある点について、あなたも同意しますか?
00:14:43 フロリアン・ブランド:実は、昨年の総括に基づいて今年に関する予測を再確認したのですが、その中で「格差は数ヶ月以内に縮まる」と述べていました。この予測は概ね的中しているようです。幸いにも、私たちが具体的な数字(3 ヶ月か 6 ヶ月か 9 ヶ月か)を提示していなかったため、その点ではセーフです。
しかし、私が中国で現地の人々と話し合い、研究者たちと接した際に感じたのは、彼らのチーム編成が 200〜300 人規模の若手(20 代半ば)で構成されており、全員が「1 つのモデルを本当に良くする」ことに集中しているということです。派生したタスクや余計な作業には着手していないように見えます。
計算資源(コンピュート)の観点については、中国製のチップが本格的に稼働し始めた現在、私たちが答えを出すのが極めて難しい問題です。実際、過去 6〜9 ヶ月でチップの密輸が大幅に増加したと私は考えています。密輸されたチップがようやくオンライン上で利用可能になり始めている状況です。
00:15:58 Nathan Lambert:密輸は輸出規制を回避する行為の総称です。チップがマレーシアにあり、そこで利用されている場合も同様にカウントされます。過去 6〜9 ヶ月でこの動きが劇的に増加しており、これが背景の一つにあると考えられます。ただ、前世代モデルを訓練していた頃と比べて、彼らが利用可能な計算リソースは格段に増えているのは事実です。
00:16:24 Florian Brand:はい。文脈として補足すると、先週 LongCat が中国製チップだけで完全に訓練されたモデルを発表しました。発表内容や我々の知る限り、これはほぼ間違いなく真実です。具体的な機種名は公開されていませんが、Huawei の Ascend シリーズではないかと推測されています。国内生産能力が高まるにつれ、これらのチップは主に訓練に使用される一方で、訓練の一部である推論処理にも特に有用です。
おそらく、学習段階では NVIDIA チップと他のチップを組み合わせて使用しているでしょう。しかし、推論フェーズにおいては、使用するハードウェアの割合が徐々に大きくなっており、この段階こそが非常に重要になります。彼らの全体的な計算リソースは増加していると見ています。また、実際のユーザー数は多くありません。ChatGPT が 10 億人のユーザーを支える必要があったり、Anthropic が数百から数千の企業をサポートする必要があるような規模ではありません。つまり、そのような大量の顧客がいないためです。
Nathan Lambert: はい。さらに、有料顧客についても言及できます。特にエンタープライズ向けでは、サポートを行う際に社内の時間やチャットリソースを消費します。研究部門に所属していない場合でも、業務外のタスクであっても、会社の注目を集めることになります。もし SSI が優れたモデルを発表すれば、それが「気晴らしが問題である」ということの究極的な証明となるでしょう。ただし、これは脇道の話なので後回しにしましょう。また、データや環境に関する業界の動きについても、すでにうわさが聞こえ始めています。
具体的な事例を覚えていますか?中国を訪れた際、外部データをほとんど活用していないことに驚きました。4 月の訪問から数ヶ月後の 7 月には、中国で新しい企業が現れ、データの購入を検討しているという話を聞きました。この変化のタイムラインは、少し皮肉なものです。
00:18:40 フロリアン・ブランド:彼らが実際に教えてくれたことには、誤差範囲を設けるべきだと思います。
00:18:44 ネイサン・ランバート:それもそのはずで、発表時期があまりにも近すぎて確実なことは言えませんね。
00:18:49 フロリアン・ブランド:そうですね。その可能性は否定できません。ただ、こうした要因を特定するのは難しいものです。それでも、外部データの購入がより大きな要素になりつつあるのは確かです。もしオープンモデル側もクローズドモデルと同じデータを、あるいは割引価格で入手できれば、追いつく手助けになるでしょう。実際、データ環境へのアクセスは後回しにされがちですが、それが要因の一つであることは間違いありません。その影響力の大きさがどれほどかは不明ですが、公的な洞察が得られるとも思えません。誰からも明確な情報が得られるとは考えにくいですね。だからこそ、オープンモデルが追いつき、スコアを向上させる理由の一つとなっているのです。
00:19:47 ネイサン・ランバート:では、中国の他のモデルプロバイダーについてまとめましょう。Kimi や Zhipu(GLM)についてはすでにお話ししました。今後、非常に優れた GLM 5.5 のような新モデルも登場するでしょう。Qwen についても、その最大規模モデルの話はしました。ただ、Qwen の大規模モデルは、小規模モデルの卓越した性能と比較すると、絶対的なパフォーマンスランキングでは同じレベルに達していない傾向があります。これはおそらく集中投資の結果と言えるでしょう。クラウド企業ならではの戦略です。少し目を細めて見れば、まるで Google のようです。
アリババの Qwen は、この分野で大きな機会を掴んでいます。特に、これらの小型モデルを通じて開発者を Alibaba のエコシステムに結びつけることは、同社のクラウド事業にとって極めて大きなチャンスであり、彼らはその点で驚異的な成功を収めていると考えられます。一方で、大型モデルはこれまで小型モデルほど卓越した性能を示してこなかったため、Kimi K3 や GLM 5.2 のような画期的な突破があるとは期待していません。ただし、中国の大手企業が名を連ねる中で主要なオープンソースモデルとしてニュースで大きく取り上げられることは間違いなく、"Open Bottle"(中国における大規模モデルの象徴的な呼び方)としての地位は確立されるでしょう。ただ、DeepSeek のように長期間にわたって話題が続くかどうかは別問題です。
00:21:01 Florian Brand:興味深いのは、この件をどの程度ご存知か分かりませんが、彼らはプレビュー版を利用できるエンドポイントを持っており、これを毎日更新しています。つまり、非常に迅速なイテレーションサイクルを実現しているのです。その結果、Twitter 上のベンチマークなどにおいて、我々の進捗がすべて反映されています。特に SVG や Three.js を用いた視覚生成タスクなどでは、直近の数日間でモデルの性能が大幅に向上しました。彼らは、他の企業も採用しているような高速フィードバック機構を確立したようです。Cursor などもこの点について多くのブログを投稿していますが、まさにそのイテレーション手法が鍵となっています。
原文を表示
Exciting news! My book trying to share post-training knowledge with the world is done and shipping soon. Order on Manning or Amazon. Thanks for the support. It’s currently the #1 AI book on Amazon :).
Nathan and Florian sit down to discuss everything happening with open models. Following the Kimi K3 release last week, it feels like everything is accelerating — geopolitics of US v China, economics of open vs. closed models, security at the frontier of AI, and so on.
Chapters:
00:00 Welcome & context
04:38 Living with / using Kimi K3
08:53 GLM 5.2’s continued role
12:47 How are the Chinese models this good?
17:41 Data, environments, and a tour of the Chinese labs
19:47 Roundup of Chinese providers: Qwen, DeepSeek, MiniMax…
24:08 The US open-model ecosystem
30:25 Frontier vs. near-frontier, and the cybersecurity case against bans
34:58 Distillation and the Ben Thompson debate
44:12 Predictions and a frontier tier list
48:36 Wrap-up
Listen on Apple Podcasts, Spotify, and where ever you get your podcasts. For other Interconnects interviews, go here.
Share
For more educational post-training videos, see the course I’m putting together.
Transcript
00:00:06 Nathan Lambert: Okay, welcome back to Interconnects. We’re doing our quarterly open model roundup, which is mostly us just making fun of or explaining, not making fun, why so many distillation takes are bad and understanding the state of where things stand. I think last Thursday was when Kimi K3 was released. I think we will see much much more in the near future. It seems pretty inevitable. Like over the weekend, Xi gave his speech where he directly committed to openness and open source as a strategy. It wasn’t a detailed layout state of affairs.
Qwen announced their next big model is going to be open weight, which is a big change of things. I think there’s just so much to get into. I think Flo you kind of were already going off on some of the performance gap and distillation takes. So we could probably start there and then as I go I have a little bit a little list and we could always go through the topics and the blog that I wrote which all are very nuanced. So I think we have infinite to talk about. So continue rant kind.
00:01:17 Florian Brand: Yeah, I think, or the biggest thing at every model release at least at every open model release is how much or how many months it is behind the closed frontier. Um and people love to put a definite uh definitive number onto this uh which is really really mudding because we have so many different benchmark providers these days and such uh so many different benchmarks as well that every site and I’m not innocent in that either um pulls up their favorite benchmarks to show that the current model or the newly released model is at the frontier which is then counted by the other side pulling up another benchmark and showing oh it’s actually a year behind or something. Um and it like a lot of it seemingly hinges on that question how many months we open models are behind.
00:02:26 Nathan Lambert: Yeah. So I my provocation is that some of the benchmarks are actually reasonably correlated with what people are doing and this is agentic coding and agentic computer use tasks and some of the benchmarks are correlated with the long tail which is where I think Claude and GPT is so valuable. But if it’s it’s like what is the market for Claude Code and Codex right now and if it is software engineering then like the fact that the models are say a couple months behind on that can be a very, very big deal and then I suspect that this model will be okay disclaimer the model weights aren’t out yet supposedly on 20 July 27th and a lot of the discussion will impinge on the assumption that they come.
But like people could post-train this model to very likely match Opus and GPT in many of these kind of niche domains that people want I think watching I mean we both have different views into the post-training open model industry, but there is a ton ton of excitement in progress on making these models like fine-tuned for specific high-value tasks and this has been historically done on a mix of like Qwen and GLM and GLM 5.2 really accelerated this and I I curious on the first person that puts out a blog post like we fine-tuned Kimi K3 on our task because I bet you could get big gains. I think even the you use Kimi K3 more than I do, but my hunch is that it would be a bit of a um rough edged post-training just by how big of a scale up it is and that normally means there’s a lot of performance that could still be extracted from it. No.
00:04:03 Florian Brand: Yeah. Running, running, running and especially post-training that one will be super hard because you need like one node of B300s to just load the weights which is crazy in terms of scale. So will probably take some time and uh a lot of engineering I’ve heard to actually get this into a state where it’s fine-tunable. Um but people you want to talk about using the model like you actually signed up for the the coding program and used it. So like getting this out there is good context.
00:04:38 Nathan Lambert: Yeah. So I signed up on day after release or so uh for the $200 plan uh which is their biggest one similar to to all the others but they have um like I think $40 and $100 as well. Uh but the biggest plan has uh 1 million context and I think or at least it feels like it has also some priority in terms of the API requests because so many people um online are saying that they hit API errors constantly and so far I’ve been I’ve been uh pretty well off if I’m uh going to say that. Um and in terms of model capabilities, it at some ways aside from front end where it is really good, it in some ways it really shines and it excels.
Um even my um expectations even with things like uh some research tasks like I have or at interconnects we now have over a year of data on on open models um and I ask the frontier models to come up with some interesting analysis which we haven’t done before in uh because we do our own analysis and have this published uh and I asked them all right do something new and um surprise me, basically. And a lot of the models or the frontier models u or basically all models latch onto the things we’ve done redo the data analysis part and then do some weird esoteric parts.
Uh Kimi K3 did some more interesting things um I’ve told it explicitly to scrape Reddit um and then it found uh some subreddits I haven’t even considered and then found out for example that the Reddit discussions are um one or two months in uh more recent or they found they find the interesting models one or two months before the download numbers usually take off like they are all onto Qwen and then the people download more Qwen models like those kind of analysis is groundbreaking um but it is something that Kimi surprised me at compared to to all the other frontier um models.
A simple question like can you use this for most of the core work you do in terms of like the exp you you have a distribution of stuff you tend to do most of them are with Codex I think you’re a Codex person rather than Claude person like what percentage do you think the Venn diagram overlaps where this model would be fine
00:07:24 Florian Brand: uh it’s really depends on how much leeway I give it like, the big thing I have seen with Kimi K3 right now I’m I’m working on u the framework we are doing at uh at Prime Intellect, where I work, and the main thing I found with Kimi is its code is a lot simpler uh which makes it way more readable uh but it misses some things that Codex just or like we’re talking 56, 55 and especially 54 would be on those levels. So, I would say that Kimi K3 is like 54-55 level for these kind of tasks.
But if I like I read the code and I say all right that’s really good code and then I give it a pass over with with Codex and it finds all these niche niche cases where it doesn’t excel but for supervising runs or for running uh some experiments it is actually really usable. Um and for some other niche things like you can just let it run. The one downside is but that’s also because the API is completely swamped in terms of users and it their servers are in China. The wall clock time is significantly significantly higher than GPT. But I would say like if I was to to push it and use it in my daily workflow, I would be slower, but I wouldn’t be slowed down by so much that I would say, “All right, that’s unusable.”
00:08:53 Nathan Lambert: And how does this compare to GLM 5.2? Because GLM 5.2 was still a story unfolding in my opinion where like I would go bop around SF and people are like yeah I genuinely use this for this part of my like agentic coding and/or workflow. Um, how do you like I feel like were you in that camp using GLM at all or
00:09:20 Florian Brand: Yeah. like where do you I also use used and use uh GLM mostly because we have an internal endpoint which is really fast and we have or or before that I I also used an API which had I don’t know 200 or 300 tokens per second. Um and if you can do a lot of task at a good enough level like really fast you just use that model compared to going to Codex then selecting the lesser model then selecting the right reasoning effort then selecting fast like I just use GLM get the same result and uh and it’s uh pretty fine like it it definitely is Sonnet-ish level in terms of capabilities and for a lot of cleanup task for a task that just is grunt work. It really works. Like I I would say you could probably go really far for a lot of the work uh with Kimi K3 as the main agent and GLM for for sub agent work.
00:10:24 Nathan Lambert: Something that’s pretty different with Kimi’s announcement and the scale of models this is. I think it’ll take a bit longer for these open models to really be optimized and available across the inference providers. Like GLM 5.2 is pretty fast, but one, we don’t have the weights yet, and then two, like I don’t think it’s going to be as fast of a roll out on adoption as the like 500B, 700B MoE. like I I there’s going to be more problems there which is a very different regime where in the past the Chinese models would finish their RL run and release the model with open weights within hours to days maybe a week and then like immediately the ecosystem kind of knew how to do this.
I think there’s a lot bigger of an infrastructure kind of uplift on this next scale of open weight models which I think we have to factor in like the closed labs do this behind the scenes before announcing the models. So it’s just like that is kind of manipulating the time gap in a way where it could be like an extra month before people can actually post-train and use Kimi at scale for their workflows. And like we love to say as an open weight fan like oh it’s only when the closed model is available that you could take the time gap but like now there’s similar dynamics in open models where it’s like the Kimi API is totally broken. There’s too much supply. There’s too much demand. There’s not enough supply. So it’s not like this model is immediately diffusing like the I’m just I’m just thinking about this as it relates to the performance time gap
00:11:54 Florian Brand: because that that is true but on the other hand the open ecosystem has professionalized quite a lot in the last few months. uh like during or in your initial roll out they all come with some partners which have the weights beforehand. They have the vLLM patches out days or or even weeks before these days which is completely different from from a year ago where basically weights got dropped and the model makers were like all right you got to figure this out. So I expect like the general availability on day one will be pretty okay and then race starts of all the providers starting to optimize to get even higher and higher speeds because it’s so much prestige.
00:12:47 Nathan Lambert: Yeah. Okay. Two directions to go. Why do we think the Chinese models are able to be this good? I think I’ve wrote about there’s a debate in our Discord with in with JSD at at Epoch and I think it’s very good and I had this section in my piece that I’m like coming around to think that the Chinese labs are more capital efficient and you can turn capital into compute data and talent in a way that makes the models better and this is really I think this is super important if it actually is some structural advantage whatever the cause I think the cause could be talent is better trained for whatever their education system was to work on problems that make LLMs better.
It could just be that all the compute and talent and everything cost way less in China somehow. Whether it’s a subsidy, whether it’s just average pay being lower. But this is a very big deal as we turn the crank in the model iterations. And if a next generation model costs $10 billion for Anthropic but only $4 billion for Kimi like this this is like could be very huge but it’s not clear why this is the case. For example, I think Big Eagle the Kimi engineer replied to my tweet on this and was like it helps because we’re not trying to push the frontier. are just trying to catch up, which really could be a mindset thing where how the the goals of the labs are scoped in in China so that it cost them way less money to build these models.
But I in the last year we’ve asked a lot of questions on like will the Chinese models fall off. I have thought that the gap between closed and open bottles would grow due to this kind of capital intensity of training and it seems like it’s going the opposite direction which is just like it’s it’s hard to unpack but like do you agree that the labs are keeping up a bit more than we would have expected as in the Chinese labs and why?
00:14:43 Florian Brand: Well, I I actually looked at our uh predictions for uh for this year based on our last year’s recap and we basically said that the gap will stay with within a few months. Uh so that prediction seems to largely hold. Um luckily for us, we didn’t put a concrete number whether it’s 3 months, 6 months or 9 months. So we are safe on that side. Um but I think like the general thing we both felt when we were in China and talking to these people like they are like the researchers themselves are teams of two or 300 people all mid20s and all just want one model to be really good like they don’t seem to do any side quests.
They don’t seem to do anything that uh deviates from from these things. And um they might or in terms of compute which is a really hard question for for us to answer especially as uh these Chinese uh chips are now coming online. We have I also think chips I think chip smuggling increased substantially in the last like 6 to 9 months or the chips that have been smuggled started to become online.
00:15:58 Nathan Lambert: Smuggling is a general term for getting around export restrictions. If the chips are in Malaysia and they’re using them, I I count that similar and I think that that has massively increased in the last six to nine months, this is the partially the result of that and and you’re saying but I just wanted to put that out there of like I do think that they have a lot more compute though than they did when they were training the previous generation of models.
00:16:24 Florian Brand: Yeah. like we like or just for for context two weeks ago I think LongCat released their model which they uh claim and I we know that it is very likely true uh is trained entirely on uh on Chinese chips. uh they didn’t specify publicly which ones but people speculate that it’s uh that it’s some uh Ascends from Huawei um and as the domestic production ramps up and you can they’re probably used most or they are used for for training but they are especially useful for inference which is a huge part of training as well.
So they probably use some mix of uh of Nvidia and other chips for the training part and then an increasingly larger part for the inference part during which during the stage is is really important. So I think their overall compute is increasing and also they don’t actually have a lot of users. So they don’t need to power 1 billion users like ChatGPT has to do, hundreds or thousands of enterprises like Anthropic has to do because they don’t have that magnitude of uh of of paying customers.
00:17:41 Nathan Lambert: Yeah. And I think even those paying customers also, at least on the enterprise side, there’s just like there is company time and chatter when you’re supporting these things. Even if you’re like not a research, even if it’s not in your job, it like does change the attention of the company. And if SSI comes out with a good model, it’ll be the ultimate validation that distractions are are a problem, but that’s an aside that we we can wait on. I think the there’s also rumblings of the data and environments industry starting to appear there.
Do you remember any specific ones? Because when we were in China, it was kind of shocking how little they seem to utilize external data. So just a few months hearing a whole bunch of a month months after our trip we went in April and then just months later in July, we’re are hearing a few things of like new companies in China and them wanting to buy data and things. And that is uh like a funny timeline of how that changes.
00:18:40 Florian Brand: And I would put error bars on what they actually told us.
00:18:44 Nathan Lambert: And that cuz it’s like so close in time that I don’t know.
00:18:49 Florian Brand: Yeah. That that that might that might be true. Uh but like those things are hard to to pinpoint. I but I would say it it seems like the buying of external data is becoming more of a factor. Um which will help the open models catch up to the closed ones if they just buy the same data maybe at a discount because um they buy the the data environments later. But it is it is a factor. How big of a factor like we don’t know. we don’t have any public insights and I doubt that we will get those insights uh from from anyone b uh really uh so that’s definitely one of the parts why um why we are able to to catch up or improve their their model scores.
00:19:47 Nathan Lambert: Okay, roundup of other Chinese model providers. We’ve talked about Kimi, we talked about Zhipu / GLM. I think there will be more GLM models soon that are very good. They might call it like GLM 5.5. Um Qwen, we talked about their biggest model coming. Qwen’s biggest models I will say have tended to relative to the excellence of their small models not had the same like absolute ranking in performance which is a probably a cost of focus. I think it goes with a cloud companies. It’s it’s almost like it’s if you squint it’s almost like Google.
It’s like Qwen has Alibaba has so much opportunity here and the opportunity of getting developers associated with Alibaba Qwen with these small models is such a huge opportunity for their cloud that I think they’re succeeding wildly. But their big models have always not been as excellent as their small models. So I don’t expect their model to be as breakthrough as Kimi K3 or GLM 5.2. I expect it to be covered in the news as major open quite as the open bottle name in China drops giant bottle but I don’t think it will be as sustained as a um news story um DeepSeek you can go if chime in whatever
00:21:01 Florian Brand: the the interesting thing is don’t know how how much you follow this but they are have or they have an endpoint which you can use for a preview version and they’ve updated this endpoint daily so they have some really fast iteration cycle because we the we progress in all these um Twitter um benchmarks. So a lot of these SVG things and three.js like all these visual generation tasks the model has been improving a lot over the last few days. So they have figured out some kind of fast feedback mechanism um which other companies have as well. Uh we we know this or cursor has a lot of blogs about this how they iterate
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み