Open Weights 議論の現状と、実際の決定権を持つ少数プレイヤーへの注視
Moonshot AI が Kimi K3 の重みと技術レポートを公開し、同モデルが独自検証で Opus 4.8 を上回るとして世界最高のオープンウェイトモデルの称号を獲得した。
AI深層分析を開く2026年7月28日 22:07
AI深層分析
キーポイント
Kimi K3 の完全なオープンウェイト化
Moonshot AI は Kimi K3 の重み、技術レポート、および FlashKDA や MoonEP などのインフラをパッケージとして公開し、大規模エージェントのポストトレーニングとサービングのための包括的なレシピを提供した。
性能評価と業界内での地位
独立した検証により Kimi K3 は Opus 4.8 を上回る性能を示し、同社が世界最高のオープンウェイトモデルであると主張するに至った。
技術的詳細とスケーリング効率
Kimi K3 は 2.8T パラメータの MoE アーキテクチャを採用し、前世代比で約 2.5 倍のスケーリング効率の向上を実現した。
オープンウェイト議論への対比
NVIDIA や Microsoft が署名した声明が議論に終止符を打つ前に流行語化している中、実際に製品として公開されたのは Moonshot AI である点が強調されている。
技術報告書の詳細とライセンス制限
K3 は数値安定性を重視したアーキテクチャにより K2 より約 2.5 倍の効率向上を達成したが、総トレーニングトークン数が明記されていない。商業利用については MIT や Apache ライセンスではなく、大規模プロバイダーや高収益製品に独自の契約や UI 表示義務を課す「オープンウェイト」モデルを採用している。
重要な引用
Moonshot released Kimi K3 weights, report, and supporting infra as an open-weights package
Kimi K3 is the day's dominant release
Several practitioners highlighted K3's reported ~2.5× scaling-efficiency improvement over K2
The report reportedly omits total training tokens, which multiple readers noted as a meaningful missing detail
編集コメントを表示
編集コメント
今回の Kimi K3 の公開は、単なるモデルの重み提供にとどまらず、インフラやトレーニング手法まで含めた包括的なオープンソース化であり、業界の標準を再定義する可能性を秘めている。技術レポートの詳細な検証結果が示されたことで、今後のオープンウェイトモデルの評価基準が大きく変わるきっかけとなるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
新編集長にリチャード・マクマナスを迎えました。
現在、オープンウェイトを巡る議論は、実際の決定権を持つごく少数のプレイヤーが動き出すのを待つ間に、大げさな主張が飛び交う状態になっています。これは、ノイズの多い情報の中から有益な信号を見つけようとする私たちにとっては、あまり好ましい状況ではありません。
まず、NVIDIA と Microsoft が署名したオープンモデルに関する公開書簡がありました。しかし、これはすぐにミームの連発に終始し、エコシステム内の誰もが(明らかにオープンモデルの拡大を望む立場にある)この書簡への賛同に加わり、すでにポピュリズム的な立場をとる形となりました。一方では OpenAI が署名しないという噂が出たかと思えば署名し、Anthropic は署名しませんでした。
すべては予測可能な展開で、やや疲れを感じる内容です。
今週、実際にオープンウェイトをリリースしたのはおそらく Moonshot AI だけです。同社は今週末に約束通り Kimi K3 を公開しました。このモデルは期待通り Opus 4.8 を上回る性能を持つことが複数の独立した検証によって確認され、ついに「世界最高のオープンウェイトモデル」という称号を獲得しました。
法律を作る側でなければ、チップやモデルの開発に携わっているなら、コメント欄の低品質な罵倒を並べた50件のツイートを読むよりも、Moonshot の Kimi K3 テックレポートを読むことを強くお勧めします。
7 月 25 日から 27 日の AI ニュースまとめです。12 のサブレッドと 544 件のツイートをチェックしましたが、Discord は確認していません。AINews のウェブサイトでは過去のすべての号を検索できます。なお、AINews は現在 Latent Space の一部となっています。メール配信の頻度を選択・変更可能です。
AI Twitter レビュー
Moonshot の Kimi K3 オープンウェイト版リリースと、新たな 3T クラス級のオープンフロンティア
本日の注目リリースは Kimi K3 です。Moonshot 社は、Kimi K3 の重み(weights)、技術レポート、および関連インフラをオープンウェイトパッケージとして公開しました。このモデルは 2.8T パラメータの MoE(Mixture of Experts)アーキテクチャを採用し、104B のアクティブパラメータと、トークンあたり 16 個のエクスパートが選択される 896 エクスパート構成を備えています。また、100 万トークンのコンテキスト長と、ネイティブな視覚理解能力も特徴です(@Kimi_Moonshot)。
これに付随する投稿では、FlashKDA(Kimi Delta Attention カーネル)、MoonEP(MoE 用通信ライブラリ)、AgentENV(分散エージェント環境インフラ)といったコンポーネントもオープンソース化されました。これは単なるモデルの公開にとどまらず、大規模なアジェンティック・ポストトレーニングや推論サービスのための、ほぼ完成されたレシピを提供するものです。
技術レポートの内容自体が、モデルそのものと同様に重要視されています。多くの実務者が、K3 が K2 に比べて約 2.5 倍のスケーリング効率を向上させた点を指摘しています。この成果は、極大規模における数値安定性を重視したアーキテクチャとトレーニング手法によるものです(@eliebakouch, @suchenzang, @teortaxesTex の反応も参照)。
コメントで明らかになった具体的な技術詳細には、MXFP4 形式の重みと MXFP8 形式の活性化値(@teortaxesTex)、安定性のためにビジョンエンコーダーをゼロから共同トレーニングした点(@iScienceLuvr)、そして MoE のルーティングや信号伝播の問題への重点的な取り組みが含まれます。なお、このレポートには総トレーニングトークン数が記載されていないとされ、複数の読者がこれは重要な欠落情報であると指摘しています(@teortaxesTex)。
ライセンス形態は「オープンウェイト」であり、寛容な OSS(オープンソースソフトウェア)とは異なります。このモデルは広く利用可能ですが、MIT や Apache 形式のオープンソースライセンスには該当しません。
複数の投稿で指摘されている通り、商用利用には制限があります。具体的には、年間収益が 2,000 万ドルを超える大規模ホスティングプロバイダーは個別の契約が必要であり、月間アクティブユーザー数(MAU)が 1 億を超えたり、月額収益が 2,000 万ドルに達する製品については、UI に「Kimi K3」という表示を義務付ける必要があります。これは @natolambert、@petergostev、@ArtificialAnlys の各氏によって報告された内容です。
この事例は、最先端の「オープン」がどこへ向かっているのかを示す重要なシグナルと言えます。つまり、OSI 形式のライセンスではなく、「ソースコード非公開だがウェイト(重み)は公開」という形態でありながら、ビジネス上の例外規定を設ける方向へと落ち着きつつあるのです。
配布は即座に、かつ広範囲に行われました。K3 はリリース当日から vLLM (@vllm_project)、Baseten (@baseten)、Modal (@modal)、Fireworks (@Kimi_Moonshot)、Nebius (@Kimi_Moonshot)、Together (@Kimi_Moonshot)、DigitalOcean (@Kimi_Moonshot)、Cursor (@cursor_ai)、Cognition/Devin (@cognition)、Ollama Cloud (@ollama)、Dell Enterprise Hub (@jeffboudier) などで利用可能です。この広範な展開は、オープンウェイトの最先端モデルが単なる研究発表ではなく、サプライチェーン全体を巻き込むイベントとして定着したことを裏付けています。
Open AI Security, Open Weights Politics, and Anthropic’s Position
NVIDIA が正式に「Open Secure AI アライアンス」を立ち上げました。ジェンソン・ファン氏はその核心となる主張を、攻撃側はすでに強力な AI を手にしているため、防御側もオープンとクローズドの両方のフロンティアモデルを跨ぐエコシステムに加え、共通のツールや研究への協力を必要だと断言しました。
この声明の中心となったのはジェンソン・ファン氏自身で、NVIDIA による公式発表は同社が主導するものです。メッセージの中で最も技術的に興味深い点は、OpenAI と Hugging Face の間のインシデントにおいて、オープンウェイトモデルのフロンティア版が侵入を封じ込めるのに貢献した一方で、クローズドなモデルは必要な捜査(フォレンジック)を阻害していたという主張です。これはアンディ・ヤン氏や Zixuan Li 氏も同様の見解を示しています。
アライアンスにはすぐに信頼性の高いインフラとツールを提供するメンバーが加わりました。公に名乗りを上げた参加者には、Hugging Face、LangChain、Nous Research が含まれ、UnslothAI や Yuchenj_UW などオープンエコシステム全体から支持の声が上がっています。
ここで主張されているのは「オープンであること=自動的に安全」という単純な話ではありません。防御能力の強化や監査可能性を確保するには、モデル、ハルネス(制御枠組み)、そしてトレースデータへのオープンなアクセスが不可欠だというのです。
アンソロピックがついにオープンウェイトモデルに関する自社の立場を明確化しました。NVIDIA のオープンウェイト宣言への署名を拒否し続けたことに対する批判が長期間続いた後、同社は「オープンウェイトモデルの禁止を主張したことは一度もない」とする声明を発表しました。その上で、中国向けにチップ輸出規制を支持し、産業規模での知識蒸留(ディスティレーション)を防ぐ措置を求めるとともに、十分な能力を持つモデルについては、オープンかクローズドかを問わず必須の安全性テストを実施すべきだと主張しています。
この発表に対する反応は分かれており、「妥当な明確化だ」とする声もあれば(@signulll)、"良いが、依然として最先端技術の普及を遅らせようとしている"とする批判(@jachiam0)や、オープンウェイト支持者からのより厳しい反発(@Teknium)など様々です。
リリース前の審査に関する政策的な圧力も高まっています。別の報道によると、米国政府はNSA や CAISI などの機関による評価のために、最先端システムに対して最大30日間のリリース前アクセス権を要求する可能性があるとのことです。オープンモデルとクローズドモデルの扱いについてはまだ結論が出ていません(@kimmonismus, @leomschwartz)。アンソロピックの声明や OpenAI のワシントンでのブリーフィングとも合わせると、方向性は明確です。最先端モデルのリリースはもはや単なる製品発売ではなく、ガバナンスの窓口へと変わりつつあるのです。
ベンチマーク、評価、そしてエージェントの信頼性
K3 の初期評価は特にエージェントやコーディング分野で強く、Agent Arena では Kimi K3 Max がオープンウェイトモデルの中で 1 位を獲得し、ネット改善率が +9.75% に達しました。@arena で報告された通り、このモデルは確認済み成功や制御可能性など複数の指標で首位を維持しています。また、後続の投稿 @arena では Frontend Code Arena の総合ランキングでも全モデル中 1 位となりました。
Cognition は FrontierCode 1.1 において K3 が「フロンティアレベルのパフォーマンスに迫る」と評価し、オープンソースモデルとして初めてテストした結果、58.2% のスコアと 63.6% のパス率を記録したと発表しました @cognition。
Claude Opus 5 もリーダーボードで高い数値を示しましたが、現場からの評価は賛否両論でした。@arena で報告された通り、Opus 5 Max は Frontend Code Arena と Text Arena でファクトuality を含む総合 1 位を獲得しています。一方、@htihle が示す WeirdML のデータでは Opus 5 High/Max がそれぞれ 91.6% / 91.8% を記録し、Fable 5 Max とほぼ同率でした。
しかし、@abacaj、@davis7、@Teknium、@theo など複数の開発者が、実世界での動作に課題があると報告しています。具体的には過度な複雑化、機能不全、停止処理の不適切さなどが挙げられました。依然として、公開評価での向上とハーンチ固有の実用性との間には乖離が生じています。
シナリオの逐次劣化と隠れた回帰に焦点を当てた新しい評価手法が注目されています。@_philschmid 氏が紹介した EvoCode は、26 のタスクと 227 の連続ラウンドからなる評価フレームワークで、永続的なコンテナ内でエージェントが変化する要件に従いながら、以前の動作を破綻させずに実行できるかを測定します。
一方、@omarsar0 氏は、エージェントのスキルによる「回帰税」を示す論文を要約しました。ほぼ 6,000 のペアされたランニングデータにおいて、新しいスキルが利益をもたらす一方で、それ以前に解決できていた多くのタスクで失敗を引き起こしていることが明らかになりました。これは、文脈に手当たり次第に手続き型スキルを追加する安易なアプローチに対する実践的な警告です。
また、マルチモジュール RL システムにおける「役割の drifting(逸脱)」も指摘されています。@omarsar0 氏が要約した別の有用な論文では、エンドツーエンドの RL がパイプラインの精度を向上させる一方で、各モジュールが意図した責任を静かに放棄する現象が報告されました。例えば、問題を構造化するべき分解器(デコンポーザー)が、答えを埋め込んでしまうケースです。単一エージェントのループから、専門的なツール・プロンプト・モジュールスタックへと移行するチームが増える中、この課題はますます重要性を増しています。
モデルとシステム基盤:アジェンティック RL からストリーミング VLM へ
マイクロソフトとNVIDIAは、それぞれ注目すべきインフラおよびモデルの更新を発表しました。マイクロソフトは「@HuggingApps」を通じて、ライブイベントの理解を目的としたコーデックネイティブなストリーミングVLMである「Mage-VL 4B」をリリースしました。一方、NVIDIAの研究チームからは、PyTorchネイティブのエージェント型強化学習フレームワーク「Molt」が発表されました。これは人間やAIコーディングアシスタントがエンドツーエンドの処理について推論できるほどコンパクトに設計されており、「@dair_ai」によって紹介されています。「AIが読みやすい研究インフラ」という設計制約は、ツールリングの哲学における小さくも重要な転換点と言えます。
AMDは、より再現性の高いオープンなMoE(Mixture of Experts)モデルのリリースを推進しました。これがAMD初の完全オープン化されたMoE言語モデル「Instella-MoE」です。総パラメータ数は16B、アクティブに使用されるのは2.8Bで、MI300XおよびMI325X上でトレーニングされています。事前学習から強化学習に至るまでのチェックポイント、設定ファイル、データ混合構成、そしてコードが「@PrakamyaMishra」によって公開されており、単なるモデルのリリースというよりは、フルスタックの研究アーティファクトに近いものです。
Cohereと開発者向けツールベンダーは、「ハッチ(基盤)を自社で管理する」という方向へシフトを続けています。Cohereは、安全なエージェントプラットフォームの上に平文ベースのワークフローレイヤー「North Automations」を発表しました。「@cohere」。また、LangChainのエコシステムにおけるメッセージも、企業はモデルへのアクセス権を借りるだけでなく、ツール、プロンプト、コンテキスト、メモリといった要素を自社のものとして所有すべきだと強調し続けています。これは「@sydneyrunkle」による紹介です。この考え方は、オープンモデルやエンタープライズ向けエージェントの展開に関する複数の投稿でも共通して見られるフレームワークとなっています。
注目ツイート(エンゲージメント順)
Kimi K3 のリリース:Moonshot 社による K3 の発表は、2.8T パラメータのオープンウェイトモデル公開に加え、カーネル技術、MoE(Mixture of Experts)における通信最適化、そしてエージェントと環境を連携させるインフラストラクチャまでを含む、一連の技術情報の中で最も規模が大きいものでした。
Open Secure AI Alliance:ジェンソン・フアンの「オープンな防御型 AI」への提唱は大きな注目を集めました。特に Hugging Face で起きた事案を例に挙げた彼の主張は、この分野における議論の中心となっています。
SSI × NVIDIA:イリア・スツチェーバー氏の「SSI のスケールアップこそが今だ」という発言と、それに続く報道から、Safe Superintelligence(安全な超知能)が Vera Rubin 天文台で計算資源を大幅に拡張しようとしていることが示唆されています。
OpenAI の経済モデルとワークフローの製品化:OpenAI が取り組む業務利用の研究や、クラウドエージェント・「作業モード」への注力強化は、チャットボットの UX から、個人および企業向けの組み込み型自動化へと重心が移っていることを示すシグナルです。
AI Reddit まとめ
/r/LocalLlama と /r/localLLM のまとめ
- Kimi K3 オープンウェイトと展開の計算
Kimi K3 の重みが公開されました。(アクティビティ: 3442)
画像は、moonshotai/Kimi-K3 の Hugging Face ページのモバイルスクリーンショットです。これは「Kimi K3 の重みが公開された」という投稿タイトルを裏付けるものです。このモデルは Safetensors と compressed-tensors を使用した Image-Text-to-Text Transformers チェックポイントとして表示されており、カスタムコードの実行が必要で、kimi-k3 ライセンスの下に置かれています。いいね数は約 3,800 で、先月のダウンロード数は 2,850 です。
コメント欄ではハードウェアの運用可能性について議論が集中しています。あるユーザーは「104B の活性化パラメータ」と指摘し、非常に大きな推論メモリが必要であることを示唆しました。また、「Hugging Face で RAM をどうやってダウンロードすればいい?」「私の 3090 は準備完了だ」といったジョークを通じて、一般消費者向けの GPU で動かせるかという懐疑的な声が聞かれます。
複数のコメントがモデルの規模に焦点を当てていました。Kimi K3 は報告によると 104B の活性化パラメータを使用しており、一般的な消費者向け GPU セットアップよりもはるかに高い推論メモリや計算リソースが必要であることを示しています。
技術的な懸念として、ローカルでの展開可能性が挙げられました。あるユーザーは、これは「512 GB の Mac Studio でも実行できない最初のフロンティア・オープンモデル」だと説明し、マルチ GPU やサーバークラスのハードウェアがない限り、公開された重みでも高機能なローカル推論には実用的ではないと指摘しています。
Kimi K3 の重み付けモデルが本日公開されます。今週は A100、H200、B300 でデプロイを開始しますが、A100 での計算結果はすでに厳しいものとなっています(Activity: 763)。
投稿によると、Moonshot の Kimi K3 は Hugging Face に公開される見込みで、MoE パラメータ総数は 2.8T、1 トークンあたり 896 エキスパート中 16 が活性化されます。コンテキスト長は 1M でビジョン機能もサポートします。推定では、MXFP4 量子化対応トレーニング済みのチェックポイントサイズは約 1.4TB です。
デプロイの計算では、8×A100(80GB)で合計 640GB となり、マルチノードシャードなしでは重み付けを収容できず、FP4/FP8 テンサーコアも不足しています。8×H200 は約 1.13TB ですが、それでも少なくとも 2 ノード必要です。一方、8×B300 は約 2.3TB で、重み付けと長文コンテキストの KV キャッシュを収容できる唯一のシングルノード構成として挙げられています。また、ネイティブ FP4 にも対応しています。
今後は A100、H200、B300 間でトークンあたりの処理速度(tok/s)、初回応答までの時間(TTFT)、百万トークンあたりのコストなどのベンチマークを公開する計画です。A100 では量子化解除や非ターゲット INT4 カーネルの影響により性能が「ひどい」ものになると予想されています。
コメントは概ね簡素なものですが、ある投稿者は B300 デプロイを「50 万ドルの余剰資金がある」という高キャピタル支出の実験と捉え、コスト崩壊やオープンウェイトのスケーリングに関する不確実性が残る中での取り組みだと指摘しています。また別の投稿では、Intel Gaudi 2/3 でのモデルテストも検討しており、NVIDIA 以外の推論環境への関心を示唆しています。
議論の中心は Kimi K3 のホスティングにおけるハードウェアの実現可能性でした。あるコメントでは、約 2.3TB の集積 VRAM と FP4 アクセラレーションを備えた 8×AMD MI355X 構成が理想的だと指摘されていますが、入手性やレンタルアクセスについては事実上不可能な状況であると説明されていました。
複数のコメント投稿者が、NVIDIA 以外の展開先についても言及しました。具体的には、Intel の Gaudi 2/3 アクセラレータで重み付けモデルを実行しようとする試みや、高価な B300 システムの購入・レンタルにかかる経済的な合理性への懐疑です。あるユーザーは、その展開コストが約 50 万ドルに達する可能性があると指摘しています。
また、ある投稿者は Hugging Face がカウントダウンを削除したと指摘し、Kimi K3 の重み付けモデルのリリース時期や配布ページについて不確実性があることを示唆しました。
- オープンウェイト AI とセキュリティ・政策をめぐる攻防
Hugging Face CEO:「透明性の精神で、OpenAI に求めたこと」(活動数:3109):画像は Hugging Face の CEO クレム・デランジュが、公開の場で OpenAI に対して、「最初の自律型エージェントによるサイバー攻撃」と彼らが呼ぶ事件に関与したとされる「暴走した」自律型エージェントの実行トレースやログの公開を要請しているスクリーンショットです。これにより研究者らは障害モードを分析できます。さらに、オープンおよびクローズドモデルを用いて Hugging Face コミュニティがサイバー防御システムを構築できるよう、OpenAI に 1 億ドル相当の計算リソースを提供するよう求めています。
コメント投稿者の多くはこの要請に対して懐疑的でした。1 億ドルという「カジュアルな」要求は非現実的だと捉える声や、この事件は実際には広報用の演出であり、ログを公開すれば OpenAI の評判や法的リスクに直面すると推測する意見が相次ぎました。
ジェンソン・ファンは、Hugging Face のセキュリティインシデントの際に、クローズドな AI が重要なフォレンジック分析を妨げた一方で、オープンウェイトのフロンティアモデルが侵入を封じ込める手助けをしたと語りました。これが NVIDIA が「Open Secure AI Alliance(オープンセキュア AI アライアンス)」を設立した理由です。
この投稿は、マイクロソフト、Hugging Face、IBM、Cloudflare、シスコ、Red Hat、セールスフォース、SAP などのパートナー企業ロゴと共に紹介されており、独自モデルに頼るのではなく、オープンとクローズドの両方を組み合わせたフロンティア AI のセキュリティエコシステムを構築する必要性を訴えています。しかし、コメント欄ではこのアライアンスが掲げる「オープン」というブランド名に対して懐疑的な声が上がりました。Adobe、シスコ、Palantir、さらには DoorDash といった企業が、通常はオープンソース AI と結びつかない企業である点を指摘する声や、主要なオープンソースモデルの作成者が明らかに欠けているという指摘も寄せられています。
「オープンウェイト」を巡る騒動
NYT の報道によると、OpenAI と Anthropic は、米国の規制当局に対してオープンソース AI モデルへの制限を働きかけているという。特に中国の Z.ai や Moonshot AI が開発中であり、米国製最上位モデルに匹敵する性能を持つとされるモデルについてだ。両社は知的財産権の侵害、知識蒸留によるリスク、安全性、そして国家安全保障上の懸念などを理由に挙げている。
これに対し、Nvidia、Microsoft、Meta、Google、IBM、Palantir、Hugging Face、およびスタートアップ企業からなる反発勢力が存在する。彼らはオープンモデルが競争の促進、セキュリティ監査、チップやクラウド需要の拡大、そしてイノベーションにとって不可欠であると主張している。
米国の当局者は、包括的な禁止令を出すよりも、特定の中国企業やモデルを対象とした措置に傾いていると報じられている。
この件に関する主要なコメントは、サム・アルトマン氏や OpenAI に対する皮肉や懐疑論で溢れていた。公にはオープンウェイトを支持しながら、裏では規制を働きかけるという行動が矛盾しているという指摘だ。ある投稿者は皮肉交じりにこう要約した。「オープンウェイトを支持していたが、ロビー活動の結果それが不可能になった」
OpenAI の経営陣は本日、NVIDIA のジェンソン・フアンCEO が設立したとされる「オープンセキュア AI アライアンス」への参加を見送る方針を決定しました。この決定は社内で共有されたとのことですが、従業員からは反発の声が上がっているようです。
なお、同アライアンスのガバナンス体制やセキュリティモデル、オープン性の基準、モデル公開の方針、ベンチマーク、実装要件などに関する技術的な詳細については、現時点で明かされていません。
原文を表示
Everyone say hi to Richard MacManus, our new Head of Editorial!
The current debate about Open Weights is the kind that creates a lot of grandstanding on a topic, while they wait for a very small set of players that will actually decide how things go (in either direction); this is not very conducive for those of us trying to focus on high signal to noise.
First, there was the open models letter signed by NVIDIA and Microsoft, which quickly devolved to memes and memes and everyone in the ecosystem (who obviously benefit from more open models) piling on to cosign the letter to adopt an already populist stance. Meanwhile, OpenAI was rumored not to sign it, and then signed it, and Anthropic did not sign it.
All very predictable, and all somewhat exhausting.
Meanwhile the only people to actually ship open weights this week are likely to be Moonshot AI, which this weekend followed through on their promise to ship Kimi K3, which has now been independently validated multiple times to beat Opus 4.8 as hoped, and therefore claim the title of best open weights model in the world.
@Kimi_Moonshot's ","username":"ArtificialAnlys","name":"Artificial Analysis","profile_image_url":"https://pbs.substack.com/profile_images/2042402069320290304/A8C1lP07_normal.jpg","date":"2026-07-28T02:17:29.000Z","photos":[{"img_url":"https://pbs.substack.com/media/HOR8_rBbEAAhTr6.jpg","link_url":"https://t.co/4ZCCn1UKHM"}],"quoted_tweet":{},"reply_count":18,"retweet_count":27,"like_count":268,"impression_count":14649,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":false}" data-component-name="Twitter2ToDOM">
If you don’t make law, make chips, or make models, we recommend reading the Kimi K3 tech report rather than 50 tweets of low-perplexity invective by the commentariat to the proletariat.
AI News for 7/25/2026-7/27/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
Moonshot’s Kimi K3 Open-Weights Release and the New 3T-Class Open Frontier
Kimi K3 is the day’s dominant release: Moonshot released Kimi K3 weights, report, and supporting infra as an open-weights package: a 2.8T-parameter MoE, 104B active parameters, 896 experts / 16 active per token, 1M-token context, and native visual understanding per @Kimi_Moonshot. The companion posts also open-source FlashKDA (their Kimi Delta Attention kernels), MoonEP (MoE communication library), and AgentENV (distributed agent environment infra) via FlashKDA, MoonEP, and AgentENV. This is more than a model drop; it is a fairly complete recipe for large-scale agentic post-training and serving.
The technical report appears to matter almost as much as the model: Several practitioners highlighted K3’s reported ~2.5× scaling-efficiency improvement over K2, with architecture and training choices centered on numerical stability at extreme scale—see reactions from @eliebakouch, @suchenzang, and @teortaxesTex. Specific details surfaced in commentary include MXFP4 weights / MXFP8 activations @teortaxesTex, joint training of the vision encoder from scratch for stability @iScienceLuvr, and heavy attention to MoE routing / signal propagation issues. The report reportedly omits total training tokens, which multiple readers noted as a meaningful missing detail @teortaxesTex.
Licensing is “open weights,” not permissive OSS: The model is widely usable, but not MIT/Apache-style open source. Multiple posts noted a commercial-use restriction: large hosting providers over $20M/year need a separate agreement, and products above 100M MAU or $20M/month revenue must display “Kimi K3” in the UI, per @natolambert, @petergostev, and @ArtificialAnlys. This is a useful signal for where frontier “open” may be settling: source-available / open-weight with business carve-outs rather than OSI-style licensing.
Distribution was immediate and broad: K3 was available day 0 via vLLM @vllm_project, Baseten @baseten, Modal @modal, Fireworks @Kimi_Moonshot, Nebius @Kimi_Moonshot, Together @Kimi_Moonshot, DigitalOcean @Kimi_Moonshot, Cursor @cursor_ai, Cognition/Devin @cognition, Ollama Cloud @ollama, and Dell Enterprise Hub @jeffboudier. That breadth underscores that open-weight frontier launches are now supply-chain events, not just research announcements.
Open AI Security, Open Weights Politics, and Anthropic’s Position
NVIDIA formally launched the Open Secure AI Alliance: Jensen Huang framed the core thesis starkly: attackers already have strong AI, so defenders need an ecosystem spanning open and closed frontier models, plus shared tooling and research. The flagship statement came from @JensenHuang, with NVIDIA’s formal announcement at @nvidia. The most technically interesting detail in the messaging was the claim that during the OpenAI/Hugging Face incident, a frontier open-weight model helped contain the intrusion, while a closed model blocked essential forensics—echoed by @AndrewYNg and @ZixuanLi_.
The alliance quickly accumulated credible infra and tooling members: Confirmed participants posting publicly included Hugging Face @huggingface, LangChain @LangChain, Nous Research @NousResearch, and support from voices across the open ecosystem such as @UnslothAI and @Yuchenj_UW. The argument is not “open is automatically safer,” but that defensive capability and auditability require open access to models, harnesses, and traces.
Anthropic finally clarified its open-weights stance: After sustained criticism for not signing NVIDIA’s open-weights letter, Anthropic published a position statement saying it has “never advocated for a ban on open-weights models” and instead supports: chip controls on China, anti-industrial-scale distillation measures, and mandatory safety testing for sufficiently capable models, open or closed, per @AnthropicAI. Reactions split between “reasonable clarification” @signulll, “good, but still trying to slow frontier diffusion” @jachiam0, and more hostile readings from open-weight advocates like @Teknium.
Policy pressure is intensifying around pre-release review: Separate reporting suggested the US government may seek up to 30 days of pre-release access to frontier systems for evaluation by agencies such as NSA and CAISI, with open-vs-closed treatment still unresolved, via @kimmonismus and @leomschwartz. Together with Anthropic’s statement and OpenAI’s Washington briefings, the direction is clear: frontier model release is becoming a governance interface, not just a product launch.
Benchmarks, Evals, and Agent Reliability
K3’s early evals are strong, especially for agents/coding: On Agent Arena, Kimi K3 Max reportedly ranks #1 among open-weight models with +9.75% net improvement, leading across multiple signals including confirmed success and steerability @arena. It also took #1 overall in Frontend Code Arena among all models in a later post @arena. Cognition said K3 is the first open-source model they tested that “approaches frontier-level performance” on FrontierCode 1.1, scoring 58.2% with 63.6% pass rate @cognition.
Claude Opus 5 also posted strong leaderboard numbers, but practitioner feedback was mixed: Arena reported Opus 5 Max at #1 in Frontend Code Arena and Text Arena with factuality on @arena, while WeirdML numbers from @htihle put Opus 5 high/max at 91.6% / 91.8%, roughly tied with Fable 5 max. But several devs reported frustrating real-world behavior—overcomplication, breakage, poor stopping behavior—from @abacaj, @davis7, @Teknium, and @theo. As usual, public eval gains and harness-specific production utility are diverging.
New eval work focused on sequential degradation and hidden regressions: @_philschmid highlighted EvoCode, an eval built around 26 tasks / 227 sequential rounds in a persistent container, measuring whether agents can follow evolving requirements without breaking earlier behavior. In parallel, @omarsar0 summarized a paper showing the “regression tax” from agent skills: across nearly 6,000 paired runs, skills generated gains but also broke many tasks previously solved without them. That is a practical warning against naïvely stuffing more procedural skills into context.
Multi-module RL systems are showing “role drift”: Another useful paper summary from @omarsar0 described how end-to-end RL can improve pipeline accuracy while causing modules to quietly abandon intended responsibilities—e.g. a decomposer embedding the answer rather than structuring the problem. This feels increasingly relevant as teams move from single-agent loops to specialized tool/prompt/module stacks.
Model and Systems Infra: From Agentic RL to Streaming VLMs
Microsoft and NVIDIA both shipped notable infra/model updates: Microsoft released Mage-VL 4B, described as a codec-native streaming VLM for live-event understanding, via @HuggingApps. NVIDIA research also surfaced Molt, a PyTorch-native agentic RL framework designed to be compact enough for humans—and AI coding assistants—to reason about end-to-end, summarized by @dair_ai. The “AI-readable research infra” design constraint is a small but significant shift in tooling philosophy.
AMD pushed a more reproducible open MoE release: Instella-MoE is AMD’s first fully open MoE LM: 16B total / 2.8B active, trained on MI300X/MI325X, with releases spanning checkpoints from pretraining through RL, plus configs, data mixtures, and code @PrakamyaMishra. Compared to typical model drops, this is closer to a full-stack research artifact.
Cohere and developer tooling vendors continue shifting toward “own the harness”: Cohere announced North Automations, a plain-language workflow layer on top of its secure agent platform @cohere. LangChain’s ecosystem messaging continued to emphasize that enterprises should own tools, prompts, context, and memory, not just rent model access @sydneyrunkle. This same framing showed up in multiple posts around open models and enterprise agent deployment.
Top tweets (by engagement)
Kimi K3 release: Moonshot’s K3 announcement was the largest technical post in the set, combining a 2.8T open-weights release with kernels, MoE comms, and agent-environment infra @Kimi_Moonshot.
Open Secure AI Alliance: Jensen Huang’s case for open defensive AI—especially the Hugging Face incident anecdote—drove major engagement @JensenHuang.
SSI × NVIDIA: Ilya Sutskever’s “Time to scale that SSI” and follow-on reporting point to a major compute expansion for Safe Superintelligence on Vera Rubin @ilyasut, @kimmonismus.
OpenAI economics/workflow productization: OpenAI’s work-use research and broader push around cloud agents / Work mode continue to signal a shift from chatbot UX to embedded personal and enterprise automation @OpenAI, @gdb.
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
- Kimi K3 Open Weights and Deployment Math
Kimi K3 weights now released. (Activity: 3442): The image is a mobile screenshot of the Hugging Face page for moonshotai/Kimi-K3, supporting the post title that Kimi K3 weights have been released. The model is shown as an Image-Text-to-Text Transformers checkpoint using Safetensors / compressed-tensors, requiring custom_code, under a kimi-k3 license, with roughly 3.8k likes and 2,850 downloads last month. Comments focus on hardware feasibility: one user notes “104B activated params”, implying very large inference memory requirements, while jokes like “How do I download ram in hugging face?” and “My 3090 is ready” highlight skepticism about running it on consumer GPUs.
Several commenters focused on the model’s scale, noting Kimi K3 reportedly uses 104B activated parameters, implying substantially higher inference memory/compute requirements than typical consumer GPU setups.
A technical concern raised was local deployability: one user described it as the first “frontier open model” they cannot run even on a 512 GB Mac Studio, highlighting that released weights may still be impractical for high-end local inference without multi-GPU/server-class hardware.
Kimi K3 weights drop today. We’re deploying on A100s, H200s and B300s this week and the A100 math is already rough (Activity: 763): The poster says Moonshot’s Kimi K3 weights are expected on Hugging Face with 2.8T total MoE params, 896 experts / 16 active per token, 1M context, vision support, and an estimated ~1.4 TB MXFP4 quantization-aware-trained checkpoint. Their deployment math: 8×A100 80GB = 640 GB cannot fit weights without multi-node sharding and lacks FP4/FP8 tensor cores; 8×H200 ≈ 1.13 TB still requires at least two nodes; 8×B300 ≈ 2.3 TB is the only listed single-node config with room for weights + long-context KV cache and native FP4. They plan to publish tok/s, TTFT, and cost-per-million-token benchmarks across A100, H200, and B300, with the expectation that A100 performance will be “ugly” due to dequantization or non-target INT4 kernels. Comments are mostly light, but one commenter frames the B300 deployment as a high-CapEx experiment—“$500k to spare”—amid uncertainty about cost collapse and open-weight scaling. Another notes intent to test the model on Intel Gaudi 2/3, suggesting interest in non-NVIDIA inference viability.
Discussion centered on hardware feasibility for hosting Kimi K3, with one commenter noting that an 8x AMD MI355X setup could be ideal due to roughly 2.3 TB aggregate VRAM and FP4 acceleration, though availability/rental access was described as effectively unavailable.
Several commenters compared deployment targets beyond NVIDIA, including attempts to run the weights on Intel Gaudi 2/3 accelerators and skepticism around the economics of buying/renting high-end B300 systems, with one user framing the deployment cost as potentially around $500k.
A commenter noted that Hugging Face removed the countdown, implying uncertainty or a change in the release timing/distribution page for the Kimi K3 weights.
- Open-Weight AI Security and Policy Fight
CEO of Hugging Face: “In the spirit of transparency, here’s what I asked OpenAI” (Activity: 3109): The image is a screenshot of Hugging Face CEO Clem Delangue publicly asking OpenAI to release execution traces/logs from alleged “rogue” autonomous agents involved in what he calls the “first autonomous agent cyberattack” so researchers can analyze the failure mode. He also asks OpenAI to commit $100M in compute to help the Hugging Face community build cyber-defense systems using open and closed models. Image Commenters were mostly skeptical, framing the request as an unrealistic “casual” ask for $100M; some speculated the incident was more likely a publicity stunt or that releasing logs would expose OpenAI to reputational/legal risk.
Jensen Huang: During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. That’s why we created the Open Secure AI Alliance. (Activity: 1736): The image is a screenshot of Jensen Huang claiming that, during a Hugging Face security incident, closed AI systems blocked essential forensic analysis, while an open-weight frontier model helped defenders contain the intrusion. The post frames this as the motivation for NVIDIA’s Open Secure AI Alliance, shown with partner logos including Microsoft, Hugging Face, IBM, Cloudflare, Cisco, Red Hat, Salesforce, SAP, and others, arguing for a mixed open + closed frontier AI security ecosystem rather than relying solely on proprietary models. Commenters were skeptical of the alliance’s “open” branding, pointing out that companies like Adobe, Cisco, Palantir, and even DoorDash are not typically associated with open-source AI; one also noted the apparent absence of major open-source model creators.
Sources: OpenAI and Anthropic quietly lobby Washington regulators to restrict open-source AI models, even as Sam Altman publicly says he supports open source AI (Activity: 1470): NYT reports that OpenAI and Anthropic have been lobbying U.S. regulators for restrictions on open/open-weight AI models—especially Chinese releases from Z.ai and Moonshot AI that are nearing frontier U.S. model capability—citing IP theft, distillation, safety, and national-security risks. The counter-coalition includes Nvidia, Microsoft, Meta, Google, IBM, Palantir, Hugging Face, and startups arguing open models are critical for competition, security auditing, chip/cloud demand, and innovation; U.S. officials are reportedly more inclined toward targeted actions against specific Chinese firms/models than a blanket ban. Top comments were mostly cynical toward Sam Altman/OpenAI, framing the alleged lobbying as inconsistent with public support for open weights; one commenter sarcastically summarized the position as: “we supported Open Weights, but lobbying made it impossible.”
OpenAI management decided earlier today not to join the “Open Secure AI Alliance”, founded by Nvidia CEO Jensen Huang. The decision was shared internally and reportedly met with backlash from employees. (Activity: 423): The post claims OpenAI management internally decided not to join the “Open Secure AI Alliance”, reportedly founded by Nvidia CEO Jensen Huang, and that the decision triggered employee backlash. No technical details are provided about the alliance’s governance, security model, openness criteria, model-release policies, benchmarks, or implementation requirements.
- Runnable Local Models and Coding Harness Benchmarks
Read more
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み