Moonshot AI と Thinking Machines Lab がオープンウェイトモデルを相次ぎ公開
本文の状態
日本語全文を表示中
詳細モードで約13分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
OpenHands Engineering
Moonshot AI が巨大パラメータ数の K3 を、Thinking Machines Lab が微調整用ベースモデル Inkling を公開し、オープンモデルによるコード生成の能力格差が縮小した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 14:28
AI深層分析
キーポイント
Kimi K3 の発表と仕様
Moonshot AI は2.8兆パラメータのマルチモーダルモデル「Kimi K3」を公開し、フロントエンドコード評価でトップクローズドモデルを上回る結果を出した。同社は7月27日に重み(weights)の公開を予定しており、API 利用も可能である。
Inkling の特性と戦略
Thinking Machines Lab は元 OpenAI CTO の Mira Murati が設立した同社から、9750億パラメータの「Inkling」を Apache-2.0 ライセンスでリリースした。これは完成品ではなく微調整用ベースモデルとして位置づけられ、Tinker プラットフォームでの学習に優れている。
オープンモデルとクローズドモデルの格差縮小
以前はオープンモデルでエージェント型コーディングを行う際、能力の欠如を許容する必要があったが、今回の発表によりその格差は小さくなり、コストやデータ制御などの他の変数が意思決定の主要因となっている。
実運用上の課題と要件
Kimi K3 の完全な自己ホスティングには64基以上のアクセラレータが必要であり、Inkling は NVFP4 量子化版を備えているが、両モデルとも大規模インフラの準備が求められる。
Inklingモデルの概要と特徴
975B総パラメータ、41BアクティブのMoE構造を持つInklingは、テキスト・画像・音声・動画の45兆トークンで訓練され、最大1Mトークンのコンテキストウィンドウをサポートする。ベンチマークでは最上位ではないが、エージェント評価や微調整において優れた性能を示し、低遅延向けに小型版も用意されている。
重要な引用
"A year ago, 'use an open model for agentic coding' meant accepting a real capability gap. That gap is now small enough that the other variables — cost, data control, customization — start to dominate the decision."
"Kimi K3 is a sparse mixture-of-experts model: 2.8T total parameters, but only 16 of 896 experts fire per token."
"Thinking Machines is candid that Inkling isn't the top model on popular benchmarks; their pitch is that it scores well among open-weights models on agentic evals like Terminal-Bench 2.1 and that it responds unusually well to fine-tuning through their Tinker platform."
A competitive, Apache-2.0, US-built open model — joining Nemotron and Gemma — means an open-weights strategy no longer has a single point of geopolitical failure.
編集コメントを表示
編集コメント
業界をリードするクローズドモデルに匹敵する性能を持つオープンソースモデルが相次いで登場したことは、開発エコシステムにおける選択肢の拡大を意味する。特に K3 の重み公開と Inkling の微調整指向という異なるアプローチは、ユーザーが自社の要件に合わせて最適な戦略を選べる環境を整えたと言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
今週、非常に異なる背景を持つ 2 つのラボから、フロンティアクラスのオープンウェイトモデルがリリースされました。
Moonshot AI は、100 万トークンのコンテキストウィンドウを備えた 2.8 兆パラメータのマルチモーダルモデル「Kimi K3」を発表しました。このモデルは、7 月 27 日までに修正 MIT ライセンスの下で完全な重みが公開される予定です。
一方、元 OpenAI CTO のミラ・ムラティ氏が設立したスタートアップ、Thinking Machines Lab は、「Inkling」という 975B パラメータの Apache-2.0 ライセンスモデルをリリースしました。同社はこれを完成品としてではなく、ファインチューニングのためのベースモデルとして明確に位置付けています。
コードエージェントをあらゆる規模で運用している場合、このニュースは選択肢を広げます。1 年前、「オープンモデルを使ってエージェントによるコーディングを行う」と言えば、能力の格差を受け入れることを意味していました。しかし現在、その格差は十分に小さくなり、コストやデータ制御、カスタマイズといった他の変数が意思決定の主要な要素になり始めています。
Kimi K3: 過去最大規模のオープンウェイトリリース
Kimi K3 はスパースな混合専門家(MoE)モデルです。総パラメータ数は 2.8T ですが、トークンごとに発火するのは 896 個ある専門家のうちわずか 16 個のみです。
テキスト、画像、動画の入力を処理でき、Frontend Code Arena のリーダーボードでトップのクローズドモデルを抜いて第 1 位を獲得しました。API の料金は、入力トークン 100 万あたり 0.30 ドル(キャッシュ済み)〜3.00 ドル(未キャッシュ)、出力トークン 100 万あたり 15 ドルです。
明言しておくべき二つの注意点があります。まず、この記事執筆時点では K3 の重みは発表済みですが、まだダウンロード可能ではありません。Moonshot 社は 7 月 27 日に公開する予定としています。第三者は API を通じてモデルのベンチマークを評価できますが(後述)、チェックポイントがリリースされるまで、自己ホスト環境でのスループットやレイテンシ、ハードウェア効率を検証することは誰にもできません。
第二に、このモデルを自前でホストするのは週末のプロジェクトではありません。Moonshot 社は 64 基以上のアクセラレーターを搭載した構成でのデプロイを推奨しており、最低限必要な GPU セットアップについても公開していません。
Inkling:ファインチューニングを前提とした西方発のベースモデル
Inkling は現在利用可能です。これは 975B の総パラメータを持つ MoE(Mixture of Experts)モデルで、41B がアクティブに動作します。テキスト、画像、音声、動画からなる 45 トリリオンのトークンで訓練され、コンテキストウィンドウは最大 100 万トークンをサポートします。
transformers、SGLang、llama.cpp ではリリース初日からネイティブサポートが提供されており、推論コストを抑えるための NVFP4 量子化版も用意されています。Thinking Machines は率直に、Inkling が一般的なベンチマークでトップクラスではないと認めています。彼らの主張は、Terminal-Bench 2.1 などのエージェント評価においてオープンウェイトモデルの中で高いスコアを記録し、Tinker プラットフォームを通じたファインチューニングに対して非常に良好な反応を示す点にあります。
また、レイテンシが敏感に作用するワークロード向けに、アクティブパラメータ数 12B の軽量版「Inkling-Small」のプレビューも公開されています。
Inkling は戦略的な隙間も埋めています。今週に至るまで、ほぼすべてのフロンティア級のオープンモデルは中国の研究所からのものでした。しかし、これらのモデルを取り巻く規制状況は依然として不透明です。米国の連邦政府や州政府はすでに DeepSeek を政府システムでの使用に制限しており、今月の報道によると、より広範な規制も検討されているとのことです。Nemotron や Gemma に加わり、Apache-2.0 ライセンスで米国製のオープンモデルが参入することで、重み付きのオープン戦略における地政学的リスクが単一ポイントから分散されます。
スコアボード:オープン vs クローズド
この格差を最も明確に示すのが、SWE-bench の 2 つの変種です。1 つは人間が検証した GitHub の課題 500 件を対象とした「Verified」、もう 1 つはより難易度が高く最近公開されたセットである「Pro」です。
以下の Verified の数値は Vals AI が独立して再実行した結果(7 月 17 日更新)から引用しています。これは、すべてのモデルを同じ最小限の bash のみで動作する mini-SWE-agent ハーネスで評価したものです。ベンダーによる自己申告やハーネスの操作はありません。Pro の数値は Thinking Machines が公開した比較 と morphllm の集計データ に基づいています。
| モデル | 重み | SWE-bench Verified | SWE-bench Pro (public) | --- | --- | --- | --- | GPT-5.6 Sol | クローズド | 96.2% | 64.6% | Claude Fable 5 | クローズド | 95.0% | 80.0% | Kimi K3 | オープン(重みは 7 月 27 日公開予定) | 93.4% | 未発表 | Claude Opus 4.8 | クローズド | 88.6% | 69.2% | GLM-5.2 | オープン | 82.8% | 62.1% | Inkling | オープン | 77.6%* | 54.3% |
|---|
「Inkling の Verified スコアは、Thinking Machines による自己申告値であり、bash のみを使用するハーンチで測定されたものです。」
この K3 という数字が今回の見出しです。オープンウェイトモデルが独立して測定され、Claude Fable 5 よりも 1.6 ポイント遅れながら、ボード上のすべてのクローズドモデルを上回りました。その中には Claude Opus 4.8 も含まれていますが、K3 はそれよりも約 5 ポント上です。第三者が運営する Verified リーダーボードで、オープンリリース版がこれほどトップに迫った例はこれまでありませんでした。Vals の難易度別内訳を見ると、K3 は人間で 1〜4 時間かかる作業と推定されるタスクでも 93% を達成しており、Fable 5 と並んでいます。
正直な注釈を付け加えます。K3 は Vals の最上位クラスの中で最も遅く、タスクあたりの平均レイテンシは 619 秒でした。これは GPT-5.6 Sol の 182 秒の約 3.4 倍です。重みファイルのダウンロードは 7 月 27 日まで解禁されませんし、SWE-bench Pro でのスコアを報告した論文もまだ出ていません。Pro ではすべてのモデルが 20 ポイント以上スコアを落としますが、Fable 5 の 80.0% は、ベストなオープンモデルである GLM-5.2 の 62.1% を大きく引き離しています。つまり、すべての分野で格差が埋まったわけではありません。しかし、業界が過去 3 年間も標準として扱ってきたベンチマークにおいては、ついにオープンモデルがフロンティアモデルの誤差範囲内に到達したのです。
オープンウェイトがもたらすもの
コストは、十分な規模次第です。コーディングエージェントはトークンの大消費地です。1 つの課題解決実行で数十万トークンが消費されることも珍しくなく、CI やすべてのプルリクエストでエージェントを実行しているチームでは、月間のトークン数がすぐに 7 桁に達します。その規模になれば、専用推論クラスターのコスト計算や、オープンモデル API プロバイダーが提供する安価なトークン単価の計算は、最先端クローズドモデルの価格を何倍も上回ります。ただし、それ以下の規模では通常そうなりません。GPU を購入する前に必ず計算してください。
データ主権です。セルフホストされた重みを使用すれば、プロンプト、コードベース、エージェントのトレースが、自社の管理下にあるインフラから外れることはありません。規制産業に属するチームや、社内の法務部門が proprietary ソースコードをサードパーティの API を経由させることに懸念を持つ組織にとって、これが最大の理由です。「パケットが施設を出たことがない」という事実には、どんなデータ処理契約も勝てません。
ファインチューニングです。クローズドモデルは、プロバイダーが学習させた時点の姿で凍結されています。一方、オープンモデルなら、自社の内部フレームワークや移行パターン、コードレビューの慣習を学習させることが可能です。これが「Thinking Machines」の賭けの核心です。彼らの発表記事では、中央集権的に訓練され固定された AI は、組織が自ら形成する AI よりも劣ると主張しています。なぜなら、多くの専門知識はそれを保持している人々に固有のものだからです。この仮説全体を支持しなくても、モノレポに特化したファインチューニング済みの中規模モデルが、そのモノレポ内で発生するタスクにおいて、一般向けの最先端モデルよりも優れていることは明らかです。
書き換えなしで切り替え可能。 重みが公開され、主要な提供元すべてが OpenAI 互換 API を採用しているなら、モデルを切り替えるのは設定変更だけで済みます。しかし、ワークフローが特定の提供元に強く結びついている場合、それは大規模な移行プロジェクトになります。
コストも現実的な問題です。推論クラスターの監視(ページャー)は自社で管理し、キャパシティプランニングも自前で対応する必要があります。また、モデルが奇妙な挙動を示した際にも、ベンダーへのサポートチケットを発行することはできません。オープンウェイトのモデルは、月額利用料というコストを運用責任と引き換えにします。このトレードオフはスケールメリットがある場合こそ価値がありますが、小規模では割に合わない取引です。
ハーネス(枠組み)にも縛られないでください
エージェントの基盤となるハーネスが、裏で特定の提供元だけを前提としているなら、モデルの柔軟性は意味をなしません。提供元が自社向けに開発したコーディングツールは、自社のモデルに対してエンドツーエンドで最適化されています。プロンプト、ツールの定義、そしてモデル自体が協調して訓練されているのです。
Claude Code を Anthropic 互換 API のシェイム(中継層)を介して Kimi と連携させるようなハッキングは可能です。実際に K2 ではこれが行われましたが、Thoughtworks のエンジニアたちは その結果 を「Claude Code は Kimi にとって最適なツールではなかった」と結論付けています。実行コストが安くても、です。このハーネスは機能しますが、「よく」機能するわけではありません。最適化の一部が転用できないからです。
これは特定のベンダー製ツールの批判ではありません。これはサプライチェーンに関する観察です。もしハルネス(実行環境)とモデルが同じベンダーから提供されている場合、モデルを切り替える能力は理論上の話に過ぎません。多くのモデルでテストされ調整された「モデル非依存型」のハルネスこそが、今週のような期間を実行可能なものにし、単なる興味深い話題で終わらせないのです。
誰が作ったか気にしないエージェントでこれらのモデルを試す
OpenHands は設計上オープンソースかつモデル非依存です。同じエージェントが、クローズドなフロントティア API でもあれば、あるいは自社のクラスター上で動作するオープンモデルであっても実行されます。そのため、今週のような期間でも、移行ではなく設定の変更で済みます。
あるモデルをサービスする価値があるかどうかを判断するには、OpenHands Index を確認してください。これは、課題解決、新規アプリ開発、フロントエンド開発、テスト、情報収集といった実際のソフトウェアエンジニアリングの作業において、クローズド・オープン両方のモデルをベンチマークする、継続的に更新されるリーダーボードです。精度、コスト、実行時間の観点からスコアリングされています。新しいオープンモデルがリリースされた際、それがローテーションに値するかを知る必要がある場合、このコストと精度のパレート曲線こそが必要なグラフです。ベンチマークコードはオープンソースであり、結果に誤りが見つかった場合は 訂正が公開されます。
データが存在する場所でエージェントを実行するには、Agent Canvas が最適です。MIT ライセンスのセルフホスティング可能なコーディングワークスペースで、ローカル環境や Docker、あるいは独自の VM で実行できます。また、お好みのモデルエンドポイントを選択してバックアップすることも可能です。
OpenHands Enterprise は、SSO/RBAC、隔離されたサンドボックス、予算管理、可視化機能を追加し、クラウドやデータセンター内で完全なスタックを必要とするチーム向けです。これにより、セルフホスト型のモデルを指すことが可能になります。
次の実践的なステップとして、7 月 27 日に K3 チェックポイントがリリースされるか、あるいは今日 Inkling を利用して、エンドポイントを構築し、モデル非依存のハネスに接続して、一週間自社のバックログに対して実行してみてください。オープンソースとクローズドソースの差はすでに小さくなっています。重要なのはリーダーボードではなく、あなたのワークロードが判断を下すことです。
原文を表示
Two frontier-class models shipped open weights this week, from two very different labs. Moonshot AI announced Kimi K3, a 2.8-trillion-parameter multimodal model with a 1M-token context window, with full weights promised by July 27 under a modified MIT license. Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, released Inkling, a 975B-parameter Apache-2.0 model that they explicitly position as a base for fine-tuning rather than a finished product.
If you run coding agents at any scale, this changes your options. A year ago, "use an open model for agentic coding" meant accepting a real capability gap. That gap is now small enough that the other variables — cost, data control, customization — start to dominate the decision.
Kimi K3: the largest open-weight release yet
Kimi K3 is a sparse mixture-of-experts model: 2.8T total parameters, but only 16 of 896 experts fire per token. It takes text, images, and video, and it debuted at #1 on the Frontend Code Arena leaderboard, ahead of the top closed models. API pricing runs 0.30–0.30–3.00 per million input tokens (cached vs. uncached) and $15 per million output tokens.
Two caveats worth stating plainly. First, as of this writing the K3 weights are announced, not downloadable — Moonshot says July 27. Third parties can benchmark the model through its API (more on that below), but nobody can verify self-hosted throughput, latency, or hardware efficiency until the checkpoint ships. Second, self-hosting this thing is not a weekend project: Moonshot recommends deployment on configurations with 64 or more accelerators and hasn't published a minimum viable GPU setup.
Inkling: a Western base model built to be fine-tuned
Inkling is out now. It's a 975B-total, 41B-active MoE trained on 45 trillion tokens of text, image, audio, and video, with a context window up to 1M tokens. It ships with day-0 support in transformers, SGLang, and llama.cpp, plus an NVFP4 quantized variant for cheaper inference. Thinking Machines is candid that Inkling isn't the top model on popular benchmarks; their pitch is that it scores well among open-weights models on agentic evals like Terminal-Bench 2.1 and that it responds unusually well to fine-tuning through their Tinker platform. There's also a preview of Inkling-Small, a 12B-active sibling for latency-sensitive work.
Inkling also fills a strategic gap. Until this week, nearly every frontier-class open model came from a Chinese lab, and the regulatory picture around those models is unsettled: US federal and state governments have already restricted DeepSeek on government systems, and reporting this month suggests broader restrictions are under consideration. A competitive, Apache-2.0, US-built open model — joining Nemotron and Gemma — means an open-weights strategy no longer has a single point of geopolitical failure.
The scoreboard: open vs. closed
The cleanest way to see the gap is on the two SWE-bench variants: Verified (500 human-validated GitHub issues) and the harder, more recent Pro (public set). The Verified numbers below come from Vals AI's independent re-run (updated July 17), which evaluates every model with the same minimal bash-only mini-SWE-agent harness — no vendor self-reporting, no harness games. Pro numbers are from Thinking Machines' published comparison and the morphllm aggregate.
| Model | Weights | SWE-bench Verified | SWE-bench Pro (public) | --- | --- | --- | --- | GPT-5.6 Sol | Closed | 96.2% | 64.6% | Claude Fable 5 | Closed | 95.0% | 80.0% | Kimi K3 | Open (weights due July 27) | 93.4% | not yet published | Claude Opus 4.8 | Closed | 88.6% | 69.2% | GLM-5.2 | Open | 82.8% | 62.1% | Inkling | Open | 77.6%* | 54.3% |
|---|
**Inkling's Verified score is Thinking Machines' self-report, also with a bash-only harness.*
That K3 number is the headline. An open-weight model, independently measured, sits 1.6 points behind Claude Fable 5 and ahead of every other closed model on the board — including Claude Opus 4.8 by nearly five points. No open release has been that close to the top of a third-party Verified leaderboard before. Vals' difficulty breakdown shows K3 holding 93% even on tasks estimated at 1–4 hours of human work, right alongside Fable 5.
The honest footnotes: K3 was the slowest model in Vals' top tier at 619 seconds average latency per task, 3.4x GPT-5.6 Sol's 182 seconds. Its weights aren't downloadable until July 27, and nobody has published a SWE-bench Pro number for it — on Pro, where every model drops 20-plus points, Fable 5's 80.0% still leads the best open score (GLM-5.2's 62.1%) by a wide margin. So the gap hasn't closed everywhere. But on the benchmark the industry has treated as the standard for three years, an open model is now inside the error bars of the frontier.
What open weights buy you
Cost, at sufficient scale. Coding agents are token furnaces. A single issue-resolution run can burn hundreds of thousands of input tokens, and teams running agents in CI or on every pull request hit seven-figure monthly token counts fast. At that volume, the math on a dedicated inference cluster — or on the much cheaper per-token pricing that open-model API providers offer — starts to beat frontier closed-model pricing by multiples. Below that volume, it usually doesn't. Do the arithmetic before you buy GPUs.
Data sovereignty. With self-hosted weights, your prompts, your codebase, and your agent traces never leave infrastructure you control. For teams in regulated industries, or anyone whose legal department has opinions about proprietary source code transiting a third-party API, this is the whole argument. No data processing agreement can match "the packets never left the building."
Fine-tuning. A closed model is frozen at whatever the provider trained it to be. An open model can learn your internal frameworks, your migration patterns, your code review conventions. This is Thinking Machines' entire bet: their launch post argues that centrally trained, set-in-stone AI underperforms AI that organizations shape themselves, because so much expertise is specific to the people who hold it. You don't have to buy the full thesis to see that a mid-size model fine-tuned on your monorepo can beat a generalist frontier model on tasks that live in that monorepo.
Switching without rewrites. When weights are open and every serious serving stack speaks the OpenAI-compatible API, moving from one model to another is a config change. When your workflow is welded to one provider, it's a migration project.
The costs are real too: you own the pager for your inference cluster, you own capacity planning, and when the model does something baffling there's no vendor support ticket. Open weights trade a monthly bill for an operational responsibility. That trade is worth it at scale and a bad deal below it.
Don't marry your harness either
Model flexibility is worthless if your agent harness quietly assumes one provider. Provider-built coding tools are tuned end-to-end for their own models: the prompts, the tool definitions, and the models themselves are co-trained. You can hack Claude Code into talking to Kimi through an Anthropic-compatible API shim — people did exactly this with K2 — but engineers at Thoughtworks who tried it concluded that Claude Code wasn't the best tool for Kimi even when it was cheaper to run. The harness works; it just doesn't work *well*, because none of its tuning transfers.
This isn't a knock on provider tools. It's a supply-chain observation: if the harness and the model come from the same vendor, your ability to switch models is theoretical. A model-agnostic harness — one that's tested and tuned across many models — is what makes weeks like this one actionable instead of merely interesting.
Trying these models with an agent that doesn't care who made them
OpenHands is open source and model-agnostic by design: the same agent runs against a closed frontier API or an open model served on your own cluster, so weeks like this one are a configuration change rather than a migration.
To decide *whether* a model is worth serving, check the OpenHands Index, a continually updated leaderboard that benchmarks both closed and open models on real software engineering work — issue resolution, greenfield apps, frontend development, testing, and information gathering — scored on accuracy, cost, and runtime. The cost-accuracy Pareto curve is exactly the graph you need when a new open model drops and you want to know if it earns a spot in your rotation. The benchmarking code is open source, and when results turn out to be wrong, the corrections are published.
For running agents where your data lives: Agent Canvas is the MIT-licensed, self-hostable coding workspace — run it locally, in Docker, or on your own VMs, backed by whatever model endpoint you choose. OpenHands Enterprise adds SSO/RBAC, isolated sandboxes, budgeting, and observability for teams that need the whole stack in their own cloud or datacenter, pointed at self-hosted models.
The practical next step: when the K3 checkpoint lands on July 27, or with Inkling today, stand up an endpoint, wire it into a model-agnostic harness, and run it against your own backlog for a week. The gap between open and closed is now small enough that your workload — not the leaderboard — should make the call.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み