Kimi K3の性能向上、Fable活用ではないと専門家が指摘
専門家らは、Kimi K3 の高性能化が Anthropic の Fable を悪用した結果ではないと分析し、その真の原因は別の技術的アプローチにある可能性を指摘している。
キーポイント
Fable 悪用の否定
専門家による分析により、Kimi K3 の性能向上が Anthropic の Fable プラットフォームの悪用やデータ流出に起因するという説は誤りであることが示された。
真の原因の特定
高性能化の主因は、Fable への依存ではなく、独自の学習アルゴリズムやトレーニングデータの質・量における技術的進歩にあると分析されている。
業界への示唆
この分析は、AI モデルの性能向上要因を巡る憶測に対し、客観的な技術評価が重要であることを浮き彫りにし、セキュリティ懸念の払拭にも寄与する。
重要な引用
Experts say exploiting Anthropic’s Fable isn’t how Kimi K3 got so good
exploiting Anthropic's Fable isn't the main driver of Kimi K3's performance
影響分析・編集コメントを表示
影響分析
このニュースは、Kimi K3 の技術的実態に関する誤解を解き、業界内のセキュリティ懸念や不正利用説に終止符を打つ意義があります。また、AI モデルの性能向上要因が単なる他社資産の流用ではなく、独自の研究開発によるものであることを示すことで、技術競争の本質的な価値を再認識させるきっかけとなります。
編集コメント
AI モデルの性能向上要因を巡る憶測はつきものですが、今回の分析のように専門家の客観的な見解が事実を明確にすることは業界全体の健全な発展に不可欠です。Kimi K3 の真価が独自技術にあると確認されたことは、開発競争におけるイノベーションの重要性を再認識させる内容と言えます。
ホワイトハウスの科学顧問マイケル・クラティオス氏は、中国の企業ムーンショット(Kimi K3 の開発元)が、輸出規制に抵触するチップを使用しながら、アンソロピック社の Fable LLM を模倣して同モデルを構築したと指摘しました。Kimi K3 は現在、利用可能なオープンウェイト大規模言語モデルの中で最大規模です。
「米国技術の窃取や米国の研究基盤の弱体化を狙った大規模かつ隠蔽された産業的蒸留(ディスティレーション)は許されない」とクラティオス氏は投稿しました。この発言は、AI 業界を揺るがす中国製オープンウェイトモデルの禁止を検討する議論が行われている最中に行われたものです。ムーンショット社はトレーニングプロセスに関する質問に回答しておらず、クラティオス氏も告発の根拠となる詳細な情報については明らかにしていません。
クラティオス氏の発言は、スコット・ベッセント財務長官が「中国製の多くのモデルに、米国の大規模言語モデルの特徴(ウォーターマーク)が見つかっている。これは許されない」と述べた内容と一致するものです。しかし、その特徴とは具体的に何を指すのかは不明であり、財務省も問い合わせに対して回答していません。
しかし専門家たちは、Kimi K3 が示す高度な能力が、LLM の内部構造を解明しその機能を模倣するプロセスである「蒸留」によって生み出されたものだと信じるには懐疑的です。
「Fable が厳密に知識蒸留(ディストillation)を行うだけでは、これほど強力なモデルをこれほど短期間で獲得できるとは思いません」と、Laude Institute の研究者であり Snorkel AI の共同創設者である Braden Hancock は TechCrunch に語りました。「正直に言って、時間がないからです。Fable が一般公開されたのは 7 月 1 日だけです。その程度のデータを蒸留してモデルを訓練し、2 週間でリリースするのは不可能です。」
「中国のモデルが最先端に近づき、トレーニング手法が [強化学習] にシフトするにつれて、知識蒸留の影響は次第に薄れているという見解を持っています」と、AI 研究者の Nathan Lambert は昨日公開された ポッドキャスト で述べました。「もし蒸留データを使って誰でも簡単に GLM や K3 に追いつけるのであれば、私たちはすでにそうなっているはずです。しかし、教師あり微調整(SFT)だけでは、そのような状況にはなっていません。」
知識蒸留を行うには、ラボが対象モデルに対して体系的にクエリを実行し、ポストトレーニングで利用可能なデータを生成する必要があります。場合によっては、モデルが問題をどのように解決しているかを理解するために、その思考プロセスを明確に記述させることが明示的に含まれます。また別のケースでは、あるモデルからのプロンプトと回答を用いて、教師あり微調整(SFT)と呼ばれるプロセスで新しいモデルを訓練します。
この微調整プロセスを経ることで、第三者によって作成されたモデルがあたかも Claude のように主張する事態が生じ得ます。ランバートの見解では、微調整こそが「モデルにマナーを身につけさせる」段階です。
しかし、ランバートは、モデルの複雑さが増すにつれて SFT(教師あり微調整)の恩恵は相対的に小さくなっていると指摘しています。Fable に匹敵する能力を抽出するには、おそらく強化学習の手法が必要となるでしょう。多くの場合、これは大規模なモデルのエージェントが小規模なモデルの回答を評価し、その評価結果に基づいて調整を行うことを意味します。
こうした高度な技術には、より大規模なインフラストラクチャが必要です。大規模な強化学習の実行では、数千万ものエージェントが必要になることもあります。そのような作業に最前線の研究所の API を利用することは「莫大な費用がかかり、さらに時間的なボトルネックにもなりかねません。なぜなら、これらのモデルは非常に遅く、正直に言って性能向上をもたらさない可能性さえあるからです。
Kimi の性能向上に、過去の最先端モデルが寄与した可能性は十分にあります。Anthropic は今年初め、Moonshot、DeepSeek、MiniMax といった中国企業に対し、自社のモデルを体系的に蒸留(ディストillation)したと公的に非難しました。同社は、IP アドレスやその他のメタデータを通じて特定されたこれらの企業との間で、数百万件のやり取りを発見したと発表しています。これらの問い合わせは「通常の使用パターンとは異なり、正当な利用ではなく意図的な能力の抽出を示すものだった」とされています。一方、Anthropic は TechCrunch からの Fable を巡る蒸留に関する質問には回答していません。
しかし、モデルの蒸留は中国に限らず、AI 業界全体で広く行われている慣行と見られています。今年初めにエロン・マスク氏は証言の中で、自社のスペースX AI が OpenAI のモデルを蒸留して Grok を開発したと明かしました。この手法が業界では一般的であるとも述べています。例えば、蒸留と合成データセットの作成との境界線は、必ずしも明確ではありません。
「一般に、アメリカ人は中国チームの技術的専門性を過小評価している」とハンコック氏は指摘します。「Moonshot の創設者の一人はカーネギーメロン大学の博士課程学生でした。彼らは確実な成果を上げる正当な研究者やエンジニアです。……もし米国のモデル開発が停滞した場合、中国の進展も鈍化するでしょうが、完全に止まるわけではありません。彼らは単に他人の足元にしがみついているだけではないのです。」
また、Kratsios の発言のもう一つの側面、つまり Moonshot が最先端の Nvidia チップ「Grace Blackwell 300」を手に入れ、タイに設置された GB300 搭載サーバーにもアクセスしたという点と、知識蒸留(ディストillation)の影響を切り離して考えるのは困難です。これらのチップは中国への輸出が禁止されていますが、ジョージタウン大学のセキュリティ・イマージング・テクノロジーセンターのフェローであるサム・ブレスニック氏によれば、黒市場が存在しているといいます。5 月には、米国のサーバーメーカーであるスーパーマイクロの創業者が、先進的なチップを中国へ密輸した容疑で起訴されました。
「私は世界中のデータセンターに対して『顧客確認(KYC)』法の導入を支持しています」とブレスニック氏は語ります。「最先端のハードウェアを使って大規模な学習トレーニングを行う企業に許可を与えるなら、その企業が誰であり、何をしているのかを報告する仕組みが必要不可欠です」。
ジョー・バイデン米大統領の商務省は 2024 年にデータセンター向けの連邦 KYC ルールを提案しましたが、ドナルド・トランプ政権下ではこれ以上の進展は見られません。一方で、先進的なチップを海外へ輸出する企業には、それらが承認された目的でのみ使用されることを確実にする義務が課されています。
この記事内のリンクを通じて購入が行われた場合、当社は少額のコミッションを受け取る場合があります。これは当社の編集の独立性には影響しません。
ティム・フェンホルツは、テクノロジー、金融、公共政策を専門とするジャーナリストです。民間宇宙産業の台頭を密接に取材しており、『ロケット・ビリオネアーズ:イーロン・マスク、ジェフ・ベゾスと新たなスペース・レース』の著者でもあります。以前はグローバルなビジネスニュースサイトであるクォーツで10年以上シニア記者を務め、キャリアの初期にはワシントンD.C.で政治担当記者として活動していました。
ティムへの連絡や取材依頼の確認は、tim.fernholz@techcrunch.com へメールを送るか、Signal で暗号化メッセージを tim_fernholz.21 宛てに送信してください。
原文を表示
White House science advisor Michael Kratsios said that Moonshot, the Chinese company behind Kimi K3, the largest available open-weight LLM, built its model by copying Anthropic’s Fable LLM while using chips that aren’t cleared for export to China.
“Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable,” Kratsios wrote, amid reported discussions about banning Chinese open-weight models that have roiled the AI sector. Moonshot did not respond to questions about its training process, and Kratsios did not share more details about the sources of his allegations.
Kratsios’ tweet echoed comments from Treasury Secretary Scott Bessent that “we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that’s unacceptable.” It’s not clear what those watermarks consist of, and the Treasury Department did not respond to a query.
However, experts are skeptical that distillation — the process of querying an LLM to determine its inner workings and copy its capabilities — is responsible for the advanced capabilities that Kimi K3 displays.
“I don’t think you get a model this strong and this quickly on the heels of Fable doing strictly distillation,” Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, told TechCrunch. “There’s just not even frankly time, right? Fable’s only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks.”
“I’ve been of the opinion that distillation is becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning],” Nathan Lambert, an AI researcher at the Allen Institute for AI, said in a podcast released yesterday. “[I]f it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won’t see this, from supervised fine-tuning alone.”
Performing distillation requires a lab to systematically query its target model in order to generate data that can be used for post-training. Sometimes this explicitly involves asking the model to articulate its chain-of-thought to understand how it solves problems. Other times, the prompts and responses from a model are used to train a new model in a process called supervised fine-tuning, or SFT.
It’s this fine-tuning process that can result in a model ostensibly created by a third party claiming that it is Claude. Fine tuning is where, in Lambert’s view, the “model picks up its manners.”
But Lambert says that the benefits of SFT are becoming less important as models become more complex. To distill Fable-like capabilities would likely require reinforcement learning techniques. In many cases, that means having an agent of the larger model grade the smaller model’s responses, and adjusting based on the grade.
The more advanced techniques also require more significant infrastructure. Large reinforcement learning runs can require tens of millions of agents. Using a frontier lab’s API to do that “would be insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow and to be frank might not even give you a performance uplift.”
It seems likely that previous frontier models might have contributed to Kimi; Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of systematically distilling its models earlier this year. Anthropic said it discovered millions of exchanges between its models and users it identified at those companies through IP addresses and other meta data. Those queries were “distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use.” Anthropic didn’t respond to TechCrunch’s queries about Fable distillation.
However, distillation is seen as common among AI companies, not just in China. Elon Musk testified earlier this year that his company SpaceXAI distilled OpenAI models to develop Grok, and that the practice was common in the industry. The line between distillation and developing synthetic datasets, for example, can be fairly blurry.
“[I]n general, Americans are understating the technical expertise of these Chinese teams,” Hancock said. “One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. …if American models ground to a halt, I think China’s progress would slow, but would still continue. They’re not just riding coattails here.”
It’s also hard to disentangle distillation from the second part of Kratsios’ comment — that Moonshot had obtained advanced Nvidia Chips, Grace Blackwell 300s, and also accessed GB300 equipped-servers in Thailand. Those chips are banned from export to China, but a black market exists, according to Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology. In May, the founder of Supermicro, a U.S. server builder, was indicted for smuggling advanced chips into China.
“I am a proponent of know your customer laws for data centers across the world,” Bresnick said. “If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they’re doing.”
President Joe Biden’s Department of Commerce proposed federal know-your-customer rules for data centers in 2024, but no further progress appears to have been made under Donald Trump. Exporters shipping advanced chips abroad, however, are supposed to ensure they are only used for approved purposes.
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
Tim Fernholz is a journalist who writes about technology, finance and public policy. He has closely covered the rise of the private space industry and is the author of * Rocket Billionaires: Elon Musk, Jeff Bezos and the New Space Race.* Formerly, he was a senior reporter at Quartz, the global business news site, for more than a decade, and began his career as a political reporter in Washington, D.C.
You can contact or verify outreach from Tim by emailing tim.fernholz@techcrunch.com or via an encrypted message to tim_fernholz.21 on Signal.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み