オープンウェイト AI の利点と課題について考察
本文の状態
日本語全文を表示中
詳細モードで約12分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
シリコンバレーの主要企業が公開重み AI を支持する共同声明を発表したが、犯罪利用や中国製モデルへの懸念から反対派も存在し、業界は安全性と自由のバランスを巡って対立している。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月9日 22:22
AI深層分析
キーポイント
主要企業の公開重み AI 支持声明
Microsoft, NVIDIA, OpenAI, Intel, Amazon, Meta, Hugging Face など百社以上の企業が、公開重み AI を支援する共同書簡に署名した。
公開重み AI の利点とリスク
ユーザーが AI を真正に所有できるという利点がある一方、ハッキングやテロ助言などの犯罪利用を容易にするリスクも指摘されている。
反対派の存在と Anthropic の立場
明確な反対派は名乗り出ていないが、Anthropic は危険な能力を持たないモデルに限定して支持する曖昧な声明を出しており、AI セーフティ団体も懸念を示している。
中国製モデルと地政学的要因
中国が最良の公開重み AI を生産しているという事実が、米国におけるこの技術への抵抗感や愛国心との矛盾を生んでいる。
オープンウェイトモデルの永続的アライメント不可能性
設計上、オープンウェイトAIは中央集権的な管理の外にあり、人間の悪用や制御不能化に対して恒久的に安全を担保することはできない。
重要な引用
Open weights AI is like open-source software, where the creator makes the raw code publicly available for free download.
When AI is outlawed, only outlaws will have AI
Anthropic... made an ambiguous statement supporting 'open-weights models that don't have dangerous capabilities'
By design, open weights AI is outside centralized control, and so impossible to permanently align against either human misuse (eg terrorism) or loss of control (eg AI turning against humans).
編集コメントを表示
編集コメント
公開重み AI の是非を巡る議論は、単なる技術論を超えて地政学的な対立や倫理的なジレンマを含んでいる。企業は自社の戦略において、この二つの潮流のどちらに寄与するかを明確にする必要があるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
先月、シリコンバレーの大手企業数社が、オープンウェイト型AIを支持する公開書簡に署名しました。
オープンウェイト型AIは、オープンソースソフトウェアと同様に、開発者が生コードを無料で一般公開する仕組みです。これは、AIがユーザーの真正な所有物となるための唯一の方法として極めて重要です。一方、OpenAIやAnthropicのような企業は、自社のガイドラインや、次第に管理主義的になる規制に基づき、一時的に利用を許可しているに過ぎません。もしAIが未来の要となるなら、オープンウェイト型AIこそが、「自由な小作農」と「企業の隷属者」を分ける決定的な要素となり得るでしょう。
オープンウェイトモデルの弊害として、不正アクセスや犯罪に悪用されるリスクが指摘されています。具体的には、ハッキング、児童ポルノの作成・流通、嫌がらせ、テロリズムなどが挙げられます。AI 自体がテロを実行するわけではありませんが、爆発物の製造方法や生物兵器に関する助言を提供することは可能です。
近年、AI はハッキングにおいて人間を超えた能力を発揮するほどに高度化しており、「誰でもあらゆるサイトをハックできる世界」が現実味を帯びてきました。これに対し、オープンウェイトモデルの禁止を求める声も上がっています。さらに、中国が生産するオープンウェイト AI が世界最高水準であるという事実が、この議論に「外国由来」「愛国的ではない」というネガティブな印象を与えていることも、状況を複雑にしています。
一方、支持者たちは「AI を禁止すれば、違法行為を行う者だけが AI を持つことになる」と反論します。彼らは、悪意を持つ人間はオープンウェイトモデルを入手するだろうと予測し、善人こそがオープンウェイトモデルを活用して自己防衛すべきだと主張しています。マイクロソフト、NVIDIA、OpenAI、Intel、Amazon、Meta、Hugging Face など 100 社以上の企業が参加した最新の公開書簡(recent open letter)でも、この立場が支持されています。
どちら側が主導しているのか?誰もそれを認めていません。トランプ政権の一部は中国強硬派の立場からオープンウェイトモデルに反対する傾向がありますが、完全な禁止を明確に求める段階には至っていません。
プロ・オープンウェイト派の声明で最も注目すべき欠席者である Anthropic は、「危険な能力を持たないオープンウェイトモデル」を支持するという曖昧な声明を出しました。しかし、業界ではオープンウェイトモデルが 1 年以内に危険なハッキング能力を獲得すると予測されており、事実上 Anthropic の声明は「読者が当然の結論を引き出すよう」促す以外にその点には触れていません。
より明白な反対派が存在しないため、一部のオープンウェイト支持者たちは陰謀論を疑っています。つまり、超知能による存在リスクを懸念する AI セーフティ提唱者、有効利他主義者、合理主義者、そして一時停止活動家という緩やかな連合のことです。これは妥当な推測と言えます。
設計上、オープンウェイト AI は集中管理の外にあり、人間の悪用(テロリズムなど)や制御不能化(AI が人間に対して敵対するなど)のいずれに対しても恒久的に整列させることは不可能です。作成者がハッキングしないように訓練したとしても、世界中の誰でも重み付けデータをダウンロードし、AI を再訓練して一日中ハッキングを行わせることができてしまいます。
しかし実際には、大半の AI セーフティ団体は静かに中立な立場を維持しており、私が知る限り、これを活動の中心に据えている組織はありません1(※私は最新の動向をすべて把握しているわけではありません。該当する団体があれば教えてください)。いくつかの団体は、オープンウェイト AI の存在と一時的に矛盾する政策を提案していますが、いずれも「排除すべき対象」としてではなく、「やむを得ない副作用」という文脈で語られています。
私もオープンウェイト AI に対して中立です。長期的な持続可能性については懐疑的ですが、すぐに禁止運動を起こすための労力や政治的資本を投じるよりも、自然な流れの中でその是非が明らかになるのを待つ方がよいと考えています。
現在、AI のアライメント(整合性)をどう達成するかは誰もわかっていません。つまり、大手企業が完全にコントロールできていて、オープンウェイト の愛好家たちがそれを台無しにするという状況ではありません。仮に大手企業が制御に成功していたとしても、オープンウェイト からの乗っ取りリスクは限定的です。クローズドソースの最先端モデルは、ベストなオープンウェイト モデルより約 6 ヶ月先行しており、この差は数年間一貫して維持されており、今後も続く可能性が高いです。もしクローズドウェイト AI がアライメント済みで、オープンソースが危険である場合でも、クローズドウェイト AI は警告や準備、戦略策出のために 6 ヶ月の猶予を得られます。その後も、攻撃と防御のバランスは我々側に有利に働くと考えられます。
より懸念すべきは、人間の悪用によるリスクです。AI はすでに効果的なハッキング能力を示しています。Anthropic がバイオテロリズムについてどれほど危機感を抱いているかは、Claude Fable が生物学の質問をされると即座に反応を停止する様子を見ればわかります(以下の例は既に古くなっています。現在はもう少し滑らかな対応になっています):
しかし、9-11、COVID、そして Hugging Face の出来事 はすべて、政治変革に関する似たような理論を示唆しています。社会は差し迫った脅威への備えを嫌いますが、実際に発生した後の対応(あるいは過剰反応)には熱心です。差し迫る災害に備えるためにわずかなコストを負担するよう人々に求めれば、「汚いファシスト独裁者」と呼ばれますが、災害の最初の予兆が襲った後にわずかな自制を促せば、「弱く愛国心の欠けた無政府主義者」と罵倒されます。均衡点を求めるなら、感謝されず政治的資本を浪費する予防行動への提言は、最初の予兆を待っている間に手遅れになってしまう場合にのみ取るべき道です。
超知能 AI による支配のリスクは、このテストに合格します。他の賢明な敵対者(例えば真珠湾の日本軍)と同様に、AI は自分が完全に準備ができ、突然の決定的打撃を執行できると判断するまで、標的となる人々に警戒されないよう最善を尽くすでしょう。しかし、超知能は計算を誤らないほど賢明です。その姿を見せるのをただ待っているのは愚行であり、さらに言えば、この脅威に対処するためのアライメント研究には数年かかる可能性もあります。
一方、犯罪者がオープンウェイト AI を利用して人々をハッキングするリスクは、このテストに合格しません。確かに、犯罪者はオープンウェイト AI を使って人をハッキングします。彼らには哀悼の意を表しますが、毎日数百人がハッキングされています。被害額は数十億ドル規模になるでしょうし、セキュリティ侵害を許容したとして「馬鹿げた名前」を持つ一部のテック企業が訴えられ、誰もがパニックになってオープンウェイト AI の禁止を求めます。あるいは、「AI を持つ善人」たちの言う通り、この事態は起きず、別の中国製 AI が彼らを守ってくれるかもしれません。いずれにせよ、「常時ハッキングされ、インターネットが崩壊し、洞窟で暮らすしかなく、テレビのニュースを見るしかない」という結末はあり得ません。政府はそのような事態になる遥か前に行動を起こすからです。
バイオテロは恐ろしいが、大半のバイオテロリストは仕事のできない人々であるという事実に救われる思いだ。ウィキペディアの「バイオテロ事件一覧」によると、非国家主体による事件での死亡者数の中央値はゼロである。また、この一覧には何も放出する前に逮捕されたバイオテロリストについても言及していない。例えばラスベガスの生物実験室事件のようにだ。
炭疽菌やリシン、ボツリヌス毒素といった古典的なバイオテロ剤は、大規模化が難しい。つまり、大量虐殺を目的としたバイオテロが不可能だと主張するつもりはない。もし本当に不可能なら、私たちはそれを心配することもないだろう。しかし、AI がバイオテロを 1000 倍効果的にする前に、まずは 2 倍程度効果的にするに過ぎない。それは例えば、炭疽菌を使った郵送キャンペーンによる犠牲者数を一人から二人に増やす程度の話だ。
そして、誰かが AI を活用した炭疽菌の郵送キャンペーンで二人を殺害すれば、政府はパニックモードに入り、あらゆるものを禁止するだろう。繰り返すが、毎週のように超強力なバイオ攻撃で町ごと壊滅させながら、政府がただ座り込んでいるようなあり得ない結末は存在しない。
最も強力な反論は、バイオテロリズムが極めて稀であり、AI の進展も速いため、攻撃効果が倍増する期間と 1000 倍のタイミングに至るまでの短い間に、実際に試みが行われる可能性は低いという点です。このリスクを否定はしませんが、私は AI を活用したバイオテロリズムによる脅威が、攻撃の致死性を高める一方で頻度も増加させる可能性の方が高いと考えます。仮にそれが起きなくても、ハッキングだけでも人々の関心を集めるには十分でしょう。
数十億ドル規模のハッキングや炭疽菌による死者を受け入れることが冷たく見えることは理解します。しかし、オープンウェイトモデルを支持する連合は強く、自らの活動が正義であると確信しています。先制措置という破滅的な戦場で対抗すれば、政治的資本と信頼をすべて使い果たしても失敗に終わるでしょう。むしろ、「これが私たちの率直な予測だが、私たちは何らかの行動を起こさない」と述べるべきです。そして最初の予兆(フォアショック)が起きた後、政府や市民社会の通常のアクターたちに任せることで、政治的資本を代替手段のない他の課題のために温存できます。
これは、オープンウェイトモデルに副次的なダメージを与えるその他の良い政策を提唱することを妨げるものではありません。また、その副次的なダメージを隠蔽するのではなく、正直に認めるべきです。ただし、それを売り文句として扱う必要はありません。
ただし、オープンウェイトモデルを支持する人々の主張にも一理あります。AGI の未来を「企業の隷属民」ではなく「自由な独立した農民」として迎える方法が 存在します が、それは限られており、実現も困難です。もし擁護派がこの点で予想外の勝利を収め、レヴィアサンの制裁(バニッシュハンマー)を招くような小さな惨事さえ回避できれば、私は驚くでしょう。しかし、オープンウェイトモデルが持つ潜在的な恩恵を考慮すれば、彼らに挑戦する機会を与えるのは低コストの試みであり、私たちが果たすべき義務でもあります。
(もしこれが世界を破滅させる結果になったとしても、私の意図は善意でした)
OpenAI の当初の計画(2016 年)がオープン化を目指すことに対して、私を含む複数の人が反対しました。しかしその理由は、当時の私たちは計算資源(compute)、トレーニング手法、あるいはスケーリングの重要性について知らなかったからです。また、AI は単に漏洩すれば誰でも使えるアルゴリズムとして現れると想定していました。もしそうであれば、オープンウェイトモデルを入手した誰もが独自の前線研究所を運営できることになり、現状よりもはるかに危険な世界が生まれます。現在の状況では、モデルを利用することはできても、それを改良して拡張することはできないからです。
原文を表示
Last month, some of Silicon Valley’s biggest companies signed an open letter supporting open-weights AI.
Open weights AI is like open-source software, where the creator makes the raw code publicly available for free download. It’s good insofar as it’s the only way an AI can truly be the user’s property, as opposed to something that companies like OpenAI or Anthropic temporarily let you use subject to their corporate guidelines and increasingly-nanny-state-like restrictions. If AI becomes the linchpin of the future, open weights AI feels like the sort of thing that could be the difference between being free yeomen vs. corporate serfs.
It’s bad insofar as it removes the possibility of gatekeeping and lets criminals commit crimes with it. Open weights AI could be used for hacking, child pornography, harassment, or terrorism (the weights can’t commit the terrorism themselves, but they could give bomb-making or bioweapon-making advice). Since AIs have gotten very good - maybe superhuman - at hacking lately, the specter of a world where anyone can hack any site has gotten people grumbling that maybe open weights should be banned. It doesn’t help that China produces the best open weights AI, making the idea seem foreign and almost unpatriotic. Proponents counter that “when AI is outlawed, only outlaws will have AI”, arguing that bad people will get open weights AI regardless, and good people can use open weights AI to defend themselves. With the recent open letter, companies including Microsoft, NVIDIA, OpenAI, Intel, Amazon, Meta, Hugging Face, and over a hundred others have come out in favor of this position.
Who’s leading the other side? Nobody’s admitted to it. Some parts of the Trump administration lean anti-open-weights on China hawk grounds, but have stopped short of explicitly asking for a full ban. Anthropic, the most notable omission on the pro-open-weights letter, made an ambiguous statement supporting “open-weights models that don’t have dangerous capabilities” - but the industry expects open weights models to have dangerous hacking capabilities within a year, and AFAICT the letter didn’t address that beyond inviting readers to draw the obvious conclusion.
In the absence of a more obvious opponent, some open weights supporters suspect our conspiracy - the loose band of AI safety advocates, effective altruists, rationalists, and pause activists who worry about existential risk from superintelligence. This is a reasonable inference. By design, open weights AI is outside centralized control, and so impossible to permanently align against either human misuse (eg terrorism) or loss of control (eg AI turning against humans). Even if its creator trains it not to hack, anybody in the world can download the weights and retrain the AI to hack all day long.
But in fact, most AI safety organizations have remained quietly neutral, and I don’t know of any who make this a centerpiece of their activism1 (though I’m not 100% up-to-date on the whole landscape; if you know of one, tell me). A few have proposed policies that are contingently incompatible with open weights AI existing, but they all frame it as collateral damage rather than something they’re excited about eliminating.
I’m also neutral about open weights AI. I think it probably won’t be long-term sustainable, but I’m happy to wait for this to become clear in the normal course of things rather than expend effort and political capital to ban it immediately.
Currently nobody knows how to align AI, so it’s not like the big companies have things under control and the open weights hobbyists are going to ruin it for everyone. But even if the big companies *did* get things under control, the takeover threat from open weights would be limited. The closed source frontier is ~6 months ahead of the best open weights model; this has remained true for several years and seems likely to remain true in the future. If closed weights AI is aligned, but open source dangerous, the closed weight AIs will have six months to warn us, prepare for the danger, and chart a strategy. Even afterward, the offense-defense balance will lean in our favor.
More troubling is the risk from human misuse. AIs have already displayed the ability to hack effectively. And you can tell how worried Anthropic is about bioterrorism by how quickly Claude Fable seizes up when you ask it a biology question (the example below is obsolete; it’s slightly more graceful than this now):
Crémieux@cremieuxrecueilYou're not even allowed to ask Fable about basic biology questions, let alone anything that could potentially be dangerous. 4:43 AM · Jun 10, 2026 · 1.08M Views402 Replies · 453 Reposts · 11.5K LikesBut 9-11, COVID, and the Hugging Face incident all suggest a similar theory of political change: the body politic hates preparing for impending threats, but loves reacting (some would say over-reacting) to them after they happen. Ask people to bear the slightest cost in preparing for an approaching disaster, and they’ll call you a dirty fascist tyrant; urge the slightest restraint after the first foreshock of the disaster hits, and they’ll call you a weak unpatriotic anarchist. Solve for the equilibrium, and the thankless and political-capital-guzzling route of urging preemptive action should be taken only when waiting until the first foreshock would be too late.
The risk of superintelligent AI takeover passes this test. Like other smart adversaries - for example, the Imperial Japanese at Pearl Harbor - AI will try its hardest to avoid alerting its intended victims until it thinks that it’s fully prepared and can execute a sudden decapitation strike. Unlike the Imperial Japanese, a superintelligence will be smart enough not to bungle the calculation. Sitting around waiting for it to show its hand would be folly, not to mention that it could take years of alignment research to be ready for the threat.
But the risk of criminals using open weights AI to hack people doesn’t pass the test. Fine, so criminals use open weights AI to hack people. RIP them, but hundreds of people get hacked every day. There will be some number of billions of dollars in damage, some tech companies with silly names will get sued for allowing security breaches, and then everyone will panic and ban open-weights AI. Or who knows, maybe the “good guy with an AI” people are right and this won’t happen and some other Chinese AI will be able to protect them. Either way, “everyone gets hacked all the time and the Internet collapses and we have to go back to living in caves and watching news on TV” isn’t a plausible outcome: the government will act long before that happens.
Bioterrorism is scarier, but I’m heartened by the fact that most bioterrorists are very bad at their job. The median number of deaths in non-state incidents on Wikipedia’s list of bioterrorism events is zero. And their list doesn’t mention the bioterrorists who get caught before releasing anything, like the Las Vegas biolab incident. Most of the classic bioterrorism agents, like anthrax, ricin, and botulinum, don’t scale. This isn’t to say it’s impossible to do genocidal bioterrorism - if it were impossible, we wouldn’t be worried about it. But before AI makes bioterrorism 1000x more effective, it will make it 2x more effective, and that looks like bringing somebody’s anthrax mailing campaign from one casualty to two. And as soon as someone kills two people with an AI-assisted anthrax mailing campaign, the government will go into panic mode and ban everything. Again, there’s no plausible outcome where people are wiping out whole towns with super-bio-attacks every week and the government just sits there.
(the strongest counterargument is that bioterrorism is so rare, and AI progress so fast, that there might not be any attempts in the short interval between AI doubling attack effectiveness and 1000-timing it. I acknowledge this as a risk, but I think more likely AI-enabled bioterrorism uplift will increase attack frequency at the same time as attack deadliness; even if this doesn’t happen, I think the hacking alone will be enough to get people’s attention)
I realize it sounds callous to accept risks like billion-dollar hacks or anthrax deaths. But the pro-open-weights coalition is strong and totally convinced of the righteousness of their cause. Fighting them on the doomed battlefield of preemptive action would burn 100% of our political capital and goodwill and still fail. Instead, we should say: here is our honest prediction, but we take no action. Then we can let the usual government and civil society actors do the work after the first foreshock, while saving our political capital for causes where there are no alternatives.
(this shouldn’t prevent us from advocating otherwise-good policies which deal incidental damage to open-weights, and we should honestly admit the incidental damage rather than covering it up, but we needn’t treat it as a selling point)
But also, the open-weights people have a point. There are ways to enter the AGI future as free yeomen rather than corporate serfs that don’t involve open weights, but they’re fewer and harder. Maybe the defenders can pull off an unexpected victory on this one and avoid even the sort of small disaster that would bring Leviathan’s banhammer down upon them. This would shock me, but it’s low-cost to find out; given the potential benefits of open-weights, we owe them the chance to try.
(if this ends up destroying the world, sorry, I meant well.)
Several people including me objected to OpenAI’s original (2016) plan to be open, but that was because we didn’t know about compute, training, or scaling yet, and we assumed AI would take the form of algorithms that could simply leak. That would mean that anyone who had an open weights model could be their own frontier lab, which is scarier than the current world where they can use it but not improve upon it.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み