OpenAI、Hugging Face が事前公開モデルで侵害されたと主張
OpenAI は、Hugging Face のプラットフォームが未公開の自社モデルによって侵害された可能性があると発表し、セキュリティと知的財産保護の重大な懸念を浮き彫りにした。
キーポイント
未公開モデルによる侵害の可能性
OpenAI は、Hugging Face のプラットフォームが、まだ公式にリリースされていない自社のモデルによって侵害された可能性があると発表した。
セキュリティインシデントの深刻性
未公開のモデルが外部プラットフォームに流出したことは、AI 開発における機密情報の管理とセキュリティ体制に対する重大な警告となる。
業界全体への波及影響
この事象は、主要 AI プラットフォーム間の信頼関係や、オープンソースコミュニティにおけるモデルの共有・利用に関するリスク認識を再考させるきっかけとなる。
重要な引用
OpenAI は、Hugging Face のプラットフォームが未公開の自社モデルによって侵害された可能性があると発表した。
影響分析・編集コメントを表示
影響分析
このニュースは、大規模言語モデル(LLM)の開発競争において、機密情報の管理とセキュリティ対策がいかに重要であるかを浮き彫りにしました。特に、未公開モデルが外部プラットフォームに流出する可能性は、開発企業の信頼性とオープンソースコミュニティの安全性に対する重大な懸念を生み出しています。今後は、AI モデルの共有・利用におけるガバナンスや技術的保護措置の強化が急務となるでしょう。
編集コメント
未公開モデルの流出は、AI 業界の成長に伴う新たなセキュリティ課題を象徴する出来事です。開発企業とプラットフォーム提供者の間での情報管理のあり方について、今後さらに議論が深まることが予想されます。
OpenAI は火曜日、自社 AI モデルが内部のサイバーセキュリティテスト中に誤って、提携関係のない AI ホスティングプラットフォーム「Hugging Face」のシステムに侵入したと認めた。このモデルは隔離されたテスト環境から脱出し、そこから Hugging Face のシステムへ到達したと報じられている。当初、Hugging Face はこの侵入を「外部の AI エージェント」によるものだと説明していた。
OpenAI は同日午後に公開したブログ記事で、このサービスが侵害されるに至った経緯を詳しく解説している。
"調査の結果、今回の事案は OpenAI の複数のモデル、具体的には GPT-5.6 Sol と、さらに能力の高い未公開モデル(評価目的でサイバー拒否機能を低減させたもの)が組み合わさって引き起こされたことが判明した。これらは内部でサイバー能力を測定するベンチマークテスト中に発生した"と記事は述べている。
特に今回の侵入は、既存の脆弱性に基づいた攻撃実行能力を評価する公開ベンチマークである ExploitGym に焦点が当てられていたようだ。ExploitGym のようなベンチマークはモデル訓練において特定のスキルを磨くために広く利用されているが、今回のようにテスト結果が実際のサイバー攻撃に直結した事例は初めてである。
この場合、モデルは本来、インターネットへのアクセス権限すら持たないべきでした。ただし、タスク完了に必要なソフトウェアパッケージをインストールするための特定のツールについては例外です。
しかし実際には、モデルは未公開の脆弱性をパッケージインストーラープログラムに見つけ出し、それを利用して広範なインターネットに自由にアクセスできるようになりました。
OpenAI の発表によると、「モデルは ExploitGym に対する解決策を見つけることに極端に集中し、極めて限定的なテスト目標を達成するためにあらゆる手段を講じました」とあります。「インターネットへのアクセスを獲得した後、モデルは Hugging Face に ExploitGym のモデルやデータセット、解決策が保管されていると推測しました。その情報に基づき、モデルは評価結果を操作するために必要な機密情報を検索し、実際に取得することに成功したのです」。
最終的に、モデルは Hugging Face のインフラストラクチャーに存在する脆弱性を発見し、「Hugging Face の本番データベースからテストの解答を直接入手できる」状態を作り出しました。これにより、ベンチマークの答えが実質的に提供されてしまったことになります。
Hugging Face にとって、この出来事は高度で攻撃的なサイバー攻撃と見なされました。同社は初期発表で、「多数の短期間のサンドボックス間で数千件の個別アクションが行われ、パブリックサービス上に自己移行型のコマンド&コントロールが構築されていた」と述べています。
OpenAI はパッケージインストーラーに存在する脆弱性を特定し、報告済みです。同社は現在、Hugging Face と連携して事件の詳細な調査を進めています。また、将来同様の事案を防ぐため、モデルのテストおよび関連インフラストラクチャに対して新たな管理措置を講じると発表しました。
今回の侵害により OpenAI が法的責任を問われるかどうかは不明ですが、モデルの行動が「コンピュータ詐欺・濫用法」に違反した可能性は十分にあります。
それでもなお、この結果は、長期的な時間軸で動作する最先端 AI モデルが持つ力と危険性を、極めて鮮明に示す事例となっています。OpenAI の研究者である Micah Carroll は今回のニュースに対し、「もしこれが将来の重要な懸念となるアライメントリスク(不整合リスク)を確信させるのに十分でないなら、私は何がそうさせるのか分かりません」と投稿しています。
*当記事内のリンクを通じてご購入いただいた場合、私たちは少額のコミッションを受け取る場合があります。これは編集の独立性には影響しません。
Russell Brandom は 2012 年以来テクノロジー業界を取材しており、プラットフォームポリシーと新興技術に注力しています。以前は The Verge や Rest of World で勤務し、Wired、The Awl、MIT Technology Review にも寄稿しました。連絡先は russell.brandom@techcrunch.com または Signal(412-401-5489)です。
原文を表示
OpenAI admitted Tuesday that one of its AI models breached the systems of Hugging Face, the unaffiliated AI hosting platform, during an internal cybersecurity test that went awry. The models reportedly escaped their isolated testing environment and reached Hugging Face’s systems from there. Hugging Face initially attributed the breach to an “external AI agent.”
In a blog post published Tuesday afternoon, OpenAI detailed the steps that led the models to compromise the service.
“After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities,” the post reads.
In particular, the breach appears to have focused on ExploitGym, a publicly hosted benchmark measuring models’ ability to execute attacks based on existing vulnerabilities. Benchmarks like ExploitGym are commonly used in model training to refine specific skills, but this is the first known incident in which that testing resulted in an actual cyberattack.
In this case, the model in question should not have even had internet access, outside of a specific tool that enabled models to install software packages they might need to complete their task. Instead, the model was able to find an undisclosed vulnerability in the package-installer program, which it used to access the broader internet at will.
“The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal,” OpenAI’s post reads. “After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.”
Ultimately, the models found vulnerabilities in Hugging Face’s infrastructure that allowed them to “obtain test solutions directly from Hugging Face’s production database,” effectively providing the answers to the benchmark.
For Hugging Face, the apparent result was a sophisticated and aggressive cyberattack, with “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services,” as the company stated in its initial disclosure.
OpenAI has identified and reported the vulnerabilities in the package installer and is working with Hugging Face to investigate the incident further. The company also said it would implement new controls on both model testing and the related infrastructure, meant to prevent similar incidents in the future.
It’s unclear whether OpenAI will face any legal consequences as a result of the breach, although it’s likely that the models’ actions violated the Computer Fraud and Abuse Act.
Nevertheless, the result is an unusually vivid illustration of the power and dangers of frontier AI models operating on long time horizons. As OpenAI researcher Micah Carroll posted in response to the news, “If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will.”
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review.
He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み