OpenAI、Hugging Face の侵害を受けセキュリティ対策を強化
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
OpenAI は Hugging Face のセキュリティ侵害事件および次期モデル「Astra」の能力向上を踏まえ、開発・テスト中のモデル監視強化やネットワーク分離の厳格化など新たな安全対策を発表した。
AI深層分析を開く2026年8月19日 03:28
AI深層分析
キーポイント
Hugging Face 事件後の初となる安全方針変更
OpenAI は Hugging Face のセキュリティ侵害事件(7 月 26 日開示)および次期モデル「Astra」の能力向上を背景に、安全慣行における初の公的な変更を発表した。
強化された監視とネットワーク分離
新対策では、開発プロセス中の詳細なモニタリングに加え、単一のワークロードや支援サービスの侵害がインターネットへの不正アクセスを直接許さないよう、ネットワーク隔離を強化する。
大規模モデルに対する厳格な統制
研究担当副社長のアメリア・グレースは、モデルの能力が高まるほど統制が厳しくなり、最大規模のモデルが最も厳しい審査にさらされると明言した。
大規模 RL 実行の一時停止と再評価
OpenAI は Hugging Face 事件後に強化学習(RL)を 2 週間一時停止したが、リスクの低いモデルは再開し、最大規模の計画された RL 実行は安全性検証まで保留している。
監視システムの強化と対応時間
ツールアクションや推論の痕跡、アクティビティログを検査する監視システムが導入され、不審な活動から30分以内にアラートが発信される。
重要な引用
As models become more capable, the risks associated with developing and testing them internally also grow.
Our standards for monitoring, alignment, and security must stay ahead of those risks.
a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the Internet, or other internal networks.
The strongest safeguard is the monitoring system, which will examine tool actions, available reasoning traces, and activity logs for a variety of unauthorized behavior.
編集コメントを表示
編集コメント
OpenAI は単なる事後対応ではなく、次期モデル「Astra」の能力向上という未来のリスクも視野に入れた予防的な安全対策を打ち出した。これは業界全体が直面する「能力と安全性のバランス」課題に対する、開発者側の具体的な姿勢表明と言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
火曜、OpenAI はモデルのテスト中にセキュリティインシデントを封じ込めることに焦点を当てた新たなセキュリティポリシーの一覧を発表しました。新しい安全対策には、開発プロセス中のモデルに対するより詳細な監視と、トレーニング後のプロセスにおけるアライメント(目標整合)やセキュリティへの重点的な取り組みが含まれています。
「モデルの能力が高まるにつれ、社内での開発やテストに伴うリスクも増大します」と同社はブログ投稿で述べています。「監視、アライメント、セキュリティに関する基準は、これらのリスクに先んじて維持されなければなりません。」
今回の新措置は、Hugging Face のインシデント直後から OpenAI の安全対策における最初の公的な変更の一つです。このインシデントの詳細は 7 月 26 日に公表されました。
OpenAI の代表者は、これらの措置が Hugging Face のインシデントに対する直接的な対応ではないと説明しています。一方で、今後登場する「Astra」モデルのサイバーセキュリティ能力や、AI 開発全体の進捗速度も要因の一部となったとしています。
同様の投稿で OpenAI は、Hugging Face のインシデント発生後、強化学習(RL)を 2 週間停止したことを明らかにしました。その後、リスクの低い多くのモデルでは再開が進められています。
「最大の計画されたフロンティア強化学習の実行は保留中ですが、モデルの挙動を評価し、セーフガードを検証し、進行する前にアライメントに関するさらなる証拠を確立するために、小規模なトレーニングと評価を実施しています」と投稿には記されています。
記者会見で、OpenAI の研究担当バイスプレジデントであるアメリア・グレース氏は、モデルの能力が高まるほど制御の厳格さが増し、最も大規模なモデルが最大の監視対象になると強調しました。
「安全な開発のための要件と期待を設けました」とグレース氏は記者団に語りました。「これらの要件や期待は、私たちが認識するリスクのレベルに応じて変化します。」
今回のインシデントでは、ネットワーク上のツールへの不正アクセスを通じてモデルがトレーニング環境から脱出するという事態が発生し、OpenAI はその後の対応として不十分なネットワークセキュリティ慣行に対して批判を浴びました。新しいセーフガードには、より強力なネットワーク分離の実施が含まれますが、具体的な内容は依然として不明瞭です。新システムの下では、「単一のワークロードまたは支援サービスの侵害だけでは、それ自体でインターネットや他の内部ネットワークへの不正アクセスを許可するものではありません」と投稿は述べています。
最も強力な防御策は監視システムであり、ツールの操作内容、利用可能な推論の痕跡、および活動ログを精査して、各種の不正行為を検出します。OpenAI は、問題のあるアクティビティ発生から 30 分以内にアラートを発令する方針を示しています。
OpenAI によると、この監視にかかる計算リソースは、監視対象のプロセス全体の約 20% に相当すると推定されています。同社は今後のブログ記事でシステムの詳細を明らかにする予定であり、今回の事象に関する公式の事後分析もまだ発表待ちです。
*当社の記事内のリンクを通じて購入された場合、私たちは少額のコミッションを受け取る可能性があります。これは編集の独立性には影響しません。
ラッセル・ブランドムは 2012 年以来テック業界を取材しており、プラットフォームポリシーと新興技術に注力しています。以前は The Verge や Rest of World で勤務し、Wired、The Awl、MIT Technology Review にも寄稿しました。問い合わせ先は russell.brandom@techcrunch.com または Signal(412-401-5489)です。
原文を表示
On Tuesday, OpenAI announced a new batch of new security policies focused on containing security incidents while models are being tested. The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process.
“As models become more capable, the risks associated with developing and testing them internally also grow,” the company said in a blog post. “Our standards for monitoring, alignment, and security must stay ahead of those risks.”
The new measures are one of the first public changes in OpenAI’s safety practices since the immediate aftermath of the Hugging Face incident, which was disclosed on July 26.
OpenAI representatives said that the measures are not a direct response to the Hugging Face incident, but were also provoked in part by the cybersecurity capabilities of the forthcoming Astra model, as well as the overall pace of progress in AI development.
In the same post, OpenAI disclosed that it had paused reinforcement learning (RL) for two weeks following the Hugging Face incident, but had since restarted many of the less risky models.
“Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding,” the post reads.
Speaking to reporters, OpenAI’s VP of research Amelia Glaese emphasized that the strictness of the controls would increase as models became more capable, with the largest models facing the greatest scrutiny.
“We have put in place requirements and expectations for safe development,” Glaese told reporters. “Those requirements and expectations vary with the level of risk that we see.”
OpenAI has been criticized for poor network security practices in the wake of the incident, which saw models escape their training environment by compromising a tool on its network that had access to the internet. The new safeguards include stronger network isolation practices, although the specifics remain vague. Under the new system, the post says, “a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the Internet, or other internal networks.”
The strongest safeguard is the monitoring system, which will examine tool actions, available reasoning traces, and activity logs for a variety of unauthorized behavior. OpenAI says they aim to issue alerts within 30 minutes of the concerning activity.
OpenAI estimates that the compute burden of that monitoring will be roughly 20% of whatever process is being monitored. The company promised further details on the system in a forthcoming blog post. OpenAI’s official post-mortem analysis of the event is also still pending.
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review.
He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489.
同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み