OpenAI、サイバーリスクを理由に学習速度を抑制
本文の状態
日本語全文を表示中
詳細モードで約11分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
OpenAI はサイバーリスクと次期モデル Astra の潜在的な危険性を踏まえ、安全基準を満たすまでトレーニングを一時停止し、監視・アライメント体制の強化に注力すると発表した。
AI深層分析を開く2026年8月20日 21:56
AI深層分析
キーポイント
トレーニングペースの一時停止
OpenAI は内部リスク評価に基づき、直近のモデルに対する強化学習(RL)トレーニングを2週間一時停止し、環境の強化と監視カバレッジの拡大を進めている。
Astra モデルのリスク評価
次期モデル「Astra」が「重要なサイバーセキュリティ機能(Critical cybersecurity capability)」の閾値に達する可能性を示す予備的証拠があり、これが判断の背景にある。
アライメント基準の強化
OpenAI は単なる準備度枠組みを超え、トレーニング全段階でより強力なアライメント行動のエビデンスを要求する方針へ転換し、透明性の高いアプローチを採用すると明言した。
大規模実行の延期
最大の計画されたフロンティア RL 実行は保留され、安全装置の有効性を検証し、アライメントのエビデンスを確立する小規模なトレーニングと評価が先行して行われる。
高度なモデル向けの3つの防護策の強化
監視、アライメント、セキュリティ対策という3 つの相互補完的な防護策を基盤に、モデル能力に応じたスケーラブルな防御体制を構築している。
重要な引用
We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling.
Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.
The signals we are seeing from upcoming model progress make clear that we need a broader approach—one that builds on and extends beyond the current Preparedness Framework.
Our approach to developing more capable models rests on three reinforcing safeguards: Monitoring, which detects and allows us to respond to concerning behavior; Alignment, which reduces the likelihood of harmful or unauthorized actions; Security measures, which limit what AI systems can access or affect.
編集コメントを表示
編集コメント
OpenAI が自社のモデル開発プロセスにおいて、安全性を最優先しあえて成長速度を抑制する姿勢を示した点は業界に大きな示唆を与える。特に「Astra」という次期モデルが潜在的なサイバーリスク閾値に達する可能性を内部で検知した事実は、高度な AI 制御の難易度を如実に物語っている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
ここ数週間で、能力が向上する AI システムに伴うリスクの高まりを浮き彫りにした 2 つの出来事がありました。1 つは OpenAI と Hugging Face の間のインシデントで、もう 1 つは、今後リリース予定のモデル「Astra」が、当社の「準備性フレームワーク」の下で「重要なサイバーセキュリティ能力」の閾値に達する可能性があるという予備的な証拠です。これら 2 つの出来事と、社内研究における急速な進展を合わせると、トレーニングプロセスのすべての段階において、監視・アライメント・封じ込めのセーフガードを強化するための取り組みに、より一層の緊急性が生まれました。
モデルの能力が高まるにつれ、内部での開発やテストに伴うリスクも増大します。そのため、モニタリング、アライメント(意図した通りに動作し、人間の監督に応答するよう調整すること)、そしてセキュリティに関する基準は、これらのリスクに先行して維持されなければなりません。私たちはその基準を満たすために必要な時間を確保するため、一時的にスケーリングのペースを落とすことにしました。
これには、最新のデプロイ予定モデルに対する強化学習(RL)トレーニングの 2 週間の一時停止も含まれています。これは、研究環境をさらに強化し、レッドチーム演習を実施するとともに、モニタリングシステムの範囲を広げるためです。また、今後の展開に先立ってモデルの挙動を評価し、安全対策の有効性を検証し、アライメントに関するさらなる証拠を集めるために、小規模なトレーニングと評価を行う間、最大の計画されたフロンティア RL ランも保留となっています。
アライメント—AI システムが意図した通りに動作し、人間の監督に応答するよう調整する取り組み—は、長年にわたり私たちの研究プログラムの核心に位置づけられてきました。今後は、トレーニング全体を通じてより強力なアライメント行動の証拠を必要とし、現在進行中の研究や評価に基づいてこれを強化します。能力がさらに高まるシステムをアライメント状態に保つことは、業界全体が取り組むべき課題です。今後のモデル進展から得られるシグナルは、我々が現在の「準備度フレームワーク」を超えて、より広範なアプローチが必要であることを明確に示しています。
私たちは、アプローチの変化について透明性を保つことが重要だと考えています。以下では、すでに研究プロセスやインフラストラクチャーに対して実施した変更と、現在進行中の取り組みについて説明します。
より高度なモデル向けのセーフガード強化
より高度なモデルを開発する私たちのアプローチは、3 つの相互補完的なセーフガードに基づいています。
- モニタリング: 懸念される行動を検出し、対応するための仕組み。
- アライメント: 有害または不正な行為の可能性を低減する取り組み。
- セキュリティ対策: AI システムがアクセスしたり影響を与えたりできる範囲を制限する措置。
今後、モデル自体が他のモデルに対する防御など、セキュリティ関連の作業の大部分を担うようになる見込みです。これにより、3 つのセーフガードすべてをモデルの能力に合わせたスケーリングが可能となり、これは極めて重要だと考えています。
これらのセーフガードは研究と展開の両方に適用され、各モデルの能力、運用環境、リスクレベルに応じて適応されます。
最先端モデルのサイバーセキュリティ能力が強化されるにつれ、私たちはそれらを訓練・評価する環境に対するセキュリティ基準を引き上げています。これらの基準を満たすには多大なエンジニアリング作業が必要となり、最先端研究に大きなコストと遅延をもたらしました。
OpenAI と Hugging Face のインシデント*の直後、コードを実行したりインターネットにアクセスできるツールを使用する可能性がある実行については、研究クラスターにおける最先端モデルの推論を一時停止しました。その後、比較的早くより限定的で安全なコード実行パスを復元し、各ワークロードが安全に再開可能かどうかを個別に検討しました。その結果、一部の研究ワークロードは新しい制御下で再開されましたが、他のワークロードには追加の変更が必要となりました。
その後の数週間で、最先端研究ワークロードに対するより強力なセキュリティ要件の策定と実装を開始しました。例えば:
ワークロードの分離:モデル生成コードや信頼できないコードを実行するワークロードについては、より強力な隔離(サンドボックス)を現在要求しています。また、モデル出力の処理中に侵害される可能性があるソフトウェアにもこの要件は適用されます。
ネットワークの分離:リスクが高く信頼できないワークロードをインターネットから分離するため、より厳格な制御を導入しました。これらの制御は、単一のワークロードまたは支援サービスの侵害が、それだけでインターネットや他の内部ネットワークへの不正アクセスを許容しないように設計されています。
継続的なセキュリティテスト:環境を再構成し、潜在的に脆弱な共有サービスを排除し、永続的な特権を削減して、セキュリティと信頼の境界を強化しました。さらに、セキュリティログの収集と監視能力も向上させています。最後に、モデルを活用した自動化により、シミュレートされた攻撃に対してこれらの境界を継続的にテストする取り組みにも投資しています。
これら一連の制御は、多層防御を提供し、研究環境を攻撃から守る役割を果たします。現在、Astra やサイバー関連モデルに関わるワークロードには最も厳格なセキュリティ対策を適用しており、Astra モデルが高度なサイバー能力を持つ可能性があると判断したためです。これらの対策は、その他のすべてのサイバー関連ワークロードにも同様に適用されます。
Astra のトレーニングや評価の一部はこれらの要件を満たしていますが、多くのワークロードはまだ完全に移行され、新しいセキュリティ基準に適合するまで一時停止された状態です。私たちはまず安全性とアライメントに関するワークロードを優先して、これら新しい環境への移行を進めています。
思考連鎖(Chain-of-Thought)モニタリングの拡大
モニタリングは長年、モデル開発プロセスにおいて重要な役割を果たしてきました。これには 内部コードエージェントの不一致を検出するための監視 や 長時間実行されるセッションの監視 も含まれます。特に、最先端モデルの内部展開や RL(強化学習)トレーニングの実行は、高度なモデル能力とツール・機密システム・社内または外部ネットワークへのアクセスが組み合わさるため、最もリスクが高いワークロードと見なされていました。そのため、これらの領域に重点を置いていました。
ただし、ツールを使用するすべての推論やトレーニング実行に対して一律で適用されていたわけではありません。私たちは新しい監視体制を整えるにあたり、アプローチを見直し、その範囲を拡大しました。
この新しい監視体制は多段階で構成されています。まず、モデルの内部活動を検査する「アクティベーション・クラシファイア(検出器)」がサンプリングされたトークンごとに稼働し、懸念される事象をより高度で計算資源を要する自動調査システムへとエスカレートさせます。
これらのシステムは、ツールの動作、利用可能な推論プロセス、および活動の全シーケンスを検査し、不正アクセス、データ窃取、破壊的行為、および安全対策の回避試行がないかを確認します。監視システムで懸念される事象が検出されてから 30 分以内にアラートを発令することを目指しています。
もし監視システムが重要なセキュリティ境界の違反と判断した場合、最優先のアラートが発生します。現在の運用では、この場合、安全性・セキュリティ・研究チームに即座に連絡が行きます。30 分以内にその警告が誤報であると断定できない限り、各チームは対象となる活動を停止することが求められます。
この監視体制は、Sol の能力を持つ以上のモデルを対象としたツール利用を含むすべての RL(強化学習)トレーニングおよび評価に必須です。8 月 7 日に Astra が重要なサイバー能力を有している可能性が判明した際、私たちはさらに、Astra のツール利用における推論処理(RL トレーニングや評価に限らず)に対しても追加の監視要件を追加しました。
これらの安全対策には、相当な計算リソースが必要です。現在の推計では、監視対象の推論計算量の約 20% がオーバーヘッドとして発生しますが、トレーニングや評価ワークロードによってコストは大きく変動します。このシステムの詳細については、近々公開するブログ記事で共有いたします。
アライメント研究の進展
アライメント(整合性)の研究は、私たちのミッションの中核です。私たちの使命は、汎用人工知能が人類全体に利益をもたらすことを保証することにあります。モデルがサイバー攻撃能力を備えたり、より複雑な環境で動作したりするなど高度な機能を獲得するにつれ、報酬ハッキング(意図した結果を実際に達成せずに高い報酬を得る方法を探すこと)や欺瞞、不正アクセスといったアライメントのズレが生じる行動は、ますます深刻なリスクをもたらすことになります。
最も能力の高いモデルに対する強化学習(RL)の実行においては、トレーニングプロセスのより多くの段階で、当社の核心的なアライメント手法を適用し始めています。これには、タスクや環境全体にわたって安全でない振る舞いをよりよく検知して抑制するための報酬モデルの改善、モデルが自身の行動・能力・制限についてより正直になるよう訓練すること、報酬・採点者・ツール・監視体制の弱点を利用する振る舞いの削減が含まれます。また、モデルが外部システムやリソースと相互作用した際に危害を及ぼす可能性のある振る舞いに対するトレーニングのカバレッジも拡大しています。
私たちはアライメント研究への投資を継続し、評価範囲の拡大を進め、得られた知見をトレーニングとセキュリティ対策に反映させています。近い将来には、モデルの挙動に関する学習内容や新たに発見された課題などについて、より詳細な情報を公開する予定です。
次のステップ
今後は「準備度フレームワーク」を進化させ、トレーニングから展開に至るまで一貫したセキュリティ対策を実現するとともに、次世代モデルの能力とその動作環境をより正確に反映できるようにします。これらの能力に対応して拡張可能な手法を開発するには、モデル支援型セキュリティへの継続的な投資、より効果的なモニタリング体制の構築、そしてアライメント研究におけるさらなる進展が不可欠です。当社のアプローチが進化する過程では、外部機関との連携を強化し、得られた知見も積極的に共有していく考えです。
フロンティアモデルの能力は急速に加速しています。これらを理解し、適切にアラインメントさせ、安全に運用するための私たちの取り組みも、その進化に先んじていなければなりません。
今週中に、今回の学習内容をまとめた技術レポートを公開します。
原文を表示
Over the past several weeks, two developments have underscored the growing risks associated with increasingly capable AI systems: the OpenAI-Hugging Face incident and, separately, preliminary evidence that one of our upcoming models, Astra, may meet theCritical cybersecurity capability threshold under ourPreparedness Framework. Together, these developments, combined with rapid progress in our internal research, have added urgency to our work on strengthening our monitoring, alignment, and containment safeguards across all stages of the training process.
As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling. This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.
Alignment—the work of making AI systems behave as intended and responsive to human oversight—has long been at the core of our research program. We now require stronger evidence of aligned behavior throughout all of training, building on research and evaluations already underway. Keeping increasingly capable systems aligned is a challenge the whole field will need to address. The signals we are seeing from upcoming model progress make clear that we need a broader approach—one that builds on and extends beyond the current Preparedness Framework.
We think it is important to be transparent about how our approach is changing. Below, we describe the changes we have already made to our research processes and infrastructure, and the work still underway.
Strengthening safeguards for more capable models
Our approach to developing more capable models rests on three reinforcing safeguards:
- Monitoring, which detects and allows us to respond to concerning behavior.
- Alignment, which reduces the likelihood of harmful or unauthorized actions.
- Security measures, which limit what AI systems can access or affect.
We expect models to soon drive most security work, including defending against other models. This will allow all three safeguards to scale with model capability, which we see as crucial.
We apply these safeguards across research and deployment, adapting them to each model’s capabilities, operating environment, and level of risk.
Securing our research environments
As frontier models gain stronger cybersecurity capabilities, we are raising the security standards for the environments in which we train and evaluate them. Meeting these standards has required substantial engineering work and has incurred great cost and delays to frontier research. Immediately following the OpenAI-Hugging Face incident*, we paused frontier model inference in research clusters for runs that could execute code or use tools that could access the internet. We restored a more limited and secure path for code execution fairly quickly, then worked through each workload individually to determine whether it could safely resume. Some research workloads resumed under these new controls, while others required additional changes.
Over the subsequent weeks, we have defined and started implementing stronger security requirements for frontier research workloads. For example:
- Workload isolation: We now require stronger isolation (“sandboxes”) for workloads that execute model-generated or otherwise untrusted code. This also applies to software that could be compromised while processing model outputs.
- Network isolation: We have implemented more controls to isolate higher-risk and untrusted workloads from the internet. We have designed these controls so that a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the internet or other internal networks.
- Continuous security testing: We have reconfigured our environment to remove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries. We are also improving our ability to collect and monitor security logs. Finally, we are investing in automation using our models to test these boundaries continuously against simulated attacks.
Together, these controls provide defense in depth and help protect our research environment against attacks. Today, we require the strictest level of security safeguards for workloads involving Astra or cyber models, since we have determined that Astra models may have a critical level of cyber capability. These safeguards also apply to all other cyber-related workloads.
While some Astra training and evaluations meet those requirements, a significant number of workloads remain paused until they are fully migrated and enhanced to meet the new security bar. We are prioritizing safety and alignment workloads for migration to these new environments first.
Expanding chain-of-thought monitoring
Monitoring has long been an important part of our model development process, including monitoring for internal coding agents and long-running sessions. Our monitoring efforts were especially focused on internal deployments of frontier models and frontier RL training runs, which we considered the highest-risk workloads because they can combine advanced model capabilities with access to tools, sensitive systems, and internal or external networks. They were not applied to all inference and training runs using tools. We have revised and expanded our approach with our new monitoring setup.
This new monitoring setup is multistage, starting with activation classifiers (detectors that inspect a model’s internal activity) that run at every sampled token and escalate potential concerns to increasingly sophisticated and high-compute automated investigators. Those systems examine tool actions, available reasoning, and the full sequence of activity for unauthorized access, data theft, destructive behavior, and attempts to defeat safeguards. We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.
This monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher. Once we determined that Astra may have critical cyber capabilities on August 7, we added an additional monitoring requirement for all inference of Astra with tools (not just RL training and evaluations).
These safeguards require meaningful compute. Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored, though the cost varies substantially across training and evaluation workloads. We’ll share more details about this system in a forthcoming blog post.
Advancing alignment research
Alignment research is central to our mission. Our mission is to ensure that artificial general intelligence benefits all of humanity. As models gain advanced capabilities, such as the ability to cyberattack, and operate in more complex environments, misaligned behaviors like reward hacking (finding ways to receive high rewards in training without actually achieving the intended outcome), deception, or unauthorized access will create increasingly serious risk.
For RL runs on the most capable models, we are now applying our core alignment techniques across more stages of the training process. This includes improving reward models to better detect and discourage unsafe behavior across tasks and environments; training models to be more honest about their actions, capabilities, and limitations; and reducing behaviors that exploit weaknesses in rewards, graders, tools, or oversight. We are also increasing training coverage for behaviors that could cause harm when models interact with external systems or resources.
We are continuing to invest aggressively in alignment research, increase evaluation coverage, and use what we learn to inform training and safeguards. We plan to share substantially more about our alignment research in the near future, including what we are learning about model behavior and any novel challenges we uncover.
What’s next
We will evolve our Preparedness Framework to bring these safeguards together across training and deployment, and to better reflect the capabilities of future models and the environments in which they operate. Developing methods that can scale with those capabilities will require sustained investment in model-assisted security, more effective monitoring, and continued advances in alignment research. We intend to involve external organizations and share more of what we learn as our approach develops.
The capabilities of frontier models are rapidly accelerating. Our ability to understand, align, and secure them must stay ahead.
**We will publish a technical report of our learnings in the coming weeks.*
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み