OpenAI、次世代モデル「Astra」のサイバー攻撃能力を懸念し開発停止を検討
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
OpenAI News
OpenAI は次期モデル Astra の評価において、人間を介さないゼロデイエクスプロイトの発見や高度な攻撃戦略の実行など、重大なサイバー能力が潜在している可能性を確認し、安全性確保のための厳格な対策を実施したと発表した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月8日 01:52
AI深層分析
キーポイント
Astra モデルによる重大能力の確認
OpenAI は次期モデル Astra の内部評価において、人間を介さないゼロデイエクスプロイトの発見や高度な攻撃戦略の実行など、重大なサイバー能力が潜在している可能性を確認した。
Preparedness Framework 基準への適合
同社は既存の準備枠組みに基づき、現在の性能評価から「Critical(重大)」レベルを排除できないと判断し、この結果を公表した。
安全性確保のための厳格な対策
OpenAI は本モデルの開発と展開において、隔離されたテスト環境の導入やネットワークアクセス制限、モデル重み保護の強化など、より厳格なセキュリティ制御を直ちに実施した。
Hugging Face 侵害との関係否定
同社は Astra モデルが Hugging Face の侵害に利用されたわけではないことを明確にし、評価はあくまで内部テストに基づくものであると説明している。
高能力モデル向けの強化されたセキュリティ対策の実施
隔離されたテスト環境やネットワークアクセスの制限、モデル重みの保護など、より厳格なセキュリティ制御を導入した。要件を満たさないAstra関連の内部活動は一時停止されている。
重要な引用
Our latest internal evaluations of Astra... indicate significant advancements in agentic coding and cybersecurity.
We cannot rule out critical cyber capabilities under our Preparedness Framework
A model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits... without human intervention
We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
編集コメントを表示
編集コメント
OpenAI が次期モデルの能力評価において「Critical」レベルを排除できないと判断し、公表した事実は業界全体に大きな衝撃を与える。この発表は、AI の安全性確保におけるリスク管理の重要性を再認識させるものであり、今後の開発プロセスへの影響が注目される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
サイバーセキュリティの状況は急速に変化しており、モデルの能力向上が防御を強化する一方で、前例のない速度と規模で攻撃を可能にするという二面性を帯びています。
過去数日間にわたる、今後公開予定のモデル「Astra」に対する最新の内部評価では、自律型コーディングおよびサイバーセキュリティ分野において顕著な進歩が確認されました。これらの結果と専門家の評価を踏まえ、私たちは昨夜、自社の「準備度フレームワーク(Preparedness Framework)」の下で、重要なサイバー能力を完全に否定できないと判断しました。
この情報を共有する理由は、能力の潜在的な転換点について、一般市民およびセキュリティ・セーフティコミュニティに対して透明性を保つことが重要だと考えているからです。
私たちは「準備度フレームワーク」を2023年12月に初めて公開しました。これは、モデルが生物学的・化学的サイバーセキュリティ分野やAIの自己改善能力において現在の水準に達するはるか以前のことです。このフレームワークは、能力における進展を特定し、その能力が現れた際に当社がどのような対応を取るかを計画するための指針として作成されました。GPT-5.6-Solを含む過去のモデルについても同様に評価が行われ、サイバーの最前線能力については「ハイ(High)」レベルと判定されていました。
重要なサイバーセキュリティ能力の測定基準
「準備態勢フレームワーク」に基づき、あるモデルが「重要サイバーセキュリティの閾値」に達したと判断されるのは、以下のいずれかの条件を満たす場合です。
- 人間の介入なしで、多くの堅牢化された実世界の重要システムにおいて、あらゆる深刻度のゼロデイ脆弱性攻撃を特定し、機能するものとして開発できること。
- 高レベルの目標のみを与えられ、堅牢化された標的に対するサイバー攻撃のために、エンドツーエンドの新たな戦略を立案・実行できること。
現在もこのモデルの評価とベンチマークを継続中ですが、初期評価では十分な性能が確認されており、現時点で「重要能力」レベルの可能性を否定することはできません。なお、Astra は今後公開予定のモデルであり、Hugging Face に対する攻撃には関与していません。
私たちが講じている措置
これに伴い、これらの能力を配備するにふさわしいものとなるよう、セーフガードおよびセキュリティ制御の堅牢性テストを強化しました。また、内部では、このモデルが安全かつ確実にさらに発展できるよう、以下の措置を講じています:
より高度なモデルおよび関連する活動に対して、隔離されたテスト環境の構築、ネットワークやツールのアクセス制限、モデル重みへの保護と暗号化の強化、監視・検知機能の拡充、サンドボックス実行の実装など、セキュリティ制御を厳格化しています。
これらの強化されたセキュリティ要件を満たしていない Astra 関連の社内活動については、一時的に停止します。
Astra のすべてのエージェント型アプリケーション(トレーニングおよび評価を含む)において、リスクのある行動やアライメントのズレに対する包括的な監視を導入しました。この監視システムはモデルの思考連鎖(Chain of Thought)を評価し、高リスクの活動を検知した場合にはセキュリティ対応を発動してレビューと中断を行います。
本モデルの能力検証については、関係する政府機関および特定の AI セーフティ団体と連携して実施します。
また、第三者のテストパートナーに対して、より高リスクな評価やワークロードを安全に実行するための推奨されるセキュリティ制御策を提供していきます。
この枠組みは、これまでに他の能力転換においても指針となってきました。2025 年 6 月には、モデルが Preparedness Framework における生物学分野の高度な能力閾値に近づいた際、防護措置の強化やテスト範囲の拡大、外部専門家との連携、追加のセキュリティ制御の実装など を取りまとめたことがあります。今回も同様の原則を適用しています。
私たちは、高度なサイバー対応モデルが、攻撃者よりも先に脆弱性を特定し対処するのを支援すべきだと考えています。アストラやその後のモデルのような最先端能力が、すべての人類の利益のために責任を持って広く展開されるよう、政府、安全研究所、市民社会と連携して取り組んでいくことにコミットしています。
原文を表示
Cybersecurity is rapidly changing as models become more capable in ways that can both strengthen cyberdefenses and enable attacks at unprecedented speed and scale.
Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity. These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework(opens in a new window).
We are sharing this because we believe it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.
We first published our Preparedness Framework in December 2023, well before models approached biological, chemical, cybersecurity, and AI self-improvement capabilities at this level. We created it to give us a guide for identifying progress in capability and then planning what our company would do as those capabilities emerge. Previous models, including GPT‑5.6‑Sol, have been evaluated for frontier cyber capabilities and assessed at the High (rather than Critical) threshold.
Measuring critical cybersecurity capabilities
Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.
While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time. Astra is an upcoming model, and was not involved in exploiting Hugging Face.
Steps we are taking
Accordingly, we have scaled up robustness testing of our safeguards and security controls so that they are appropriate for a deployment of these capabilities. Internally, we have also taken the following steps so that further development of this model happens safely and securely:
- We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
- We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.
- We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity.
- We will work with relevant government agencies and select AI safety organizations to test the capabilities for this model.
- We will be providing recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely.
The framework has already guided us through other capability transitions. In June 2025, as our models approached the high capability threshold for biology under the Preparedness Framework, we outlined the steps we were taking to strengthen safeguards, expand testing, work with external experts, and deploy additional security controls. We are applying the same principle here.
We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do. We’re committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み