OpenAI、セキュリティリスクを理由に次期モデル「Astra」の開発を一時停止
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
OpenAI は新モデル「Astra」の内部評価において、人間を介さずに高度なゼロデイ脆弱性を発見・開発する能力が確認されたため、安全性確保のため開発を一時的に停止し、厳格なセキュリティ対策を講じたと発表した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 22:21
AI深層分析
キーポイント
Astra の危険な能力評価
OpenAI は直近の評価で、Astra が複数の深刻度レベルのゼロデイ脆弱性を人間を介さずに特定・開発する能力を持つ可能性があり、重大なサイバーセキュリティリスクと判断した。
開発停止と対策強化
同社は「Critical」レベルの能力を排除できないとして開発を一時停止し、堅牢なテスト環境やアクセス制限など、より厳格なセキュリティ制御を実装している。
評価基準の再定義
今回の判断は、2023 年 12 月に策定された「Preparedness Framework」に基づき、生物・化学・サイバー分野での自律的な攻撃能力を厳格に測定した結果である。
ハッキング事案との区別
OpenAI は、Astra が Hugging Face に対する実際の攻撃や悪用に関与したことはないと明確に声明し、今回の措置は将来のリスク予防のためのものだと強調している。
Astra開発の安全対策と一時停止
OpenAIは高能力モデル向けに隔離環境や暗号化など厳格なセキュリティ制御を導入し、要件を満たさないAstra関連活動を一時的に停止した。
重要な引用
Our latest internal evaluations of Astra... indicate significant advancements in agentic coding and cybersecurity.
We cannot rule out critical cyber capabilities under our Preparedness Framework.
A model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits... without human intervention.
We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
編集コメントを表示
編集コメント
OpenAI が自社のモデルに対して最も厳しい評価基準を適用し、開発を停止した事実は、自律型 AI の安全性確保がいかに喫緊の課題かを如実に示している。業界全体として、能力の向上とリスク管理のバランスを取るためのガバナンス体制の構築が急務となっている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
サイバーセキュリティの状況は急速に変化しており、モデルの能力向上がサイバー防御を強化する一方で、前例のない速度と規模で攻撃を可能にするという両面を持っています。
今後公開予定のモデル「Astra」に対する最新の社内評価(過去数日間のもの)では、エージェント型コーディングおよびサイバーセキュリティ分野において顕著な進歩が確認されました。これらの結果と専門家の評価を踏まえ、私たちは昨夜、当社の 準備状況フレームワーク の基準において、重大なサイバー能力を排除できないと結論付けました。
私たちはこの判断を公表します。これは、能力の潜在的な変化について一般市民およびセキュリティ・セーフティコミュニティに対して透明性を保つことが重要だと考えているためです。
当社の「準備状況フレームワーク」は、モデルが生物学的・化学的サイバーセキュリティや自己改善型 AI のような高度な能力に到達する以前、2023 年 12 月に初めて公開しました。これは、能力の進展を特定し、その能力が現れた際に当社が取るべき行動を計画するための指針として作成されたものです。GPT‑5.6‑Sol を含む過去のモデルについても同様に評価が行われ、最前線のサイバー能力については「高(High)」レベルと判定されました。
重大なサイバーセキュリティ能力の測定基準
「準備態勢フレームワーク」に基づき、モデルが人間の介入なしに多くの堅牢化された実世界の重要システムにおいてあらゆる深刻度のゼロデイ脆弱性を特定・開発できる場合、あるいは高レベルの目標のみを与えられて堅牢な標的に対するサイバー攻撃のための新規戦略を立案・実行できる場合は、「クリティカル(重大)」セキュリティリスク閾値に達したと判断します。
現在も同モデルのベンチマーク評価を継続中ですが、予備的な評価結果から、現時点で「クリティカル」な能力レベルを否定することはできません。なお、Astra は今後公開予定のモデルであり、Hugging Face への攻撃には関与していません。
私たちが講じている対策
これに伴い、これらの機能を展開するにふさわしいよう、セーフガードおよびセキュリティ制御の堅牢性テストを強化しました。また、内部では同モデルが安全かつ確実に開発されるよう、以下の措置を講じています:
より高能力なモデルおよび関連する活動に対して、隔離されたテスト環境の構築、ネットワークやツールへのアクセス制限、モデル重み保護と暗号化の強化、監視・検知機能の拡充、サンドボックス実行の実施など、セキュリティ制御を厳格化しています。
これらの強化されたセキュリティ基準を満たしていない Astra に関わる社内活動については、一時的に停止します。
Astra のすべてのエージェント型アプリケーション(トレーニングおよび評価を含む)において、危険な行動やアライメントのズレに対する包括的な監視を導入しました。この監視システムはモデルの思考連鎖(Chain of Thought)を評価し、リスクの高い活動を検知した際にはセキュリティ対応としてレビューと中断をトリガーします。
政府関係機関や特定の AI セーフティ団体と連携し、本モデルの能力テストを実施していきます。
また、第三者のテストパートナーに対しては、より高リスクな評価やワークロードを安全に実行するための推奨されるセキュリティ制御策を提供する予定です。
このフレームワークは、これまでに他の能力転換においても指針となってきました。2025 年 6 月には、モデルが Preparedness Framework に基づく生物学分野の高能力閾値に近づいた際、 safeguards の強化、テストの拡大、外部専門家との連携、追加のセキュリティ制御の実装など を実施した手順を明らかにしました。今回も同様の原則を適用しています。
私たちは、高度なサイバー能力を持つモデルが、攻撃者よりも先に脆弱性を特定し、対応する防御者を支援すべきだと考えています。アストラやその後のモデルのような最先端の能力が、すべての人類の利益のために責任を持って広く展開されるよう、政府、安全研究所、市民社会と連携して取り組んでいくことを約束します。
原文を表示
Cybersecurity is rapidly changing as models become more capable in ways that can both strengthen cyberdefenses and enable attacks at unprecedented speed and scale.
Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity. These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework(opens in a new window).
We are sharing this because we believe it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.
We first published our Preparedness Framework in December 2023, well before models approached biological, chemical, cybersecurity, and AI self-improvement capabilities at this level. We created it to give us a guide for identifying progress in capability and then planning what our company would do as those capabilities emerge. Previous models, including GPT‑5.6‑Sol, have been evaluated for frontier cyber capabilities and assessed at the High (rather than Critical) threshold.
Measuring critical cybersecurity capabilities
Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.
While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time. Astra is an upcoming model, and was not involved in exploiting Hugging Face.
Steps we are taking
Accordingly, we have scaled up robustness testing of our safeguards and security controls so that they are appropriate for a deployment of these capabilities. Internally, we have also taken the following steps so that further development of this model happens safely and securely:
- We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
- We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.
- We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity.
- We will work with relevant government agencies and select AI safety organizations to test the capabilities for this model.
- We will be providing recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely.
The framework has already guided us through other capability transitions. In June 2025, as our models approached the high capability threshold for biology under the Preparedness Framework, we outlined the steps we were taking to strengthen safeguards, expand testing, work with external experts, and deploy additional security controls. We are applying the same principle here.
We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do. We’re committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity.
同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み