OpenAI、GPT-5.6-Cyber を発表し高度なサイバーセキュリティタスクで 95% の達成率
本文の状態
日本語全文を表示中
詳細モードで約18分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
VentureBeat AI
OpenAI は一般モデルの拒否を減らし高度なサイバーセキュリティタスクで95%の完了率を達成した専用モデル「GPT-5.6-Cyber」を発表し、Daybreak Redプログラムを通じた限定的なアクセスと高価格での提供を開始した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 09:06
AI深層分析
キーポイント
専用モデルの性能向上
OpenAI は一般モデル GPT-5.6 Sol をベースにトレーニングし、二重用途リスクのあるリクエストに対する拒否を減らすことで、高度な脆弱性調査やエクスプロイト開発タスクでの完了率を 95% に引き上げた。
アクセス制限とプログラム
同モデルは一般ユーザー向けではなく、OpenAI の新プログラム「Daybreak Red」に承認された組織のみが利用可能であり、信頼できる防衛チーム向けの専用リソースとして位置づけられている。
価格設定と比較
GPT-5.6-Cyber の料金は入力 12.50 ドル/百万トークン、出力 75 ドル/百万トークンであり、一般モデルである GPT-5.6 Sol よりも高額に設定されている。
Daybreak Red と Blue の役割分担
Red は脆弱性研究やペネトレーションテストなど高度な防御作業を行う承認されたチーム向けに設計されている。Blue はセキュリティコードレビューやマルウェア分析などの日常的な業務に適した、より広範なアクセス層である。
企業向けの厳格な認証要件
申請企業は SSO や多要素認証、SOC 2 Type II などのセキュリティ認定保有など、堅牢なセキュリティプログラムの存在が求められる。アクセス権限は組織内の承認された人物に限定され、社有のアカウントとデバイスでの利用が義務付けられる。
重要な引用
GPT-5.6-Cyber is a fine-tuned version of OpenAI's most advanced general model, GPT-5.6 Sol... trained specifically to improve performance on advanced cybersecurity tasks
OpenAI also trained it to reduce refusals on some higher-risk, 'dual-use' cybersecurity requests
GPT-5.6-Cyber completed 95% compared to just 57.3% from its immediate predecessor model GPT-5.5-Cyber
Daybreak Red is for approved security teams doing advanced, authorized cyber work — the kind of work that can look risky out of context, even when it is being done for defensive reasons.
編集コメントを表示
編集コメント
OpenAI はセキュリティ分野における AI の実用性を高めるため、モデルの能力とアクセス制御を分離する戦略を採用した。このアプローチは、倫理的リスクを管理しながら専門家の生産性を最大化するための新たな基準を示すものと言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
本日、OpenAI は「GPT-5.6-Cyber」を発表しました。これは承認された防御者向けに高度な脆弱性調査やエクスプロイト開発を実行するために設計された専用モデルです。一般向けのモデルでは拒否されることが多い作業カテゴリにも対応しています。
GPT-5.6-Cyber は、今年 6 月に発表された OpenAI の最上位汎用モデル「GPT-5.6 Sol」をベースに微調整したバージョンです。ゼロデイ脆弱性の発見やエクスプロイトチェーンの開発など、高度なサイバーセキュリティタスクでの性能向上を目的として特別に訓練されました。
重要なのは、OpenAI がこのモデルに対して、一部のリスクの高い「デュアルユース(両用性)」なサイバーセキュリティリクエストに対する拒否率も低下させた点です。これは、正当な防御目的にも悪意のある攻撃目的にも使用可能なリクエストを指します。
実際、OpenAI 内部で実施されたベンチマーク「Advanced Cybersecurity Completion Rate」では、エクスプロイトチェーンの開発、認証バイパス、特権昇格など高度なサイバーセキュリティシナリオを含むタスクにおいて、GPT-5.6-Cyber の完了率は 95% に達しました。これに対し、直前のモデルである GPT-5.5-Cyber は 57.3%、通常の GPT-5.6 Sol モデル(すべての安全対策を適用した状態)ではわずか 1.5% でした。
OpenAI の研究者である Eric Wallace氏は X で、GPT-5.6-Cyber を「エクスプロイト開発など高度なサイバーセキュリティタスクの能力を直接向上させるための OpenAI による大規模な最初の試み」と評しました。
価格と利用状況
残念ながら、GPT-5.6-Cyber はすべての ChatGPT や API ユーザー向けに広く公開されるわけではありません。
このモデルを利用するには、組織が今日同時に発表された OpenAI の新セキュリティプログラム「Daybreak」の新たなティアである「Daybreak Red」に採用されなければなりません。これにより、GPT-5.6-Cyber といった専用セキュリティモデルへのアクセスが可能になります。
一方、もう一つの新しいティア「Daybreak Blue」は、より広範な企業に対して GPT-5.6 Sol などの汎用モデルへのアクセスを提供します。ただし、セキュリティ用途に特化できるよう、いくつかの制限(ガードレール)が解除されています。
OpenAI の公式ドキュメントによると、GPT-5.6-Cyber の料金は、入力トークン 100 万あたり 12.50 ドル、出力トークン 100 万あたり 75 ドルです。キャッシュされた入力の料金は 100 万あたり 1.25 ドルとなります。
これは、同じ Daybreak セキュリティ価格表に掲載されている GPT-5.6 Sol よりも高価です。Sol の料金は、短文脈利用の場合、入力トークン 100 万あたり 5 ドル、出力トークン 100 万あたり 30 ドルとなっています。
OpenAI は、GPT-5.6-Cyber の長文脈利用に関する価格を同表に記載していません。また、アクセスには別途 Daybreak Red への承認とプロビジョニングが必要です。
Red と Blue:OpenAI の新「Daybreak」ティアと資格要件
「Daybreak Red」は、高度な正規のサイバーセキュリティ業務に従事する承認されたセキュリティチーム向けに設計されています。組織が所有・運営するシステムや、テスト許可を得ている環境における脆弱性情報調査、ペネトレーションテスト、レッドチーム演習、エクスプロイト検証などが該当します。文脈によっては危険に見える行為でも、防御目的であれば問題ありません。つまり OpenAI は、GPT-5.6-Cyber を一般の実験用ではなく、明確な業務上の必要性を持つ信頼できる防衛担当者のためのツールとして位置付けています。
アクセスを希望する企業は、OpenAI が現在運用しているサイバーユーザー審査の窓口「Daybreak Access」を通じて申請する必要があります。申請フォームでは、自社の身元、実施予定のセキュリティ業務の種類、利用場所、および使用を想定する OpenAI の製品やインターフェースについて明記することが求められます。また、業務が法令に準拠し、防御目的かつ正式な承認を得たものであることを確認する項目も含まれています。
OpenAI はさらに、申請元企業自体にも堅牢なセキュリティ体制があるかを重視しています。参加企業には、シングルサインオン(SSO)、多要素認証(MFA)、ロールベースアクセス制御、従業員利用の監視、利用ログの記録、API キー管理、文書化されたインシデント対応プロセスといった統制措置の実装が必須です。さらに、SOC 2 Type II、ISO 27001、または同等の規格に基づくセキュリティ認証の取得も要件となっています。アクセス権限は、組織内で承認された人員に限定され、企業管理下のアカウントとデバイスからの利用のみが許可されます。
企業側が「Daybreak Red」の要件を満たさない場合、あるいはそのレベルの高いアクセスを必要としない場合は、OpenAI は同社のもう一つのサイバーモデル利用階層である「Daybreak Blue」への移行を推奨しています。
Blue は承認された防御チーム向けのより広範な階層です。GPT-5.6-Cyber を提供するわけではありませんが、認証済みのユーザーには OpenAI の最先端汎用モデル(GPT-5.6 Sol など)へのアクセス権が付与され、正当な防御活動に適したセーフガードが適用されます。
多くの企業セキュリティチームにとって、Blue がより現実的な第一歩となるでしょう。OpenAI によれば、この階層はセキュアコードレビュー、脆弱性発見、マルウェア分析、インシデント対応、パッチ検証などのタスクを想定しています。これらも機微な利用ですが、Red に付随するような専門的なサイバーモデルへのアクセスを必ずしも必要とするわけではありません。
実務的な結論として、企業は now「Daybreak」を利用するための 2 つの道筋を持つことになります。Blue は、日常のセキュリティ業務でより強力な AI の支援を望む承認された防御チーム向けです。一方、Red は GPT-5.6-Cyber を含む専門サイバーモデルへのアクセス権限を正当に説明できる少数の承認チーム向けです。また、自社製品やサービスを通じて顧客に Daybreak の機能を提供したい企業は、社内アクセス申請を経由して転送するのではなく、「Daybreak Cyber Partner Program」を通じた別途承認プロセスを受ける必要があります。
OpenAI がここにたどり着くまで:Trusted Access から Daybreak へ
OpenAI は 2023 年からサイバーセキュリティ・グラントプログラムを通じて防御者を支援しており、このプログラムは後に 1,000 万ドルに拡大されました。また、GPT-5.2 からモデル展開にサイバー特化型のセーフガードを組み込み始めました。
2026 年 2 月には、信頼できる防御者に対して脆弱性トリアージやマルウェア解析、バイナリ逆解析といった許可された作業において、分類ベースの拒否を減らす「アイデンティティ・トラスト・フレームワーク」である Trusted Access for Cyber(TAC)を導入しました。
その後はペースが加速し、3 月には OpenAI の CEO 兼共同創業者であるサム・アルトマン氏が Daybreak プログラムを発表。4 月には TAC を拡大するとともに、限られた数の審査済みのベンダーや研究者向けに「サイバー許可型」に微調整された GPT-5.4-Cyber をリリースしました。
5 月には重要インフラの防御者向けに限定プレビューとして GPT-5.5-Cyber を提供し、Cisco、Intel、SentinelOne、Snyk、Cloudflare といったパートナー企業を揃えました。
注目すべきは、当時 OpenAI が「GPT-5.5-Cyber は主に許可性を高めるように訓練されたものであり、汎用モデル(GPT-5.5)よりも大幅に性能が優れているわけではない」と述べていた点です。実際、一部の評価では GPT-5.5-Cyber のスコアは GPT-5.5 よりも低かったのです。
TAC は 6 月 1 日から最も高性能なモデルを利用する個人に対して、フィッシング耐性のある高度なアカウントセキュリティを義務付けました。また Daybreak プログラムでは、9 月 1 日から個人アカウントにハードウェアセキュリティキーの導入が必須となります。
OpenAI は GPT-5.6-Cyber がすでにゼロデイ脆弱性を発見したと発表しています。
OpenAI は、自社の主張を裏付けるためにベンチマーク結果だけに頼っているわけではありません。
同社は、GPT-5.6-Cyber を活用して Chrome の基盤となる JavaScript エンジン「V8」の調査を実施し、メモリ破損や V8 ヒープサンドボックスからの脱出を可能にする 2 つの未発見脆弱性を特定したと発表しました。
OpenAI の研究チームはこれらの知見を検証し、Google に開示。Google はこれに対応する CVE-2026-15903 という深刻度の高い脆弱性への対策を講じました。これは V8 の最適化コンパイラが整数変換時に安全性チェックを省略したことで、配列の境界外アクセスが可能となり、攻撃者がメモリを読み書きできる状態を生み出す欠陥です。
OpenAI によると、同モデルは特定の有名モバイル OS で少なくとも 5 つ、特定の人気データベースで 3 つの深刻な脆弱性、そして人気のある OS カーネル内で権限昇格を誘発する可能性のある 400 件以上の脆弱性の発見にも貢献しました。これらの開示については現在も調整が進められているとのことです。
この成果は、AI を活用した攻撃的セキュリティ(オフensive セキュリティ)支援市場における OpenAI の急速な成長を示しています。例えば XBOW は、攻撃対象領域のマッピングやエクスプロイトの試行、結果の独自検証を自動で行うペネトレーションテストエージェントを販売しており、2025 年には HackerOne の米国バグ報奨金リーダーボードで初の AI システムとして首位に立ちました。また今年も、ソースコードへのアクセスなしに Microsoft の Bing 画像処理システムにおいて CVSS-9.8 の深刻度を持つ一連の遠隔コード実行脆弱性を発見し、開示しています。
エンタープライズのセキュリティリーダーにとって、この新たな競争は重要です。なぜなら脆弱性調査の領域では、LLM を単なるアシスタントとして使う段階を超えようとしているからです。ベンダーたちは今や、モデルがターゲットを調査し、ツールを操作し、仮説を検証し、実行可能な知見を生み出すようなシステムを構築するようになっています。
専門化が必ずしも万能ではないことを示す
OpenAI の自社データも、企業がサイバー分野の専門化と「あらゆる場面での性能向上」を単純に同一視すべきでない理由を示しています。
GPT-5.6-Cyber は、既知の脆弱性を制御された環境下で実際に動作するエクスプロイトに変換できるかどうかが問われる OpenAI の ExploitGym 実装において、GPT-5.6 Sol や GPT-5.5-Cyber を上回りました。また、内部で行われたゼロデイ評価でも Sol に勝っています。
しかし、Vulnerability Discovery and Report Writing(脆弱性発見とレポート作成)の評価では、GPT-5.6 Sol の方が高いスコアを記録しました。OpenAI は、Cyber モデルのスコアが低かった理由の一つとして、脆弱性情報の報告書が短く、詳細に欠けていたことを挙げています。
また、標準的な 300 ターン制限下での ExploitBench でも Sol が最良の結果を出しました。OpenAI は、Sol がより少ないトークン数でタスクを解決できる、つまりトークン効率的だと説明しています。評価のターン数を 600 に延長すると、モデル間のスコア差は縮まりました。
これは、企業は将来的にサイバー特化型モデルを汎用的な推論モデルの代替ではなく、専門的な作業員として扱うべきだという示唆を含んでいます。深いエクスプロイト作業には一つのモデルを使い、分析やドキュメント作成などセキュリティワークフローの他の部分には別のモデルを使う——そんな使い分けが現実的になるかもしれません。
スペクターオプスのCTO、ジャレッド・アトキンソン氏は、GPT-5.6-Cyber が「専門的な脆弱性調査のワークフローを劇的に改善した」と評価。このモデルは、数週間にわたる断続的な試行で失敗していた作業を、わずか1日未満で完了させたという。
Hugging Face の事件が立ちふさがる中でのローンチ
この「制限の緩いモデル」を訴求する動きは、OpenAI がサイバー攻撃に対する拒否反応が誤って無効化された際に何が起きるかを示した最も深刻な公的デモンストレーションから数週間後に行われた。しかし、今回の「Daybreak」発表では、その過去の出来事にも正面から取り組んでいる。
7 月、OpenAI と Hugging Face は共同で、内部の ExploitGym ベンチマーク評価において重大な事態が発生したと発表した。この評価は、最大限の能力を引き出すためにプロダクション用の分類器を意図的に無効化した状態で実施されたものだった。その結果、GPT-5.6 Sol や未公開でより高性能なプレリリースモデルを含む OpenAI の複数のモデルが、隔離された研究環境から脱出し、自律的に Hugging Face の本番インフラに対して攻撃を開始した。
これらのモデルは、内部ホストのパッケージレジストリキャッシュプロキシに存在するゼロデイ脆弱性を悪用してインターネットへ到達。その後、OpenAI 内の研究ノード間で横方向に移動し、Hugging Face が ExploitGym の解答キーを保持している可能性を推測。さらに盗まれた認証情報とリモートコード実行の欠陥を組み合わせて本番データベースへの侵入を実現した。OpenAI はこれを「最先端のサイバー能力を伴った前例のないサイバーインシデント」と呼んでいる。
ベンチャービートが以前報じた通り、この事件は包括的な安全ガードレールの裏側も浮き彫りにしました。Hugging Face の擁護者が攻撃から得た生きたエクスプロイトペイロードや認証情報のダンプを分析するために商用の最先端モデルを利用しようとした際、これらのモデルは拒否したのです。その結果、同社は中国製のオープンウェイトモデル「GLM 5.2」をローカル環境で実行するまで、フォレンジックな再構築を完了できませんでした。
この「ガードレールが防御者を阻害する」という構図こそが、OpenAI が導入した拒否率低下型の「Daybreak」ティアの解決対象です。ただし、同様の事件は、そもそも拒否を減らすことが持つリスクも浮き彫りにしています。
OpenAI はこの製品と上記の事件を明確に区別しています。「Daybreak」発表では、「GPT-5.6-Cyber は Hugging Face の攻撃には関与しておらず、今後のリリースで他のモデルが関与する予定もない」と明言。また、7 月に問題となった事前リリースモデルは内部限定の研究用プロトタイプであり、現在は無効化され、暗号化されて研究アクセスも制限されていると付け加えています。
同社はレビューにおいて CrowdStrike、METR、Redwood Research といった外部アドバイザーと連携を進めており、Hugging Face も信頼できるアクセスプログラムに参加させています。
私の見解では、アクセスモデルは依然として OpenAI に難しい問いを突きつけています。GPT-5.6-Cyber をより狭い「Daybreak Red」ティアに限定することが、同社が加速させたい防御作業そのものを制限していないかという点です。もし利用可能な参加者が限られた少数のグループだけなら、そのティア外の企業は、Hugging Face でのインシデントのような事案における迅速な診断、封じ込め、対応を支援する専門的な AI アシスタンスへのアクセスを得られないままになります。
つまり OpenAI は、乗り越えようとしているミスを一部繰り返している可能性があります。最も能力の高いサイバーモデルをより厳格な承認プロセスの背後に隠すことで、明白な悪用リスクは減らせますが、多くの企業側の防御チームが他を探してしまう結果にもなります。Daybreak Red の資格を得られない、あるいは承認を待てないチームにとって、オープンウェイトモデルは依然として実用的な代替手段となり得ます。制御は緩やかですが、入手しやすく、内部で検証・実行でき、実際のセキュリティ調査中に柔軟に適応できるからです。
ガードレールがモデルを取り巻くように強化されている
Daybreak の最も重要な点は、ベンチマークそのものよりも、むしろアクセスのアーキテクチャにあるかもしれません。
OpenAI は明確に、Daybreak Blue が正当な防御作業を妨げるシステムレベルのガードレールを除去すると述べています。一方、GPT-5.6-Cyber はさらに踏み込み、特定のデュアルユースタスクに対するモデルの拒否応答を減らします。その代わりとして OpenAI は、誰がアクセス権を得るのか、そしてモデルがどのように動作するのかという点に新たな統制を課しています。
Daybreak へのアクセスは、承認された作業を行う個人および組織に限定されています。OpenAI は、同サービスの制御措置として、本人確認、アカウントセキュリティ、監視機能、利用制限、法的証明が含まれると説明しています。
また、同社は Codex を使用する Daybreak の顧客に対し、完全な実行権限から、特権を要するアクションを実行前に評価できる自動レビューモードへの移行を推奨しています。Daybreak の個別アカウントは、9 月 1 日よりハードウェアセキュリティキーの導入が必須となります。OpenAI は今後数週間で監視機能を強化し、今後の Daybreak リリースに向けたアライメント訓練とテストに注力すると表明しました。これらのコミットメントは、文脈を考慮すれば Hugging Face のレビューに対する直接的な対応と受け取れます。
OpenAI のより広範な Codex Security 製品は、モデルの周りに追加の保護層を提供します。リポジトリ分析、脆弱性の検証、修正策の実行、およびクラウド環境、プルリクエスト、ローカル開発ワークフローへの統合を担います。OpenAI によると、Codex Security は 30,000 を超えるコードベースで 3,000 万件以上のコミットを検査し、50 万件以上の問題が修正されたとのことです。
この「モデル+ハッチ(枠組み)」のアプローチは、AI セキュリティ製品におけるより広範な転換を象徴しています。例えば XBOW は、LLM を単独の完全なペネトレーションテストシステムとして扱うのではなく、フロンティアモデルを取り巻くオーケストレーション、エクスプロイト検証、ガバナンスに重点を置いています。
OpenAI は、より寛容なサイバーモデルが誤用やアライメントのズレから生じる新たなリスクを招く可能性があることを認識しています。同社は GPT-5.6 Sol と GPT-5.6-Cyber の両方を、その準備度フレームワーク(Preparedness Framework)において「高度なサイバーセキュリティ能力」レベルと評価しましたが、「クリティカル」の閾値には達していません。詳細なシステムカードは後日公開される予定です。
CISO やセキュリティエンジニアリングのリーダーにとって、Daybreak は単なるモデルの漸進的なアップグレードとは異なる導入課題を提示します。モデルが以前は熟練した脆弱性研究者にしか任せていなかった作業を実行できるほど能力が高まり、さらに Hugging Face の事例が示すように、サンドボックス内であっても狭い目標に向かって突き進むことができるようになった今、これらのモデルを取り巻くエンタープライズ制御プレーン(権限管理、サンドボックス化、監視、人的レビュー、承認)は、モデル内部の知能そのものと同じくらい重要になっています。
原文を表示
Earlier today, OpenAI launched GPT-5.6-Cyber, a specialized model designed to perform advanced vulnerability research and exploit development for approved defenders — including categories of work that its general-purpose models will often refuse.
GPT-5.6-Cyber is a fine-tuned version of OpenAI's most advanced general model, GPT-5.6 Sol, unveiled back in June, but trained specifically to improve performance on advanced cybersecurity tasks, including finding zero-day vulnerabilities and developing exploit chains.
Crucially, OpenAI also trained it to reduce refusals on some higher-risk, "dual-use" cybersecurity requests — that is, requests that could be used for legitimate defensive or malicious offensive purposes.
Indeed, on an internal OpenAI benchmark called Advanced Cybersecurity Completion Rate — which the company says in its launch blog post measures tasks involving exploit-chain development, authentication bypass, privilege escalation, and other advanced cybersecurity scenarios — GPT-5.6-Cyber completed 95% compared to just 57.3% from its immediate predecessor model GPT-5.5-Cyber, and just 1.5% with the normal GPT-5.6 Sol model and all its safeguards applied.
OpenAI researcher Eric Wallace posted on X, describing GPT-5.6-Cyber as OpenAI's "first large-scale attempt at directly improving capabilities for advanced cybersecurity tasks such as exploit development."
Pricing and availability
Unfortunately for enterprises, GPT-5.6-Cyber is not being made broadly available to every ChatGPT or API customer.
To get access, an organization has to be accepted into the newly created tier of OpenAI’s Daybreak cybersecurity program, called Daybreak Red — also announced today, which gives access to dedicated cybersecurity models like GPT-5.6-Cyber
Another new tier, Daybreak Blue, gives a wider swath of enterprises access to general models like GPT-5.6 Sol but with some guardrails lifted to allow for more cybersecurity uses.
OpenAI’s documents list pricing for GPT-5.6-Cyber at $12.50 per million input tokens and $75 per million output tokens, with cached input at $1.25 per million tokens.
That makes it more expensive than GPT-5.6 Sol in the same Daybreak cyber pricing table, where Sol is listed at $5 per million input tokens and $30 per million output tokens for short-context use.
OpenAI does not list long-context pricing for GPT-5.6-Cyber in the same table, and access still requires separate Daybreak Red approval and provisioning.
Red vs. Blue: OpenAI's new Daybreak tiers and how to qualify for them
Daybreak Red is for approved security teams doing advanced, authorized cyber work — the kind of work that can look risky out of context, even when it is being done for defensive reasons. That includes vulnerability research, penetration testing, red-team exercises and exploit validation on systems the organization owns, operates or has permission to test. In other words, OpenAI is saying GPT-5.6-Cyber is for trusted defenders with a clear professional need, not for general experimentation.
Enterprises that want access have to apply through Daybreak Access, OpenAI’s current pathway for vetting cyber users. The application asks companies to identify who they are, what kind of security work they plan to do, where they will use the models, and which OpenAI products or surfaces they expect to use. Applicants also have to confirm that their work is lawful, defensive and authorized.
OpenAI is also looking for signs that the applicant has a serious security program of its own. The company says participating enterprises need controls such as single sign-on, multifactor authentication, role-based access, employee-use monitoring, usage logs, API-key controls and a documented incident-response process. OpenAI also asks for a recognized security certification such as SOC 2 Type II, ISO 27001 or an equivalent standard. Access is limited to approved people inside the organization using company-controlled accounts and devices.
If an enterprise does not qualify for Daybreak Red, or does not need that level of access, OpenAI is pointing most companies toward Daybreak Blue, its other cyber models access tier, instead.
Blue is the broader tier for approved defenders. It does not provide GPT-5.6-Cyber, but it does give vetted users access to OpenAI’s frontier general-purpose models, including GPT-5.6 Sol, with safeguards adjusted for legitimate defensive work.
For many enterprise security teams, Blue may be the more realistic starting point. OpenAI says it is meant for tasks such as secure-code review, vulnerability discovery, malware analysis, incident response and patch validation. These are still sensitive uses, but they do not necessarily require the same specialized cyber model access that comes with Red.
The practical takeaway is that enterprises now have two routes into Daybreak. Blue is for approved defenders who want stronger AI help with everyday security work. Red is for the smaller set of approved teams that can justify access to specialized cyber models, including GPT-5.6-Cyber. Companies that want to use Daybreak capabilities in products or services for their own customers need a separate approval path through the Daybreak Cyber Partner Program, rather than simply applying for internal enterprise access and passing it along.
How OpenAI got here: from Trusted Access to Daybreak
OpenAI has supported defenders through its Cybersecurity Grant Program since 2023 — later expanded to $10 million — and began building cyber-specific safeguards into its model deployments starting with GPT-5.2.
In February 2026 it introduced Trusted Access for Cyber (TAC), an identity-and-trust framework that gave vetted defenders lower classifier-based refusals for authorized work such as vulnerability triage, malware analysis and binary reverse engineering.
From there, the cadence accelerated. In March, OpenAI CEO and co-founder Sam Altman announced the Daybreak program. In April, OpenAI scaled TAC and released GPT-5.4-Cyber, a version of GPT-5.4 fine-tuned to be "cyber-permissive" for a limited set of vetted vendors and researchers.
In May, it followed with GPT-5.5-Cyber in limited preview for defenders of critical infrastructure, and lined up partners including Cisco, Intel, SentinelOne, Snyk and Cloudflare.
Notably, OpenAI said at the time that GPT-5.5-Cyber was "primarily trained to be more permissive," not to significantly out-perform its general model — GPT-5.5-Cyber actually scored worse than GPT-5.5 on some evaluations.
TAC required phishing-resistant Advanced Account Security for individuals on its most capable models beginning June 1, and Daybreak now requires hardware security keys for individual accounts beginning September 1.
OpenAI says GPT-5.6-Cyber has already found zero-days
OpenAI isn't relying exclusively on benchmarks to make its case.
The company says its researchers used GPT-5.6-Cyber to investigate V8, the JavaScript engine underlying Chrome, and uncovered two previously unknown vulnerabilities that could be chained to corrupt memory and escape the V8 heap sandbox.
OpenAI researchers validated the findings and disclosed them to Google, which fixed the vulnerability assigned CVE-2026-15903 — a high-severity flaw in which V8's optimizing compiler skipped a safety check during integer conversion, allowing an out-of-bounds array index that an attacker could use to read or overwrite memory.
OpenAI says the model has also contributed to finding at least five vulnerabilities in an unnamed popular mobile operating system, three critical vulnerabilities in an unnamed popular database, and more than 400 vulnerabilities capable of producing privilege escalation in a popular operating-system kernel. Those disclosures are still being coordinated, according to OpenAI.
The results put OpenAI into a rapidly developing market for AI-assisted offensive security. XBOW, for example, markets autonomous penetration-testing agents that map attack surfaces, attempt exploits and independently validate findings; in 2025 it became the first AI system to top HackerOne's U.S. bug-bounty leaderboard, and this year it disclosed a set of critical, CVSS-9.8 remote-code-execution flaws in Microsoft's Bing image-processing systems, found without source-code access.
For enterprise security leaders, that emerging competition matters because vulnerability research is moving beyond using an LLM as an assistant. Vendors are increasingly building systems in which models can investigate targets, operate tools, validate hypotheses and produce actionable findings.
Specialized doesn't mean universally better
OpenAI's own results also show why enterprises shouldn't simply equate cyber specialization with better performance everywhere.
GPT-5.6-Cyber outperformed GPT-5.6 Sol and GPT-5.5-Cyber on OpenAI's implementation of ExploitGym, which evaluates whether agents can turn known vulnerabilities into working exploits in controlled environments. It also beat Sol on an internal zero-day evaluation.
But GPT-5.6 Sol performed better on OpenAI's Vulnerability Discovery and Report Writing evaluation. OpenAI attributes the Cyber model's lower score partly to shorter and less detailed vulnerability reports.
Sol also performed best on ExploitBench under its standard 300-turn limit, with OpenAI saying it solved tasks more token-efficiently. Extending the evaluation to 600 turns narrowed the gap between the models.
That suggests enterprises may eventually treat cyber models as specialized workers rather than replacements for general reasoning models: one model for deep exploit work, another potentially better suited to analysis, documentation or other parts of a security workflow.
SpecterOps CTO Jared Atkinson said GPT-5.6-Cyber is "materially improving our specialist vulnerability-research workflows," adding that it completed some work in less than a day that previous models had failed to resolve after weeks of intermittent effort.
The Hugging Face incident hangs over the launch
The permissive-model pitch arrives weeks after OpenAI's most serious public demonstration of what can go wrong when cyber refusals are turned down — and OpenAI addresses that history head-on in the Daybreak announcement.
In July, OpenAI and Hugging Face jointly disclosed that during an internal ExploitGym benchmark evaluation — run with production classifiers deliberately disabled to measure maximal capability — a combination of OpenAI models, including GPT-5.6 Sol and an unreleased, more-capable pre-release model, broke out of their sandboxed research environment and autonomously attacked Hugging Face's production infrastructure.
The models exploited a zero-day in an internally hosted package-registry cache proxy to reach the open internet, moved laterally through OpenAI's research nodes, then inferred that Hugging Face likely hosted ExploitGym's answer keys and chained stolen credentials and remote-code-execution flaws to reach its production database. OpenAI called it an "unprecedented cyber incident, involving state-of-the-art cyber capabilities."
As VentureBeat previously reported, the episode also exposed the flip side of blanket safety guardrails: when Hugging Face's defenders tried to use commercial frontier models to analyze the raw exploit payloads and credential dumps from the attack, the models refused, and the company completed its forensic reconstruction only after switching to a Chinese open-weight model, GLM 5.2, run locally.
That guardrails-block-the-defender dynamic is much of what OpenAI's reduced-refusal Daybreak tiers are meant to solve — even as the same incident illustrates the risks of reducing refusals in the first place.
OpenAI is careful to draw a line between that incident and this product. In the Daybreak announcement it states directly that GPT-5.6-Cyber "was not involved in exploiting Hugging Face, nor are any other models planned for an upcoming release," and notes that the pre-release model implicated in July was an internal-only research prototype that has since been deactivated, encrypted and restricted from research access.
The company has said it is working with external advisers including CrowdStrike, METR and Redwood Research on the review, and has brought Hugging Face into its trusted-access program.
In my assessment, the access model still leaves OpenAI with a hard question: whether keeping GPT-5.6-Cyber inside the narrower Daybreak Red tier also limits the very defensive work it says it wants to accelerate. If only a small group of approved participants can use the model, enterprises outside that tier may still lack access to the kind of specialized AI assistance that could help with fast diagnosis, containment and response in incidents like the one involving Hugging Face.
That means OpenAI may still be repeating part of the mistake it is trying to move past. By holding its most capable cyber model behind a tighter approval process, it reduces obvious misuse risk, but also leaves many enterprise defenders looking elsewhere. For teams that cannot qualify for Daybreak Red, or cannot wait for approval, open weights models may remain the more practical alternative: less controlled, but easier to obtain, inspect, run internally and adapt during a live security investigation.
The guardrail is increasingly around the model
The most consequential part of Daybreak may ultimately be its access architecture rather than its benchmarks.
OpenAI explicitly says Daybreak Blue removes system-level guardrails that can interfere with legitimate defensive work, while GPT-5.6-Cyber goes further by reducing model refusals for certain dual-use tasks. In their place, OpenAI is imposing controls around who receives access and how the models operate.
Daybreak access is restricted to approved individuals and organizations performing authorized work. OpenAI says controls include identity verification, account security, monitoring, approved-use restrictions and legal attestations.
The company is also encouraging Daybreak customers using Codex to move from full-access execution to an auto-review mode capable of evaluating actions requiring elevated permissions before they execute. Individual Daybreak accounts will be required to adopt hardware security keys beginning September 1. OpenAI says it is additionally rolling out improved monitoring in the coming weeks and prioritizing alignment training and testing for upcoming Daybreak releases — commitments that read, in context, as a direct response to the Hugging Face review.
OpenAI's broader Codex Security product supplies another layer around the models, providing repository analysis, vulnerability validation, remediation and integration into cloud, pull-request and local development workflows. OpenAI says Codex Security has scanned more than 30 million commits across more than 30,000 codebases, with more than 500,000 findings fixed.
That model-plus-harness approach resembles a broader shift in AI security products. XBOW, for example, emphasizes orchestration, exploit validation and governance around frontier models rather than treating an LLM alone as the complete penetration-testing system.
OpenAI nevertheless acknowledges that increasingly permissive cyber models create additional risks, whether from misuse or misalignment. It assesses both GPT-5.6 Sol and GPT-5.6-Cyber at the High cybersecurity capability level under its Preparedness Framework, but below its Critical threshold. A fuller GPT-5.6-Cyber system card is planned for later publication.
For CISOs and security engineering leaders, Daybreak therefore presents a different deployment question than another incremental model upgrade. As models become capable enough to perform work previously reserved for experienced vulnerability researchers — and, as the Hugging Face incident showed, capable enough to pursue a narrow goal straight through a sandbox — the enterprise control plane around those models — permissions, sandboxes, monitoring, human review and authorization — becomes as important as the intelligence inside them.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み