GPT-5.5 がサイバーセキュリティテストで Mythos Preview に匹敵する性能を示す
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
英国の AI セキュリティ研究所(AISI)が実施した新たなサイバーセキュリティ評価において、先週公開された OpenAI の GPT-5.5 が、Anthropic の Mythos Preview と同程度の性能を達成したことが判明しました。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
先月、Anthropic は、Mythos Preview モデルが示す supposedly 異常に大きなサイバーセキュリティ脅威について大々的に取り上げ、同社が初期リリースを「重要な産業パートナー」に限定するに至ったと報じました。しかし、英国の AI セキュリティ研究所(AISI)による新たな調査では、先週一般公開された OpenAI の GPT-5.5 が、「先月当グループが評価した Mythos Preview と同程度のサイバー評価におけるパフォーマンス水準に達している」ことが示されました。
2023 年以来、AISI は逆エンジニアリング、ウェブエクスプロイト、暗号化など、サイバーセキュリティタスクにおける能力をテストするために設計された 95 の異なる Capture the Flag チャレンジを通じて、さまざまな最先端 AI モデルを検証してきました。最高レベルの「エキスパート」タスクにおいて、GPT-5.5 は平均 71.4 パーセントを達成し、Mythos Preview が記録した 68.6 パーセント(ただし誤差範囲内)をわずかに上回りました。特に困難なタスクの一つである、Rust バイナリをデコードするためのディスアセンブラを構築する課題については、AISI は「GPT-5.5 は API 呼び出しに 1.73 ドルのコストをかけ、人間の支援なしで 10 分 22 秒でこの課題を解決した」と指摘しています。
GPT-5.5 はまた、企業ネットワークに対する 32 ステップのデータ抽出攻撃をシミュレートする AISI テスト範囲である「The Last Ones」(TLO)における進捗においても、Mythos Preview と同等の結果を示しました。GPT-5.5 は TLO で 10 回の試行のうち 3 回で成功しましたが、Mythos Preview は 10 回中 2 回でした。これまでにテストされたどのモデルも、このテストに一度も成功したことはありませんでした。しかし、GPT-5.5 もまた、発電所の制御ソフトウェアの妨害を試みるという、AISI のより困難な「Cooling Tower(冷却塔)」シミュレーションでは失敗しており、これはこれまでテストされたすべての AI モデルが同様に示してきた結果です。
記事全文を読む
コメント
原文を表示
Last month, Anthropic made a big deal about the supposedly outsize cybersecurity threat represented by its Mythos Preview model, leading the company to restrict the initial release to “critical industry partners.” But new research from the UK's AI Security Institute (AISI) suggests that OpenAI's GPT-5.5, which launched publicly last week, reached "a similar level of performance on our cyber evaluations" as Mythos Preview, which the group evaluated last month.
Since 2023, the AISI has run a variety of frontier AI models through 95 different Capture the Flag challenges designed to test capabilities on cybersecurity tasks, such as reverse engineering, web exploitation, and cryptography. On the highest-level "Expert" tasks, GPT-5.5 passed an average of 71.4 percent, slightly higher than the 68.6 percent achieved by Mythos Preview (though within the margin of error). In one particularly difficult task that involved building a disassembler to decode a Rust binary, AISI notes that "GPT-5.5 solved the challenge in 10 minutes and 22 seconds with no human assistance at a cost of $1.73" in API calls.
GPT-5.5 also matched Mythos Preview in its progress on "The Last Ones" (TLO), an AISI test range set up to simulate a 32-step data extraction attack on a corporate network. GPT-5.5 succeeded in 3 of 10 attempts on TLO, compared to 2 of 10 for Mythos Preview—no previous model had ever succeeded at the test even once. But GPT-5.5 still fails at AISI's more difficult "Cooling Tower" simulation of an attempted disruption of the control software for a power plant, as every previously tested AI model also has.
Read full article
Comments
同じ出来事を3媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み