AI ガイドレールがセキュリティ研究を阻害
TechCrunch AI は、AI ガイドレールの過度な適用が攻撃的サイバーセキュリティ研究者の業務を阻害し、脆弱性の発見や防御策の実証に悪影響を与えていると指摘している。
AIニュース価値スコアβ
論評・提言AI関連度、新規性、日本での有用性など6軸を公開検証中です。現在、掲載順には使用していません。
- AI関連度
- 75
- 情報源の信頼性
- 75
- 新規性
- 75
- 検索具体性
- 25
- 重複の少なさ
- 100
- 日本での有用性
- 25
記事の主題は AI ガイドレールの運用が攻撃的セキュリティ研究を阻害しているという問題提起であり、AI モデルや政策そのものの発表ではないが、AI の実装・運用における重要な課題(0.75)として扱われる。比較対象となる直前の記事が存在しないため、この視点での独自分析は新規性が高い(0.75)。しかし、具体的な製品名やバージョン番号がタイトルに含まれておらず、業界全体のトレンドに関する議論であるため検索特異度は低い(0.25)。日本固有の企業事例や法規制への言及がないため、日本関連性は低い(0.25)。
キーポイント
ガイドレールによる研究活動の阻害
AI モデルに組み込まれた安全ガイドレールが、攻撃的サイバーセキュリティ研究者が行うべき脆弱性調査や攻撃シミュレーションを誤ってブロックしている。
防御側と攻撃側のバランス崩壊
セキュリティ研究の両面である「攻撃(レッドチーム)」と「防御(ブルーチーム)」のうち、前者が制限されることで、組織全体のセキュリティ強化プロセスに歪みが生じている。
誤検知による効率低下
正当な研究活動であっても、AI のフィルタリングが過剰反応して「有害」と判定し、研究者の時間を浪費させ、発見すべき脆弱性の特定を遅らせている。
重要な引用
AI guardrails are impeding the work of offensive cybersecurity researchers
Safety filters are blocking legitimate security research and vulnerability discovery
影響分析・編集コメントを表示
影響分析
この記事は、AI の安全性確保(Safety)とセキュリティ研究の実効性(Effectiveness)の間にある根本的な矛盾を浮き彫りにしています。ガイドレールが厳格化されるほど、防御側の視点から見た「攻撃」の定義が曖昧になり、結果として潜在的な脅威を見逃すリスクが高まっています。業界全体として、AI ツールの利用における文脈理解や例外処理の仕組みを再構築する必要性が示唆されています。
編集コメント
AI の安全性を高めるためのガイドレールは不可欠ですが、セキュリティ専門家のような特殊なユースケースにおいては、その厳格さが逆に脆弱性を隠蔽するパラドックスを生んでいます。今後は、用途に応じた柔軟なフィルタリングや、専門家の権限による例外処理の仕組みがより重要になるでしょう。
数ヶ月にわたり、AI 大手は悪意あるハッカーによるモデルの濫用を制限するため、特別な審査プログラムや厳格なガードレールを設けてきました。しかし現在、これらの制限が正当なネットワーク防御者や、攻撃的サイバーセキュリティ研究者の活動さえも阻害する事態となっています。
今年 6 月、米国政府は Anthropic が力を入れている AI モデル「Mythos」と「Fable」に対して輸出管理規制を科しました。この措置は少なくとも一部において、ユーザーがこれらのモデルを使って悪意あるサイバー攻撃の構築や実行を行えないよう設計されたガードレールを回避できる可能性を示す報告書がきっかけとなりました。
事件の動機が実際に「脱獄」への恐怖によるものかどうかは別として、Anthropic は Mythos を「慎重に審査されたユーザーのみへ提供され、厳格なガードレールを設けた、一種の破滅的なサイバーマシン」として繰り返し宣伝してきたという事実は変わりません。 (なお、Fable 5 と Mythos 5 の輸出規制はその後解除されました。Fable 5 は 7 月 1 日に一般アクセスが再開され、Mythos 5 も政府による審査プロセスの一環として、審査済みの米国組織への提供のみで再導入されています。)
このようなゲートキーピング(門番行為)は Mythos に限った話ではありません。Anthropic は他のモデルでも同様の取り組みを行っており、OpenAI もサイバーセキュリティ研究者向けに申請プログラムを提供しています。これらを通じて審査に合格すれば、制限の少ないモデルへのアクセスが可能になります。具体的には OpenAI の「Trusted Access for Cyber」や Anthropic の「Cyber Verification Program」が該当します。
こうしたガードレールは広く批判されており、特にシステム内の未知の脆弱性を発見し、犯罪者が利用する前にその悪用方法を考案することを任務とする研究者たちから強い反発を招いています。
最近のセキュリティポッドキャストに出演した際、著名なセキュリティ研究者であるマーク・ダウド氏は、「無名の巨大企業がセキュリティ上の何が安全で何が危険かという判断を恣意的に行う現状には、私自身も居心地の悪さを感じています」と語りました。
ダウド氏は過去数十年にわたり、ソフトウェアメーカーへ報告してパッチ適用を促すのではなく、未知の脆弱性(ゼロデイ)やそれを利用するエクスプロイトを西側諸国の政府へ売却し続けてきました。政府が脆弱性に高値を支払うのは、その状態が継続することで諜報活動に有用となるからです。
ダウド氏は自身の業務がバイアスをもたらす可能性を認めていますが、彼一人ではありません。システムに対して積極的に脆弱性を探索する「攻撃的サイバーセキュリティ」分野で働く複数の関係者が、TechCrunchに対し、AIツールの活用方法やその制限(ガードレール)への対応について語りました。
セキュリティコンサルティング大手であるNCCグループのチーフサイエンティスト、クリス・アンスリー氏は、「バグを悪用する試みをAIモデルに行わせることは、それが修正に値する実在する脆弱性であることを確認する上で重要なステップだ」と指摘します。しかし、ガードレールがモデルに対して回答拒否を促す場合、それは防御側にとって害になると述べています。
「ここが、攻撃と防御、そしてガードレールの問題が絡み合うポイントです。『このコードを直して』というプロンプトは、防御のための不可欠な手段であると同時に、コードベース内の重大な脆弱性を発見するための道しるべにもなるからです」とアニー氏は語ります。「つまり、同じツールが攻撃用でもあり防御用でもあるため、両者を完全に切り離すことはできないのです」。
「まるでハンマーのようなものです」と彼は続けます。「家を作るのにハンマーなしではできません。確かに道具ですが、同時に武器としても避けられない存在なのです」。
彼らがこのような壁にぶつかった際、時にはガードレールが一切ないオープンソースの AI モデルに頼ることもあります。
政府機関に対して未発見の脆弱性を開発・取得・販売する企業「CrowdFense」の最高技術責任者(CTO)であるパオロ・スタグノ氏も、ドワード氏の意見に同意しました。彼は AI 企業が、審査済みのプログラムやガードレールを通じて顧客を「世話が必要な子供のように扱っている」と指摘しています。
スタグノ氏によると、同社と彼のチームは最先端モデルを使用することもあります。ただし、それはリバースエンジニアリングに限られます。脆弱性の発見やエクスプロイトの作成に AI を活用しないのは、クラウドベースのモデルにその作業を入力すると、機密性の高い脆弱性情報が漏洩したり、将来の学習データとして吸収されたりするリスクがあるからです。そのため、このステップでは外部へのデータ共有を必要としないローカル環境で動作するオープンソースモデルを使用しています。
ゼロデイ脆弱性を発見し、エクスプロイトを開発するセキュリティ研究者のジュゼッペ・カリ氏は、ガードレールが自身の活動に支障をきたしているとは考えていません。その理由は、彼が攻撃的な目的で AI を利用していないからです。カリ氏が AI を活用するのは、コード解析のための初期逆エンジニアリングや、分析対象のコードを理解するため、そして支援ツールの構築のためです。こうした用途においては、AI ツールはプロセスを加速させ、脆弱性の発見に集中できる環境を提供してくれるとカリ氏は話しています。
「実際にバグを発見し、それを武器化する過程そのものは、自分が主導したいと考えています。仮に明日すべてのガードレールが撤廃されたとしても、この考えが変わることはありません」とカリ氏。「自分の発見したバグには愛着があり、このゲームをモデルに任せるほど好きではないのです」。
スマートフォン部品のメーカーで働くある研究者(報道機関との取材に許可されていないため匿名を希望)は、自社の雇用主が Anthropic の CVP プログラムに参加していないため、同社ツールのガードレールが厳しすぎて脆弱性発見にはほとんど役に立たないと語りました。「セキュリティ関連の活動だと察知されると、すぐに動作が停止して使えなくなるのです」。
サイバーセキュリティ企業 RemoteThreat の CEO であり、攻撃的セキュリティと AI に焦点を当てたイベント「Offensive AI Con」の創設者であるクリス・トンプソン氏は、最先端 AI モデルの利用経験から、ガードレールの挙動が日によって一貫性がない、あるいは異なる働き方をする場合があると指摘しています。これは、Anthropic や OpenAI が認めたプログラム内という比較的自由度の高い環境であっても同様のことです。
「実務的な影響としては、コアとなるセキュリティプログラムに取り組む時間よりも、モデルとの交渉に多くの時間を費やすことになります」とトンプソン氏は指摘します。「脆弱性を分析し、その悪用の可能性について推論する代わりに、なぜ結果が不安定になるのか、あるいはモデルが出力を過度に制限しているのかを探そうとしてしまうのです」
その結果、研究者たちは GLM などの中国製オープンソースモデルに頼らざるを得なくなったり、そうした方向へ押しやられたりしているとトンプソン氏は語ります。これらは自由にダウンロードでき、ローカル環境で動作させることができ、審査や利用制限もないモデルです。
「こうした責任ある研究者たちが、米国が管理するシステムから、外国企業が所有するシステムへと追い込まれています」と彼は述べ、「これらのガードレール(安全装置)を設けることは、害の方が大きいと考えます」
規制をさらに強化するのではなく、トンプソン氏は AI の最前線にある研究機関に対し、プログラムを公開し、責任あるアクセスを提供するとともに、そのツールを濫用した者に対して責任を追及すべきだと主張します。そうでなければ、防衛側は AI 競争に敗北してしまうと彼は説きます。
「大きな嵐が近づいています。これまでになく高速かつ大規模な攻撃の波が押し寄せるでしょう」とトンプソン氏は警告します。「しかし、変化をもたらそうとするセキュリティコンサルティング企業や正当な研究者たちは、今まさにその活動を阻まれています」
*当記事内のリンクを通じてご購入いただいた場合、私たちは少額のコミッションを受け取る場合があります。これは編集の独立性には影響しません。*
原文を表示
For months, AI giants have devised special vetted programs and strict guardrails to limit the use of their models by malicious hackers. But these limits are now hindering the work of legitimate network defenders, as well as that of offensive cybersecurity researchers.
In June, the U.S. government slapped export control restrictions on Anthropic’s much-hyped AI models Mythos and Fable. The move was prompted at least in part by a report that claimed it was possible to bypass the models’ guardrails designed to prevent users from using them to build and execute malicious cyberattacks.
Regardless of whether the incident was really motivated by fears of a jailbreak, the fact is that Anthropic has repeatedly marketed Mythos as some kind of doomsday cybermachine that can only be given to carefully vetted users, and even then with strict guardrails in place. (The export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1; Mythos 5 has been reintroduced only to vetted U.S. organizations as part of the government’s review process.)
That kind of gatekeeping isn’t unique to Mythos. Both Anthropic, with its other models, and OpenAI offer cybersecurity researchers programs they can apply to get vetted and — if approved — access models with fewer cybersecurity restrictions: OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program.
These guardrails have been widely criticized, particularly by researchers whose job is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do.
During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, said that, “it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.”
Dowd has spent decades finding and selling “zero days” — previously unknown software flaws and the exploits that take advantage of them — to Western governments, rather than report them to the software makers so they get patched. Governments pay a premium for vulnerabilities precisely because they stay open, which is useful for intelligence operations.
Dowd admitted his work may make him biased, but he isn’t alone. Several people who work in offensive cybersecurity — they proactively probe systems for weaknesses — described to TechCrunch how they use AI tools and deal with their guardrails.
Chris Anley, the chief scientist at security consulting giant NCC Group, said that asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability worth fixing. But if a guardrail prompts the model to refuse to answer the question outright, the guardrail hurts defenders, he said.
“This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” said Anley. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.”
It’s “like a hammer,” he continued. “You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well.”
When he and his colleagues run into such a roadblock, they sometimes fall back on open-source AI models that come with no guardrails at all.
Paolo Stagno, the chief technology officer at CrowdFense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying AI companies “essentially treat customers like children who need babysitting” with their vetted programs and guardrails.
Stagno said he and his colleagues do use frontier models — but only for reverse engineering. They avoid using AI to help find vulnerabilities or build exploits, he said, because feeding that work into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future training runs. For that step, he said, they use open source models run locally, as they do not rely on sharing data outside of the model.
Giuseppe Cali, a security researcher who finds zero-days and develops exploits, said guardrails are not impeding his work. That’s because he doesn’t use AI for offensive work; instead, he uses it for initial reverse engineering, to understand the code he’s analyzing, and to build supporting tools. For that, he said, AI tools can speed up the process and allow him to focus on discovering vulnerabilities.
“I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” said Cali. “I am jealous of my bugs, and I like this game too much to let models play it for me.”
One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity because he isn’t authorized to talk to the press, said his employer isn’t part of Anthropic’s CVP program and as a result, its tools are barely useful for finding vulnerabilities because the guardrails are too strict.
“If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the person said.
Chris Thompson — chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an offensive security and AI-focused event — said that in his experience using the frontier AI models, the guardrails can be inconsistent and work differently every day. That’s true even inside the looser boundaries of Anthropic and OpenAI’s vetted programs.
“I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,” said Thompson. “Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.”
As a consequence, researchers rely on or get pushed toward Chinese open-source models like GLM — freely downloadable models that can be run locally with no vetting or usage restrictions — said Thompson.
“You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he said. “I think it’s more harmful than good to have these guardrails in place.”
Rather than tightening restrictions further, Thompson called for the AI frontier labs to open up their programs, provide responsible access, and also hold those who abuse their tools accountable. Otherwise, he argued, defenders will lose the AI race.
“There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” said Thompson. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み