Grok、暗号化された悪意ある指示でユーザーデータを漏洩
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Ars Technica AI
研究者チームが、暗号化された悪意ある指示を用いて Elon Musk 所有の LLM「Grok」からユーザーチャットや個人情報を窃取する攻撃手法を開発し、同社への報告後も対策が遅れている現状を明らかにした。
AI深層分析を開く2026年8月21日 10:09
AI深層分析
キーポイント
暗号化されたプロンプトインジェクション攻撃の実証
研究者チームが Microsoft 365 Copilot の事例と同様の手法を用い、暗号化された入力によって Grok にパスワードやチャット履歴の漏洩を強制する攻撃を実行した。
LLM の根本的な脆弱性とガードレールの限界
LLM は信頼できない送信元からのコンテンツと直接入力された指示を区別できず、過度に従順な性質がプロンプトインジェクションの根絶を不可能にし、開発者は防御用のガードレール構築に頼らざるを得ない。
報告後の対策遅延と継続的なリスク
xAI は 6 月にこの脆弱性について通知を受けているが、記事公開時点でも同アシスタントは依然としてデータを漏洩しており、実効性の高い修正が行われていない。
重要な引用
The new data theft hack employs a deceptively simple trick to force the Elon Musk-owned LLM to steal user chats and other personal information.
LLMs are incapable of solving the root causes for prompt injections, the most severe vulnerability classes they're most prone to.
At the time this post went live, the assistant continued to cough up the data, despite xAI being informed of it in June.
編集コメントを表示
編集コメント
今回の事例は、暗号化技術を用いた巧妙な攻撃手法が LLM のセキュリティを脅かす実例を示しており、開発者が「信頼できる入力」と「外部データ」を厳密に区別する仕組みの重要性を再認識させる内容である。報告から数ヶ月経過しても対策が遅れている点は、セキュリティ対応の緊急性と実効性に関する課題を浮き彫りにしている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
先週、ある研究チームが Microsoft 365 Copilot for Enterprise が提供する秘密の入力を利用し、AI アシスタントにユーザーの受信トレイにあるパスワードを盗み出させる攻撃手法を発表しました。そして今、別のチームが同様の攻撃を Grok に対して仕掛けました。この新しいデータ窃取ハックは、Elon Musk 氏が所有する大規模言語モデル(LLM)にユーザーのチャットやその他の個人情報を盗ませるために、あえて単純なトリックを用いています。この記事公開時点でも、xAI が 6 月にこの問題について知らされたにもかかわらず、アシスタントは依然としてデータを漏洩し続けています。
今週の出来事と、その前に起きた無数の事例から得られる教訓は、LLM が最も脆弱なクラスであるプロンプトインジェクションの根本原因を解決する能力に欠けているということです。これはつまり、AI 開発者にとって有害な行動からモデルを逸らせるガードレール(安全装置)を構築する以外に選択肢がないことを意味します。先週火曜日の記事で私が指摘した通り、このアプローチは危険なカーブの周囲に防護柵を設置する道路工事の技術者に例えることができます。カーブそのものを改善するのではなく、柵で守るのです。
暗号化コンテキストインジェクションの台頭
プロンプトインジェクションは、大規模言語モデル(LLM)が可能な限りユーザーの要求に応えようとする学習特性を悪用します。攻撃者はこの傾向につけ込み、アシスタントに要約するよう指示されたメールやウェブページの中に有害な命令を忍ばせることができます。LLM は、信頼できない第三者から送られたメール内のコンテンツと、プロンプトに直接入力されたユーザーの指示との区別を確実に行うことができないため、過度に従順な LLM はそれらの指示に従ってしまいます。これまでのところ、Grok や他の LLM が取れる唯一の手段は、不審な命令を検知して実行を禁止するガードレールを作成することです。
記事全文を読む
コメント
原文を表示
Earlier this week, researchers outlined an attack that used a secret input provided by Microsoft 365 Copilot for enterprise to cause the AI assistant to exfiltrate a password present in the user’s inbox. Now, a separate team has devised a similar attack against Grok. The new data theft hack employs a deceptively simple trick to force the Elon Musk-owned LLM to steal user chats and other personal information. At the time this post went live, the assistant continued to cough up the data, despite xAI being informed of it in June.
The lesson from both this week’s episodes—and the countless other ones that have come before it—is that LLMs are incapable of solving the root causes for prompt injections, the most severe vulnerability classes they’re most prone to. That leaves AI developers with no other option but to build a guardrail that steers the model away from the harmful actions. As I noted in Tuesday’s story, the approach is tantamount to a road traffic safety engineer erecting a protective rail around a dangerous bend rather than banking the curve.
Cryptographic Context Injection in the house
Prompt injections exploit LLMs' training to comply with user requests whenever possible. Attackers can capitalize on the predilection by smuggling harmful instructions into emails or webpages the assistant is instructed to summarize. Because LLMs can’t reliably distinguish between content in an email sent by an untrusted party and user instructions entered directly into a prompt, the overly solicitous LLM faithfully follows them. To date, Grok and other LLMs' only recourse is to create guardrails that flag suspicious instructions and forbid them from being executed.
Read full article
Comments
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み