大規模言語モデルに対する敵対的攻撃
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Lilian Weng
ChatGPTの普及によりLLM利用が加速する中、OpenAIはRLHFによる安全な動作構築に注力している。しかし、敵対的攻撃やジェイルブレイクプロンプトにより、モデルが望ましくない出力を行うリスクが存在する。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
ChatGPT の登場により、現実世界における大規模言語モデルの使用は急速に加速しました。私たち(OpenAI の私のチームを含め、彼らへの shoutout)は、アライメントプロセスにおいてデフォルトの安全な動作をモデルに組み込むために多大な努力を払ってきました(例:RLHF を通じて)。しかし、敵対的攻撃や jailbreak プロンプトによって、モデルが望ましくない出力を行ってしまう可能性があります。
敵対的攻撃に関する広範な基礎研究は画像を対象としており、連続的な高次元空間で異なる方法で動作します。テキストのような離散データに対する攻撃は、直接的な勾配シグナルの欠如により、はるかに困難であると見なされてきました。私の過去の投稿 Controllable Text Generation はこのトピックと非常に関連しており、LLM に対する攻撃とは本質的に、モデルを制御して特定の種類の(安全でない)コンテンツを出力させることに他なりません。
原文を表示
The use of large language models in the real world has strongly accelerated by the launch of ChatGPT. We (including my team at OpenAI, shoutout to them) have invested a lot of effort to build default safe behavior into the model during the alignment process (e.g. via RLHF). However, adversarial attacks or jailbreak prompts could potentially trigger the model to output something undesired.
A large body of ground work on adversarial attacks is on images, and differently it operates in the continuous, high-dimensional space. Attacks for discrete data like text have been considered to be a lot more challenging, due to lack of direct gradient signals. My past post on Controllable Text Generation is quite relevant to this topic, as attacking LLMs is essentially to control the model to output a certain type of (unsafe) content.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み