Qwen3Guard:トークンストリームのリアルタイム安全性確保
Qwenチームは安全分類用に微調整した「Qwen3Guard」を発表しました。同モデルはプロンプトと応答の安全性をリアルタイム検出し、リスクレベルと分類を提供してAI対話の安全確保を実現します。
Tech Report GitHub Hugging Face ModelScope DISCORD
導入 Qwen3Guard をご紹介します。これは Qwen ファミリーにおける初の安全性ガードレールモデルです。強力な Qwen3 基盤モデルをベースに、安全分類のために特別にファインチューニングされており、プロンプトとレスポンスの両方に対して正確な安全性検出を提供し、リスクレベルとカテゴリ分類を付与することで、責任ある AI インタラクションを実現します。
Qwen3Guard は主要な安全性ベンチマークにおいて最先端のパフォーマンスを達成しており、英語、中国語、および多言語環境におけるプロンプト分類およびレスポンス分類タスクの両方で強力な能力を示しています。
原文を表示
Tech Report GitHub Hugging Face ModelScope DISCORD
Introduction We are excited to introduce Qwen3Guard, the first safety guardrail model in the Qwen family. Built upon the powerful Qwen3 foundation models and fine-tuned specifically for safety classificatoin, Qwen3Guard ensures responsible AI interactions by delivering precise safety detection for both prompts and responses, complete with risk levels and categorized classifications for accurate moderation.
Qwen3Guard achieves state-of-the-art performance on major safety benchmarks, demonstrating strong capabilities in both prompt and response classification tasks across English, Chinese, and multilingual environments.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み