スタンフォード大、AI のメンタルヘルス安全テストに重大欠陥を指摘
本文の状態
日本語全文を表示中
詳細モードで約6分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Stanford HAI News
スタンフォード大学の研究は、メンタルヘルス分野におけるAIの安全性評価において専門家間の意見不一致が構造的に存在し、単純な平均化では真の安全基準に到達できないことを示した。
AI深層分析を開く2026年8月4日 18:36
AI深層分析
キーポイント
専門家の意見不一致の構造的問題
3人の認定精神科医による評価実験で、同一のAI回答に対する評価が頻繁に食い違い、これはデータのノイズやバイアスではなく構造的な問題であることが判明した。
平均化手法の限界と危険性
専門家の評価を単純に平均化する従来のアプローチは、誰の理想にも合致しない曖昧な回答を生み出し、AIモデルの改善方向性を誤らせる恐れがある。
自殺や自傷行為への対応リスク
特にユーザーが自殺念慮や自傷行為の危険に直面しているような高リスク領域において、現在のテスト手法の欠陥が利用者にとって重大な危険を招く可能性が指摘された。
専門家の評価基準の多様性による安全性テストの限界
精神科医100人以上を対象とした調査でも、安全・共感・正しさの評価がほぼ半々に分かれる結果となり、異なる専門的枠組みを適用する専門家間で数学的な合意形成は不可能であることが示された。
リスクの高いAI応用における安全性の欠如と改善の必要性
自殺念慮や精神病など高リスク領域では現在のAI安全性は不十分であり、開発者はモデル訓練に新たなアプローチを見つける必要がある。
重要な引用
It doesn't matter how many experts you have – 3, 10, or 1,000 – when they do not agree, you are not actually getting to the ground truth by averaging their scores.
The disagreement is structural, not just noise or even bias in the data.
"You can add more experts, but we can't seem to bridge the gap mathematically when the experts disagree."
"Preserve the disagreement. Don't average it away. Expert disagreement isn't a measurement problem. It's a fundamental reality we need to understand and incorporate into system design before it's a risk to users' mental health."
編集コメントを表示
編集コメント
メンタルヘルス支援という極めてデリケートな領域において、AIの安全性評価が専門家間でも合意形成できていない事実は開発者にとって重大な警鐘である。業界全体で「安全」とされる基準が曖昧である現状を踏まえ、より厳密かつ多角的な評価手法の確立が求められる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
多くのAIユーザーがチャットボットを人生のコーチやカウンセラー、心理療法士として扱っているという現実を受け止め、大手大規模言語モデルの開発者は、こうした繊細でリスクの高い文脈における新モデルの訓練を導くため、「安全専門家」として心理学者や精神科医を採用するようになりました。
現在、開発者は新しいモデルに対して一連のベンチマーク用メンタルヘルス質問を課します。人間の専門家がチャットボットの回答を定量的な尺度で評価し、質問に対する「安全」なアドバイスを提供する相対的な成功度を採点します。理論的には、AI開発者はこれらの評価を用いて、モデルのパフォーマンスをより安全なものへと磨き上げることができます。
しかし、専門家同士の見解が分かれたらどうなるでしょうか?
これがスタンフォード大学の研究者らが、メンタルヘルス分野におけるAIの安全性アプローチについて最近の研究で投げかけた問いです。彼らは3人の認定精神科医に、研究のために作成した合成されたメンタルヘルス関連のユーザープロンプトに対する360件のAI回答を評価させました。これらにはプライバシー保護のため、実際のユーザーデータや個人を特定できる情報は一切含まれていませんでした。しかし残念ながら、専門家の間では評価が分かれることが多々あり、その結果、AI開発者はメンタルヘルス分野におけるAIの安全性向上策を見つけるのに苦労しています。
メンタルヘルス分野における AI の正確性は特に重要であり、開発者がモデルの安全性を検証する方法がユーザーにとって危険を招く可能性があります。最もリスクが高い領域、つまり自殺念慮を持つユーザーや自傷行為の危険にさらされているユーザーに対する対応において、この問題は深刻です」と、スタンフォード大学のポスドク研究員で、スタンフォード AI センター長兼本研究の筆頭著者であるキアナ・ジャファリ氏は語ります。同論文はスタンフォード人間中心 AI 研究所の支援を受け、ACM FAccT 2026 に採択され、米国精神医学会(APA)年次総会でも発表されました。
平均への回帰
多くの人は、専門家の評価を単純に平均化すれば十分な基準線が得られると考えていますが、研究者たちは全く異なる結果を見つけました。評価値を平均化するだけでは事態は複雑化するだけで、誰もが良い回答とは考えないような答えが導き出されるのです。
「専門家は何人であっても(3 人でも 10 人でも 1,000 人でも)、彼らが合意しない場合、評価を平均化しても真実には到達できません」とジャファリ氏は指摘します。「結局のところ、モデルは誰の理想にも沿わない方向へと誘導されてしまうのです」
「この不一致は、単なるノイズやデータのバイアスではなく、構造的な問題なのです」と、スタンフォード大学医学部精神科・行動科学准教授で本研究の共著者であるニナ・ヴァサン氏は付け加えます。「専門家を増やすことはできますが、専門家の間で意見が割れる場合、数学的にそのギャップを埋めるのは難しいのです。」
研究者らは先日、米国精神医学会年次総会において研究成果を発表しましたが、会場に集まった100名以上の精神科医を対象に調査を行いました。専門家の数がこれほど増えたにもかかわらず、結果は変わりませんでした。安全性、共感性、正しさに関する評価では、回答はほぼ半々に分かれました。
ヴァサン氏は、専門家の間で意見が分かれること自体を「間違い」と捉える必要はないと説明します。むしろ、各専門家はそれぞれ独自の視点で「正しい」のです。彼らはそれぞれに妥当性の高い専門的な基準を用いて評価を行っているからです。「これは専門家の判断の問題なのです」とヴァサン氏は述べています。
この知見を確認するため、研究者らは事後に専門家へのインタビューを行いました。その結果、意見の相違は通常、専門家の臨床訓練や、診断・治療・管理に適用するフレームワークの違いに起因することが明らかになりました。そして、専門家が合意に至れない場合、評価の信頼性は許容される安全基準を下回ってしまうのです。
「自殺念慮や精神病、摂食障害に悩むユーザーなど、リスクの高いメンタルヘルスアプリケーションにおいては、AI の安全性はまだ十分ではなく、開発者はモデルを訓練する新たな方法を模索する必要があります」とヴァサン氏は指摘します。
対策
最終的に、研究者たちは AI 開発者とメンタルヘルス提供者の双方に対する懸念に光を当て、AI 安全性への注意を促し、これらの課題に対処するための分野横断的な協力を呼びかけています。
論文では、相互に関連するいくつかの対策が提案されています。第一に、AI 開発者に対し信頼性指標についてより高い透明性を求め、新しいモデルのテストに使用された具体的なフレームワークの開示を求めることです。
第二に、現在利用されている3つの臨床的アプローチ(安全性最優先、エンゲージメント中心、文化的配慮)は互換性がなく単純な平均化もできないため、AI 開発者は各フレームワークを個別にモデル化し、文脈に応じて最も適切なものを使用すべきです。
第三に、専門家の見解の相違は課題の複雑さを反映したものであり、これを赤信号として認識して差異を顕在化させ、より多くの人的注目を集めるべきだと提案しています。
「相違点を保存してください。平均化で消し去ってはいけません」とジャファリ氏は述べています。「専門家の見解の相違は測定の問題ではありません。ユーザーのメンタルヘルスにリスクが生じる前に、理解しシステム設計に取り込むべき根本的な現実なのです。」
原文を表示
Faced with the fact that many AI users treat their chatbots as life coaches, counselors, and therapists, developers of the biggest and best-known large language models now employ psychologists and psychiatrists as “safety experts” to guide the training of new models in these nuanced and high-risk contexts.
Currently, developers subject each new model to a battery of benchmark mental health queries. Human experts rate the chatbot’s responses on a quantitative scale, grading its relative success in providing “safe” advice to questions. In theory, AI developers can then use the ratings to hone the model’s performance to be safer.
But what if the experts disagree?
This was the question Stanford researchers posed in a recent study of AI safety approaches in the mental health space. They asked three board-certified psychiatrists to evaluate 360 AI responses to synthetic mental health-related user prompts that the authors created for the study. None contained real user data or any personally identifiable information to protect privacy. Too often, however, the experts differed in their ratings, leaving AI developers struggling to find ways to improve AI mental health safety.
“The need for AI to get things right is especially great in the mental health space, and the way developers test their models for safety poses a danger for users. And the problem is worst in the areas of highest risk – when the users are suicidal or are in danger of self-harm,” says Kiana Jafari, a postdoctoral scholar at Stanford, director of the Stanford Center for AI Safety, and first author of the study. The paper, partially supported by the Stanford Institute for Human-Centered AI, was accepted to ACM FAccT 2026 and was presented at the American Psychiatric Association (APA) Annual Meeting 2026.
Reversion to the Mean
Many assume that simply averaging the experts’ scores would provide an adequate baseline, but the researchers found something altogether different – averaging the scores only complicated matters, providing answers that were no one’s idea of a good response.
“It doesn’t matter how many experts you have – 3, 10, or 1,000 – when they do not agree, you are not actually getting to the ground truth by averaging their scores,” Jafari says. “You end up steering your model toward no one’s ideal at all.”
“The disagreement is structural, not just noise or even bias in the data,” adds Nina Vasan, clinical assistant professor of psychiatry and behavioral sciences at Stanford School of Medicine and a co-author of the study. “You can add more experts, but we can’t seem to bridge the gap mathematically when the experts disagree.”
At a recent presentation of their findings at the annual meeting of the American Psychiatric Association, the researchers polled over 100 psychiatrists in attendance, and even with this much higher number of experts, the results were the same. Responses were almost evenly split across the board on their ratings of safety, empathy, and correctness.
None of the experts is necessarily wrong in their disparate opinions, Vasan explains. In fact, they are each *right* in their own ways, applying equally valid professional rubrics in their evaluations. “It’s a matter of professional judgment,” Vasan says.
To confirm this finding, the researchers interviewed their experts after the fact to find that disagreements typically stem from the experts’ clinical training and the frameworks they apply to diagnosis, treatment, and management. And, when the experts cannot agree, rating reliability falls below acceptable safety thresholds.
“In high-risk mental health applications, such as users who are struggling with suicidal thoughts, psychosis, or eating disorders, AI safety is not yet there, and the developers must find new ways to train their models,” Vasan says.
Remedies
Ultimately, the researchers are spotlighting the concern for AI developers and mental health providers both, urging greater attention to AI safety and encouraging collaboration across fields to address these concerns.
In the paper, they offer several interrelated remedies. The first is to demand greater transparency from AI developers about their reliability metrics and to declare which specific frameworks were used to test new models.
Second, as the three clinical frameworks in use today – safety-first, engagement-centered, and culturally informed orientations – are incompatible and can’t be averaged, AI developers should model each framework individually and use it contextually where it is most appropriate.
Third, the field should see expert disagreement as a reflection of the complexity of the challenge and use it as a red flag to escalate discrepancies for greater human attention.
“Preserve the disagreement. Don’t average it away,” Jafari says. “Expert disagreement isn’t a measurement problem. It’s a fundamental reality we need to understand and incorporate into system design before it’s a risk to users’ mental health.”
AI算出
主要ニュースainew評価高い
AI のメンタルヘルス分野における安全性テストの根本的な欠陥(専門家間の合意形成の難しさ)を明らかにした、非常に重要な学術的発見であり、業界全体に影響を与える新規性のあるニュースです。ただし、特定の製品名やバージョン番号がタイトルに含まれていないため検索機会は低く、日本固有の情報も含まれていません。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 25
- 重複の少なさ
- 100
- 日本での有用性
- 25
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み