MIT 研究、医療 AI の効果はユーザーの専門性で異なることを示す
本文の状態
日本語全文を表示中
詳細モードで約8分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MIT ML News
MIT とスタンフォード大学の研究は、医療 AI の説明可能性が非専門家の診断精度を向上させる一方で、専門家には誤った信頼を生むリスクがあることを示し、ユーザーの知識レベルに応じた設計が必要であると結論付けた。
AI深層分析を開く2026年8月4日 18:06
AI深層分析
キーポイント
ユーザーの知識レベルによる効果の違い
非専門家は AI の説明に依存して診断精度が向上したが、これは AI への盲従(デフェレンス)によるものであり、専門家はそのような影響を受けなかった。
説明可能性の逆説的効果
非専門家は曖昧な説明ほど信頼しやすくなる一方、専門家は AI の予測のみを提示された場合に最も高いパフォーマンスを発揮した。
自動化バイアスとアルゴリズムへの依存
AI とその説明方法の両方が人間の自動化バイアスを誘発し、特に知識が少ないユーザーが誤った出力に導かれるリスクがあることが示された。
説明可能性のあるAIは非専門家の診断精度を向上させる
すべての説明可能なAIアプローチが非専門家の正確性を改善し、特に良性のほくろの診断に寄与した。
バイアス抑制モデルが皮膚色による診断格差を縮小
暗い肌色に対するバイアスを防ぐために設計された公平性制約付きモデルを使用すると、精度が大幅に向上し、診断の格差が減少した。
重要な引用
"Good AI systems can improve performance in some health settings, but this has to be balanced carefully with algorithmic deference that can lead to more error."
"These findings are important as patients increasingly turn to AI to help with their health care. Our findings show that those with the least medical knowledge are most likely to be led astray when explainable AI models give an erroneous output."
"Often the people who could benefit most from AI are the ones most likely to be led astray by it, so how we present a recommendation matters as much as whether it's correct," says lead author Orson Xu, an assistant professor in the Department of Biomedical Informatics at Columbia University.
"But the reason non-expert users are better is because they are more reliant on the models. When the model is wrong, it hurts performance more than it helps performance when the model is right."
編集コメントを表示
編集コメント
医療 AI の実装において、技術的な精度だけでなく人間の認知バイアスをどう扱うかが成否を分ける重要な示唆となった。開発者は「説明可能であること」が常に良い結果をもたらすわけではないという前提で、ユーザー層に応じた柔軟な設計戦略を構築する必要がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
疾患診断を支援する人工知能(AI)システムを設計する際、画一的なアプローチが最善の戦略とは限りません。
MIT などの研究者による新しい研究では、AI の支援が皮膚疾患の診断において非専門家や臨床医の精度を全体的に向上させる一方で、AI の説明可能性手法の影響はユーザーの知識レベルによって異なることが明らかになりました。
説明可能な AI(XAI)手法は、モデルの意思決定を記述または検証することで、ユーザーがいつ予測を信頼すべきかを理解するのを助けます。例えば、モデルは診断において最も重要な画像領域を強調するためにヒートマップを使用したり、自然言語で予測を説明するために大規模言語モデル(LLM)を活用したりします。
本研究では、異なる説明可能な AI システムの有無による皮膚疾患の診断について、非専門家と一次医療提供者がテストされました。
その結果、非専門家の診断精度は向上しましたが、それは主に AI システムへの依存によるものでした。非専門家は、LLM による説明が正しいか間違っているかを問わず信頼し、曖昧または一般的な説明の方が説得力があると感じました。
一方、臨床医は誤った AI の支援に惑わされることはなく、モデルの予測のみを与えられ、説明を伴わない場合に最も高いパフォーマンスを発揮しました。
「優れたAIシステムは特定の医療現場でのパフォーマンスを向上させる可能性がありますが、誤りを招く恐れのあるアルゴリズムへの過度な依存とのバランスを慎重に取る必要があります。AIと説明可能な手法の両方が人間における自動化バイアスを引き起こすことが知られており、このアンカリング効果はAIシステムを設計する際に考慮すべき重要な要素です」と、MIT電気工学・情報工学科(EECS)准教授であり、医学工学科学研究所メンバー、情報意思決定システム研究所およびAbdul Latif Jameelヘルスケア機械学習研究所の主任研究者であるMarzyeh Ghassemi氏は述べています。
「患者が医療支援のためにAIに頼るケースが増える中、これらの知見は極めて重要です。説明可能なAIモデルが誤った出力を示した際、医学知識が最も少ない人ほどその影響を受けやすく、道迷いやすいことが私たちの研究で明らかになりました」と、スタンフォード大学生物医科学データ学および皮膚科学准教授であり共同著者のRoxana Daneshjou氏は指摘しています。
研究者たちは、これらの結果は、ユーザーを念頭に置いたAIシステムの構築と、モデルへの過度な依存ではなく批判的思考を促す説明可能性手法の開発がいかに重要であるかを浮き彫りにしていると強調しています。
「優れた AI がすべての問題を解決してくれると安易に考えるのはもう通用しなくなってきました。重要なのは、AI システムを利用するユーザーを慎重に見極めることです。同じ説明内容でも、専門家には有益であっても、初心者を誤解させる可能性があります。実は、AI から最も恩恵を受けるべき人こそが、AI に導かれてしまうリスクも最も高いのです。したがって、推奨事項をどう提示するかは、その内容が正しいかどうかと同じくらい重要なのです」と、コロンビア大学医学情報学科の准教授であり、本研究の筆頭著者である Orson Xu は述べています。
Ghassemi 氏、Xu 氏、Daneshjou 氏のほか、MIT の大学院生 Haoran Zhang さん、学部生の Reina Wang さん、Jameel クリニックの研究員で博士課程修了(2020 年)の Luis Soenksen 氏ら多くの共著者が名を連ねています。また、臨床医や研究者も参加しています。本研究の詳細は本日、『Nature Medicine』誌に掲載されました。
説明の探求
いくつかの FDA 承認済み AI インターフェースが、医療画像から皮膚疾患を特定する際、医師の支援ツールとして活用されています。これは早期診断のプロセスを効率化するための手段です。これらのツールは、画像に病気が存在するかどうかの予測結果を提供するだけでなく、モデルの判断根拠を説明するために、いくつかある手法の一つを採用していることが一般的です。
一方、非専門家の間でも、ユーザーの入力に基づいて皮膚疾患を予測する AI 搭載型検索エンジンを用いた自己診断が可能になっています。これらのシステムでは、LLM(大規模言語モデル)を活用して、モデルの予測結果をより平易な言葉で説明することがよくあります。
研究チームは、説明可能なAIツールの効果と、皮膚疾患の検出における一次医療従事者や非専門家への潜在的なメリットを探りました。被験者には医療画像とAIによる皮膚疾患の予測結果を示し、異なる説明可能なAIのアプローチを適用してテストを行いました。
これらのアプローチには、説明のないAIの予測と信頼度を示す方法、予測を補強するために類似画像を提供する手法、重要な画像領域を強調するヒートマップベースのアプローチ、そしてモデルの推論を平易な言語で解説するLLMが含まれていました。
非専門家には、皮膚のほくろが癌かどうかを判断するタスク(説明可能なAIの有無に関わらず)が与えられました。一方、臨床医にはより難易度の高い、皮膚疾患の鑑別診断を行う課題が出されました。
研究チームは、すべての説明可能なAIアプローチが非専門家の精度を向上させたことを発見しました。これは主に、ツールが非癌性のほくろの診断を支援したためです。
さらに、肌の色に対するバイアスを排除するために設計された公平性制約付きモデルを採用した場合、システムの精度が大幅に向上し、肌色に基づく診断格差も減少しました。
「非専門家の方が優れている理由は、彼らがAIモデルにより依存しているからです。モデルが間違っている場合、それが性能を低下させる影響は、正しい場合に性能を高める効果よりも大きくなります。私たちはこの設定において非常に優れたAIモデルを訓練することができました」とGhassemi氏は述べています。
この「信頼しすぎ効果」は、LLM(大規模言語モデル)による説明の場合に最も顕著に現れ、ユーザーは LLM の支援を受けた際、誤った回答に対してより自信を持つ傾向がありました。
一方、臨床医は AI からの誤った説明にも強く対応でき、すべての説明手法の中で LLM が精度を向上させる効果は最小でした。
「結局のところ、これは各グループがどのように説明を利用するかにかかっています。臨床医はすでに診断の候補を持っており、AI の出力を自身の訓練で検証するため、不十分な説明は見抜くことができます。一方、専門家ではないユーザーはその説明自体を根拠に判断を下すため、もっともらしく自信ありげな論理が誤った方向へ誘導してしまうのです。同じツールが、ある利用者には資産となり、別の利用者には負債となる」と Xu は述べています。
信頼しすぎ効果の克服
研究者がさらに深く分析したところ、AI の支援に最も従順だったユーザーこそが、AI の助けなしではタスクの成績が最悪であることがわかりました。
また、AI の説明を提示するタイミングもユーザーの行動に影響を与えることが判明しました。診断を行う前にまず説明が提示されると、ユーザーはモデルに対してより依存しやすくなります。
さらに、病状の現れ方が微妙な場合、AI システムの方が人間よりも優れた結果を出しましたが、画像に非典型的な症状や無関係な特徴が含まれている場合は、人間の判断の方が大幅に優れていました。
これらの結果を総合すると、説明可能な AI はモデルへの過度な依存を引き起こし、AI の推奨が誤っている場合でもユーザーが無条件に従ってしまうリスクがあることが示唆されます。
LLM を用いてより詳細な説明を生成するのではなく、まずユーザーに診断仮説を提示させ、その後に AI による提案を加えることで、検討すべき他の可能性を浮き彫りにする方が効果的かもしれません。
「私たちは、AI が創造性を高め、ユーザーが見過ごしやすい微妙な症候を見逃さないようスキルを向上させるか、不足部分を補完することを強く望んでいます。そうでなければ、自動化バイアスに陥り、モデルが誤った場合でもユーザーが回復できなくなる恐れがあります」と、Ghassemi 氏は述べています。
本研究は、全米科学財団(NSF)、シュミット・サイエンシズ、国立経済研究局、およびコロンビア大学の支援により実施されました。
原文を表示
A one-size-fits-all approach likely isn’t the best strategy when designing artificial intelligence systems that assist users in disease diagnosis.
A new study by researchers at MIT and elsewhere found that, while AI assistance generally improved the accuracy of non-experts and clinicians in diagnosing skin diseases, AI explainability methods had different impacts depending on the users’ knowledge level.
Explainable AI methods help users know when to trust a model’s predictions by describing or validating the model’s decision-making. For instance, a model might use a heat map to highlight image regions that were most important in its diagnosis or a large language model (LLM) to explain the prediction in plain language.
In this study, researchers tested non-experts and primary care providers in skin disease diagnosis, with and without the help of different explainable AI systems.
They found that non-experts’ diagnostic accuracy improved, but it was largely due to deference to the AI system. Non-experts trusted LLM-based explanations whether they were right or wrong, and found explanations more convincing when they were vague or generic.
By contrast, clinicians were not tripped up by incorrect AI assistance and performed best when given only a model’s prediction, with no accompanying explanation.
“Good AI systems can improve performance in some health settings, but this has to be balanced carefully with algorithmic deference that can lead to more error. We know that both AI and explainability methods can engage automation bias in humans, and this anchoring effect is something that must be accounted for when we design AI systems,” says Marzyeh Ghassemi, an associate professor in MIT’s Department of Electrical Engineering and Computer Science (EECS), a member of the Institute for Medical Engineering and Science, and a principal investigator at the Laboratory for Information and Decision Systems and the Abdul Latif Jameel Clinic for Machine Learning in Health.
“These findings are important as patients increasingly turn to AI to help with their health care. Our findings show that those with the least medical knowledge are most likely to be led astray when explainable AI models give an erroneous output,” says Roxana Daneshjou, a co-author and assistant professor of biomedical data science and dermatology at Stanford University.
These results underscore the importance of building AI systems with users in mind and of developing explainability methods that encourage critical thinking rather than overreliance on the model, the researchers say.
“It’s getting obvious that we cannot just assume a good AI will solve all problems. We need to pay careful attention to the users who will be using the AI system, because the same explanation can help an expert and mislead a beginner. Often the people who could benefit most from AI are the ones most likely to be led astray by it, so how we present a recommendation matters as much as whether it’s correct,” says lead author Orson Xu, an assistant professor in the Department of Biomedical Informatics at Columbia University.
Ghassemi, Xu, and Daneshjou are joined on the paper by many authors, including MIT graduate student Haoran Zhang, undergraduate Reina Wang, and Luis Soenksen PhD ’20, a research affiliate at the Jameel Clinic, along with clinicians and researchers. A description of the work appears today in *Nature Medicine*.
Exploring explanations
Several FDA-approved AI interfaces are being used to help clinicians identify skin conditions in medical images, as a way to streamline early diagnosis. In addition to providing a prediction of whether disease is present in the image, these tools often use one of several methods that explain the model’s decision-making.
At the same time, non-experts can perform digital diagnosis on their own using AI-powered search engines that predict skin diseases based on user prompts. These systems often use LLMs to explain the model’s prediction in simpler terms.
The researchers explored the effects and potential benefits of these explainable AI tools on primary care physicians and non-experts in dermatological disease detection. They tested users by showing them medical images plus an AI prediction of skin disease, employing different explainable AI approaches.
These approaches included: an AI prediction and confidence level with no explanation, a method that provides similar images to reinforce its prediction, a heat map-based approach that highlights important image regions, and an LLM that explains the model’s reasoning in plain language.
Non-experts were tasked with deciding whether an image of a skin mole was cancerous, with and without the help of explainable AI. Clinicians were given the more challenging task of providing a differential diagnosis of dermatological disease.
The researchers found that all explainable AI approaches improved the accuracy of non-experts, mostly because the tools helped users diagnose non-cancerous moles.
In addition, when they employed a fairness-constrained model designed to combat bias against darker skin tones, the system significantly improved accuracy and reduced diagnostic disparities based on skin tone.
“But the reason non-expert users are better is because they are more reliant on the models. When the model is wrong, it hurts performance more than it helps performance when the model is right. We were just able to train very good AI models for this setting,” Ghassemi says.
This deference effect is largest with LLM explanations, and users were more confident about their wrong answers when aided by an LLM.
On the other hand, clinicians were resilient to incorrect AI explanations and, of all the explainability methods, LLMs boost their accuracy the least.
“It really comes down to how each group uses the explanation. A clinician already has a diagnosis in mind and checks the AI against their own training, so a bad explanation gets caught. Meanwhile, a non-expert can use that exact same explanation to form an opinion in the first place, so a plausible, confident-sounding rationale can pull them toward the wrong answer. The same tool ends up being an asset for one user and a liability for another,” Xu says.
Overcoming the deference effect
When the researchers dug deeper, they found that users who were most deferential to AI assistance were the worst performers on the task without the help of AI.
They also found that the time at which* *users were presented with AI explanations influenced their behavior. If an explanation is given first, before the user can perform the diagnosis on their own, they tend to become more deferential to the model.
In addition, AI systems outperformed humans when the presentation of disease was subtle, but humans performed much better if there are atypical symptoms or unrelated features in an image.
Taken together, these results indicate that explainable AI can cause overreliance on models and lead users to blindly follow AI recommendations even when they are wrong.
Rather than using LLMs to generate more detailed explanations, it might be more effective to force users to give a diagnostic hypothesis first, then provide an AI-based suggestion to highlight other possible conditions for consideration.
“We really want AI to improve creativity and either upskill or fill in gaps where users are missing subtle presentations. Otherwise, we risk engaging automation bias and then, when the model is wrong, users can’t recover,” Ghassemi says.
This research was funded, in part, by the National Science Foundation, Schmidt Sciences, the National Bureau of Economic Research, and Columbia University.
AI算出
主要ニュースainew評価標準
MIT などの研究者による医療 AI の説明可能性とユーザー専門性の関係に関する新規の実証研究(Nature Medicine 掲載)であり、AI の効果やリスクが利用者の知識レベルによって異なるという具体的な知見を提供しているため、新規性と関連性は高い。ただし、特定の製品名やバージョン番号ではなく業界全体への示唆であるため検索機会スコアは低め、また日本固有の企業・規制情報がないため日本関連性も低い。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 75
- 新規性
- 75
- 調べる価値
- 25
- 重複の少なさ
- 100
- 日本での有用性
- 25
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み