Google Research、臨床相談の専門家レベル達成へ向けた AMIE の進展を発表
本文の状態
日本語全文を表示中
詳細モードで約12分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Google Research Blog
Google は研究用 AI システム「AMIE」の能力を拡張し、音声・視覚情報を統合した専門家レベルの臨床相談や、腫瘍学・心臓病学などの専門領域における診断支援を実現したと発表した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 02:32
AI深層分析
キーポイント
マルチモーダル診断能力の向上
AMIE は患者の歩行、呼吸、表情といった非言語的な視覚・聴覚情報を統合し、専門家レベルの臨床相談をシミュレートする能力を獲得した。
専門領域への拡大と長期管理
腫瘍学、心臓病学、眼科における専門評価に加え、診断から治療・疾病管理までの長期にわたる対応へと機能が拡張された。
臨床実装に向けた枠組みの構築
研究段階での成果を臨床現場へ移行するため、医師中心の監督体制(physician-centered oversight)を含む実用化の枠組みが整備され始めた。
テキストベースの限界と動画コンファレンスの導入
従来のテキストベースシステムは視覚・聴覚情報を欠き、患者が症状を記述する負担や診断情報の損失という課題があった。これに対し、AMIE (Video) はリアルタイムのビデオ通話で非言語的臨床シグナルを感知し、仮想身体検査を案内することでこれらの限界を克服した。
Gemini と Project Astra を基盤とした多エージェントアーキテクチャ
AMIE (Video) は Gemini および Project Astra を基盤とし、非言語的臨床シグナルの知覚、患者への仮想身体検査の案内、リアルタイムでの診断推論を同期して実行する。
重要な引用
The physician observes the patient's gait, registers visible signs of discomfort, notes their breathing, and guides the patient through physical examination maneuvers.
AI systems capable of clinical reasoning and dialogue have the potential to dramatically increase access to medical expertise and care
We have also extended AMIE's capabilities towards specialist-level evaluations in oncology, cardiology and ophthalmology
Patients must translate complex physical symptoms into written descriptions, a process that discards diagnostic information and can negatively affect patients with limited digital or health literacy.
編集コメントを表示
編集コメント
AMIE の進化は、AI が医療現場で「聞く」「見る」能力を獲得し、医師の補助としてより自然な対話を実現する重要な転換点である。研究段階から臨床実装への移行プロセスが明記されている点は、技術の実用化に向けた現実的なアプローチを示している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
医師が患者と対面したとき、診療は交わされた言葉だけにとどまりません。医師は患者の歩行を観察し、目に見える不快感の兆候を把握し、呼吸の状態を確認し、身体検査の手順を患者に導きます。この絶え間ない視覚・聴覚情報の流れは、話される臨床歴とシームレスに統合されます。これらの非言語的な視覚・聴覚の手がかりは、効果的な診断、患者との信頼関係の構築、そして臨床コミュニケーションにおいて中心的な役割を果たしています。
臨床推論と対話を実現する AI システムは、医療専門知識やケアへのアクセスを劇的に拡大し、医師が患者とのやり取りにおいて最も意味のある部分に集中できる未来を築く可能性があります。初期の研究では、臨床推論と対話を目的とした当社の研究用 AI システム「Articulate Medical Intelligence Explorer(AMIE)」が、テキストベースの診断対話において専門家レベルのパフォーマンスを示し、医師に対する鑑別診断支援ツールとして有効であることを実証しました。最近では、AMIE の能力を診断を超えて拡張し、時間経過に伴う疾患の治療と管理へと進化させました。
また、AMIE の能力を専門医レベルの評価へと拡張し、腫瘍学、心臓病学、眼科における評価や、画像および臨床文書に基づく多モーダル診断推論の実現にも取り組んでいます。これらはすべて、患者役の俳優を用いたシミュレーション環境で検証されたものです。
並行して、これらの研究成果を実際の臨床現場へ移行する取り組みも開始しました。その一環として、「医師中心の監督」を実現するための枠組みを構築し、ベス・イスラエル・ディアコネス医療センターとの共同による最初の現実世界臨床試験(会話型診断 AI の実用性を検証した feasibility study)や、Included Health と提携して実施中の全国規模の無作為化比較試験などを通じて、その実現を目指しています。
これらの進展にもかかわらず、本研究には根本的な制約が残っていました。それは、テキストベースのインターフェースが臨床実践における視覚的・聴覚的側面を排除してしまう点です。患者は複雑な身体的症状を文章で記述する必要があり、この過程では診断に重要な情報が失われ、デジタルリテラシーや健康リテラシーが低い患者には悪影響を及ぼす可能性があります。テキストのみのシステムでは、臨床推論を補完する視覚的・聴覚的手がかりを独立して観察することもできず、鑑別診断を形成する身体診察の手順を患者に案内することもできません。
本日、「*Towards expert-level medical AI for real-time video consultations**」において、これらの課題に対応したリアルタイム動画構成の AMIE(AMIE Video)を発表します。Gemini と Project Astra を基盤とする AMIE(Video)は、非言語的な臨床シグナルを感知し、患者役の人々にバーチャルな身体診察の手順を案内し、リアルタイムで診断的推論を行うことで、同期型の臨床動画診療を実現します。100 のシナリオと 300 のライブ診療、そして 30 名の認定一次保健医(PCP)による多腕ランダム化比較試験において、私たちはリアルタイムの臨床動画診療において専門家のレベルに達する AI システムを示す初のデモンストレーションを発表しました。
AMIE (Video): 非同期型マルチエージェントアーキテクチャ
ビデオを通じた効果的な臨床対話を行うには、競合する要件のバランスが求められます。システムは患者に対して自然な会話速度で応答しつつ、慎重な臨床推論を行い、視覚・聴覚ストリームを継続的に処理する必要があります。しかし現在、単一のエージェントではこれらの要求すべてを満たすことはできません。
深い推論には時間がかかりますが、会話の沈黙が続けば患者との信頼関係やラポールは損なわれます。
この課題に対処するため、AMIE (Video) は非同期マルチエージェントアーキテクチャを採用し、3 つの専門化されたエージェントが並列して継続的に作業を行うことで役割を分担しています。
- *Talker エージェント:* 患者と対話するフロントエンドのエージェントで、応答性が高く低遅延な音声対話を担います。他のエージェントからのガイダンスを取り入れつつも、自然な会話の流れを維持します。
- *Planner エージェント:* バックグラウンドで動作し、システムの臨床推論を継続的に洗練させます。鑑別診断や管理計画を更新しながら情報ギャップを特定し、多様な臨床目標の優先順位を再設定します。
- *Perception エージェント:* 音声と映像ストリームを絶えずレビューし、臨床的に意味のある非言語的シグナル(明らかな苦痛の兆候、身体所見、聴覚信号など)を識別します。これらの観察結果は、進行中の会話の文脈の中で解釈・統合されます。
この分離型設計により、AMIE (Video) は診断推論や音声・視覚知覚を実行しながらも、自然な会話の遅延を維持できます。これらは通常、許容できないほどの遅延を引き起こす要因となります。
自動評価では、この 3 エージェント構成における各エージェントが、病歴聴取や臨床推論、治療提案といった臨床指標の向上に重要な貢献をしていることが確認されました。また、患者中心のコミュニケーションスキルや応答遅延など、対話品質に関連する指標についても同様に改善が見られました。
自動評価による開発ガイダンス
音声・視覚医療 AI を構築する際の主要な課題は、システムの知覚能力と推論能力を大規模に定量化することです。開発を導くため、私たちは遠隔医療に関連する臨床の音声・視覚コンピテンシーに関する分類体系を医学文献から導き出し、非言語的な視覚シグナル、聴覚信号、身体診察の手技を網羅しました。その後、この分類体系を基盤とした自動評価スイートを構築しました。
本評価フレームワークは、対象を絞った単発の音声・視覚アセスメントと、複数回にわたるシミュレーションによる音声相談を組み合わせています。単発の音声・視覚アセスメントでは、臨床的な知覚や推論の特定の側面(例えば、解剖学的な左右の識別や呼吸困難の兆候の認識など)がテストされます。一方、複数回のシミュレーション音声相談では、一連の対話パフォーマンス全体を評価しつつ、視覚情報をテキスト記述としてシナリオに組み込みます。具体的には、パーキンソン病のシナリオで手書きを見せるよう指示された AI 患者シミュレーターが、「[カメラに向かって紙を持ち上げ、ぎこちなく小さな文字を書いている様子]」という音声による説明を挿入するケースなどが該当します。
これらの相補的な評価手法により、人間による評価を行う前に、システム設計の迅速な反復と、AMIE (Video) の能力および失敗モードの詳細な分析が可能となりました。
ランダム化されたビデオ研究を通じた評価
エンドツーエンドの音声・視覚臨床相談という、より困難で現実的なシナリオにおける臨床的有能さを評価するため、同期型ビデオ相談インターフェースを用いた大規模なランダム化 客観的構造化臨床試験 (OSCE) 研究を実施しました。
本研究では、評価対象となる医療疾患の範囲を広くカバーするため、心臓・呼吸器系、腹部、頭部・眼・耳・鼻・喉(HEENT)、神経・精神科、筋骨格系の5つの身体システムにまたがる100の臨床シナリオを対象としました。訓練を受けた患者役俳優15名が、3つの研究グループ間で合計300件の標準化された診療を行いました。
- *AMIE (Video):* AMIE の動画構成によるリアルタイムビデオ診療の実施。
- *AMIE (Text):* 音声・映像機能を切り離し、その貢献度を特定するためのベースラインとして機能するテキストのみのバージョン。
- *PCP (Video):* 同じビデオインターフェースを通じて診療を行う、10名の認定プライマリケア医(PCP)。
20名の経験豊富なプライマリケア医師からなる独立した評価パネルが、一般的な臨床能力スケールに加え、各シナリオに特化した詳細なケース別評価基準を含む確立された臨床ルブリックを用いて、すべての診療を評価しました。
主な結果
*専門家レベルの臨床パフォーマンス:* 診察の徹底度、診断精度、治療方針の適切さ、コミュニケーションの質といった中核的な臨床能力において、AMIE (Video) は PCP と同等の評価を得ました。また、これらの次元において AMIE (Text) を上回るか、少なくとも匹敵する結果を示しました。
身体観察と診察における強み:AMIE(Video)は、平均して一般開業医(PCP)やテキスト版の AMIE よりも有意に高い評価を受け、身体的徴候をより多く引き出し、患者役の人に対して仮想検査の手順を能動的に案内しました。この優位性は、症例ごとの認識と診察基準におけるスコアにも反映されています。
患者役は動画体験を好む:患者役はテキストベースのチャットよりも同期型ビデオインターフェースを強く好み、使いやすさと健康上の懸念を伝える効果の高さを有意に高く評価しました。また、共感力、信頼関係の構築、治療への自信という点でも、AMIE(Video)は一般開業医やテキスト版 AMIE よりも高い評価を得ています。
限界と責任ある開発
本研究には重要な限界があり、これらの結果を限定的な文脈の中で解釈することが極めて重要です。本調査は、実際の患者が自身の健康状態を抱えて来院するのではなく、専門的な俳優による患者役とシミュレーションされた臨床環境において完全に実施されました。俳優がいかに熟練していても、実際の臨床現場の複雑さや予測不可能さを完全に再現することはできず、また演技を通じて誠実に表現可能な病状にのみ対象を限定しました。その結果、聴覚・視覚的な知覚が診断的に決定的な役割を果たす重要な臨床症例は除外されています。
本研究の範囲を超えた自動評価では、全体的な会話の質や診断精度が高いにもかかわらず、時折、知覚や推論における誤りが確認されました。また、システムには依然として会話の自然さを阻害する一時的な技術的問題も存在します。プロジェクト・アストラがプロトタイプである点を踏まえると、これらは本研究で探求した特定の医療応用を超え、将来の開発においてシステムレベルで解決されるべき技術的な課題を含んでいます。
現実世界での有用性について結論を導き出す前に、実際の患者と実際の臨床状況を用いた研究でこれらの知見を検証することは不可欠な次のステップです。
今後の展望
本研究は、テキストベースから音声・映像を扱う臨床 AI への移行が、専門家レベルの品質で実現可能であることを示しています。AMIE (Video) は、臨床現場の知覚的な豊かさに応じ、非言語的シグナルを観察し、身体検査を誘導し、自然な会話を通じて対話を行います。これらは、遠隔医療におけるビデオ診察の体験に、より近い能力です。
責任ある実世界エビデンスの確立に向けた道筋には、まだ重要な課題が残っています。今回の知見は、実際の患者を用いた検証が必要であり、シミュレーションが困難な臨床症状にも対象を拡大し、堅牢な安全性フレームワークによる裏付けが必要です。
私たちはすでにこの方向へ一歩踏み出しています。ベス・イスラエル・ディアコネス医療センターとの「実世界における可行性研究」では、テキストベースの AMIE が臨床現場で安全かつ有用であるという初期のエビデンスが得られました。また、Included Health と共同で実施中の全国無作為化試験では、リアルワールドのバーチャルケアにおける AI の有効性をさらに評価しています。
これらの研究成果は、音声・映像機能を実際の医療現場にどのように責任を持って統合すべきかを示す指針となるでしょう。まだやるべきことは山積ですが、今回の結果は、臨床現場の感覚的な複雑さに応答することで将来的には医療を補完する可能性のある AI システムへの重要なマイルストーンです。
謝辞
本稿で取り上げる研究は、Google Research と Google Deepmind の多数のチームによる共同作業です。共著者である Mahvish Nagda 氏、Jihyeon Lee 氏、Matthew Thompson 氏、CJ Park 氏、Tim Strother 氏、Valentin Liévin 氏、Roma Ruparel 氏、Akshay Goel 氏、Teya Bergamaschi 氏、Suhana Bedi 氏、Meet Shah 氏、Pavel Dubov 氏、Liviu Panait 氏、Toshiyuki Fukuzawa 氏、Sam Schmidgall 氏、Craig Schiff 氏、Joseph Xu 氏、Aliya Rysbek 氏、Yana Lunts 氏、Jan Freyberg 氏、Rebecca Hemenway 氏、Sunny Virmani 氏、David Racz 氏、Carey Radebaugh 氏、Joëlle Barral 氏、Kavi Goel 氏、Dale R. Webster 氏、Katherine Chou 氏、Avinatan Hassidim 氏、Yossi Matias 氏、James Manyika 氏、Gregory Wayne 氏、Tao Tu 氏、Yun Liu 氏、Ethan Goh 氏、Christina Chen 氏、Ryutaro Tanno 氏、Cameron Chen 氏の皆様に心より感謝申し上げます。
原文を表示
When a physician meets a patient, the consultation extends far beyond the words exchanged. The physician observes the patient's gait, registers visible signs of discomfort, notes their breathing, and guides the patient through physical examination maneuvers. This continuous stream of visual and auditory information is seamlessly integrated with the spoken clinical history. These non-verbal visual and auditory cues are central to effective diagnosis, patient trust, and clinical communication.
AI systems capable of clinical reasoning and dialogue have the potential to dramatically increase access to medical expertise and care, fostering a future where physicians can focus their time on the most meaningful aspects of patient interactions. In early work, the Articulate Medical Intelligence Explorer (AMIE), our research AI system for clinical reasoning and dialogue, demonstrated expert-level performance in text-based diagnostic dialogue and proved effective as a differential diagnosis aid for clinicians. Recently, we advanced AMIE’s capabilities beyond diagnosis towards treating and managing disease over time.
We have also extended AMIE's capabilities towards specialist-level evaluations in oncology, cardiology and ophthalmology, and multimodal diagnostic reasoning over images and clinical documents, in simulated settings with patient actors. In parallel, we have begun translating these research advances towards clinical practice, through a framework for physician-centered oversight, as well as our first real-world clinical studies including a clinical feasibility study with Beth Israel Deaconess Medical Center, and an ongoing nationwide randomized study in partnership with Included Health.
Despite these advances, a fundamental constraint in our research remained that text-based interfaces discard the visual and auditory dimensions of clinical practice. Patients must translate complex physical symptoms into written descriptions, a process that discards diagnostic information and can negatively affect patients with limited digital or health literacy. Text-only systems cannot independently observe the visual and auditory cues that inform clinical reasoning, nor can they guide patients through the physical examination maneuvers that shape differential diagnosis.
Today, in “Towards expert-level medical AI for real-time video consultations*”*, we present AMIE in a real-time video configuration, AMIE (Video), that addresses these limitations. Built on Gemini and Project Astra, AMIE (Video) conducts synchronous clinical video consultations, perceiving non-verbal clinical cues, guiding patient actors through virtual physical examinations, and reasoning diagnostically, all in real time. In a multi-arm randomized study with 100 scenarios, 300 live consultations, and a group of 30 board-certified primary care physicians (PCPs), we present the first demonstration of an AI system exhibiting expert-level performance in real-time clinical video consultations.
AMIE (Video): An asynchronous multi-agent architecture
Conducting an effective clinical conversation over video requires balancing competing demands: the system must respond to patients at natural conversational speed while simultaneously performing careful clinical reasoning and continuously processing visual and auditory streams. Currently, a single agent cannot satisfy all these requirements. Deep reasoning takes time, but conversational pauses erode patient trust and rapport.
To address this challenge, AMIE (Video) uses an asynchronous multi-agent architecture that divides labor across three specialized agents working continuously in parallel:
- Talker agent: The patient-facing agent drives responsive, low-latency spoken interaction. It maintains natural conversational flow while incorporating guidance from the other agents.
- Planner agent: Operating in the background, this agent continuously refines the system's clinical reasoning, updating differential diagnoses and management plans while identifying information gaps, and re-prioritizing diverse clinical goals.
- Perception agent: This agent continuously reviews the audio and visual streams, identifying clinically relevant non-verbal cues (such as visible signs of distress, physical findings, or auditory signals) and contextualizing these observations within the ongoing conversation.
This decoupled design allows AMIE (Video) to maintain natural conversational latency while performing diagnostic reasoning and audio-visual perception that would otherwise introduce unacceptable delays. Automated evaluations confirm that each agent in this three-agent architecture makes important contributions towards improvements on clinical metrics, such as competency in history-taking, clinical reasoning and treatment recommendations, as well as on metrics related to dialogue quality, including patient-centered communication skills and response latency.
Guiding development with automated evaluation
A key challenge in building audio-visual medical AI is characterizing a system's perceptual and reasoning capabilities at scale. To guide development, we derived a taxonomy of clinical audio-visual competencies relevant to telehealth from the medical literature, covering non-verbal visual cues, auditory signals, and physical examination maneuvers. We then built an automated evaluation suite structured around this taxonomy.
This evaluation framework combines targeted single-turn audio-visual assessments with multi-turn simulated audio consultations. The single-turn audio-visual assessments test specific instances of clinical perception and reasoning (e.g., correctly identifying anatomical laterality or recognizing signs of respiratory distress). And the multi-turn simulated audio consultations assess end-to-end conversational performance while injecting visual cues as textual descriptions into the simulation (for example, an AI patient simulator for a Parkinson’s scenario prompted to show their handwriting may inject a verbal description of “[holding up paper to camera showing cramped, tiny script]”). Together, these complementary evaluations enabled rapid iteration on system design and richly characterized capabilities and failure modes of AMIE (Video) prior to human evaluation.
Evaluation through a randomized video study
To evaluate clinical competence in the more challenging and realistic setting of an end-to-end audio-visual clinical consultation, we conducted a large-scale, randomized Objective Structured Clinical Examination (OSCE) study with a synchronous video consultation interface.
To cover a breadth of medical conditions in our evaluation, the study spanned 100 clinical scenarios covering five body systems, including cardiopulmonary, abdominal, head/eyes/ears/nose/throat (HEENT), neurological/psychiatric, and musculoskeletal conditions. Fifteen trained patient actors carried out 300 standardized consultations across three study arms:
- AMIE (Video): The video configuration of AMIE conducting real-time video consultations.
- AMIE (Text): A text-only version serving as a baseline to isolate the contribution of audio-visual capabilities.
- PCP (Video): Ten board-certified PCPs consulting via the same video interface.
An independent panel of 20 experienced primary care physicians evaluated all consultations using established clinical rubrics, including both general clinical competency scales and detailed case-specific scoring criteria tailored to each scenario.
Key results
*Expert-level clinical performance:* Across core clinical competencies, history-taking thoroughness, diagnostic accuracy, management appropriateness, and communication quality, clinical evaluators rated AMIE (Video) on par with PCPs. AMIE (Video) also matched or exceeded AMIE (Text) on these dimensions.
*Strength in physical observation and examination:* AMIE (Video) was rated significantly higher, on average, than both PCPs and AMIE (Text) eliciting physical signs and proactively guiding patient actors through virtual examination maneuvers. This advantage was also reflected in case-specific perception and examination rubric scores.
*Patient actors preferred the video experience:* Patient actors strongly preferred the synchronous video interface over text-based chat, rating it as significantly easier to use and more effective for communicating health concerns. They also rated AMIE (Video) favorably on empathy, rapport, and confidence in care compared to both PCPs and AMIE (Text).
Limitations and responsible development
This research has important limitations and it is critical to interpret these results within the context of these limitations. This study was conducted entirely with professional patient actors in simulated clinical settings, not with real patients presenting with their own health conditions. Patient actors, however skilled, cannot fully replicate the complexity and unpredictability of real clinical encounters, and the scenarios were limited to conditions that can be authentically portrayed through acting, omitting important clinical presentations where audio-visual perception would be diagnostically consequential. Beyond the scope of the study, targeted automated evaluations revealed occasional perceptual and reasoning errors, despite overall high-quality conversation and diagnostic accuracy, and the system still exhibits intermittent technical issues that can disrupt conversational naturalness. Given the prototype nature of Project Astra, this includes technical considerations that future development may address at a system level that go beyond the specific medical application explored in this work. Assessing these findings in studies with real patients and real clinical conditions is an essential next step before any conclusions about real-world utility can be drawn.
Looking ahead
This work demonstrates that the transition from text-based to audio-visual clinical AI is achievable at expert-level quality. AMIE (Video) engages with the perceptual richness of clinical practice, observing non-verbal cues, guiding physical examination, and conversing naturally through spoken dialogue — capabilities that more closely approximate the experience of a telehealth video encounter.
Important questions remain on the path towards responsible real-world evidence. Our findings need to be validated with real patients, expanded to encompass clinical presentations that cannot be enacted, and supported by robust safety frameworks. We have already taken early steps in this direction: a real-world feasibility study with Beth Israel Deaconess Medical Center provided initial evidence for the safety and utility of text-based AMIE in clinical practice, and our ongoing nationwide randomized study with Included Health is further evaluating AI in real-world virtual care. Together, these research experiences will help inform how audio-visual capabilities might be responsibly integrated into clinical practice. While much remains to be done, these results mark an important milestone towards AI systems that could one day augment care by engaging with the sensory complexity of clinical practice.
Acknowledgements
*The research described here is joint work across many teams at Google Research and Google Deepmind. We are grateful to all our co-authors - Mahvish Nagda, Jihyeon Lee, Matthew Thompson, CJ Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemenway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, and Cameron Chen.*
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み