Jump to Content

Google's
research
medical AI

play silent looping video pause silent looping video
unmute video mute video

AMIE (Articulate Medical Intelligence Explorer) is Google’s flagship research project for AI in clinical reasoning and conversations. Our research mission is to strive for universal access to expert-level medical care enabling all to live their healthiest lives. AMIE is a collaborative effort across Google Research, Google DeepMind, Google Platforms and Devices, and Google for Health.

Clinical studies

Trust in clinical AI cannot be benchmarked into existence. It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself. Therefore, we partner with clinical research collaborators to study AMIE in real-world care settings.

Single-center feasibility study

Video preview image

Watch the film

In partnership with Beth Israel Deaconess Medical Center, we explored the feasibility of conversational diagnostic AI in ambulatory primary care through a first-of-its-kind real-world clinical study, published in The Lancet.

Nationwide randomized study

In partnership with Included Health, we launched a first-of-its-kind nationwide study to evaluate conversational AI within real-world virtual care workflows. Moving beyond simulation and retrospective data, this research aims to gather rigorous, prospective evidence on how AI performs in clinical settings at scale.

Publications

A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic

Authors: Peter Brodeur, Jacob M. Koshy, Anil Palepu, Khaled Saab, Ava Homiar, Roma Ruparel, Charles Wu, Ryutaro Tanno, Joseph Xu, Amy Wang, David Stutz, Wei-Hung Weng, Hannah M. Ferrera, David Barrett, Lindsey Crowley, Jihyeon Lee, Spencer E. Rittner, Ellery Wulczyn, Selena K. Zhang, Elahe Vedadi, Christine G. Kohn, Kavita Kulkarni, Vinay Kadiyala, Sara Mahdavi, Wendy Du, Jessica Williams, David Feinbloom, Renee Wong, Tao Tu, Petar Sirkovic, Alessio Orlandi, Christopher Semturs, Yun Liu, Juro Gottweis, Dale Webster, Joelle Barral, Kat Chou, Pushmeet Kohli, Avinatan Hassidim, Yossi Matias, James Manyika, Rob Fields, Jonathan X. Li, Marc L. Cohen, Vivek Natarajan, Mike Schaekermann, Alan Karthikesalingam, and Adam Rodman.

arXiv (2026)

 

Prospective evidence for conversational medical AI is hard, but non-negotiable

Authors: Mike Schaekermann, Anil Palepu, Adam Rodman, Ami Parekh, Ethan Goh, Shek Azizi, Yun Liu, Sunny Virmani, Christina Chen, Dale Webster, Joelle Barral, Avinatan Hassidim, Yossi Matias, and Cameron Chen.

Nature Medicine (2026)

Art of the possible

Explore our “art of the possible” research spanning advancements from text-based diagnostic settings toward multimodal medical reasoning, disease management over time, and AMIE in specialty care.

Multimodal medical reasoning

Medicine is an inherently multimodal discipline that requires interpreting audiovisual cues, ranging from observing a patient's gait or noting their breathing during a live clinical encounter, to reviewing a photo of a skin concern or a medical document. Explore our research on AMIE for multimodal medical reasoning below.

Video preview image

Watch the film

Publications

Towards expert-level medical AI for real-time video consultations

Authors: Mahvish Nagda, Jihyeon Lee, Matthew Thompson, CJ Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemenway, Sunny Virmani, David Racz, Carey Radebaugh, Joelle Barral, Kavi Goel, Dale Webster, Kat Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Cameron Chen, Mike Schaekermann, and Anil Palepu.

arXiv (2026)

 

Advancing conversational diagnostic AI with multimodal reasoning

Authors: Khaled Saab, CJ Park, Tim Strother, Jan Freyberg, David Barrett, Yong Cheng, Wei-Hung Weng, David Stutz, Nenad Tomašev, Anil Palepu, Valentin Liévin, Yash Sharma, Roma Ruparel, Abdullah Ahmed, Elahe Vedadi, Kimberly Kanada, Cían Hughes, Yun Liu, Geoff Brown, Yang Gao, Sean Li, Sara Mahdavi, James Manyika, Kat Chou, Yossi Matias, Avinatan Hassidim, Dale Webster, Joelle Barral, Ali Eslami, Pushmeet Kohli, Adam Rodman, Vivek Natarajan, Mike Schaekermann, Tao Tu, Alan Karthikesalingam, and Ryutaro Tanno.

Nature Medicine (2026)

From diagnosis to disease management

Receiving a diagnosis is often just the first step in a long journey. It is the effective treatment and management of disease over time that ultimately allows people to resume good health or manage symptoms. Our Nature-published research goes beyond diagnostic reasoning and dialogue and explores AMIE’s capabilities in disease management, including reasoning about disease progression, therapeutic response, safe medication prescription, and the appropriate use of accepted guidelines or evidence.

play silent looping video pause silent looping video
unmute video mute video

Leveraging Gemini’s long-context capabilities, AMIE supports longitudinal disease management, grounding its reasoning in clinical guidelines and adapting to patient needs across multiple visits.

Publications

Towards conversational artificial intelligence for disease management

Authors: Valentin Liévin, Anil Palepu, Wei-Hung Weng, Khaled Saab, David Stutz, Yong Cheng, Kavita Kulkarni, Sara Mahdavi, Joelle Barral, Dale Webster, Avinatan Hassidim, Yossi Matias, James Manyika, Ryutaro Tanno, Vivek Natarajan, Adam Rodman, Tao Tu, Alan Karthikesalingam, and Mike Schaekermann.

Nature (2026)

 

Towards conversational diagnostic artificial intelligence

Authors: Tao Tu, Mike Schaekermann, Anil Palepu, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Nenad Tomašev, Shekoofeh Azizi, Karan Singhal, Yong Cheng, Le Hou, Albert Webson, Kavita Kulkarni, Sara Mahdavi, Christopher Semturs, Juro Gottweis, Joelle Barral, Kat Chou, Greg Corrado, Yossi Matias, Alan Karthikesalingam, and Vivek Natarajan.

Nature (2025)

 

Towards accurate differential diagnosis with large language models

Authors: Daniel McDuff, Mike Schaekermann, Tao Tu, Anil Palepu, Amy Wang, Jake Garrison, Karan Singhal, Yash Sharma, Shekoofeh Azizi, Kavita Kulkarni, Le Hou, Yong Cheng, Yun Liu, Sara Mahdavi, Sushant Prakash, Anupam Pathak, Christopher Semturs, Shwetak Patel, Dale Webster, Ewa Dominowska, Juro Gottweis, Joelle Barral, Kat Chou, Greg Corrado, Yossi Matias, Jake Sunshine, Alan Karthikesalingam, and Vivek Natarajan.

Nature (2025)

Specialty care

We have partnered with research collaborators at Stanford, Houston Methodist, University College London and other institutions to assess AMIE’s clinical reasoning capabilities in specialty care, including cardiology, ophthalmology, and oncology.

Publications

A large language model for complex cardiology care

Authors: Jack W O’Sullivan, Anil Palepu, Khaled Saab, Wei-Hung Weng, Daniel K. Amponsah, Evaline Cheng, Yong Cheng, Emily Chu, Yaanik Desai, Aly Elezaby, Muhammad Fazal, Tasmeen Hussain, Sneha S. Jain, Daniel Seung Kim, Roy Lan, Jiwen Li, Wilson Tang, Natalie Tapaskar, Victoria Parikh, Ryan Sandoval, Gabriela Spencer-Bonilla, Bryan Wu, Kavita Kulkarni, Philip Mansfield, Dale Webster, Juro Gottweis, Joelle Barral, Mike Schaekermann, Ryutaro Tanno, Sara Mahdavi, Vivek Natarajan, Alan Karthikesalingam, Euan Ashley, and Tao Tu.

Nature Medicine (2026)

 

Exploring large language models for specialist-level oncology care

Authors: Anil Palepu, Vikram Dhillon, Polly Niravath, Wei-Hung Weng, Preethi Prasad, Khaled Saab, Ryutaro Tanno, Yong Cheng, Hanh Mai, Ethan Burns, Zainub Ajmal, Kavita Kulkarni, Philip Mansfield, Dale Webster, Joelle Barral, Juro Gottweis, Mike Schaekermann, Sara Mahdavi, Vivek Natarajan, Alan Karthikesalingam, and Tao Tu.

NEJM AI (2025)

 

Complementary human-AI clinical reasoning in ophthalmology

Authors: Mertcan Sevgi, Fares Antaki, Abdullah Zafar Khan, Ariel Yuhan Ong, David Adrian Merle, Kuang Hu, Shafi Balal, Sophie-Christin Kornelia Ernst, Josef Huemer, Gabriel T. Kaufmann, Hagar Khalid, Faye Levina, Celeste Limoli, Ana Paula Ribeiro Reis, Samir Touma, Anil Palepu, Khaled Saab, Ryutaro Tanno, Valentin Liévin, Tao Tu, Yong Cheng, Mike Schaekermann, Sara Mahdavi, Elahe Vedadi, David Stutz, Vivek Natarajan, Alan Karthikesalingam, Pearse Keane, and Wei-Hung Weng.

arXiv (2025)

In the press

Video preview image

Watch the film

Joëlle Barral on “AI and the future of health” | Google DeepMind Podcast with Hannah Fry.

Selected publications

We believe that an evidence-based approach backed by rigorous science is essential for ensuring that AI in medical settings can be explored safely and responsibly while building trust with patient and clinical communities. AMIE research has been published in high-impact clinical and scientific journals including The Lancet, Nature, Nature Medicine and NEJM AI.
Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study
Peter Brodeur
Jacob M. Koshy
Khaled Saab
Ava Homiar
Roma Ruparel
Charles Wu
Ryutaro Tanno
Joseph Xu
Amy Wang
David Stutz
Hannah M. Ferrera
David Barrett
Lindsey Crowley
Jihyeon Lee
Spencer E. Rittner
Selena K. Zhang
Elahe Vedadi
Christine G. Kohn
Kavita Kulkarni
Vinay Kadiyala
Sara Mahdavi
Wendy Du
David Feinbloom
Renee Wong
Petar Sirkovic
Alessio Orlandi
Juro Gottweis
Joelle Barral
Kat Chou
James Manyika
Rob Fields
Jonathan X. Li
Marc L. Cohen
Adam Rodman
The Lancet (2026)
Preview abstract Background: Artificial intelligence (AI)-based systems show promise for assisting primary care providers (PCPs) with patient care. We aimed to evaluate the safety and quality of clinical conversations of a patient-facing conversational AI system, which engaged in real-world urgent primary care appointments. Methods: In this prospective, single-centre, single-arm feasibility study, English-speaking patients aged at least 18 years interacted with the Articulate Medical Intelligence Explorer (AMIE) up to 5 days before a single-complaint urgent primary care appointment. Physician safety supervisors monitored all interactions and were trained to intervene on the basis of predefined safety criteria. AMIE transcripts and summaries were shared with PCPs before the visit. Primary outcomes were the number of supervised conversation safety stops, AMIE’s conversation quality assessed by clinical evaluators, and patient and PCP experiences per surveys. This study is registered with ClinicalTrials.gov (NCT06911398). Findings: From April to November, 2025, 114 patients were enrolled with 98 completing both the AMIE interaction and the PCP appointment. Zero conversation safety stops were required on the basis of prespecified criteria. Safety supervisors noted one hallucination and added clinical information in five interactions. AMIE’s conversations were rated favourably in 87–100% of cases (17 criteria) by clinical evaluators, and 48–96% (16 criteria) by patients. Patient attitudes towards AI improved after interacting with AMIE and remained elevated after the patient’s visit with their physician. PCPs completed post-surveys in 60 of 98 cases, including 44 cases in which they reviewed the AMIE transcript before the visit. PCPs found AMIE helpful for visit preparation in 33 of 44 cases and reported that it might have changed their behaviour in 25 of 44 cases. Interpretation: Although further research is needed, this study shows the initial feasibility of conversational AI in a real-world setting—assessed via conversation safety and quality, as well as user acceptance—and represents a crucial step towards clinical translation. Funding: Alphabet. View details
Preview abstract Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence. It must be earned through rigorous prospective studies in real-world clinical settings, where the hardest lessons often concern the humans and systems around the AI, not the technology itself. View details
Towards expert-level medical AI for real-time video consultations
Mahvish Nagda
Jihyeon Lee
Matthew Thompson
CJ Park
Tim Strother
Roma Ruparel
Teya Bergamaschi
Suhana Bedi
Meet Shah
Pavel Dubov
Toshiyuki Fukuzawa
Sam Schmidgall
Craig Schiff
Joseph Xu
Aliya Rysbek
Yana Lunts
Jan Freyberg
Rebecca Hemenway
David Racz
Carey Radebaugh
Joelle Barral
Kavi Goel
Kat Chou
James Manyika
Gregory Wayne
Yun Liu
Ethan Goh
Christina Chen
Ryutaro Tanno
arXiv (2026)
Preview abstract Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility, but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice. View details
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
Sam Schmidgall
Tim Strother
Alex Bijamov
CJ Park
Vahid Balazadeh
Min Woo Sun
Marius Guerard
Justin Chen
Vikram Dhillon
Kat Chou
Quoc Le
Raia Hadsell
Joelle Barral
Carey Radebaugh
Aleksandra Faust
David Racz
Lin Yang
Arxiv (2026)
Preview abstract In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy. Much of clinical reasoning relies on the patient encounter, a dialogue in which a clinician elicits history, refines diagnostic hypotheses, and decides management under uncertainty. While large language models (LLMs) excel on static medical benchmarks, methods to optimize the full sequence of clinical decisions remain underdeveloped. We present ResidencyRL, a reinforcement learning (RL) method for training clinical artificial intelligence (AI) agents through simulated multi-turn clinical encounters (up to 60 dialogue turns and 8 tool calls per trajectory). ResidencyRL pairs the policy agent with LLM simulators capable of complex, adversarial behaviors, training against a structured reward aligned to diagnostic accuracy, management quality, communication, documentation, and safety. On held-out evaluations, the ResidencyRL agent improves diagnostic accuracy by 7.0% under adversarial conditions (88.0% vs. 81.0%) and reduces missed red flag rates by 31%, demonstrating rigorous mitigation of premature closure. Blinded expert clinicians validated these gains, preferring the trained agent in 87.6% of side-by-side comparisons. The procedural competencies transfer to unseen benchmarks: the agent outperforms the base model across all six clinical axes of the AMIE multi-visit benchmark, and shows consistent directional improvements on AgentClinic and CRAFT-MD. Our findings demonstrate that sequential clinical decision-making can be effectively learned through multi-turn RL in simulation, yielding robust, generalizable capabilities, paving the way towards clinical mastery. Prospective validation with real-world workflows remains necessary to establish clinical utility. View details
Toward a test of medical AI superintelligence
Ethan Goh
David Wu
Chase Walton
Liam McCoy
Anastasia Perez
Laura Wegner
Fateme Nateghi Haredasht
Luyang Luo
Kathleen Lacar
Thomas Buckley
Austin Schoeffler
Peter Brodeur
Kameron C. Black
John Havlik
John Rumsfeld
Daniel Lopez-martinez
Paxton Maeder-York
Karan Singhal
David Gunning
Bon Ku
Haider Warraich
Shantanu Nundy
Vishnu Ravi
Arnold Milstein
Jason Hom
Kevin Schulman
Pranav Rajpurkar
Arjun Manrai
Robert Wachter, MD
Eric Topol
Eric horvitz
Adam Rodman
Jonathan Chen
Nature Medicine (2026)
Preview abstract Researchers urgently need a rigorous, task-based framework to define and measure medical AI ‘superintelligence’, because existing benchmarks are misleading and insufficient. View details
Preview abstract Although large language models have shown promise in diagnostic dialogue, their capabilities for effective management reasoning, including disease progression, therapeutic response and safe medication prescription, have remained underexplored. We have advanced the previously demonstrated diagnostic capabilities of the Articulate Medical Intelligence Explorer (AMIE) using a new large-language-model-based agentic system optimized for multivisit clinical management and dialogue. To ground the reasoning of AMIE in authoritative clinical knowledge, we leveraged the long-context capabilities of Gemini, combining in-context retrieval with structured reasoning to align its output with up-to-date clinical practice guidelines and drug formularies. In a randomized, blinded virtual Objective Structured Clinical Examination study, AMIE was compared to 21 primary care physicians (PCPs) across 100 multivisit case scenarios designed to reflect the guidance of the UK National Institute for Health and Care Excellence and BMJ Best Practice guidelines. AMIE was non-inferior to PCPs in management reasoning, as assessed by specialists, and scored better both with respect to preciseness of treatment and investigation, and in terms of its alignment with and grounding in clinical guidelines. To benchmark medication reasoning, we developed RxQA, a multiple-choice question benchmark that was derived from two national drug formularies (from the USA and UK) and validated by board-certified pharmacists. Although AMIE and PCPs both benefited from the ability to access external drug information, AMIE outperformed PCPs on higher-difficulty questions. Although further research will be needed before real-world translation of AMIE, its strong performance across evaluations marks a significant step towards use of conversational artificial intelligence as a tool in disease management. View details
Advancing conversational diagnostic AI with multimodal reasoning
Khaled Saab
CJ Park
Tim Strother
Jan Freyberg
David Barrett
Yong Cheng
David Stutz
Nenad Tomašev
Yash Sharma
Roma Ruparel
Abdullah Ahmed
Elahe Vedadi
Kimberly Kanada
Cían Hughes
Geoff Brown
Yang Gao
Sean Li
Sara Mahdavi
James Manyika
Kat Chou
Joelle Barral
Ali Eslami
Adam Rodman
Ryutaro Tanno
Nature Medicine (2026)
Preview abstract Real-world clinical practice is inherently multimodal, relying on the synthesis of patient history with visual information such as medical imagery and clinical documents. Although large language models (LLMs) have shown promise in diagnostic dialogue, their evaluation has been largely restricted to text-only interactions, failing to capture the complexity of modern remote care delivery. Here we introduce a multimodal extension of the Articulate Medical Intelligence Explorer (multimodal AMIE), capable of gathering, interpreting and reasoning about multimodal data within a diagnostic conversation. To achieve this, we developed a state-aware dialogue framework that dynamically guides history-taking based on diagnostic uncertainty and evolving patient states, emulating the structured reasoning of experienced clinicians. We evaluated this updated, state-aware version of multimodal AMIE against primary care physicians (PCPs) in a randomized, blinded exploratory study comprising 105 simulated telehealth consultations, which included dermatology photographs, electrocardiograms and clinical documents. As assessed by 18 specialist physicians, multimodal AMIE outperformed PCPs not only in diagnostic accuracy but also in conversation quality, including history-taking and empathy. Specifically, multimodal AMIE demonstrated superior performance on 29 of 32 evaluation axes, including seven of nine metrics that assess multimodal reasoning. These results validate the efficacy of state-aware reasoning in bridging the gap between text and visual information and demonstrate the potential for artificial intelligence (AI) systems to augment clinicians in complex, multimodal diagnostic settings. View details
SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment
Joe Breda
Fadi Yousif
Beszel Hawkins
Marinela Cotoi
Miao Liu
Ray Luo
Sam Schmidgall
Girish Narayanswamy
Samuel Solomon
Max Xu
Longfei Shangguan
Bhavna Daryani
Buddy Herkenham
Cara Tan
Mark Malhotra
Shwetak Patel
Zach Wasson
Dimitrios Antos
Bob Lou
Matthew Thompson
Jonathan Richina
Anupam Pathak
Nichole Young-Lin
Jake Sunshine
Daniel McDuff
Arxiv preprint, 2605.040 (2026) (to appear)
Preview abstract Language models excel at diagnostic assessments on curated medical case-studies and vignettes, performing on par with, or better than, clinical professionals. However, existing studies focus on complex scenarios with rich context making it difficult to draw conclusions about how these systems perform for patients reporting symptoms in everyday life. We deployed SymptomAI, a set of conversational AI agents for end-to-end patient interviewing and differential diagnosis (DDx), via the Fitbit app in a study that randomized participants (N=13,917) to interact with five AI agents. This corpus captures diverse communication and a realistic distribution of illnesses from a real world population. A subset of 1,228 participants reported a clinician-provided diagnosis, and 517 of these were further evaluated by a panel of clinicians during over 250 hours of annotation. SymptomAI DDx were significantly more accurate (OR = 2.56, p < 0.001) than those from independent clinicians given the same dialogue in a blinded randomized comparison. Moreover, agentic strategies which conduct a dedicated symptom interview that elicit additional symptom information before providing a diagnosis, perform substantially better than baseline, user-guided conversations (p < 0.001). An auxiliary analysis on 1,509 conversations from a general US population panel validated that these results generalize beyond wearable device users. We used SymptomAI diagnoses as labels for all 13,917 participants to analyze over 500,000 days of wearable metrics across nearly 400 unique conditions. We identified strong associations between acute infections and physiological shifts (e.g., OR > 7 for influenza). While limited by self-reported ground truth, these results demonstrate the benefits of a dedicated and complete symptom interview compared to a user-guided symptom discussion, which is the default of most consumer LLMs. View details
A large language model for complex cardiology care
Jack W O’Sullivan
Khaled Saab
Daniel K. Amponsah
Evaline Cheng
Yong Cheng
Emily Chu
Yaanik Desai
Aly Elezaby
Muhammad Fazal
Tasmeen Hussain
Sneha S. Jain
Daniel Seung Kim
Roy Lan
Jiwen Li
Wilson Tang
Natalie Tapaskar
Victoria Parikh
Ryan Sandoval
Gabriela Spencer-Bonilla
Bryan Wu
Kavita Kulkarni
Philip Mansfield
Juro Gottweis
Joelle Barral
Ryutaro Tanno
Sara Mahdavi
Euan Ashley
Nature Medicine (2026)
Preview abstract The scarcity of subspecialist medical expertise poses a considerable challenge for healthcare delivery. This issue is particularly acute in cardiology, where timely, accurate management determines outcomes. We explored the potential of Articulate Medical Intelligence Explorer (AMIE), a large language model-based experimental medical artificial intelligence system, to augment clinical decision-making in this challenging context. We conducted a randomized controlled trial comparing large language model-assisted care with the usual care of complex patients suspected of having a genetic cardiomyopathy, and we curated a real-world dataset of complex cases from a subspecialist cardiology practice. Nine participating general cardiologists were provided with access to both clinical text reports and raw diagnostic data—including electrocardiograms, echocardiograms, cardiac magnetic resonance imaging scans and cardiopulmonary exercise testing—and were randomized to manage these cases, either with or without assistance from AMIE. We developed a ten-domain evaluation rubric used by three blinded subspecialists to evaluate the quality of triage, diagnosis and management. In our randomized controlled trial with retrospective patient data, subspecialists favored large language model-assisted responses overall, and for the management plan and diagnostic testing domains, with the remaining domains considered a tie. Overall, subspecialists preferred AMIE-assisted cardiology assessments 46.7% of the time, compared with 32.7% for cardiologists alone (P = 0.02), with 20.6% rated as a tie. Subspecialists also quantified errors, extra and missing content, reasoning and potential bias. Cardiologists alone had more clinically significant errors (24.3% versus 13.1%, P = 0.033) and more missing content (37.4% versus 17.8%, P = 0.0021) than cardiologists assisted by AMIE. Lastly, cardiologists who used AMIE reported that AMIE helped their assessment more than half the time (57.0%) and saved time in 50.5% of cases. View details
Complementary human-AI clinical reasoning in ophthalmology
Mertcan Sevgi
Fares Antaki
Abdullah Zafar Khan
Ariel Yuhan Ong
David Adrian Merle
Kuang Hu
Shafi Balal
Sophie-Christin Kornelia Ernst
Josef Huemer
Gabriel T. Kaufmann
Hagar Khalid
Faye Levina
Celeste Limoli
Ana Paula Ribeiro Reis
Samir Touma
Khaled Saab
Ryutaro Tanno
Yong Cheng
Sara Mahdavi
Elahe Vedadi
David Stutz
Pearse Keane
arXiv (2025)
Preview abstract Vision impairment and blindness are a major global health challenge where gaps in the ophthalmology workforce limit access to specialist care. We evaluate AMIE, a medically fine-tuned conversational system based on Gemini with integrated web search and self-critique reasoning, using real-world clinical vignettes that reflect scenarios a general ophthalmologist would be expected to manage. We conducted two complementary evaluations: (1) a human-AI interactive diagnostic reasoning study in which ophthalmologists recorded initial differentials and plans, then reviewed AMIE's structured output and revised their answers; and (2) a masked preference and quality study comparing AMIE's narrative outputs with case author reference answers using a predefined rubric. AMIE showed standalone diagnostic performance comparable to clinicians at baseline. Crucially, after reviewing AMIE's responses, ophthalmologists tended to rank the correct diagnosis higher, reached greater agreement with one another, and enriched their investigation and management plans. Improvements were observed even when AMIE's top choice differed from or underperformed the clinician baseline, consistent with a complementary effect in which structured reasoning support helps clinicians re-rank rather than simply accept the model output. Preferences varied by clinical grade, suggesting opportunities to personalise responses by experience. Without ophthalmology-specific fine-tuning, AMIE matched clinician baseline and augmented clinical reasoning at the point of need, motivating multi-axis evaluation, domain adaptation, and prospective multimodal studies in real-world settings. View details
Exploring large language models for specialist-level oncology care
Vikram Dhillon
Polly Niravath
Preethi Prasad
Khaled Saab
Ryutaro Tanno
Yong Cheng
Hanh Mai
Ethan Burns
Zainub Ajmal
Kavita Kulkarni
Philip Mansfield
Joelle Barral
Juro Gottweis
Sara Mahdavi
NEJM AI (2025)
Preview abstract Large language models have shown rapid progress in encoding clinical knowledge and demonstrating clinical reasoning. However, their capabilities in subspecialty or complex medical settings remain underexplored. In this work, we probe the performance of Articulate Medical Intelligence Explorer (AMIE), a conversational diagnostic AI system in the subspecialty of breast oncology care without specific fine-tuning to this challenging domain. To perform this evaluation, we curated a set of 60 synthetic breast cancer vignettes representing a range of treatment-naive, treatment-refractory, and rare histology cases encountered in a community-based breast oncology clinic. We developed a detailed clinical rubric for evaluating management plans, including axes such as the quality of case summarization, safety of the proposed care plan, and recommendations for treatment (i.e., chemotherapy, radiotherapy, surgery, and hormonal therapy). To improve performance, we enhanced AMIE with the inference-time ability to perform web search retrieval to gather relevant and up-to-date clinical knowledge and refine its responses with a multistage, self-critique pipeline. We compare the response quality of AMIE with that of internal medicine trainees, oncology fellows, and general oncology attendings under both automated and specialist clinician evaluations. Although our evaluations were limited to a few physicians, AMIE outperformed trainees and fellows, demonstrating the potential of the system in this important domain. However, AMIE’s performance was overall inferior to that of attending oncologists, suggesting that further prospective research is needed. View details
Towards physician-centered oversight of conversational diagnostic AI
Elahe Vedadi
David Barrett
Shashir Reddy
Roma Ruparel
Tim Strother
Ryutaro Tanno
Yash Sharma
Jihyeon Lee
Cían Hughes
Jan Freyberg
Khaled Saab
Nenad Tomašev
Kavita Kulkarni
Sara Mahdavi
Kelvin Guu
Joelle Barral
James Manyika
Kat Chou
Adam Rodman
David Stutz
arXiv (2025)
Preview abstract Recent work has demonstrated the promise of conversational AI systems for diagnostic dialogue. However, real-world assurance of patient safety means that providing individual diagnoses and treatment plans is considered a regulated activity by licensed professionals. Furthermore, physicians commonly oversee other team members in such activities, including nurse practitioners (NPs) or physician assistants/associates (PAs). Inspired by this, we propose a framework for effective, asynchronous oversight of the Articulate Medical Intelligence Explorer (AMIE) AI system. We propose guardrailed-AMIE (g-AMIE), a multi-agent system that performs history taking within guardrails, abstaining from individualized medical advice. Afterwards, g-AMIE conveys assessments to an overseeing primary care physician (PCP) in a clinician cockpit interface. The PCP provides oversight and retains accountability of the clinical decision. This effectively decouples oversight from intake and can thus happen asynchronously. In a randomized, blinded virtual Objective Structured Clinical Examination (OSCE) of text consultations with asynchronous oversight, we compared g-AMIE to NPs/PAs or a group of PCPs under the same guardrails. Across 60 scenarios, g-AMIE outperformed both groups in performing high-quality intake, summarizing cases, and proposing diagnoses and management plans for the overseeing PCP to review. This resulted in higher quality composite decisions. PCP oversight of g-AMIE was also more time-efficient than standalone PCP consultations in prior work. While our study does not replicate existing clinical practices and likely underestimates clinicians' capabilities, our results demonstrate the promise of asynchronous oversight as a feasible paradigm for diagnostic AI systems to operate under expert human oversight for enhancing real-world care. View details
Towards conversational diagnostic artificial intelligence
Khaled Saab
Jan Freyberg
Ryutaro Tanno
Amy Wang
Brenna Li
Nenad Tomašev
Karan Singhal
Yong Cheng
Le Hou
Albert Webson
Kavita Kulkarni
Sara Mahdavi
Juro Gottweis
Joelle Barral
Kat Chou
Nature (2025)
Preview abstract At the heart of medicine lies physician–patient dialogue, where skillful history-taking enables effective diagnosis, management and enduring trust. Artificial intelligence (AI) systems capable of diagnostic dialogue could increase accessibility and quality of care. However, approximating clinicians’ expertise is an outstanding challenge. Here we introduce AMIE (Articulate Medical Intelligence Explorer), a large language model (LLM)-based AI system optimized for diagnostic dialogue. AMIE uses a self-play-based simulated environment with automated feedback for scaling learning across disease conditions, specialties and contexts. We designed a framework for evaluating clinically meaningful axes of performance, including history-taking, diagnostic accuracy, management, communication skills and empathy. We compared AMIE’s performance to that of primary care physicians in a randomized, double-blind crossover study of text-based consultations with validated patient-actors similar to objective structured clinical examination. The study included 159 case scenarios from providers in Canada, the United Kingdom and India, 20 primary care physicians compared to AMIE, and evaluations by specialist physicians and patient-actors. AMIE demonstrated greater diagnostic accuracy and superior performance on 30 out of 32 axes according to the specialist physicians and 25 out of 26 axes according to the patient-actors. Our research has several limitations and should be interpreted with caution. Clinicians used synchronous text chat, which permits large-scale LLM–patient interactions, but this is unfamiliar in clinical practice. While further research is required before AMIE could be translated to real-world settings, the results represent a milestone towards conversational diagnostic AI. View details
Towards accurate differential diagnosis with large language models
Daniel McDuff
Amy Wang
Karan Singhal
Yash Sharma
Kavita Kulkarni
Le Hou
Yong Cheng
Sara Mahdavi
Sushant Prakash
Anupam Pathak
Shwetak Patel
Ewa Dominowska
Juro Gottweis
Joelle Barral
Kat Chou
Jake Sunshine
Nature (2025)
Preview abstract A comprehensive differential diagnosis is a cornerstone of medical care that is often reached through an iterative process of interpretation that combines clinical history, physical examination, investigations and procedures. Interactive interfaces powered by large language models present new opportunities to assist and automate aspects of this process. Here we introduce the Articulate Medical Intelligence Explorer (AMIE), a large language model that is optimized for diagnostic reasoning, and evaluate its ability to generate a differential diagnosis alone or as an aid to clinicians. Twenty clinicians evaluated 302 challenging, real-world medical cases sourced from published case reports. Each case report was read by two clinicians, who were randomized to one of two assistive conditions: assistance from search engines and standard medical resources; or assistance from AMIE in addition to these tools. All clinicians provided a baseline, unassisted differential diagnosis prior to using the respective assistive tools. AMIE exhibited standalone performance that exceeded that of unassisted clinicians (top-10 accuracy 59.1% versus 33.6%, P = 0.04). Comparing the two assisted study arms, the differential diagnosis quality score was higher for clinicians assisted by AMIE (top-10 accuracy 51.7%) compared with clinicians without its assistance (36.1%; McNemar’s test: 45.7, P < 0.01) and clinicians with search (44.4%; McNemar’s test: 4.75, P = 0.03). Further, clinicians assisted by AMIE arrived at more comprehensive differential lists than those without assistance from AMIE. Our study suggests that AMIE has potential to improve clinicians’ diagnostic reasoning and accuracy in challenging cases, meriting further real-world evaluation for its ability to empower physicians and widen patients’ access to specialist-level expertise. View details

Contact us

If you are a research-focused health organization interested in novel research partnerships with the AMIE team, please reach out to the project leads, Mike Schaekermann, Sunny Virmani and Cameron Chen.

×