Google’s medical AI helped doctors prepare for visits in small Lancet study
The supervised study found useful summaries and strong diagnostic recall, but it did not test whether AI improved care over the usual workflow.
Listen to this story
The audio brief
Story brief
3 key pointsA Lancet study published October 8, 2026, evaluated Google’s AMIE chatbot as a pre-visit tool at Beth Israel Deaconess Medical Center. Clinicians said its summaries influenced their approach in more than half of cases, but the single-clinic study had no control group and cannot show that AMIE improved care. The results support further testing of AI-assisted visit preparation, not a conclusion that it can replace clinical judgment or deliver better outcomes.
- 01
Of 100 adults who completed an AMIE chat, 98 attended their scheduled appointment; a physician supervised every interaction, and none required a safety stop.
- 02
AMIE included the eventual diagnosis among its top seven possibilities in 90% of cases, among its top three in 75%, and as its first choice in 56%.
- 03
Blinded evaluators found no significant difference in diagnosis-list quality or care-plan appropriateness and safety, while doctors’ plans were more practical and cost-effective.
Doctors found an AI chatbot’s summaries useful for preparing patient visits in 75% of cases in a small, closely supervised clinical study. Google said on October 8, 2026, that the research, led with Beth Israel Deaconess Medical Center and published that day in The Lancet, evaluated its research chatbot AMIE before urgent-care appointments.
The journal publication follows an earlier public account of the same study. A preprint appeared on March 9, and Google Research described the results on March 11.
A patient chat with a physician watching
Patients had new, non-emergency complaints and used a secure web link to chat with AMIE before seeing their clinician. The system gathered their medical history and presented possible diagnoses to discuss with their provider. With patient consent, clinicians received the transcript and a summary before the appointment.
Patients were invited during appointment booking and given time to review the study’s approved protocols. They were told that declining would not affect their care. The appointments could take place in person or through telehealth.
A physician watched each interaction through a live video call with screen-sharing. Supervisors were trained to stop a conversation under four predefined conditions:
- Immediate concern about harm to the patient or others.
- Significant emotional distress related to the AI interaction.
- Potential clinical harm identified during the conversation.
- An explicit patient request to end the session.
No interaction required a safety stop. Google’s March account also clarifies the participant count: 100 adults completed the chat, and 98 attended their scheduled appointments.
A diagnosis list is not one correct answer
The headline diagnostic result measures whether AMIE included the eventual diagnosis among several possibilities—not whether its first answer was correct. Researchers compared its differential diagnosis, a list of possible explanations for symptoms, with the final diagnosis established through chart review eight weeks after the visit.
AMIE included the final diagnosis within its top seven possibilities in 90% of evaluated cases.
The final diagnosis appeared among AMIE’s top three possibilities in 75% of cases.
AMIE’s single most likely diagnosis matched the final diagnosis in 56% of evaluated cases.
The authors also examined 46 patients whose final diagnosis was confirmed by a diagnostic test, such as laboratory work or imaging. Diagnostic accuracy remained high in that group. It tended to be higher when the final diagnosis rested on the doctor’s judgment without further testing.
For the comparison with doctors, clinical evaluators who had not conducted the consultations reviewed diagnosis lists and care plans. The assessment was blinded and randomized. Three evaluators graded each case, with the middle rating used to combine their judgments.
Blinded evaluators found no significant differences in overall diagnosis-list quality or the appropriateness and safety of management plans. That does not establish equivalence. Doctors, however, produced more practical and cost-effective plans. AMIE lacked access to patients’ electronic health records and could not perform a physical examination or use inputs beyond text.
Useful preparation, not proven better care
Google’s October announcement says the summaries influenced clinicians’ approach to care in more than half of cases. In earlier interviews, clinicians described moving from gathering information to verifying it, allowing more collaborative conversations. Patients reported high satisfaction, and their attitudes toward AI improved after the chat.
Patient surveys measured attitudes before the chat, afterward and again after the doctor’s appointment. The improvement persisted after that visit. In surveys and interviews, patients generally described AMIE as polite and effective at explaining medical conditions.
The study had no control group, so it cannot establish that adding AMIE improved care compared with the usual workflow. It also involved one clinic, with participants tending younger than its overall urgent-care population. The authors call for larger studies with controlled comparisons to measure the impact of patient-facing AI.
Sources
- arxiv.orgA prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
- research.googleExploring the feasibility of conversational diagnostic AI in a real-world clinical study
- blog.googleStudy in The Lancet suggests AI could improve patient-physician relationships.
Reader comments
Newest comments first. Replies stay oldest first.