AI News · Good news ·

Google and BIDMC study AMIE chatbot with 98 primary care patients

Google and BIDMC study AMIE chatbot with 98 primary care patients

Researchers at Google and Beth Israel Deaconess Medical Center published a real-world study of Google's AMIE diagnostic AI chatbot in The Lancet. At BIDMC's ambulatory primary care clinic, 98 patients consulted AMIE before urgent care visits. Supervising physicians monitored the interactions in real time. The study examined patient-facing AI in a clinical setting rather than relying only on simulations.

Key points

  • 98 patients completed AMIE consultations and urgent primary care appointments in a real-world BIDMC study.
  • Physicians monitored every conversation; none required a safety-stop intervention.
  • Supervisors identified one hallucination and provided clinical clarification in five cases.
  • Among 44 reviewed cases, 75% of clinicians said AI-generated records helped them prepare.
  • The study evaluated feasibility, not improved health outcomes or reduced workload.

What happened: Researchers at Google and Beth Israel Deaconess Medical Center (BIDMC) have published a real-world study of Google’s AMIE diagnostic AI chatbot in The Lancet. At BIDMC’s ambulatory primary care clinic, 98 patients completed a chatbot consultation before an urgent primary care visit, with physicians monitoring every conversation in real time. The research moves beyond simulated encounters to examine how patient-facing AI works in routine care, but it tested feasibility rather than whether the technology improves health outcomes.

The details: AMIE, short for Articulate Medical Intelligence Explorer, is a research system that asks patients about symptoms, gathers medical histories and offers possible diagnoses to discuss with a doctor. Patients used a secure text-chat interface from home after booking an appointment. The study ran from April through November 2025 and enrolled 114 patients, of whom 98 completed both the AI interaction and their appointment. The system also generated notes for clinicians to review before seeing patients.

The details: According to BIDMC’s announcement on Newswise, a board-certified internal medicine physician monitored each conversation and could intervene over potential harm, emotional distress, a request to stop or other safety concerns. None of the 98 completed encounters required a safety-stop intervention. That does not mean the conversations were error-free: supervising physicians identified one hallucination and offered additional clinical clarification in five cases. Google also reported that AMIE’s differential diagnoses, the possible explanations for a patient’s symptoms, matched doctors’ final diagnoses 90% of the time.

Who it affects: For healthcare teams assessing AI for appointment preparation, the workflow findings offer an early indication of usefulness. Clinicians reviewed an AI-generated transcript or notes before the visit in 44 cases. Among those cases, 75% said the review helped them prepare, and 57% said it may have influenced their clinical approach, BIDMC reported. One provider rated the interaction somewhat harmful, citing concern that a patient may have experienced anxiety after AMIE listed lymphoma as a possible diagnosis.

Background: BIDMC framed the research against pressure on primary care from an aging population, workforce changes and increasingly complex medical needs. Although chatbots have performed well in simulations, the hospital said little is known about their performance with real patients. Participants generally rated AMIE favorably for listening, explaining information and helping them feel at ease. Their attitudes toward AI improved after using it and remained elevated after the clinician visit. Concerns persisted about confidentiality and the chatbot’s honesty and trustworthiness.

What to watch: The central deployment condition is physician oversight: these findings describe monitored conversations, not independent AI care. For healthcare buyers, the study offers evidence about patient engagement and supervised workflow feasibility, rather than proof of better outcomes or reduced workload. Google said larger clinical trials are needed to assess patient-facing AI at scale. Whether the approach improves health outcomes or eases staffing pressure remains to be established.

Our take

Real-patient studies offer more relevant evidence for healthcare buyers than simulated benchmarks. The physician supervision is an important condition when assessing how these findings might translate into deployment.

Sources