AMIE matches doctors in a simulated clinical follow-up study
Google's research system produced plans comparable to those of primary care doctors in simulated chat consultations. The study measured reasoning and guideline alignment, not health outcomes in patients.
A study published in Nature on June 17, 2026 compared AMIE, Google's conversational medical research system, with 21 primary care physicians in simulated follow-up scenarios. AMIE's plans were non-inferior across fifteen evaluation axes over three visits and received higher scores on several measures of precision and guideline alignment. The experiment did not treat real patients or measure improvement, complications or mortality.
That distinction changes the headline's scope. AMIE matched or exceeded physicians inside an objective structured clinical examination conducted by chat, with trained patient actors and specialist raters. The study did not establish equal quality in real disease management or safe unsupervised use. The authors explicitly conclude that the system is not ready for clinical care and requires prospective studies with patients, safety controls and ethical oversight.
Google framed the work as a move from one-time diagnostic conversation toward longitudinal management. The official Google Research post describes two agents: one maintains dialogue with the patient, while another reasons in the background over the case, drug formularies and hundreds of pages of authoritative clinical knowledge.
The question the study answered
The question was not whether an AI cures disease more effectively. It was whether AMIE could sustain successive consultations, incorporate information revealed at each visit and produce management plans comparable to those of certified doctors in a controlled environment. Researchers adapted a virtual OSCE, a clinical examination format based on standardized scenarios.
The study was randomized and blinded. It involved 21 primary care doctors and 21 validated patient actors distributed across India and Canada. Every actor completed the same scenario with AMIE and with a physician in randomized order. Thirty specialists rated the consultations. The paper documents 120 scenarios: twenty for validation and one hundred for evaluation.
Both comparators had access to a corpus of 627 guidelines: 527 NICE documents and one hundred BMJ Best Practice documents. Fifty from each collection were directly relevant to the scenarios, while the rest provided less relevant material. This detail prevents a simplistic explanation in which AMIE had documents and doctors did not. The protocol instructed both to use the same corpus.
What “non-inferior” means
A non-inferiority analysis asks whether a new system falls no more than a defined margin below a comparator. It does not mean the two are identical or that the tool is safe for general use. The result depends on the margin, axes, raters and setting. In this study, AMIE's plans were rated at least as well as physicians' plans across all fifteen axes and three visits.
For overall plan appropriateness, AMIE scored 95%, 96% and 98% across the visits, compared with 72%, 80% and 81% for physicians; the paper reports the values and statistical tests. Treatment recommendation precision also favored AMIE. These figures describe ratings of written plans, not shares of patients cured.
Guideline alignment contains another qualification. Selection of applicable guidelines was similar for AMIE and physicians. The advantage appeared in whether recommended treatments aligned with guidance and included explicit references. The system was evaluated on a task where it could consult the full corpus at test time. That is valuable evidence, but it cannot be extended to a consultation with unsuitable documents or incomplete information.
The architecture separates conversation from deliberation
AMIE combines a dialogue agent, responsible for real-time conversation and persistent state across visits, with a management reasoning agent, or Mx, which spends more computation analyzing the case and guidelines. Before sending a response, the dialogue agent revises it against quality criteria. Both agents are built on Gemini models.
The state retains a patient summary, differential diagnosis and current plan and is updated during the dialogue. This division lets conversation remain responsive while another component performs deeper deliberation. It also creates separate audit points: which facts enter state, what is lost during summarization, which source supports a recommendation and how disagreement between agents is resolved.
The architecture does not remove the risk of automating an error. Incorrect information in the summary can shape later visits. An unsuitable retrieved guideline can make a recommendation well-cited but still wrong. Traceability must connect each decision to case data, source document, version and reviewing professional.
Surrogate outcomes and clinical outcomes
The study measured conversation and plan quality through specialist ratings. These are surrogate outcomes: they estimate a relevant capability without directly observing effects on health. A precise plan may not be followed, a guideline-aligned recommendation may require adaptation, and a patient may respond unexpectedly.
There was no real continuity over months or years. The design revealed information incrementally across simulated visits, testing whether reasoning could update, but compressing the complexity of follow-up. Real care includes adherence, medication access, delayed tests, other professionals, guideline changes, comorbidity and shared decisions.
Actors allow repeatable cases and reduce variation, but they do not represent all clinical communication. Text chat does not reproduce physical examination, nonverbal signals, emergencies or every language and access barrier. The next step is therefore not a broader claim; it is a feasibility, safety and workflow study with real patients and clinicians.
A record for reading medical AI studies
Readers can use six fields. First, population: real patients, actors or records. Second, intervention: system version, data and tools. Third, comparator: who competed and with which resources. Fourth, outcome: plan rating, diagnosis, time or health. Fifth, horizon: one visit, simulated follow-up or real follow-up. Sixth, oversight: who reviews and remains accountable.
Applied here, the record reads: trained actors; a two-agent AMIE with a clinical corpus; certified doctors using the same corpus; specialist ratings; three simulated visits; and no care deployment. This sentence preserves the achievement without turning it into medical authorization.
Reproducibility also has documented limits. Nature publishes the 120 scenarios and AMIE outputs, but physician responses are shared for only two sample scenarios. Part of the RxQA set comes from OpenFDA and is available, while the British National Formulary portion cannot be distributed directly because of licensing. AMIE's code and weights are not open because of the risks of unmonitored medical use.
What a clinical study would need to show
A prospective evaluation must measure errors that matter to patients: contraindicated recommendations, missed tests, delayed referral, adverse events, comprehension, equity and extra clinician workload. It also needs an escalation protocol, version records, monitoring and a way to withdraw the system. Time comparisons are meaningful only when safety and quality are preserved.
AMIE is not a service from which a patient should request a treatment plan. It is a research system showing capability in a demanding simulation with meaningful transparency. The transferable skill is to separate population, intervention, comparator, outcome, horizon and oversight. With those six fields, “matches doctors” stops being a general promise and becomes what the evidence supports: a bounded experimental result that still has to pass a clinical test.
This article was produced with artificial intelligence under human editorial oversight.