AGI in Medicine: A Promise That Requires Clinical Evidence
Medical AGI remains hypothetical. This guide shows how to demand intended use, external validation, clinical utility, human factors, and monitoring.
On July 30, 2026, discussing artificial general intelligence (AGI) in medicine requires separating two planes. In the present, AI systems are authorized or studied for specific healthcare uses: analyzing one imaging modality, estimating a risk, or supporting a defined decision. On the hypothetical plane, AGI might transfer knowledge across specialties, combine heterogeneous data, and handle unforeseen clinical tasks. Presenting the first as proof of the second confuses a bounded tool with general capability—and in healthcare, that confusion can affect decisions about real people.
The useful question is not whether a demonstration “looks medical,” but which clinical claim it supports and which evidence accompanies it. This guide teaches an evidence ladder: intended use, technical performance, external validation, interaction with professionals, patient benefit, and post-deployment monitoring. The AGI label does not allow any rung to be skipped.
Medicine evaluates uses, not abstract intelligence
The FDA’s AI-enabled medical device list links products to their public records and explains that review considers safety and effectiveness for the intended use. The agency also says the list is not comprehensive. The key is not the number of entries but their form: product, company, specialty, product code, and regulatory decision. Authorization does not certify “general intelligence”; it concerns a particular device and use.
That principle forces broad promises to be translated. “Detects disease” must become something like: identifies a specified finding, in a defined type of study, within a stated population and workflow. “Personalizes treatment” must specify the outcome, alternatives, clinical moment, and decision-maker. If that sentence cannot be written, there is not yet an intervention that can be evaluated.
First test: which population did the model see?
An internal result answers how a system performed on data similar to those used for development. Performance may fall when the hospital, device, protocol, prevalence, language, or patient profile changes. A training-test split is therefore necessary but insufficient: both sets may share the same institutional biases.
External validation uses genuinely independent data and should report who was included and excluded. TRIPOD+AI updates reporting guidance for diagnostic and prognostic prediction models, whether they use machine learning or regression. It directs attention to details that a headline number hides: participant source, outcome definition, available predictors, handling of missing data, and performance evaluation.
Prevalence and clinical spectrum also matter. Sensitivity and specificity alone do not tell a hospital how many alerts will be correct; predictive values change with the frequency of the condition. A model tested on clear cases versus healthy controls may perform worse among consecutive patients with ambiguous symptoms. The sample should resemble the difficult decisions the system will encounter, not merely the cases that were easiest to label.
Discrimination, calibration, and utility are not synonyms
A tool may rank higher-risk patients correctly while assigning poorly calibrated probabilities. If it predicts 30% for a group, calibration asks whether roughly that share experiences the outcome under the studied conditions. Discrimination asks whether those who will experience it are ranked ahead of those who will not. These are different properties, and both can change outside the development center.
Clinical utility adds another layer: which decision changes, and with what consequences? A threshold produces false positives and false negatives with different costs. Ordering a test earlier may help, create anxiety, consume resources, or trigger an unnecessary procedure. The metric must connect to that trade-off. A statistical improvement over an unsuitable reference does not show that patients live longer, suffer less, or receive safer care.
From the laboratory to the human team
A retrospective study can measure predictions without observing what happens when a tool enters a clinic. Live use introduces response times, ignored alerts, automation bias, duplicated work, and cases in which the clinician lacks data expected by the model. DECIDE-AI provides reporting guidance for early clinical evaluations of AI decision-support systems, including real-world performance, safety, and human factors.
This changes the unit of analysis. Comparing an algorithm and a physician as isolated rivals is not enough; the human-AI team must be evaluated inside the workflow. Who receives the output? Can it be challenged? What explanation does the user need? What happens if the system is unavailable? Is its influence on the decision recorded? The transparency principles from the FDA, Health Canada, and MHRA explicitly emphasize human-AI team performance and information relevant to device users.
A trial must describe the whole intervention
When the hypothesis is that AI improves care, a prospective comparison must describe both the model and how it is introduced. SPIRIT-AI extends guidance for protocols of trials involving an AI intervention. Its companion, CONSORT-AI, guides reporting of trial results. These are transparency guidelines, not automatic seals of quality, but they help reveal whether the setting, version, interaction, errors, or planned analysis is missing.
The distinction prevents misleading headlines. “Comparable to specialists” may describe a test outside clinical context, whereas “improves care” requires observing a decision or health outcome. And “works across several tasks” is still not AGI if each task was selected, trained, and validated within a closed repertoire. Transfer must be tested on tasks and conditions that are more than variants of the same exam.
A changing model needs change control
Conventional software can keep the evaluated version fixed. A system updated with new data raises a further question: does the modification preserve safety and performance? The FDA’s final guidance on predetermined change control plans, issued in August 2025, gives recommendations for planned modifications to AI-enabled medical device software functions. It does not authorize unconstrained learning; it asks what will change and how that change will be validated.
Monitoring should look for drift in inputs, outcomes, and use. A hospital may replace equipment, a clinical practice may alter the referred population, or a new policy may change the measured outcome. Owners, alert thresholds, review frequency, and a process for withdrawing or rolling back a version should be defined. “It keeps learning” is not a clinical advantage if nobody can reconstruct which version produced a recommendation.
Governance: broader capability widens the radius of harm
The more cross-cutting a system becomes, the more data, decisions, and teams it may affect. The WHO guidance on ethics and governance of AI for health sets out principles covering autonomy, safety, transparency, accountability, inclusion, and sustainability. Applying them requires mechanisms: consent and legal basis, restricted access, documentation, complaint pathways, subgroup evaluation, and a person or institution responsible for stopping the system.
Combining health records, genomics, images, and behavior may enrich an assessment, but it also multiplies surfaces for error and exposure. More modalities do not guarantee causal inference or personalized recommendations. Each source’s contribution must be tested against a simpler baseline, along with whether the gain justifies cost, inequity, brittleness, and privacy risk.
The checklist that survives the word AGI
For any project, readers can ask for seven elements: intended use; population and site; clinical comparator; discrimination and calibration; effect on decisions or outcomes; human role and failure handling; and version, monitoring, and change control. Each element should then be located in the protocol, study, regulatory decision, or manufacturer documentation—not in promotional copy.
The transferable skill is to distinguish technical breadth from clinical utility: an AI system may combine many modalities and still fail to demonstrate benefit for a single patient. If a system with genuinely general transfer ever arrives, medicine will need more evidence, not less, because every added capability creates additional failure conditions. The promise may be broad; the test must always be concrete.
This article was produced with artificial intelligence under human editorial oversight.