AI in Medicine and Biology: Follow the Evidence Chain
Several rungs separate an image or variant from a clinical decision: distinguish prediction, validation, and real benefit.
On July 30, 2026, saying that machine learning “is used in medicine” explains very little. A model may classify an image, flag a genetic variant, or propose a protein structure, but none of those outputs is automatically a diagnosis, treatment, or drug. A rigorous reading of any biomedical claim follows the entire chain: intended use, input data, reference standard, external validation, human intervention, and the outcome that actually matters to the patient or experiment.
Start with intended use and the action it triggers
The same score can serve different functions. A system might prioritize studies for earlier professional review, provide a second read, recommend another test, or issue an alert. Each position in the workflow changes the risk. If it merely orders a queue, an error changes timing; if it rules cases out, a false negative may prevent someone from reaching review. “Detects disease” is insufficient. A meaningful description names the population, user, moment, input, output, and subsequent action.
This distinction also separates research from a medical device. The FDA’s page of [AI-enabled medical devices] gathers products identified through United States marketing-authorization documents. Appearance on that list does not mean an agency certified general intelligence or approved the device for uses outside its indication. The specific authorization must be opened to learn which population, modality, and task were evaluated.
The FDA, Health Canada, and the United Kingdom’s MHRA published [Good Machine Learning Practice principles] spanning the product life cycle: multidisciplinary expertise, representative data, independence between training and test sets, the human-system interaction, clear information for users, and monitoring of deployed performance. The durable lesson is that the algorithm is only one component of a sociotechnical intervention.
Data and the reference standard determine what “correct” means
A large dataset can still be narrow. Images from one hospital, scanner manufacturer, or population contain regularities a model may exploit without learning the intended clinical signal. A random split does not prevent studies from the same patient, duplicates, or nearly identical acquisitions from landing on both sides. Separation must occur at the unit that could leak information; a transportability claim requires testing across other sites, periods, equipment, and groups.
What supplies the answer also matters. A label may come from pathology, a molecular test, follow-up, specialist consensus, or a routine report. These references are not interchangeable, and each can contain disagreement or error. A model trained to imitate historical reports learns documented practice, not a noise-free clinical truth. A responsible reader asks who labelled cases, what information they had, whether they were blinded to the model output, and how long it took for the outcome to become known.
Accuracy, sensitivity, and specificity describe different mistakes; prevalence changes the predictive value of an alert. Beyond discrimination, a clinical probability needs calibration: a 20% risk should correspond to roughly twenty events per hundred comparable cases. The [TRIPOD+AI] guidance calls for descriptions of data sources, participants, outcomes, development, evaluation, and availability. Following a guideline improves reporting transparency; it does not automatically turn a weak design into sound evidence.
Missing data are part of the system. An ordered test, a repeated image, or an empty field may reveal an earlier clinical decision, and the model can learn that trace. The signal disappears if a hospital changes its protocol. A report should therefore explain how missing values were handled, whether every input was available at the intended moment, and what the product does when quality is insufficient. Refusing an input can be safer than manufacturing false confidence.
Imaging and genomics share a method, not a meaning
In medical imaging, a network may segment a region, classify a study, detect a finding, or predict an outcome. Comparing it with “the doctor” erases tasks and conditions. The professional may have received medical history and earlier studies while the model saw only one image, or the reverse. A fair comparison keeps information equal, specifies whether it measures a human alone or a human-model team, and records time, referrals, unevaluable cases, and the effects of automation.
In genomics, DeepVariant provides a concrete example. It converts aligned reads around candidate variants into representations that a convolutional network classifies as genotypes. The [original Nature Biotechnology paper] describes calling single-nucleotide polymorphisms and small insertions or deletions. That does not mean the network diagnoses disease. Calling a variant is a technical stage that comes before annotation, evidence assessment, and interpretation alongside phenotype and family context.
Keeping those stages separate prevents scope inflation. Agreement with a benchmark evaluates variant calling under particular technologies, genomic regions, and filters. Clinical utility raises additional questions: was the variant detectable in that sample type, does it change classification, can it alter a decision, and does performance survive across populations and platforms? The chain should preserve uncertainty and provenance instead of turning every model output into a binary answer.
A predicted structure is not a discovered drug
AlphaFold showed how far a well-defined task can advance. In the blind CASP14 assessment, the system predicted protein structures with accuracy competitive with experimental structures in a majority of cases, according to the [Nature paper]. The precise wording matters: it predicts a structure from a sequence and supplies confidence estimates. It does not automatically reproduce dynamics, function, interaction in every cellular environment, or therapeutic response.
The EMBL-EBI [AlphaFold Protein Structure Database] exposes models and confidence signals. Low-confidence regions should not be read as observed coordinates, and a high-confidence prediction remains a computational hypothesis when an experiment depends on alternative states, ligands, or complexes. Its value is in prioritizing questions, designing assays, and supplying structural context—not abolishing crystallography, cryo-electron microscopy, or biochemical validation.
Drug discovery adds more links: representing molecules, predicting affinity or properties, screening candidates, synthesizing them, measuring activity, toxicity, and pharmacokinetics, and eventually studying safety and benefit in people. A graph neural network can learn patterns over atoms and bonds, but its evaluation depends on how chemical families were split and whether genuinely new structures were tested. A retrospective hit or docking score is not a medicine.
Evidence climbs from technical performance to real benefit
An offline result shows that an output agrees with a reference on held-out data. External validation asks whether it transports. A prospective study observes future inputs and workflows. An early clinical evaluation examines safety, human factors, and integration. Finally, a comparative trial may measure whether the intervention improves decisions, patient outcomes, or resource use against an alternative. Skipping rungs turns a possibility into a claim the evidence does not yet support.
The [DECIDE-AI] consensus guideline addresses early clinical evaluation and asks authors to describe the system and version, participants, users, workflow, modifications, errors, safety, and human factors. For randomized trials, [CONSORT-AI] extends the reporting standard with elements specific to AI interventions. These are maps of what must remain visible for the evidence to be judged.
Any biomedical headline can be tested with seven questions: what exactly was the task; who would use the output; which reference defined a correct result; how were data separated and represented; where was it validated; what happened when it failed; and whether the outcome was technical performance, a decision, or a benefit. If the story offers only an architecture and a percentage, most of the path is missing. Learning to identify the rung makes it possible to value genuine progress without turning a useful prediction into an imaginary cure.
This article was produced with artificial intelligence under human editorial oversight.