From model to scanner: making an AI explanation falsifiable
Generative causal testing turns opaque patterns into hypotheses and new fMRI experiments. Its value lies in closing the loop; its limits rule out calling it an autonomous scientist.
A model can predict which part of the brain will respond to a story without explaining why. It can also produce a persuasive phrase that does not describe the pattern it learned. On 13 June 2026, a team from the University of Texas at Austin, Microsoft Research, UC San Francisco and UC Berkeley updated its study of generative causal testing (GCT), a method that turns patterns from a brain-encoding model into short explanations and subjects them to a new functional magnetic resonance imaging experiment.
This is not an AI system independently “discovering how the brain works.” Researchers chose the data, models, voxels, comparisons and statistical criteria, then measured real people in a scanner. The contribution is a loop: prediction, explanation, intervention and test. Recognising those four parts helps distinguish an AI-assisted scientific hypothesis from an attractive interpretation added after the result.
Part one: a model that predicts outside its training data
The work begins with encoding models. A person listens to stories while fMRI records changes in the BOLD signal from small volumes called voxels. Internal representations from a language model are aligned in time with that signal, and a regression learns to predict each voxel’s response. The important evaluation uses held-out stories: fitting the training observations does not show that the model has captured a general pattern.
The authors used roughly twenty hours of listening data per participant for three people and built models from two different families, LLaMA and OPT. Two model families do not eliminate error, but they make it possible to ask whether explanations remain stable under a consequential modelling choice. The team continued only with voxels whose predictive performance passed a threshold and selected seventeen per person for the follow-up experiment. The relevant denominator is not “the whole brain”; it is regions deliberately chosen because their models already worked relatively well.
That selection teaches a transferable lesson. When a tool explains only its best cases, the success rate cannot be extended to the discarded ones. Before celebrating a percentage, reconstruct the funnel: how many items existed, which were excluded, under what rule and over which set the final result was calculated.
Part two: turn a black box into an operational sentence
For each voxel model, GCT searched for short fragments that produced high predictions. GPT-4 summarised those fragments into candidates such as “food preparation” or “location names.” It then generated synthetic phrases associated with each candidate and retained the explanation that drove the predictive model most strongly. That tests consistency with the model, not yet with the brain.
Such a label is an operational hypothesis, not the final name of a brain function. A voxel aggregates signals from many neuronal populations and may respond to overlapping semantic dimensions. The paper stresses that a successful explanation shows a feature can modulate a response, not that it is the only feature capable of doing so. In extended analyses, models built from a single question captured about half the performance of full models, and most voxels improved when multiple features were allowed.
Wording matters. “This region responds to food” is too strong; “text about preparing food should raise the measured response in this voxel relative to comparison passages” can fail. A scientific explanation specifies an observation, an intervention, an expected direction and a reference against which it will be judged.
Part three: intervene with a new story
The language model wrote coherent narratives in which each paragraph was designed around a different explanation. The same participants returned to the scanner and heard those stories. For each voxel, researchers compared its signal during the target paragraph with its mean signal during the other paragraphs. Of 51 selected voxels — 17 in each of three people — 41 showed an increase; the reported average effect was 0.198 standard deviations per measurement interval above baseline.
“Causal” here refers to that controlled intervention: changing the semantic content of the stimulus produced a measurable change consistent with the hypothesis. It does not mean the study identified the entire neuronal chain, that the phrase is the only possible cause or that a region has one function. fMRI records a BOLD signal, a haemodynamic change linked to blood oxygenation, rather than the activity of each neuron. The authors themselves identify the technique’s spatial and temporal granularity as a limit on explanation.
Nor is it enough for a paragraph to activate its target. Related concepts can activate several voxels. The study found off-diagonal structure: directions, measurements and locations, for example, share content. In an adversarial test of three place-related regions, the method separated two with more specific stimuli but failed to isolate the third. That failure is informative: it marks where the hypothesis or measurement still lacks discrimination.
Which result belongs to the brain and which to the model?
GCT contains two different tests. The first checks whether generated text activates the encoding model; it may fail because of a poor explanation or unsuitable stimulus. The second checks whether the text activates the measured brain signal; it may fail because the predictor is inaccurate, fMRI is noisy or the hypothesis does not describe selectivity. Conflating them turns internal validation into a biological discovery.
In the failures the team analysed, paragraphs generally matched their explanation and raised the model’s prediction, yet they did not always raise the measured signal. The best indicator of experimental success was stability: agreement between rankings produced by the LLaMA- and OPT-based models. High prediction and cross-model stability together offered more confidence than a one-off generated explanation.
The pattern applies outside neuroscience. If a model proposes a rule about proteins, weather or materials and that same model scores its own generated example, the loop remains closed around its assumptions. A strong test requires a measurement that does not depend on the score: a physical assay, newly observed data or an independent instrument.
Limits contained in the primary source
The sample comprised three healthy participants listening to English-language stories. The models were trained within a particular domain, and the authors warn that stimuli outside that distribution are harder to predict. They also note that different encoding models can yield different, equally valid explanations. A phrase should therefore not be read as the unique description of a voxel or region.
The method favours one concise explanation and may miss polysemanticity. Accuracy depends on the starting model, prompt wording and story type: the appendices show variation between narratives and better performance from one of the two generation templates tested. One participant moved more in the scanner than the others, another reminder that biological measurement introduces noise after the automated stages.
The work provides materials for inspection. The project page traces the path from question-based models to GCT; the repository publishes code and reproduction routes; and the passive-listening dataset used to fit models is available through OpenNeuro. Availability is not an independent replication, but it enables auditing and makes replication possible.
A template for judging AI-generated hypotheses
The first question is predictive: does the model work on held-out data, and are selected cases disclosed? The second is semantic: does the explanation become a prediction that could be false? The third is experimental: is the proposed factor manipulated and the outcome measured outside the model? The fourth is comparative: do controls separate nearby concepts, and are failures reported? The fifth is cumulative: can another sample, method or laboratory repeat the test?
GCT passes several of those steps and leaves others open. It produces readable hypotheses, designs new stimuli, returns to the scanner, and documents limitations and code. It does not establish a complete theory of language in the brain, generalise from three people to a population or make the LLM an autonomous author. Its enduring value is methodological: an AI explanation deserves scientific attention when it leaves the model’s loop, risks a prediction against the world and preserves the outcome even when it contradicts the most appealing story.
Sources for this piece
This piece draws on 5 primary source(s), gathered during reporting.
- Microsoft Research: Understanding the brain with AI-driven explanations and experiments
- Microsoft Research publication: Generative causal testing
- ArXiv: Generative causal testing to bridge data-driven models and scientific theories in language neuroscience
- bioRxiv/PubMed: Evaluating scientific theories as predictive models in language neuroscience
- Datos y límites que se mantendrán en la pieza
This article was produced with artificial intelligence under human editorial oversight.