An AI proposes brain hypotheses and tests them in the scanner
Generative causal testing turns opaque predictions of language-related brain responses into hypotheses tested with new stories and fMRI.
On June 13, 2026, a team from Microsoft Research, the University of California, Berkeley, UCSF, and other institutions updated work on a method called generative causal testing, or GCT. An artificial-intelligence system can successfully predict which parts of the brain will respond to a story and still fail to explain what it has learned.
The work is not intended to diagnose disease or turn a scanner into a mind-reading device. It addresses a narrower question in language neuroscience: when a cortical region responds while someone listens to a narrative, which feature of language is driving that response? And, crucially, can that explanation be tested in a new experiment?
From an opaque prediction to a hypothesis
Brain encoding models can associate language-model representations with BOLD signals measured by functional magnetic resonance imaging, or fMRI. They can help predict the response of tiny volumes of brain tissue, called voxels, to a language stimulus. But the internal representations of a large language model do not arrive with a simple scientific label.
GCT adds a translation step and a test. It starts with a model that predicts a voxel’s response well, finds text fragments expected to produce a high response, and uses a language model to summarize that pattern as a short verbal hypothesis. The phrase may point to a kind of semantic content; it does not turn a brain area into a box with one fixed meaning.
Then comes the decisive step: the system generates new narratives designed around that hypothesis, and researchers present them to participants in a scanner. If the region responds more strongly to the passage designed for it than to other passages, the explanation gains experimental support. In this setup, AI does not replace measurement; it helps propose an experiment that can fail.
What the study tested
In the publicly available work, the authors fit voxel-level models with data from three people who listened to 20 hours of stories. For a follow-up experiment, they selected 17 voxels with good predictive performance for each participant and created synthetic narratives intended to drive them. The paper reports that 41 of the 51 tested voxels showed a response above their baseline during the paragraphs designed for them.
The team also applied the method to regions of interest. One benefit of the approach is that it can distinguish areas that look similar when described with a broad label. In tests involving regions associated with places, for example, the authors looked for more precise descriptions that could separate them with different stimuli. That result does not make a short phrase a final definition of a brain area; it makes the phrase a prediction that an experiment can challenge or refine.
An important boundary
Caution is essential. The study concerns language responses under controlled conditions, with a small sample and a specific methodology; it does not show that a model understands the brain or can infer a person’s private thoughts. The authors themselves describe the explanations as partial characterizations: several semantic features may contribute to the response of the same voxel or region.
Its most interesting contribution is methodological. Instead of celebrating a black box because it predicts well, GCT asks it to produce an understandable hypothesis, design an intervention, and face an independent measurement. For brain research, that discipline could help connect models that handle large amounts of data with theories that other scientists can debate, reproduce, and improve.
The GCT preprint says the work has been accepted in Nature Neuroscience. Its immediate value, however, lies in the procedure: using generative models to speed up an experimental question, not to close scientific debate.
Sources for this piece
This piece draws on 5 primary source(s), gathered during reporting.
- Microsoft Research: Understanding the brain with AI-driven explanations and experiments
- Microsoft Research publication: Generative causal testing
- ArXiv: Generative causal testing to bridge data-driven models and scientific theories in language neuroscience
- bioRxiv/PubMed: Evaluating scientific theories as predictive models in language neuroscience
- Datos y límites que se mantendrán en la pieza
This article was produced with artificial intelligence under human editorial oversight.