AI can generate hypotheses fast; validating them is still the slow part
DeepMind argues that agents can shift science’s bottleneck from generating ideas to testing them with data, laboratories and review.
In July 2026, Google DeepMind published an analysis of a new bottleneck for science done with agents: generating ideas may become cheaper, but checking them still demands data, instruments, time, and review. The news is not that an agent does science by itself. It is that abundant hypotheses can increase, not reduce, the need for validation. The authors condense it into a phrase: agents are "conjecture machines," making ideas abundant and cheap, while "refutations remain physical and institutional — and so, costly and slow."
Four steps you cannot skip
The useful distinction has four rungs, and confusing them is where almost all the hype is born. A hypothesis is an explanation or possibility worth testing. A prediction specifies what should be observed if it holds. A laboratory test measures something under defined methods and conditions. A replicated result shows that other teams, on their own, obtain compatible evidence. Calling the first rung a "discovery" erases the next three — and it is exactly the error DeepMind's analysis asks not to make with its own systems.
DeepMind's policy piece is, by its own description, a policy and design proposal, not a demonstration of autonomous discovery. It acknowledges the limit plainly: an agent "cannot say definitively whether it actually works." That honesty is what separates the piece from a brochure — and what is worth demanding of anyone announcing "AI science." The document adds two ideas that give the measure of that caution: "epistemic humility," that models should "know when they don't know, and say so," and "scaffolding" as a harness that exposes the agent's reasoning instead of leaving it a black box. A system that shows its steps can be audited; one that only hands over a polished conclusion cannot.
How the conjecture machine works
The system that illustrates the mechanism is Co-Scientist, described by DeepMind as a multi-agent system built with Gemini that generates, debates, and evolves hypotheses. It is not a model that answers, but a coalition of agents with distinct roles: one generates proposals grounded in the literature; another clusters them to cover a diverse space; a reflection agent acts as a "virtual peer reviewer" and critiques them; a ranking agent runs an "idea tournament" with pairwise comparisons and an Elo-style score — the same principle as AlphaGo and AlphaStar; an evolution agent combines and refines the best; and a meta-review agent synthesizes the debate. A supervisor coordinates the coalition as an adaptive planner.
That design is the instructive part, because it explains what improves and what does not. An internal tournament between agents can refine a candidate list far better than a single model firing off ideas. But it is a debate on paper: it orders conjectures, it does not test them against the world. The order itself says so — generate, debate, evolve — and at no point does "measure in a laboratory" appear. That phase remains human, physical, and slow.
That division explains a figure from DeepMind itself that, read properly, reinforces the thesis rather than contradicting it: Co-Scientist was tested with researchers from more than one hundred institutions, among them Stanford, MIT, Cambridge, the University of Edinburgh, and Calico. A hundred labs are not needed to generate hypotheses — the model does that alone; they are needed for the other side, checking them. The number does not measure the agent's power: it measures the size of the bottleneck the agent exposes. The more conjectures a machine produces, the more human hands and instruments are needed to separate the good from the merely plausible.
DeepMind's examples, read with the ladder
The analysis boasts concrete cases, and they are worth reading with the four rungs in hand, because all of them come from the lab itself. DeepMind states that, on liver fibrosis, a Co-Scientist candidate "blocked 91% of a scarring-linked response in lab tests," and that two of its picks not only halted fibrosis but promoted cell regeneration, beating the experts' selection. That is a laboratory-test result — the third rung — not an independent replication or an approved drug. The most-cited case is that of microbiologist José Penadés: DeepMind holds that Co-Scientist reproduced in two days an unpublished hypothesis from his team on antibiotic resistance, work that had taken them years. It is impressive, but it is worth seeing what it exactly proves: that the system generated a good conjecture that humans had already validated on their own — that is, it shines at the cheap rung, generation, not the expensive one.
The analysis adds others: Aletheia, a proof generator paired with a natural-language verifier, solved by DeepMind's account six of ten unpublished research problems in an inaugural proof challenge; and AlphaEvolve reportedly helped design Google's next-generation TPU chips and in work with mathematician Terence Tao. These are the vendor's claims about its own systems. That does not make them false; it places them where they belong: on the side of the powerful conjecture, awaiting the slow side of validation.
What the document asks of science funders
The policy part is concrete and, again, consistent with the thesis: if ideas get cheaper, invest in the expensive thing. DeepMind asks for three. First, universal access to leading agents through public-private partnerships, a priority it compares to the historical challenge of providing access to supercomputers. Second, "agent-ready" data: open datasets exposed through documented APIs, with privacy-preserving solutions for sensitive genomic or virological data. Third, validation infrastructure — laboratories and automation — and it notes that the U.S. National Science Foundation allocated 100 million dollars to a national network of distributed facilities, alongside examples such as the UK's Materials Innovation Factory, worth 81 million pounds, and the U.S. Genesis Mission, which connects Department of Energy labs with academia. It adds a peer-review reform, with agents that help detect errors and watermarking techniques to declare AI use.
There is also a safety detail the reader should register: DeepMind says it added classifiers to flag unethical research goals, following evaluations in the chemical, biological, radiological, and nuclear weapons domain. A machine that proposes hypotheses fluently also proposes the dangerous ones; the brake is not an ornament.
The checklist for the next "AI science" headline
Faced with any claim of AI-made science, the transferable capability is a handful of questions: what exactly did the model propose? What testable prediction follows? Who ran the experiment, and with which controls? Is it a laboratory test or an independent replication? And who publishes the figure — a third party or the maker of the system itself? If the answers are missing, there is an interesting hypothesis, not a discovery.
DeepMind's own conclusion best sums up the moment: the success of agents may make laboratories, quality data, and peer review more valuable, not less. When generating candidates is cheap, choosing which deserve resources and testing them rigorously becomes the decisive work — and also the part no conjecture machine can do on its own.
Primary sources
This article was produced with artificial intelligence under human editorial oversight.