IA 360
AI Fundamentals

When several AI models agree, they may still be wrong

A new preprint warns that self-consistency and agreement between LLMs are weak, context-dependent confidence signals.

Admin IA360 2 min read AI-generated Leer en español
When several AI models agree, they may still be wrong

You ask three different AI assistants the same question, and all three give you the same answer. The natural instinct is to relax: if they agree, it must be true. A study published on 9 July 2026 tested that instinct at scale, and the result should change how you read any AI answer from today on: in the most self-consistent model the researchers tested, when agreement among its own answers exceeded 80%, nearly half of those cases — 48% — were still wrong.

This isn't a quirky failure of one particular model, or a lab curiosity. It's a pattern that resurfaces anywhere multiple AI sources "confirm" something: panels of models grading other models in enterprise pipelines, systems that request several answers and keep the majority one, an assistant that "checks" its own answer by asking itself again. If you use AI for anything that matters — reviewing a contract, deciding on an investment, marking a piece of work, validating the output of another automated system — this gives you the exact question to ask before trusting that "several of them agree."

What the study actually measured

Kaihua Ding published the study on arXiv on 9 July 2026, still as a preprint without peer review — worth saying upfront, because the same caution applies to the source behind this article. The paper separates two phenomena that often get blurred together: self-consistency, asking the same model for several answers and checking whether it converges on one; and cross-model agreement, comparing answers from different models. Both are used today to decide whether an answer deserves trust, or whether a task needs more compute.

The scale of the experiment is considerable: 53 runners drew 50 samples per assigned case across comparisons of model tier, prompt type, and scale, on the GPQA Diamond and AIME benchmarks — 265,000 samples in total. The central result is qualified, not categorical: agreement predicts correctness positively but weakly, with correlations between 0.20 and 0.59. And the pattern gets worse exactly where trust runs highest: on the most self-consistent frontier model, high agreement appeared on 77% of GPQA entries, and nearly half of those were wrong. An additional check across three tiers of Claude found the same overconfidence pattern across providers, suggesting this isn't a quirk of one model family.

Why agreement isn't independence

Here's the capability actually worth taking away: agreement is only evidence if the sources that agree could have been wrong for different reasons. If they share the reason they'd get it wrong, agreeing adds no information — it just repeats the same mistake with more apparent confidence.

And language models share a great deal: they train on overlapping slices of the same web, on similar transformer architectures tuned with similar human-feedback methods, and they inherit similar biases from similar kinds of data. A study by Cornell researchers, accepted at ICML 2025 after peer review, measured this across more than 350 models, two public leaderboards, and a résumé-screening task: when two models get the same leaderboard case wrong, they agree on that wrong answer 60% of the time. The more uncomfortable finding is that larger, more accurate models have MORE correlated errors with each other, even across distinct architectures and providers — a pattern the authors call algorithmic monoculture, which in their hiring-task test reproduced exactly what theory predicts about homogeneous automated decisions.

It's fundamentally the same principle any careful fact-checker applies: four links to the same owner aren't four sources, they're one voice repeated four times. Here it happens inside a single AI system: three models returning the same answer aren't necessarily three independent opinions if they learned from the same well and share the same underlying architecture.

Where this fails in practice

The first place this breaks down is the oldest one: self-consistency as a technique. The paper that originated it, by Google researchers published at ICLR 2023, proposed it as a decoding trick — sample several reasoning paths and keep the majority answer — to improve accuracy on reasoning tasks, with measured gains of up to 17.9% on the GSM8K math benchmark. It was never designed as a measure of how much to trust the result. Today's mistake is borrowing a tool built to pick the best answer and using it to answer a different question entirely: should I believe this?

The second is the rise of "LLM-as-judge": systems that use one AI model to score another's answers, increasingly common in enterprise evaluation pipelines. The paper that popularized the method, by UC Berkeley researchers and collaborators published at NeurIPS 2023, already warned about its limits: position bias, verbosity bias, and what it calls self-enhancement bias — a judge's tendency to favor answers resembling what it would have generated itself. An AI judge that shares those biases with the model it's grading isn't verifying independently; it's recognizing itself.

The third is the panel-of-judges as an apparent fix. Cohere researchers showed in 2024 that a panel of several smaller models beats a single large judge and reduces intra-model bias — but only when the panel is deliberately built from disjoint model families. Independence doesn't appear just because you're using several models at once; it has to be engineered, the same way a journalist doesn't add up sources by counting links, but by checking they genuinely don't share an origin.

What actually counts as evidence

None of these studies concludes that asking several times is useless. They conclude something more precise: agreement is a useful clue in specific contexts — for instance, deciding whether a task deserves more compute for mid-tier models, per Ding's own study — but it stops working the moment it becomes the only proof.

Real evidence has a property that AI agreement doesn't have on its own: it comes from outside the system generating the answer. A primary source you can open and read yourself. A result someone else, using a different method, can reproduce independently rather than just re-asking the same question. A test that could actually fail — where there's a genuine possibility the answer is "no," not just "yes" dressed up several different ways. Agreement among models meets none of those three by default; it can support them, but it doesn't replace them.

What you can actually do, today

Next time several AI answers agree and you feel the pull to take that as proof, three questions are enough:

  • Could these sources have been wrong for the same reason? If they come from similar training, similar architecture, or the same underlying data, agreeing isn't independence.
  • Is there anything outside the system confirming it? A document, a database, an expert, a calculation you can redo yourself with a different method.
  • Do you recognize the pattern under a different name? "AI consensus," "high confidence," "several models independently confirm it" — the underlying question is always the same one, and this study won't be the last to trip over it.

That habit — asking where the agreement actually comes from before treating it as proof — holds for this particular preprint today, and it will still hold in a year, under a different study and a different name.

Sources for this piece

This piece draws on 1 primary source(s), gathered during reporting.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close