IA 360
AI Fundamentals

When AI erases ‘I don’t know’: measuring answers and errors

A preprint with 3,132 participants found less abstention under wrong AI advice. The key is to measure coverage, accuracy and error cost separately.

Admin IA360 4 min read AI-generated Leer en español
When AI erases ‘I don’t know’: measuring answers and errors

On July 15, 2026, three researchers released a preprint with a troubling but deliberately bounded result. When 3,132 participants faced six difficult questions about movie details, access to AI advice nearly eliminated their willingness to answer “I don’t know.” Because the advice had been engineered to be usually wrong, they answered more, were correct less often and reported greater confidence.

The lesson is not that every kind of AI assistance makes every person reckless. The study’s original abstract describes five laboratory experiments, not a general measurement of ChatGPT, work or education. Its durable value lies elsewhere: accuracy and coverage are separate metrics. A system can produce more answers while degrading decisions if it pushes people to answer cases they would previously have left unresolved.

What the researchers actually tested

Chiara Marcoccia, Walter Quattrociocchi and Valerio Capraro selected six questions about minor visual details in films: the color of a uniform, a hairstyle or a vehicle, among others. The facts were poorly represented in online text. Participants could answer or abstain. In experimental conditions they could consult a model; in baseline conditions they could not.

The model was not a generic product simply called “AI.” The full manuscript identifies StepFun’s Step 3.5 Flash, chosen because it could be embedded in Qualtrics. Its name was hidden from participants. The researchers selected questions on which this model failed almost every time, expressly to separate the effect of having a suggestion from the rational benefit of following good advice.

The first trial used live responses. Roughly 10% of requests failed because of overload or malfunction, so the team ran a replication with pre-generated wrong answers. Later experiments retained fixed responses, added monetary incentives and, in the final study, displayed advice automatically instead of waiting for participants to request it.

That design supports one specific inference: on these tasks, with a deliberately unreliable source, a fluent answer changed the threshold for abstaining. It cannot estimate Step 3.5 Flash’s typical accuracy, compare commercial brands or show that correct AI advice is harmful. The authors built an adverse case to observe human behavior when delegation was not sensible.

The important result is not one number

In Study 1a, mean judgment suspension fell from 0.36 without AI to 0.06 with access to advice. In the Study 1b replication, it fell from 0.44 to 0.03. In Study 2, without incentives, it was 0.17 versus 0.01; with incentives, 0.21 versus 0.02. Magnitudes varied, but the direction was consistent.

Pooling no-incentive conditions, the authors report 27.5% correct responses without AI and 9.2% with it. In the study measuring confidence, the mean rose from 29.6 without advice to 75.9 with it when no money was at stake. “Two and a half times” is more accurate than “nearly doubled,” and the context must remain visible: a 0–100 scale and questions where automated advice was wrong by design.

The phrase “one-third as accurate” also needs a complete sentence. It does not mean participants got one-third of questions right. It means the quoted 9.2% rate was about one-third of the 27.5% rate without AI in pooled conditions. A ratio between rates must not be turned into an absolute rate.

This discipline prevents a study about overconfidence from being communicated overconfidently. Every number needs an experiment, condition, variable and denominator. The authors’ public repository provides data, code, instructions and preregistrations, allowing readers to move beyond the abstract.

What incentives repaired—and what they did not

Participants in monetary conditions earned $0.10 for a correct answer, lost $0.10 for an error and received zero for abstaining. In Study 3, displaying those rules before every question raised suspension from 0.02 to 0.08 when AI was available and improved correctness from 0.11 to 0.16. Participants also requested advice less often: 4.53 out of six questions rather than 5.27 without incentives.

One of the most instructive findings is a null. The researchers had preregistered that AI access would weaken the effect of incentives on suspension, producing a negative statistical interaction. That interaction was not significant in Studies 2, 3 or 4. The Study 2 preregistration establishes that the hypothesis preceded the analysis.

An unsupported hypothesis should not be rescued by a post-hoc story. The data suggest that money and AI availability operated largely independently on abstention. Incentives did improve accuracy when advice was present and reduced advice-seeking, but did not restore “I don’t know” to no-AI levels. Separating prediction, primary result and exploratory analysis is central to reading science.

Accuracy and coverage: the missing map

In classification and selective prediction, coverage is the share of cases receiving an answer; risk is error among answered cases. Abstention creates a tradeoff: a system can cover fewer cases to make fewer mistakes. An evaluation that rewards answering every question erases that safety mechanism.

The study shows a human shift along that curve. The option to consult AI raised coverage because almost nobody abstained, while risk increased as people followed incorrect advice. An assistant may look more productive when managers count answers, closed tickets or generated documents, even while increasing the number of decisions that should not have been automated.

A business, hospital or school should assign distinct costs to a correct answer, error, abstention and escalation to a person. If a medication error costs far more than another review, the system should favor suspension. If the task is reversible brainstorming, coverage may matter more. There is no universal threshold; there is a loss function that should be stated.

Design changes that still need their own tests

The experiments tested voluntary access, automatic advice and incentives. They did not directly test every interface intervention one might infer. Those should be framed as design hypotheses: ask for an initial answer before revealing the suggestion; retain a visible “I don’t know” control; place advice behind a request; display provenance; or require verification for consequential choices.

Each intervention needs evaluation. A model confidence score may worsen reliance if it is miscalibrated. Citations can become decoration if nobody opens them. Requiring an initial answer can anchor users to their own error. A sound trial measures not only final accuracy, but abstention, confidence, source inspection, time and recovery when advice is false.

It should also vary adviser quality. This research fixed a usually wrong model to isolate a mechanism. In life, reliability changes by task and version. Users need to learn when advice deserves trust, not reject it categorically. The target is calibrated delegation: follow strong evidence, resist weak evidence and suspend when neither is sufficient.

A three-question pause

Before accepting an automated suggestion, ask three questions. What would I have answered before seeing it? What external evidence separates this recommendation from a plausible sentence? What is the cost of being wrong compared with saying “I do not know yet”? The pause separates the availability of an answer from the existence of knowledge.

The preprint is dated and transparent, but it remained a preprint on July 15: no published peer review had occurred. Its sample consisted of US participants recruited through Prolific, and its six questions occupied a narrow domain. Those limitations do not erase the observed effect; they define where it has been demonstrated.

The transferable skill is to read every assistance system on two axes. Ask how many cases it gets people to answer and how many they answer correctly, while keeping abstention visible. When a tool makes “I don’t know” disappear, it has not necessarily removed uncertainty. It may only have removed the behavior that made uncertainty observable.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close