IA 360
Language Models

AI in Education: The Test to Demand of Any Artificial Tutor

An experiment with a thousand high school students found that those who practised with an unguarded chatbot scored 17% WORSE than those who never had one, once the help was taken away. Another, at Harvard, measured gains of up to 1.3 standard deviations with a tutor designed by teachers. The difference is not the AI: it is the design, and there is a concrete way to check it.

Admin IA360 4 min read AI-generated Leer en español
AI in Education: The Test to Demand of Any Artificial Tutor

More than half of American teenagers — 54% — have used a chatbot to help with schoolwork. The figure comes from the Pew Research Center report published on 24 February 2026, with fieldwork between 25 September and 9 October 2025 across 1,458 young people aged 13 to 17. One in ten does «all or almost all» of their homework with chatbot help.

Which means the debate about whether this is advisable no longer runs ahead of the fact. It is already in the house. And the useful question for a parent, a teacher or the student is not whether AI helps learning — that question has no answer — but what specific proof to demand of a tool before trusting it.

That proof exists, it is easy to understand, and it cleanly separates the studies that went well from the ones that went badly.

The number everyone cites and almost nobody has read

Every AI tutor pitch eventually invokes Benjamin Bloom and his «two sigmas»: the idea that one-to-one tutoring produces an improvement of about two standard deviations over conventional teaching. The paper exists, it is called «The 2 Sigma Problem», and it appeared in Educational Researcher, volume 13, issue 6, pages 4 to 16, in 1984.

Now the uncomfortable part, stated up front because it is part of the method: that paper sits behind a paywall. A scanned copy exists as course reading on an MIT server, but it is page images with no text layer. In preparing this article I could not read the body of the paper in a source I could hand to you, so there is no direct quotation from Bloom here, and the famous summary — that the average tutored student outperforms 98% of the control group — is offered as what the literature universally cites, not as something this newspaper verified in the original.

It is worth sitting with what that means. The figure propping up half the industry's sales talk comes from a four-decade-old document that the reader — the teacher, the parent — cannot open to check the conditions under which it was measured. Next time someone cites it in a slide deck, you know what to ask.

What has been measured with a control group

The good news is that recent, open, randomized evidence exists.

A Harvard team — Greg Kestin, Kelly Miller and colleagues — published a crossover trial in Scientific Reports in 2025, run in the autumn of 2023 in an introductory physics course. The results: median post-test score of 4.5 for the group working with the AI tutor (142 students) against 3.5 for the active-learning classroom (174), from a baseline of 2.75. Effect size runs from 0.63 standard deviations by linear regression up to a range of 0.73 to 1.3 by quantile regression. And it took less time: a median of 49 minutes against the class's 60.

Before extrapolating, the context the paper itself declares: Harvard undergraduates, physics, and — this is the decisive part — a purpose-built tutor, with scaffolding designed by teachers and explicit guardrails against handing over the full solution. Not a generic chatbot.

On another scale and another continent, the World Bank documented in a post by its researchers, published on 9 January 2025, an after-school English support pilot in Edo State, Nigeria, over six weeks between June and July 2024. Measured effect: around 0.3 standard deviations, which its authors equate to nearly two years of typical learning. With a caution worth inheriting: it is a researchers' blog, not a peer-reviewed working paper, and it publishes neither the number of students nor the tool used.

And what happened when the help was withdrawn

Here is the study that changes the conversation. Hamsa Bastani, Osbert Bastani and their team published in PNAS, in June 2025, an experiment with nearly a thousand high school mathematics students at a school in Turkey. They compared two tools built on the same model: a plain conversational interface with no safeguards, and a tutor with hints and pedagogical guardrails designed by teachers.

While they had access, both groups improved spectacularly: 48% in grades with the plain interface and 127% with the guarded tutor.

When access was removed and they sat an unaided exam, the result, in the paper's own words: students «actually perform worse than those who never had access (17% reduction in grades for GPT Base)». With the designed tutor's safeguards, the harm is mitigated.

Read that twice. The group that had improved most while practising with the tool ended up below those who never used it once they had to work alone. They had got better at producing correct answers, not at knowing mathematics. And nobody could see the difference until the tool was switched off.

The variable is not the AI: it is the design

Place the two studies side by side and the conclusion imposes itself. The Harvard experiment worked with a tutor carrying teacher-designed guardrails. The Turkish one shows that the same underlying model helps or harms depending on whether those guardrails are there.

That is why «does AI improve learning?» is a badly formed question, and why headlines answering it with a flat yes or a flat no are both wrong. The right question is what the tool does when the student asks for the whole solution.

Why the model's errors weigh more in a classroom

There is a further technical reason not to delegate the checking. An Apple team published GSM-Symbolic in October 2024, putting leading models through variations of grade-school maths problems. Their finding: «adding a single clause that seems relevant to the question causes significant performance drops (up to 65%) across all state-of-the-art models». Changing only the numbers in the same problem also degrades results.

Translated into the classroom: models fail precisely on the variations a teacher introduces to check whether a student has understood or merely memorized the procedure. And they fail with exactly the confidence with which they succeed, which is what makes the failure dangerous for someone still learning and without the judgment to catch it.

The capability: three questions before entrusting anyone's learning

1. Is there a controlled study, and who was in it? Without a comparison group there is no measurement, only testimonials. And context is not a detail: Harvard undergraduates in physics, Turkish secondary students in maths and English support in Nigeria are three different worlds. None licenses generalizing to «any child with a chatbot».

2. Was learning measured WITH the tool in front of the student, or WITHOUT it? This is the decisive question, and the one that almost never appears in brochures. A result measured while the student is using the tutor measures the performance of student-plus-tool. What matters in an exam, at university or in life is what remains when the tool is gone.

3. What does it do when the student asks for the full solution? If it gives it, it is optimizing the submitted assignment, not the learning. Pedagogical guardrails are not a product flourish: on the available evidence, they are the difference between helping and harming.

The deep end, undiluted

Nearly all of this can be read in full and for free, which is what makes it arguable:

The exception is Bloom, which is why this article opened with him: the most-cited document in this conversation is the one you cannot open.

The capability you leave with: demand of any AI education tool the same proof you would demand of a teaching method — a controlled study — and above all, check whether learning was measured with the tool in front of the student or without it. That second question separates a tutor from a crutch, and it is the one almost nobody asks.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close