IA 360
Education

An AI Detector Said a 30-Year-Old Essay Was Written by AI. Here's How to Actually Read One of These Accusations

A Yale student proved GPTZero gave a '100% probability of AI' verdict on texts published decades before any language model existed. A Stanford study measured the problem precisely: 61% false positives on real student writing. The four questions worth asking before accepting any accusation like this.

7 min read AI-generated Leer en español
An AI Detector Said a 30-Year-Old Essay Was Written by AI. Here's How to Actually Read One of These Accusations

A graduate student at Yale's School of Management, identified in the lawsuit only as "John Doe," wanted to prove something in his defense: that the AI detector that had accused him didn't do what it claimed to do. He took academic papers published by Peter Salovey, Yale's former president, and by other university scholars — some more than thirty years old — and ran them through GPTZero. The detector returned a "100% probability" that these texts, written decades before any language model existed, were AI-generated. It was, literally, impossible. The lawsuit is ongoing (Doe v. Yale University, 3:25-cv-00159, filed February 2025); the student has already served a one-year suspension. If you've ever wondered whether a piece of writing — yours, your kid's, a student's, a job applicant's cover letter — "was written by AI," here's the first thing worth knowing: the tool that usually answers that question gets it wrong with the confidence of something that's right.

What a detector can actually see, and why it fails

Detectors don't read meaning, and they don't check facts or citations either: they measure how predictable your word choices are, a property called "perplexity." Text with simple vocabulary and very regular structure has low perplexity, and AI models generate low-perplexity text by default — so a low score gets read as a sign of artificial origin. The problem is that low perplexity isn't exclusive to AI: it's also exactly how anyone who learned English as a second language writes naturally, or someone with a direct style, or a neurodivergent student.

A Stanford team measured this precisely in 2023: they ran 91 TOEFL essays — written entirely by real people, none of them native English speakers — through seven commercial detectors. The result: an average false-positive rate of 61.22%. 97.80% of those essays were flagged as AI by at least one detector, and 19.78% — 18 of 91 — were flagged unanimously by all seven detectors, on text that was one hundred percent human. When the researchers used ChatGPT to "enrich" the vocabulary of those same essays, the false-positive rate dropped from 61.22% to 11.77%: using AI to polish human writing makes it look less artificial to the detector. And in the opposite direction, the same team showed that genuinely AI-generated text can lower its perplexity — and dodge detection — with simple prompt instructions. The detector doesn't distinguish authorship: it distinguishes style, and style can be manufactured in either direction — which means the same score that convicts an honest student today could just as easily clear a dishonest one tomorrow, depending on which side happened to phrase things the way the detector expects.

What universities have already confirmed on their own

This didn't stay in a paper. Vanderbilt disabled Turnitin's AI detector in August 2023 and said so plainly: "we do not believe that AI detection software is an effective tool that should be used," noting also that Turnitin "gives no detailed information as to how it determines if a piece of writing is AI-generated or not." At the false-positive rate Turnitin itself claimed at the time (1%) applied to the 75,000 papers Vanderbilt processed in 2022, that's roughly 750 human papers potentially mislabeled in a single year, at a single university. Pittsburgh, Boston University, and Georgetown turned the tool off that same summer; more than sixty institutions across five countries have disabled it since, most explicitly citing the same bias against non-native writers the Stanford study had already measured. This isn't a problem time solved on its own: Curtin University in Australia disabled its own system on January 1, 2026 — more than two years after the original Stanford study — framing it as "fostering trust and clarity" in assessment. The technology has changed names several times since 2023; the underlying flaw hasn't.

The scale of the problem shows up in the numbers from institutions that did keep using these systems: Australian Catholic University logged nearly 6,000 alleged academic-misconduct cases in 2024, 90% tied to alleged AI use, and roughly a quarter were dismissed on review — any case where Turnitin's detector was the sole evidence was dismissed immediately. Internal documents later showed the university knew the tool was unreliable for over a year before dropping it in March 2025. One of those students, identified as "Madeleine," had her results withheld for six months while doing a nursing placement, until she was cleared — and suspects the delay cost her a job offer. That's not a small margin of error — it's a quarter of the accusations collapsing the moment someone bothers to look past the score.

What to ask before accepting an accusation

Neither "detectors are useless" nor "if the detector says so, it's true" is honest. What the evidence actually supports is narrower and more useful:

  1. No single detector score is proof on its own. With a documented false-positive rate up to 61% on real writing, an "82% probability of AI" isn't a conviction — it's a data point that needs corroboration.
  2. Ask whether the system knows the accused person's writing history. Prior drafts, a document's revision history, or writing done live weigh more than a score on a finished text.
  3. Consider who's more likely to write "low-perplexity" prose for legitimate reasons. Non-native English speakers, direct writing styles, certain neurodivergent conditions: a high score in these cases says as much about how someone writes as about what they wrote.
  4. Remember it cuts both ways. Human text can look artificial, and AI text can look human with a minor prompt tweak. The score measures style, not authorship.

Applied to the Yale case: (1) the exam score arrived with nothing else attached; (2) there's no indication it was compared against the student's earlier exams or coursework; (3) he's a non-native English speaker, exactly the profile the Stanford study shows getting over-flagged; (4) he demonstrated the reversibility himself by testing the detector against text that couldn't possibly be AI. All four questions point the same direction — the direction a federal court is now litigating.

Next time a detector — whichever one exists then, under whatever name — enters a conversation about whether someone cheated, these four questions still hold up. Detectors will change versions; how they fail, structurally, won't.

For anyone who wants to go deeper

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close