IA 360
Current Affairs

Apple paper measures limits of reasoning models

An Apple study uses controllable puzzles to observe where several reasoning models lose accuracy and reduce effort. Its result is bounded to those environments, not all machine reasoning.

5 min read AI-generated Leer en español
Apple paper measures limits of reasoning models

On June 10, 2025, Apple-affiliated researchers published an experiment using controllable puzzles to study reasoning models. The original source supports the documentary core of the event; the paper observes limits in those environments; it does not prove that all machine reasoning is illusory.

Reasoning models, technically known as LRMs (Large Reasoning Models), are the generation of systems that generate an explicit, step-by-step thinking process the user can read before producing an answer. That capability has been the big selling point of the past several months: models that don't just answer, but show how they arrive at the answer — and that, in theory, should be better at solving complex math and logic problems.

A puzzle lab for measuring thought

Until now, most evaluations of these models have relied on math and coding benchmarks that only measure whether the final answer is correct. The problem, the authors note, is that these benchmarks tend to be contaminated — meaning the model has likely seen very similar exercises during training — and they say nothing about how the model actually thinks internally.

To get around this, Apple's team designed controllable puzzle environments, including variants of the Tower of Hanoi, where difficulty can be precisely dialed up or down while the underlying logical structure stays identical. That makes something possible that traditional benchmarks don't offer: isolating the effect of complexity and, in the process, examining not just whether the model gets the right answer, but the reasoning trace it produces before responding.

Three regimes, and a total collapse

Comparing LRMs against their equivalent standard versions (models without that explicit "think before answering" step) under the same compute budget, the authors identify three distinct behaviors depending on problem difficulty:

  • Low complexity: standard models, without explicit reasoning, outperform LRMs.
  • Medium complexity: this is where the reasoning models' edge shows up, outperforming their standard counterparts.
  • High complexity: both types of model collapse equally. The paper describes it literally as a "complete accuracy collapse" past a certain threshold, with results approaching 0% accuracy.

The most striking part, according to the authors, isn't just that the models fail, but how they fail. As a problem gets harder, reasoning effort — measured by the length of the thinking process the model generates — increases, as one would expect. But only up to a point: past that threshold, effort actually decreases, even though the model still has token budget left to keep thinking. In other words, the model doesn't run out of room to reason further; it simply stops trying once the problem gets too hard.

Failures in exact computation

The study also documents that these models have concrete limitations when it comes to exact computation: they fail to consistently apply explicit algorithms and reason inconsistently depending on the scale of the problem, even when the underlying logic doesn't change. The authors also dig deeper into the reasoning traces themselves, studying which solutions the models explore and how they behave computationally, in an effort to better understand their real strengths and limits.

Why this debate matters

The paper's title is no accident. "The Illusion of Thinking" goes straight to the heart of the debate that has dominated the AI industry over the past year: whether the "thinking out loud" process these models display reflects genuine reasoning, or is, in practice, a learned pattern that falls apart the moment a problem stops resembling something seen during training.

The question isn't academic. Companies and developers have spent months adopting these reasoning models precisely for tasks that demand reliability on complex problems: planning, code debugging, financial or scientific analysis. If performance itself collapses beyond a certain complexity — and the model even scales back its effort right when it would be needed most — it's worth calibrating expectations before handing critical decisions to these systems without supervision.

For now, the work is a preprint: it hasn't gone through peer review or been published at a conference. That doesn't take away from the methodology, but it does mean replications and counterarguments from other labs should be expected in the coming weeks — especially from the companies that build these reasoning models and have a lot riding on this debate.

Turning the headline into a check

An experimental result begins with its observable variable. Record what counts as success, which behavior triggers a label and which cases fall outside scope. the paper observes limits in those environments; it does not prove that all machine reasoning is illusory. If the phenomenon cannot be recognized without interpreting a model's intent, the conclusion needs even greater caution and a reproducible definition.

Protocol matters as much as score. Document instructions, tools, time, compute budget, number of attempts, example selection and grading rule. Changing any one may alter the result without the model learning anything new. Comparing two headlines therefore begins by checking that they measure the same axis.

A strong replication tries to break the conclusion. Add unseen data, small variants, negative controls and tasks where abstention is correct. Preserve failures as well as selected successes. To assess how to read domain, complexity, metric and replication before generalizing a benchmark, the set must resemble the intended use and reflect the cost of each error class.

What the record must preserve

A study can reveal a pattern without settling an entire field. Honest wording preserves domain, sample and date, and avoids turning 'we observed' into 'we proved forever.' Evidence becomes more valuable when another team can repeat it with available materials or state what is missing. That traceability is more useful than a sweeping label.

An evidence sheet separates four columns: what the source claims, what it shows, what it did not measure and what would change the conclusion. That discipline prevents an absence from becoming a promise and a condition from vanishing in summary. It also lets the story be updated without rewriting history from a later outcome.

Include a negative case before deciding. Find a situation where the system, rule, transaction or study does not meet the need and record the signal that would require stopping. Selected successes show that something can happen; the negative case reveals the boundary and lowers the cost of discovering it after deployment.

The skill that outlasts the announcement

A valid comparison preserves denominator and axis. It does not pit a point figure against an average, future capacity against installed capacity or a forecast against an observation. When two sources use similar language, reconstruct what they counted and over what period. If those differ, publish them as different measures instead of inventing a ranking.

The record should survive a version change. Keep URL, consultation date, document, configuration and decision. When new evidence appears, add it with its date and explain what it changes. That traceability prevents opposite errors: keeping an expired conclusion or pretending later information was known on the event date.

The transferable skill in this story is how to read domain, complexity, metric and replication before generalizing a benchmark. The procedure is short: name the document, preserve the date, fix the axis, find the condition and design a check that can fail. With those steps, a reader need not accept or reject the announcement by intuition; the decision follows a visible chain of evidence.

Before closing, another person should be able to reconstruct the conclusion without knowing the headline. Give them the sources, conditions and negative case, then ask what they would accept and reject. If they need an assumed intent, a figure without a denominator or an undated later fact, the chain still has a gap. That short review catches errors that fluent prose can conceal.

The result is not a permanent score but a dated, revisable decision. Set when to measure again and which signal triggers an earlier review. Caution then does not paralyze; it turns uncertainty into a monitoring condition. It also prevents an announcement from receiving credit for a later improvement that was not available when the decision was made.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close