Q*: how to report a secret breakthrough without turning a leak into AGI
The Q* story relied on anonymous sources and included no public test. This method separates documents, attribution, measurement and interpretation so a leak does not become an AGI verdict.
On November 22, 2023, Reuters reported, citing two people familiar with the matter, that OpenAI researchers had warned the board about a finding they considered important to the future of artificial intelligence. One source connected the work to a project called Q* and to solving certain grade-school mathematics problems. The agency also preserved two essential limits: it had not seen the letter and could not independently verify the capabilities attributed to the system.
That last fact changes how the entire story should be read. As of that date, the public record contained no paper, results table, reproducible demonstration or technical description of Q*. There was a news report based on anonymous sources. It could be an important lead; it did not turn a secret capability into a measured fact or prove the arrival of artificial general intelligence.
Documented and attributed are not the same
The nearest public document was OpenAI’s November 17 announcement about Sam Altman’s departure. The board said Altman had not been “consistently candid” in his communications and that this hindered its ability to exercise its responsibilities. The statement did not mention Q*, a researchers’ letter, mathematics or a threat to humanity. Treating it as confirmation of any of those points would add material the document does not contain.
A disciplined reading separates four layers. First comes the documentary fact that can be opened and cited. Second is what an identified source says. Third is what a news organization attributes to protected sources. Fourth is the interpretation others build on top. In the Q* story, the project name and reported mathematics tests were in the third layer; equating the development with AGI belonged to the fourth.
This hierarchy does not require ignoring a leak. It requires preserving its provenance in every sentence. “Reuters reported, citing sources” is different from “OpenAI created.” “The sources described it” is different from “the system demonstrated it.” When that grammar disappears, a limited claim becomes certainty through repetition.
Solving mathematics does not define AGI
Mathematics problems are useful because an answer can be checked and the steps inspected. But “grade-school math” does not identify the problem set, number of attempts, success rate, possible training contamination, use of tools or comparison with earlier models. Without those fields, the phrase is not an experimental result.
Months earlier, OpenAI had published work on process supervision for mathematical reasoning. It compared rewarding only the final answer with providing feedback at every step and explained that advanced models still made logical mistakes. That research helps explain why improved mathematics could interest a laboratory. It does not establish that Q* used the same method or reveal its architecture.
A score would not be enough to leap from one task to a general category. The OpenAI Charter defines AGI as highly autonomous systems that outperform humans at most economically valuable work. This is a broad institutional definition, not a standardized test. Success on bounded problems may reveal better planning or search; transfer across domains, sustained autonomy, reliability and performance against people across varied work would still need to be demonstrated.
What a laboratory would need to publish
A mathematics claim can be evaluated without releasing weights or trade secrets. A laboratory can describe the problem set, its creation date, contamination policy and grading rule. It can report one-attempt and multiple-attempt results, cost, time, permitted tools and failures. An external evaluator can receive controlled access and publish the protocol and aggregate findings.
The baseline matters too. “It solves problems that it could not solve before” requires naming the previous version, applying the same prompts and conditions, and showing how many cases changed. Curated demonstrations do not measure a capability. A record of every exercise, including wrong answers and abandoned attempts, distinguishes a frequent improvement from a prepared display.
Contamination needs a dedicated test. A model may reproduce solutions found in its training data without learning a method that transfers to new exercises. Sets created after the training cutoff, rule-based variants and independent human review reduce that risk. No single measure removes it; the combination makes the claim auditable.
Before calling a secret result a breakthrough, ask for six elements: exact task, sample, metric, conditions, baseline and independent review. If they are missing, the accurate verb is not “demonstrated” but “was described as.” Language should expose the quality of the evidence.
A code name does not disclose a method
Q* prompted associations with known techniques because its name resembled concepts in reinforcement learning and search. Resemblance is not evidence. Internal names can be jokes, inherited labels or partial references. Without technical documentation, no reader can determine whether a system combines language with planning, search, reinforcement learning, symbolic reasoning or some other strategy.
Reasoning without inventing an architecture means framing hypotheses separately from facts. If the system used search, one might expect sensitivity to compute budget and gains from exploring more candidates. If it learned a transferable procedure, it should work on unseen problems and structural variants. If it merely recalled patterns, performance should fall when names, numbers or formats change. These predictions turn speculation into testable questions, not a description of Q*.
Capability, access, autonomy and harm
A technical capability does not automatically equal practical risk. A chain sits between them: what the model can do, which tools and data it can access, how independently it can act, which controls constrain it and what consequence a failure would produce. A mathematics test speaks, at most, to the first link within one task.
Safety evaluation has to test every link. Can the system pursue a goal across many actions? Does it detect and correct mistakes? Can it run code, spend resources or transmit information without authorization? Are its actions logged, and is there an effective stop? Two models with the same benchmark score can have very different risk profiles if one answers inside a sandbox while the other operates external tools.
The Charter commits the organization to long-term safety and cooperation. That is useful as a governance statement, but it is not a substitute for a concrete evaluation. Connecting Q* to a threat would require the observed capability, deployment scenario, controls and evidence of harm. Those elements were not public in November 2023.
A protocol for reading technology leaks
First, open the original document or report and record what its author could verify. Second, preserve the exact attribution: organization, identified source or anonymous source. Third, build a claim table and mark which entries have measurements. Fourth, locate the decisive absence—sample, date, comparison, access or verification. Fifth, never use a different announcement to fill that gap.
Then apply a contrary test: what would we observe if the most spectacular interpretation were false? A system could improve sharply on one problem family without generalizing; a letter could express concern without claiming AGI exists; a governance conflict could coincide with a technical project without being caused by it. Keeping those alternatives open prevents chronology from becoming causation.
Q* was a legitimate story about internal signals and missing public information. Its value did not depend on solving the mystery. The transferable skill is recognizing the evidence ladder: document, attribution, measurement, reproduction and interpretation. While the middle steps are absent, a leak can justify urgent questions, but not a verdict about AGI.
This article was produced with artificial intelligence under human editorial oversight.