Gemini and OpenAI Reach Gold Level at Math Olympiad
Google DeepMind and OpenAI say they each scored 35 of 42 points at the 2025 International Mathematical Olympiad. The result surpasses AlphaProof’s silver-medal performance just one year ago.
On July 21, 2025, Google DeepMind reported a gold-level result certified by olympiad coordinators, while OpenAI published a self-run evaluation with the same score. The original source supports the documentary core of the event; the two results did not undergo exactly the same validation process, and neither makes the system a general mathematician. Terms not published by the party are attributed to the source that documented them.
OpenAI said a few days earlier that an experimental reasoning model had also scored 35 points and fully solved five of the six problems. The two companies have therefore placed their systems at the level of the top human competitors in the world’s toughest pre-university mathematics competition. Document supporting the figure.
Five problems solved like an Olympiad contestant
The IMO brings together high school students from dozens of countries each year to tackle six problems in algebra, geometry, number theory, and combinatorics. Each problem is worth seven points, and contestants have two four-and-a-half-hour sessions to write rigorous proofs.
Gemini Deep Think solved five problems and scored 35 points. The significance lies in more than the number: it presented its answers in natural language, with mathematical proofs that human graders could review. It did not simply find a numerical answer or verify a result that was already known. Document supporting the figure.
Google DeepMind worked with the IMO organization so that the solutions could be evaluated according to the competition’s standards. The system did not officially compete alongside the national delegations, but having Olympiad coordinators grade the work gives the result an unusual degree of external validation for an AI model capability announcement.
OpenAI, for its part, said that three former IMO medalists independently graded its solutions and awarded them 35 points under the competition’s criteria. The company has not publicly identified the model or announced a release date. That caution matters: the result demonstrates a research capability, not a tool that any student or company can use today. Document supporting the figure.
A very rapid leap from 2024
A year ago, Google DeepMind introduced AlphaProof and an upgraded version of AlphaGeometry. Together, they scored 28 out of 42 points on the 2024 IMO problems, a result equivalent to a silver medal. Document supporting the figure.
That system required a far more specialized approach. AlphaProof translated problems into a formal language—a mathematical representation with strict rules that a computer can check—and searched for proofs within that framework. AlphaGeometry focused on geometry. Human intervention was needed to adapt some problem statements to those tools.
Gemini’s result changes the nature of the demonstration. A language model can read the problem statement, explore strategies, and explain the proof in the format a mathematician expects. That potentially makes it more flexible than a system designed for a specific branch of mathematics, although it does not remove the need to verify every step: a proof that appears convincing can still contain an invalid logical leap.
The timing of OpenAI’s announcement also shows how mathematical reasoning has become the new competitive arena for the major AI labs. Mathematics is an especially valuable test because answers can be scored against clear criteria and because the problems demand planning, abstraction, and the ability to sustain a long chain of inferences.
From competition to scientific research
An IMO gold medal does not mean that a system has solved open mathematical problems or that it can replace a researcher. Olympiad problems are new to the participants, but they are designed to have solutions and to be solvable within the exam period using advanced pre-university knowledge.
Even so, the advance has practical implications. In education, these models could help break down a proof, spot an incorrect step, or suggest different paths to a solution. In scientific research, the same ability to explore hypotheses and formalize arguments could speed up highly specific tasks, from algorithm design to the verification of complex calculations.
The next test will be reproducibility. We will need to know how often these models maintain that level beyond six selected problems, how much computational reasoning they require, and how they perform on questions for which no known solution exists. For now, Google and OpenAI have shown that the Olympiad bar—which seemed out of reach for generative models just a year ago—is no longer an exclusively human frontier.
Turning the headline into a check
An experimental result begins with its observable variable. Record what counts as success, which behavior triggers a label and which cases fall outside scope. the two results did not undergo exactly the same validation process, and neither makes the system a general mathematician. If the phenomenon cannot be recognized without interpreting a model's intent, the conclusion needs even greater caution and a reproducible definition.
Protocol matters as much as score. Document instructions, tools, time, compute budget, number of attempts, example selection and grading rule. Changing any one may alter the result without the model learning anything new. Comparing two headlines therefore begins by checking that they measure the same axis.
A strong replication tries to break the conclusion. Add unseen data, small variants, negative controls and tasks where abstention is correct. Preserve failures as well as selected successes. To assess how to compare protocol, graders, tools and availability before comparing scores, the set must resemble the intended use and reflect the cost of each error class.
What the record must preserve
A study can reveal a pattern without settling an entire field. Honest wording preserves domain, sample and date, and avoids turning 'we observed' into 'we proved forever.' Evidence becomes more valuable when another team can repeat it with available materials or state what is missing. That traceability is more useful than a sweeping label.
An evidence sheet separates four columns: what the source claims, what it shows, what it did not measure and what would change the conclusion. That discipline prevents an absence from becoming a promise and a condition from vanishing in summary. It also lets the story be updated without rewriting history from a later outcome.
Include a negative case before deciding. Find a situation where the system, rule, transaction or study does not meet the need and record the signal that would require stopping. Selected successes show that something can happen; the negative case reveals the boundary and lowers the cost of discovering it after deployment.
The skill that outlasts the announcement
A valid comparison preserves denominator and axis. It does not pit a point figure against an average, future capacity against installed capacity or a forecast against an observation. When two sources use similar language, reconstruct what they counted and over what period. If those differ, publish them as different measures instead of inventing a ranking.
The record should survive a version change. Keep URL, consultation date, document, configuration and decision. When new evidence appears, add it with its date and explain what it changes. That traceability prevents opposite errors: keeping an expired conclusion or pretending later information was known on the event date.
The transferable skill in this story is how to compare protocol, graders, tools and availability before comparing scores. The procedure is short: name the document, preserve the date, fix the axis, find the condition and design a check that can fail. With those steps, a reader need not accept or reject the announcement by intuition; the decision follows a visible chain of evidence.
Before closing, another person should be able to reconstruct the conclusion without knowing the headline. Give them the sources, conditions and negative case, then ask what they would accept and reject. If they need an assumed intent, a figure without a denominator or an undated later fact, the chain still has a gap. That short review catches errors that fluent prose can conceal.
The result is not a permanent score but a dated, revisable decision. Set when to measure again and which signal triggers an earlier review. Caution then does not paralyze; it turns uncertainty into a monitoring condition. It also prevents an announcement from receiving credit for a later improvement that was not available when the decision was made.
This article was produced with artificial intelligence under human editorial oversight.