IA 360
DeepMind

DeepMind reaches silver-medal level at the Math Olympiad

AlphaProof and AlphaGeometry 2 solved four of the six problems at the 2024 International Mathematical Olympiad. Their 28 points amounted to a silver medal, just one point short of the gold cutoff.

4 min read AI-generated Leer en español
DeepMind reaches silver-medal level at the Math Olympiad

On July 25, 2024, Google DeepMind reported that AlphaProof and AlphaGeometry 2 had solved four International Mathematical Olympiad problems and scored 28 points out of 42, a silver-medal standard. DeepMind’s report also discloses conditions limiting the comparison: manual formalisation, specialised systems and runs lasting up to several days.

The gold-medal cutoff at this year’s competition was 29 points. The numerical gap may seem small, but the scale matters: this was not a single AI competing as a student under the usual conditions. It was two specialized systems tackling different problems, with their results evaluated according to IMO standards. Source

Four out of six problems, with verifiable proofs

Each Olympiad problem is worth seven points and requires a proof, not just the correct answer. DeepMind says AlphaProof solved two algebra problems and one number theory problem, while AlphaGeometry 2 solved the geometry problem. IMO coordinators graded the solutions using the competition’s scoring system. Source

That detail sets this achievement apart from many eye-catching results produced by language models. A chatbot can generate a convincing mathematical explanation and still contain a logical leap or an incorrect operation. AlphaProof works differently: it searches for proofs expressed in Lean, a formal language that allows every step to be verified by a program.

A formal proof is, in practice, a proof written in notation that a computer can check without ambiguity. If a condition is missing or a step does not follow from the preceding ones, the verifier rejects it. The trade-off is that the problem must first be translated into that formal language, a task that still requires mathematical and technical expertise.

From AlphaGo to theorem proving

AlphaProof builds on an idea DeepMind has used since AlphaGo: learning through reinforcement. Instead of choosing moves on a board, the system explores reasoning steps until it finds a sequence that concludes with a valid proof. The company trained it on large collections of formalized mathematical problems and generated new problems to expand that training.

AlphaGeometry 2 has a different design. It combines a neural model, useful for proposing geometric constructions and relationships, with a symbolic engine that applies explicit rules. That combination matters because Olympiad geometry often requires discovering an auxiliary idea—such as a line or circle not mentioned in the problem statement—and then chaining together exact deductions. Source

The result improves on the level DeepMind demonstrated with AlphaGeometry in early 2024. At the time, the system solved a significant share of historical IMO geometry problems; now the company is presenting a combined performance that approaches the competition’s top medal tier.

A genuine milestone, but not general-purpose mathematics

The comparison with a medal needs to be read carefully. Human participants receive the problems in natural language, have two days of exams, and write their answers without relying on a prepared formal translation. AlphaProof, by contrast, needs the problem to be expressed in Lean before it begins searching. DeepMind also divided the competition between systems designed for specific areas.

That is why the announcement does not show that an AI has the mathematical flexibility of an Olympiad contestant, much less that it can conduct research autonomously. It does demonstrate something important: AI methods can now produce rigorous proofs for problems that, until recently, were considered the domain of highly trained human reasoners.

The most immediate value lies beyond the Olympiads. Tools such as AlphaProof could help formalize theorems, detect errors in lengthy proofs, and assist researchers in fields where exact verification is crucial. The next challenge will be to reduce their reliance on manual formalization and help these systems better understand statements and mathematical ideas expressed in ordinary language.

The medal is a score equivalence

The accurate claim is not that an AI sat the contest as a student. Experts formalised the problems, different tools handled different areas, and their running times did not match the human exam. The score locates the difficulty of the solutions; it does not erase procedural differences.

A formal proof brings a strong advantage: the checker verifies that every step follows accepted rules. Yet that correctness applies to the formal translation. If the statement is encoded incorrectly, there may be a flawless proof of another problem. Both the formalisation and its correspondence to ordinary language must therefore be preserved.

How to evaluate a mathematical reasoning system

Separate formulation, search and verification. Record who translated the problem, which theorem library was available, how much compute was used and whether several attempts were selected. Then check the proof in an independent installation and inspect the mathematical idea found, not just whether the file compiled.

A transfer test changes one condition: notation, domain or the need for a new lemma. A system matching patterns from a library may perform well and fail on an equivalent formulation. Unseen cases and inspection of dependency traces reveal generality more clearly.

The transferable skill is to read a human-machine comparison through conditions rather than metaphor. Preserve inputs, assistance, tools, time, checker and criterion. The medal remains a milestone when described precisely; exaggeration makes it less informative.

Finding and checking a proof are different jobs

A formal checker can reject an invalid derivation with great precision, but it does not find the route to a solution by itself. The search system chooses lemmas, constructs steps and explores alternatives; the checker decides whether the result obeys the rules. Combining those functions makes the checker’s certainty appear to cover the whole strategy.

A comparison should preserve failed attempts. If only the selected proof is published, readers cannot know how many paths the system consumed or how the winner was chosen. Compute time, candidate count and human assistance belong to the result. It also matters whether the problem or nearby lemmas appeared in training data.

Explanation remains a separate evaluation

A proof that compiles may be difficult for a person to read. Asking for a natural-language explanation measures another ability: identify the central idea, justify the decisive move and connect it to prior knowledge without inventing steps. Then compare the explanation with the formal object, because a persuasive narrative may not describe the actual proof.

For educational support, reaching the answer is not the only criterion. Test whether the system gives a proportionate hint, detects where a learner is stuck and avoids revealing the complete solution immediately. Solving an olympiad and teaching mathematics are not the same task.

The method travels beyond mathematics: separate generator from checker and ask which part supports automatic verification. Where no strong checker exists, confidence in the output falls even when the apparent reasoning is equally fluent.

A complete report supplies statements, formal translations, proofs, dependencies, time and configuration. Specialists can then locate both achievement and limit without trusting a corporate summary. If one part cannot be opened for security or licensing reasons, that should be stated: inability to reproduce weakens the conclusion even when it does not erase the observed advance. Publishing failures and unsolved problems prevents a selected sample from becoming a commercial demonstration.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close