DeepSeek Updates R1 Model, Narrows Gap With OpenAI, Google
The new R1-0528 version lifts math performance from 70% to 87.5% on the AIME 2025 benchmark and cuts hallucinations nearly in half, bringing DeepSeek's open model closer to o3 and Gemini 2.5 Pro.
On May 28, 2025, DeepSeek released the R1-0528 update and its model card with stated results. The original source supports the documentary core of the event; benchmark figures come from the developer and do not automatically measure the reader's task or error rate.
The performance jump is concrete and measurable: on AIME 2025, a benchmark that tests the ability to solve competition-level math problems, the model's accuracy climbs from 70% to 87.5%. The new version also cuts hallucinations — answers the model presents as true when they aren't — by roughly 45-50% compared to the previous version. Document supporting the figure.
What closing the gap with o3 and Gemini 2.5 Pro means
The comparison isn't cosmetic. OpenAI's o3 and Google's Gemini 2.5 Pro are, alongside other reasoning models, the current benchmark for tasks that require step-by-step thinking before answering: advanced math, complex coding, or multi-step logic problems. An open-weight model — meaning one whose parameter file can be downloaded, run locally and freely modified — closing in on that level changes the calculus of who can afford to build on frontier AI.
Until now, accessing a reasoner of that caliber meant paying for API access from a closed company and accepting its terms of use, its rate limits and its opacity about how the model was trained. With open weights, any company, university lab or independent developer can download R1-0528, fine-tune it for their specific use case and run it on their own infrastructure, without relying on an outside provider or sending data to third-party servers.
The open-source threat to closed models
It's common to hear the argument in the industry that the capability gap between closed and open-weight models would remain wide: the latter could be useful, but not competitive at the frontier. The R1-0528 numbers complicate that argument: if an open model cuts hallucinations in half and gains nearly 18 percentage points on a demanding math benchmark in a single update, the distance to proprietary models appears to be closing fast. Document supporting the figure.
This doesn't settle the debate: closed models still offer guarantees around support, integration and, in theory, safety controls that a downloadable model doesn't come with out of the box. But the economic argument for companies with tight budgets — or for those looking to avoid dependence on external infrastructure — is tilting increasingly toward open models.
What changes for today's AI users
For a developer or company already working with reasoning models, DeepSeek's update expands the real options available with no per-token licensing cost. For the ecosystem at large, it reinforces a trend that has been building for months: leadership in open weights is no longer the exclusive domain of second-tier models — it's starting to reach the top tier of performance on complex reasoning tasks.
It remains to be seen how OpenAI and Google respond to this narrowing gap, and whether they can hold onto their edge by doubling down on capabilities a downloadable model can't easily replicate, such as deep integration with proprietary tools or large-scale fine-tuning on proprietary data.
Turning the headline into a check
The first step is to freeze the system's identity. DeepSeek released the R1-0528 update and its model card with stated results. A commercial name may cover different revisions, automatic routes and tools. A test record should preserve date, access mode, configuration, permissions and the full output. Without that snapshot, an improvement or failure observed today cannot rigorously be attributed to the version another person will use tomorrow.
Next, turn how to reproduce a comparison with version, template, budget and an independent set into cases with acceptance criteria. Build a local sample containing easy, ambiguous, long and deliberately impossible tasks. Record the input, the information the model may consult and what outcome would count as sufficient. A vendor-selected demonstration shows possibility; a test set preserved by the user measures reliability.
Autonomy needs a permission ladder. Reading and proposing are not the same as editing, sending or buying. A safer setup begins with read-only access, requires a preview and reserves execution for explicit approval. It also keeps a log and a rollback path. Judge the model by the errors the surrounding system contains, not by the confidence of its plan.
What the record must preserve
Cost and quality must be measured together. A cheap answer that must be reviewed from scratch can cost more than a slower but verifiable one. Measurement includes waiting, retries, consumption, human oversight and the consequences of failure. benchmark figures come from the developer and do not automatically measure the reader's task or error rate. That boundary turns the announcement into a testable hypothesis rather than a promise to be believed.
An evidence sheet separates four columns: what the source claims, what it shows, what it did not measure and what would change the conclusion. That discipline prevents an absence from becoming a promise and a condition from vanishing in summary. It also lets the story be updated without rewriting history from a later outcome.
Include a negative case before deciding. Find a situation where the system, rule, transaction or study does not meet the need and record the signal that would require stopping. Selected successes show that something can happen; the negative case reveals the boundary and lowers the cost of discovering it after deployment.
The skill that outlasts the announcement
A valid comparison preserves denominator and axis. It does not pit a point figure against an average, future capacity against installed capacity or a forecast against an observation. When two sources use similar language, reconstruct what they counted and over what period. If those differ, publish them as different measures instead of inventing a ranking.
The record should survive a version change. Keep URL, consultation date, document, configuration and decision. When new evidence appears, add it with its date and explain what it changes. That traceability prevents opposite errors: keeping an expired conclusion or pretending later information was known on the event date.
The transferable skill in this story is how to reproduce a comparison with version, template, budget and an independent set. The procedure is short: name the document, preserve the date, fix the axis, find the condition and design a check that can fail. With those steps, a reader need not accept or reject the announcement by intuition; the decision follows a visible chain of evidence.
Before closing, another person should be able to reconstruct the conclusion without knowing the headline. Give them the sources, conditions and negative case, then ask what they would accept and reject. If they need an assumed intent, a figure without a denominator or an undated later fact, the chain still has a gap. That short review catches errors that fluent prose can conceal.
The result is not a permanent score but a dated, revisable decision. Set when to measure again and which signal triggers an earlier review. Caution then does not paralyze; it turns uncertainty into a monitoring condition. It also prevents an announcement from receiving credit for a later improvement that was not available when the decision was made.
Finally, preserve the alternative. The question is not only whether the announcement works, but whether it improves the process compared with keeping the current approach, using another tool or waiting for evidence. A concrete baseline prevents novelty from being mistaken for benefit. The decision may be to proceed, limit scope or make no change yet, always for a checkable reason.
The same method improves discussion across different roles. A domain expert defines the costly error; an operator records conditions; a decision maker accepts the residual risk. Nobody needs to pretend to have complete certainty. It is enough for every premise to have a source, every limit to be visible and the action to stop when evidence contradicts expectations.
This article was produced with artificial intelligence under human editorial oversight.