China closes the AI gap with the US: how to read the race correctly
The 2025 AI Index shows sharp technical convergence, but not one universal 1.7% gap. A guide to separating benchmarks, models, capital and adoption.
On April 7, 2025, Stanford released its 2025 AI Index with two conclusions that can be true at once: the United States was still producing many more notable models, while China had sharply narrowed the performance gap. The useful question is not who “won” the artificial-intelligence race according to one number. It is how to read a scoreboard that mixes tests, models, capital, research and deployment, with each column answering a different question.
That distinction matters well beyond geopolitics. A procurement lead comparing vendors, a team choosing a model and a reader assessing the next grand headline all face the same problem: “best,” “leading” and “nearly equal” mean little until the measured quantity is named. The official AI Index overview does not identify one universal winner. It depicts an industry in which leadership changes with the dimension being observed.
The first correction: 1.7% was not the China-US gap
The easiest figure to misassign is 1.7%. In Stanford’s material, that number does not measure the difference between the best American and Chinese models. It measures the February 2025 gap between the leading closed-weight and open-weight models on the Chatbot Arena leaderboard. That gap fell from 8.04% in early January 2024 to 1.70%, according to the official technical-performance chapter.
For China and the United States, Stanford reports four separate differences. At the end of 2023, the US leads on MMLU, MMMU, MATH and HumanEval were 17.5, 13.5, 24.3 and 31.6 percentage points. By the end of 2024, they had fallen to 0.3, 8.1, 1.6 and 3.7 points. “Near parity” reasonably summarizes the overall direction, but not a uniform outcome: MMMU still showed an 8.1-point difference while MMLU showed 0.3.
This correction yields a reusable rule. Before repeating a percentage, turn it into a complete sentence with a subject, measure, date and comparison. “1.7%” alone says almost nothing. “A 1.70% gap between the open- and closed-weight leaders on Chatbot Arena in February 2025” can be checked. If that sentence cannot be completed with the document open, the number is not ready for publication or a decision.
A benchmark is not an all-purpose ranking
MMLU tests knowledge and question answering across many subjects; MMMU adds multimodal reasoning; MATH focuses on mathematical problems, and HumanEval on code generation. Convergence on one of them does not make two models interchangeable for customer support, legal analysis, vision, programming or Chinese-language work. Every benchmark is a designed sample, not a full reconstruction of real work.
The technical chapter itself supplies reasons to distrust an overly comfortable scoreboard. In a year, systems’ performance on SWE-bench rose from 4.4% to 71.7% of problems solved, while familiar tests such as MMLU and HumanEval were beginning to saturate. When many systems approach an exam’s ceiling, tiny differences may reveal less about usefulness than about the test’s design. Stanford also records that the Chatbot Arena gap between the first- and tenth-ranked models fell from 11.9% to 5.4%, while the top two were only 0.7% apart.
Three questions prevent an inflated conclusion. First, is the unit percentage points, a relative percentage or an Elo score? Those are not interchangeable. Second, are the objects specific models, each country’s best model or a national average? Selecting only the leader hides the depth of the field. Third, were the versions and test conditions held constant? A live leaderboard can move while an article quoting it remains unchanged.
An organization still needs to run the decisive test on its own cases. Public results can narrow a candidate list; they must be followed by representative tasks, error criteria, cost per valid result, latency, language, security and human review. Benchmark convergence creates a wider field of plausible options. It does not remove local evaluation.
Performance, production and influence are different leagues
The United States retained a clear advantage in the number of reference systems. The research-and-development chapter counts 40 notable models from US-based institutions in 2024, compared with 15 from China and three from Europe. “Notable” follows a selection methodology; it does not mean every existing model or market share.
China led in total AI research publications, while the United States led in highly influential research. China also accounted for 69.7% of AI-related patent grants in 2023. None of these measures substitutes for performance: a granted patent does not prove a superior product, a paper is not a deployed service, and a model’s institutional origin does not capture its full supply chain of chips, data, talent and finance.
A serious national scoreboard therefore keeps separate columns. One may track benchmark capability and another the number and diversity of models; others can cover papers, citations, patents, compute, capital, adoption and infrastructure. Compressing them into a word such as “dominance” destroys the information needed to understand the competition.
Money matters, but it does not decide the result alone
The financial gap was far wider than the technical one. The AI Index’s economy chapter puts US private AI investment at $109.1 billion in 2024, nearly twelve times China’s $9.3 billion and 24 times the United Kingdom’s $4.5 billion. Global private investment in generative AI reached $33.9 billion.
Those figures describe capital captured under the report’s methodology, not a country’s total spending. They do not necessarily account in comparable ways for public funds, infrastructure procurement, incentives, university research or hard-to-observe corporate investment. A dollar does not automatically buy the same capability everywhere: power prices, chip access, software efficiency, wages and the reuse of open models all matter.
The falling cost of use further complicates the relationship between spending and advantage. Stanford calculates that querying a system performing at GPT‑3.5’s level on MMLU fell from $20 per million tokens in November 2022 to $0.07 in October 2024, a decline of more than 280 times. The endpoint was Gemini 1.5 Flash 8B. This does not mean every workload became 280 times cheaper or that GPT‑3.5 turned into that model; it compares one level on one test over time.
For a technology buyer, the practical lesson is to ask not what the biggest model cost to train, but what it costs to complete one unit of the organization’s work correctly. A cheap model that forces repeated calls, error correction or escalation to a person may be expensive. A system that does not top a general leaderboard can win on total cost if it is reliable enough.
Adoption is not the same as transformation
The report says 78% of respondents reported AI use in their organizations in 2024, up from 55% in 2023; generative-AI use in at least one business function rose from 33% to 71%. Yet “use” spans a limited pilot and a core production process. Stanford supplies the caveat that headlines often remove: among organizations reporting economic effects, the most common savings or revenue gains remained modest.
Regulation deserves the same care. The policy-and-governance chapter records 59 AI-related US federal regulations introduced in 2024 by 42 agencies, more than twice the 2023 count. “Introduced” does not mean that each was a new comprehensive statute or had the same effect. Counting regulatory actions measures activity; measuring protection or burden requires reading their scope and enforcement.
A checklist for the next AI race
When the next headline says country A has caught country B, record six items: the cutoff date; models included; benchmark and version; unit of the gap; rule for assigning a model to a country; and dimensions omitted by the headline. Then separate capability, resources, scientific output, deployment and economic outcomes. Only then is a conclusion defensible.
Applied to the 2025 AI Index, the result is more exact and more revealing than a winners’ table. China greatly narrowed the technical gap on several benchmarks in 2024, though not by one universal 1.7%; the United States produced more notable models and attracted vastly more private investment; China led other measures such as publication volume and patents; and the price of using a defined capability level collapsed. The lasting skill is knowing how to stop one metric from impersonating all the others.
This article was produced with artificial intelligence under human editorial oversight.