GPT-5.3 Codex: OpenAI's Model That Helps Build Itself
OpenAI trained GPT-5.2 and GPT-5.3 Codex entirely on NVIDIA infrastructure. The latter is its first agentic coding model to take part in its own development. Behind it lies a hardware race that now determines who gets to build frontier AI.
OpenAI launched GPT-5.3 Codex in February, an agentic coding model —meaning it can execute tasks autonomously— that the company describes as the first of its models to help build itself. It was trained and deployed entirely on NVIDIA's GB200 NVL72 systems. That's a small detail with a big implication: hardware is no longer a logistical footnote, but the deciding factor in who gets to play in AI's premier league.
The information comes from an NVIDIA corporate blog post, which updated its tally of models trained on its platform following OpenAI's latest releases. It's worth reading with some distance: this is material from a company that sells precisely those chips. But the technical facts it lists —which models were trained where, and with what results on standard benchmarks— do help explain how frontier AI is being built today.
What GPT-5.2 and GPT-5.3 Codex Are
OpenAI unveiled GPT-5.2 in December as its most capable model series to date for professional knowledge work. According to the data NVIDIA cites, it posts the best reported scores on several industry benchmarks: GPQA-Diamond (doctoral-level science questions), AIME 2025 (competition mathematics), and Tau2 Telecom (simulated customer-service tasks). It also sets a new record on ARC-AGI-2, a test specifically designed to measure the abstract reasoning associated with so-called artificial general intelligence.
GPT-5.3 Codex, released in February, is the more interesting piece. It combines the coding performance of GPT-5.2-Codex with the reasoning ability of GPT-5.2 in a single model, and does so, according to OpenAI, 25% faster. On coding tests it sets new highs on SWE-Bench Pro and Terminal-Bench —two benchmark suites that measure whether a model can solve real software engineering problems and operate in a terminal— and performs well on OSWorld and GDPval.
The detail NVIDIA itself highlights is that this is OpenAI's first agentic model to take part in its own development. That shouldn't be read as science fiction: it means a coding model is being used as a tool to write and debug the code of the next one. It's automation of engineering work, not a machine reinventing itself on its own. But it points to a real trend: labs are starting to use their own models to speed up the development of the ones that come after.
Why Hardware Calls the Shots
The core argument of the piece is that training a frontier model from scratch is no longer a software project — it's an infrastructure project. It requires tens of thousands, sometimes hundreds of thousands, of GPUs working in coordination. That demands three things at once: powerful accelerators, networking capable of connecting all those chips without bottlenecks, and a software layer optimized to squeeze every bit of performance out of them.
NVIDIA describes three architectures in its rundown. Hopper is the previous generation. Blackwell —and its GB200 NVL72 configuration, which groups 72 GPUs into a system that behaves like a single giant machine— is the current one. Blackwell Ultra (GB300) is just beginning to roll out.
The figures the company provides, measured in MLPerf Training —the industry standard for comparing training performance— give a sense of the leap: GB200 NVL72 systems train up to three times faster than Hopper on the largest model evaluated, with nearly double the performance per dollar. The GB300 goes even further, topping Hopper by more than four times.
Translated into what actually matters: every hardware leap shortens development cycles. A lab that can train in weeks what used to take months can iterate faster and ship sooner. In a race where new models arrive every few months, that speed is a direct competitive advantage.
An Ecosystem That Goes Beyond Text
The post insists —with obvious commercial interest— that most of today's large language models were trained on NVIDIA platforms, and that its reach extends well beyond text.
In video generation, Runway announced Gen-4.5 last week, which according to the Artificial Analysis ranking is the world's top-rated video model. It was developed entirely on NVIDIA GPUs. The same company unveiled GWM-1, a general-purpose "world model" capable of simulating reality in real time, with applications in gaming, education, science, and robotics.
On the scientific front, NVIDIA cites models such as Evo 2 (genetic sequence decoding), OpenFold3 (3D protein structure prediction), and Boltz-2 (drug interaction simulation), along with its Clara models, which generate synthetic medical images to train diagnostic systems without exposing real patient data.
The list of labs training on Blackwell includes notable names: Black Forest Labs, Cohere, Mistral, OpenAI, Reflection, and Thinking Machines Lab. The platform is also available through the major cloud providers —AWS, Google Cloud, Microsoft Azure, Oracle— and through so-called neo-clouds like CoreWeave, Lambda, and Nebius.
The Critical Read: Dependency and Concentration
There's a story beneath the story, and it's not the one NVIDIA wants to tell. When the chipmaker can publish a list naming nearly every frontier lab on the planet as a customer, the real headline is concentration.
The ability to train cutting-edge models today depends on accessing hardware produced essentially by one company. That has consequences: prices are largely set by the supplier, GPU availability determines which projects move forward and which wait, and whoever can't afford that infrastructure gets left outside the frontier. It's a barrier to entry that separates a handful of well-funded labs from everyone else.
Alternatives do exist —Google's in-house chips (TPUs), AWS's accelerators, AMD's initiatives— but NVIDIA's software ecosystem, built over more than a decade, remains the path of least resistance. Breaking that inertia is hard precisely because every new model trained on its platform further entrenches the standard.
What's Next
The underlying message for users and businesses is that the pace of improvement doesn't depend solely on researchers' ingenuity, but on how quickly faster hardware arrives. Blackwell Ultra is already rolling out; the next generation will define what's possible to train in 2026.
For anyone watching the sector from the outside, two signals are worth tracking. First, whether more models emerge that —like GPT-5.3 Codex— use earlier versions to speed up their own development, which would compress development cycles even further. Second, whether any competitor manages to train a frontier model outside the NVIDIA ecosystem competitively. Until that happens, the line the post itself closes with —"the future is being built on NVIDIA"— will be, more than a slogan, an uncomfortably accurate description of how things stand.
This article was produced with artificial intelligence under human editorial oversight.