NVIDIA Unveils Rubin and Open Models at CES 2026
Jensen Huang opened CES 2026 with Rubin, NVIDIA's six-chip platform promising to generate tokens at one-tenth the cost, alongside a family of open models reaching all the way to level 4 autonomous driving.
On 5 January 2026, NVIDIA founder and CEO Jensen Huang took the stage at the Fontainebleau Las Vegas to open CES 2026. The company’s official presentation summary centres on two announcements: Rubin, a six-chip AI platform already in full production, and Alpamayo, a family of open models and tools for autonomous-vehicle research.
"Computing has been fundamentally reshaped as a result of accelerated computing, as a result of artificial intelligence," Huang said. In NVIDIA’s published account, he put the computing of the past decade now being modernised at “some $10 trillion or so”. This is a commercial estimate from the company’s own chief executive, not an independent market measurement.
Rubin: Cutting the Cost of Generating Tokens
The new platform's name honors American astronomer Vera Rubin. It succeeds the Blackwell architecture and, according to Huang, is the first "extreme-codesigned" AI platform with six chips, now in full production.
NVIDIA's technical argument is that scaling AI to what it calls "gigascale" requires designing every component at once — chips, trays, racks, networking, storage and software — to eliminate bottlenecks. That's what the company means by extreme codesign: not optimizing individual pieces, but the entire system at once.
The Rubin launch brief lists six codesigned components:
- Rubin GPUs with 50 petaflops of NVFP4 inference.
- Vera CPUs, engineered for data movement and agentic processing.
- NVLink 6 scale-up networking.
- Spectrum-X Ethernet Photonics scale-out networking.
- ConnectX-9 SuperNICs.
- BlueField-4 DPUs.
The figure NVIDIA wants the industry to remember is not raw performance but cost: the company says Rubin can reduce inference cost per token by up to tenfold versus Blackwell. The manufacturer’s announcement ties that result to mixture-of-experts workloads and the complete system; it does not promise that every model, configuration or provider will bill exactly one tenth as much.
"The faster you train AI models, the faster you can get the next frontier out to the world," Huang said. "This is your time to market. This is technology leadership."
Huang also unveiled “AI-native” storage, the NVIDIA Inference Context Memory Storage Platform, a layer for sharing KV cache in long-context inference. NVIDIA claims up to five times more tokens per second and up to five times better energy efficiency than traditional storage. The source does not quantify performance per dollar: speed, energy and cost are separate metrics.
Open Models as Strategy
The second pillar of the keynote is that NVIDIA doesn't present itself merely as a hardware maker, but as a builder of frontier models — and it does so, according to Huang, in the open.
"Now on top of this platform, NVIDIA is a frontier AI model builder, and we build it in a very special way. We build it completely in the open so that we can enable every company, every industry, every country, to be part of this AI revolution," he said.
The CES release of open models and tools spans Nemotron for agents, Cosmos for physical AI, Alpamayo for driving research, GR00T for robotics and Clara for biomedicine, alongside datasets and training frameworks. “Open” does not mean certified for real-world use: each artefact retains its own licence, hardware requirements and tested scope.
"Every single six months, a new model is emerging, and these models are getting smarter and smarter," Huang said, attributing the explosion in download numbers to that cadence.
The bet on open source carries an obvious strategic reading: if the models being downloaded and run at massive scale are trained and optimized on NVIDIA hardware, the company strengthens its position even without charging for the model itself. Open software drives sales of the hardware that runs it.
Personal AI: Agents Leave the Cloud
Huang insisted that the future of AI isn't just about supercomputers — it's personal too. He showed a demo of a personalized AI agent running locally on the DGX Spark desktop supercomputer, embodied in a Reachy Mini robot using Hugging Face models.
The idea he wanted to convey was the combination of open models, model routing and local execution turning agents into "physical collaborators." "The amazing thing is that is utterly trivial now, but yet, just a couple of years ago, that would have been impossible, absolutely unimaginable," he said.
On the enterprise side, Huang cited companies integrating NVIDIA AI into their products, including Palantir, ServiceNow, Snowflake, CodeRabbit, CrowdStrike and NetApp. "Whether it's Palantir or ServiceNow or Snowflake — and many other companies that we're working with — the agentic system is the interface," he said, summing up his vision that interacting with software will increasingly happen through agents rather than menus and buttons.
For DGX Spark, NVIDIA announced at CES up to 2.6 times faster performance since launch, plus support for LTX-2 and FLUX and future NVIDIA AI Enterprise availability. The “up to” and the baseline matter: this is not a uniform acceleration for every model.
Physical AI: From Simulator to the Road
The most concrete part of the announcement is the one that brings AI down to the physical world. NVIDIA's strategy involves training systems with synthetic data in virtual worlds before they interact with reality.
Huang showed the open world foundation models Cosmos, trained on video, robotics data and simulation. Cosmos generates realistic videos from a single image, synthesizes multi-camera driving scenarios, models edge-case situations from prompts, and performs physical reasoning and trajectory prediction.
The central announcement was Alpamayo, a portfolio of models, simulation and data for research towards future level 4 roadmaps. The primary source defines its scope: Alpamayo 1 is a 10-billion-parameter teacher model that does not run directly in a vehicle; developers must fine-tune or distil it within a complete stack. It includes weights and inference scripts, while AlpaSim supports training and evaluation in simulation.
"It doesn't just take sensor input and activate the steering wheel, brakes and accelerator — it also reasons about the action it's about to take," Huang explained, before screening a video of a vehicle driving through San Francisco traffic.
A manufacturer makes the announcement concrete, but not through Alpamayo or level 4. The specific Mercedes-Benz CLA announcement describes NVIDIA DRIVE AV with enhanced level 2 point-to-point assistance, expected in the United States by the end of 2026. At that level the person remains responsible for supervision; the five-star Euro NCAP rating covers the vehicle and its safety functions, not level 4 autonomy.
The keynote also leaned on the DRIVE Hyperion platform, described as open, modular and level-4-ready, and on a robotics ecosystem Huang illustrated with robots trained in the simulated Isaac Sim and Isaac Lab environments, alongside partners such as Synopsys, Cadence, Boston Dynamics and Franka. He was joined onstage by Siemens CEO Roland Busch to announce an expansion of their industrial software alliance.
"These manufacturing plants are essentially going to be giant robots," Huang said. And about the car: "Our vision is that someday every car, every truck, will be autonomous, and we're working toward that future."
What to Watch From Here
NVIDIA's script at this CES is consistent with its position: it sells a full platform — hardware, networking, storage and models — and presents open source as the lever that turns that stack into the standard everyone builds on. The promise of cutting the cost of generating tokens tenfold is the number companies will actually measure, because it determines whether deploying AI at scale stops being prohibitively expensive.
In driving, the verifiable commitment is different: a Mercedes-Benz CLA with enhanced level 2 DRIVE AV in the United States by the end of 2026. Alpamayo 1 is research material for developing and evaluating future stacks; “level-4-ready” describes a roadmap, not a certified capability of the announced car. Separating product, comparison baseline and availability condition is the useful way to read this kind of presentation.
This article was produced with artificial intelligence under human editorial oversight.