NVIDIA Unveils Rubin at CES 2026 and the Alpamayo Research Model
NVIDIA introduced Rubin with a claim of up to 10x lower inference cost per token than Blackwell and Alpamayo, an open teacher model for research into planning and reasoning traces in autonomous driving.
On January 5, 2026, at the opening of CES in Las Vegas, Jensen Huang introduced Rubin and Alpamayo. Rubin succeeds Blackwell as a data-center platform; Alpamayo combines a research model, simulation, and datasets for autonomous-driving development. The announcement placed them under one commercial thesis: extend accelerated computing and AI from data centers into physical machines.
Huang said “some $10 trillion or so” of the previous decade's computing was being modernized toward this new form of computing. NVIDIA's official presentation summary reproduces the statement but supplies no methodology, calculation period, or breakdown. Without those elements it cannot be treated as a market measurement: it is a vendor estimate and should be read with that limitation.
Rubin: six chips designed as a single piece
The announcement with the most technical weight is Rubin, named after American astronomer Vera Rubin and, according to the company, already in production. NVIDIA calls it its first “extreme-codesigned” AI platform: six chips and system software conceived as a rack, network, and storage platform. That description comes from the supplier; it is not yet an independent evaluation of a deployed system.
The technical list published by NVIDIA contains six components:
- Rubin GPU, delivering 50 petaflops of inference in NVFP4 format (a low-precision numerical representation that allows models to run faster while using less memory).
- Vera CPU, geared toward data movement and agent processing.
- NVLink 6 for internal server interconnection.
- Spectrum-X Ethernet Photonics for the network linking racks together.
- ConnectX-9 SuperNICs.
- BlueField-4 DPU, processors dedicated to offloading networking and storage tasks.
The logic behind designing everything together is to avoid bottlenecks. When a model is trained or run across thousands of chips at once, performance isn't set by the fastest GPU but by the weakest link in the chain: the network, storage, memory. Huang argues that integrating every layer drastically cuts the cost of training and inference.
The announcement's central number is narrower: NVIDIA claims up to 10x lower inference cost per token than Blackwell and four times fewer GPUs to train mixture-of-experts models. A token is a unit of model input or output, but a service's final cost also includes utilization, memory, networking, energy, software, and provider margin. Without configuration, workload, latency target, and total cost, the vendor multiplier cannot be converted into a specific bill.
NVIDIA also introduced a key-value-cache storage platform intended to reuse inference context. The announcement attributes five times more tokens per second, performance per dollar, and power efficiency to it. The summary does not provide the comparison's complete configuration; these are declared performance figures, not universal results for every model or context length.
A strategic pivot: NVIDIA as a model builder
The most interesting move from a business standpoint isn't in the hardware, but in Huang's insistence on presenting the company as a "frontier AI builder." NVIDIA has spent years making money selling the picks and shovels of the gold rush; now it wants to mine too.
The presentation grouped six families under NVIDIA's open-model label: Clara for healthcare, Earth-2 for climate, Nemotron for reasoning and multimodality, Cosmos for robotics and simulation, GR00T for robotics, and Alpamayo for autonomous driving. “Open” does not by itself confer identical rights across them: weights, code, data, and license must be checked for each release.
Huang said the company built these models “completely in the open.” The specific card adds an important qualification: Alpamayo 1's weights were released under a non-commercial license, while its inference code uses Apache 2.0 and commercial licensing is available on request. Openness is measured separately by what may be downloaded, modified, redistributed, and used in a product—not by the umbrella label.
Cars that reason and explain their decisions
The Alpamayo announcement combines models, simulation, and data for autonomous-driving research and development. NVIDIA connects the package to Level 4 roadmaps, but that objective does not mean the published model is approved to drive a vehicle without human intervention.
The centerpiece, introduced as Alpamayo R1 and later renamed Alpamayo 1, is a roughly 10-billion-parameter vision-language-action model. Its model card lists multi-camera video, text, and motion history as inputs, and a future trajectory plus a textual chain-of-causation trace as outputs. That trace shows the justification the system generates; by itself, it does not prove that the text exposes the internal process that caused the trajectory.
NVIDIA connected the presentation to the new Mercedes-Benz CLA and its DRIVE platform, but the Alpamayo release itself says the large models do not run directly in the vehicle: they are teacher models that developers may fine-tune and distill into a complete stack. The model card also places Alpamayo 1 in the cloud for research and requires testing with use-case-specific data before safe deployment.
Rounding out the package is AlpaSim, an open simulator for evaluating driving policies in closed loop with configurable sensors, vehicle dynamics, and traffic. Simulation can repeat difficult cases without creating the danger on public roads, but fidelity is part of the test: a result transfers to the physical world only to the extent that the simulator represents the relevant conditions.
From supercomputer to desktop
Huang showed an agent running locally on DGX Spark, represented by a Reachy Mini robot using Hugging Face models. The CES demonstration establishes that prepared configuration; it does not show that every agent, model, or device can run locally with the same ease or without external services.
Among the companies integrating its technology, NVIDIA cited Palantir, ServiceNow, Snowflake, CrowdStrike, and NetApp. "The agentic system is the interface," Huang summed up, pointing to a future where people interact with software through agents rather than menus and buttons.
There was also room for the business that gave rise to the company in the first place: video games. NVIDIA announced DLSS 4.5, featuring a new frame-generation mode and more than 250 games compatible with its DLSS 4 technology, along with G-SYNC Pulsar monitors and the arrival of GeForce NOW on Linux and Amazon Fire TV.
What it all means
The underlying message of CES 2026 is one of continuity and ambition. NVIDIA no longer just sells chips; it sells full platforms —from the data center to the robot and the car— and wants a presence at every layer of the AI value chain, including building the models themselves.
NVIDIA's own published limitations are part of the announcement. The Rubin release says future performance, availability, and partner arrangements are forward-looking and subject to change, and places partner products in the second half of 2026. The Alpamayo card requires iterative validation with use-case-specific data and does not present the teacher model as a road-ready autonomous product.
The transferable skill is to separate platform, metric, and deployment. For Rubin: which configuration produces “up to 10x,” against which Blackwell system, and at what total cost? For Alpamayo: which part is a teacher model, what runs in the vehicle, what license applies, and which tests bridge simulation and roads? Those questions outlast CES and keep the next vendor number from being mistaken for an operational result.
This article was produced with artificial intelligence under human editorial oversight.