Nvidia unveils Vera Rubin CPU-GPU system for AI data centers
If the Vera Rubin system lives up to its claims, Nvidia could redefine AI data-center economics with near one CPU per two GPUs, higher memory bandwidth, and simpler rack installation, influencing costs and energy use in AI workloads.
At a glance
- Vera Rubin NVL72 system pairs 36 Vera CPUs with 72 Rubin GPUs for a 1:2 CPU-to-GPU ratio.
- Nvidia claims Vera Rubin racks deliver ~10x tokens per watt versus Grace Blackwell and offer nearly triple the memory bandwidth.
- The system is designed to be more plug-and-play, with reduced cabling and 100% liquid cooling for energy efficiency.
- OpenAI is reported to already have one Vera Rubin rack in use.
The story
Nvidia introduced the Vera Rubin NVL72 system as the centerpiece of its next-generation AI data-center strategy, positioning Vera Rubin as a key platform to run large-scale AI workloads. The configuration pairs 36 Vera CPUs with 72 Rubin GPUs, delivering a CPU-to-GPU ratio designed for orchestration of AI tasks across devices.
Nvidia executives claimed the Vera Rubin racks offer dramatically higher efficiency, including an estimated tenfold improvement in tokens per watt versus the Grace Blackwell system. The company also highlighted localized memory subsystems and higher memory bandwidth, which are intended to improve performance for agentic AI workloads.
The hardware is marketed as highly integrated and easier to deploy, with significant reductions in cabling and a fully liquid-cooled design. Nvidia described Vera Rubin as providing “cable-free compute” and “hot-swappable” components to speed up rack installation.
A hardware deployment detail cited by the article notes that OpenAI already operates a Vera Rubin rack, underscoring the system’s readiness for production-scale AI tasks.