NVIDIA Emphasizes AI Factory Efficiency: Optimizing 'Tokens/Watt' with Vera Rubin and DSX
Key point
At the AI Infra Summit, NVIDIA introduced the Vera Rubin and DSX platforms, highlighting a shift in key AI factory metrics toward 'validated agentic tokens per megawatt,' while partner Lambda achieved a 23% improvement in performance per watt using DSX MaxLPS.
Details
At the AI Infra Summit, NVIDIA unveiled its full-stack AI factory strategy centered on the Vera Rubin system and the DSX platform, emphasizing that the core metric for infrastructure is shifting from peak performance to validated agentic tokens per megawatt. To achieve this, code signing from silicon to grid is essential, and they explained that power optimization across the entire factory can be achieved through DSX MaxLPS.
DSX MaxLPS and Power Optimization
DSX MaxLPS monitors GPU and rack power consumption in real time and dynamically reallocates it to recover unused capacity caused by static allocation. AI cloud provider Lambda released the first validation results based on NVIDIA Blackwell servers, showing that they could run 19 nodes within a budget of 16 full-power nodes. This resulted in a 24% increase in total cluster token throughput (from approximately 4 million to 5 million tokens per second) and a 23% improvement in performance/watt. The next-generation Vera Rubin NVL72 AI factory can secure up to 40% more GPU capacity within the same megawatt budget in the right deployment environment.
Agentic AI Performance and Benchmarks
NVIDIA Groq 3 LPX adds deterministic ultra-low-latency inference to Vera Rubin to address latency and context length issues in agentic AI. According to the SemiAnalysis AgentX benchmark, Vera Rubin NVL72 achieved up to 30x higher throughput per megawatt on the DeepSeek V4 Pro model compared to GB300 NVL72, with token costs reduced by up to 45x. This is a key factor determining the revenue-generating capability and margins of AI factories in power-constrained environments.
Grid Flexibility and Partnerships
Emerald AI demonstrated a flexible load program using NVIDIA DSX Flex that automatically adjusts power for AI workloads based on grid signals. Emerald AI and NVIDIA collaborated with Silicon Valley Power to successfully respond to hundreds of demand signals, demonstrating automatic load shedding that protects AI workload performance. Additionally, Amazon Annapurna Labs is co-developing NVHBM custom high-bandwidth memory technology with NVIDIA, and d-Matrix is integrating the NVLink Fusion platform into its Raptor XPUs.
Vera CPU and NVLink 6
The NVIDIA Vera CPU demonstrated superior performance in agent workloads and data-intensive tasks. Perplexity reported a 1.9x improvement in sandbox startup speed, while Redpanda reported a 5.5x reduction in latency and a 73% increase in throughput. NVLink 6 applies a multi-layer resilience architecture for reliability in large-scale AI factory environments, isolating faults locally through forward error correction (FEC) at the physical layer and credit-based flow control at the network layer.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.