AI Briefing
KO
Pick

Why 'Performance per Watt' Is the Ultimate Metric for AI Infrastructure Efficiency

·2026.07.15 00:00

Key point

NVIDIA highlights performance per watt as the key metric determining the profitability of AI infrastructure in power-constrained environments, showcasing the overwhelming efficiency of the Blackwell platform.

1 / 2

Details

In AI infrastructure, power is an unavoidable constraint. This is because the amount of token that an AI factory can generate within a fixed power budget is directly tied to a company's profitability. Therefore, performance per watt, a real, unmanipulable outcome, becomes the core foundation of AI infrastructure.

Most of the latest frontier AI models have adopted the Mixture-of-Experts (MoE) architecture. To serve these models smoothly at rack scale, sophisticated codesign across the entire system and software stack is essential.

The NVIDIA Blackwell NVL72 platform is designed to meet these requirements, delivering overwhelming efficiency compared to the previous Hopper generation.

  • DeepSeek V4 Pro: Delivers up to 25x performance per watt compared to Hopper
  • GLM5.1: Delivers up to 20x performance per watt compared to Hopper
  • Kimi K2.6: Delivers up to 10x performance per watt compared to Hopper

Instead of a single figure, NVIDIA presents Pareto curves to show the optimal operating points for various workloads (latency-optimized vs. throughput-optimized), and supports finding the optimal point before verification through tools such as DynoSim.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.