Garp Independent AI & technology journalism
Tuesday, July 21, 2026 Sign In · Join Subscribe
Latest Wristband enables wearers to control a robotic hand with their own movements

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency

AI News

Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency

Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency

NVIDIA Corp. — power is AI infrastructure’s inescapable constraint. How many tokens an AI factory can generate within a fixed power budget determines its revenue and profitability.

Because of this, performance per watt — a metric that can’t be gamed, only earned through real-world results — is the foundation for AI factories.  As agentic AI drives token demand higher, the infrastructure decisions organizations make today will determine who scales and who doesn’t in a power-constrained world. Virtually every frontier AI model today runs on a mixture-of-experts (MoE) architecture. Serving MoE at rack scale demands codesign across every layer of the system and software stack, plus the operational depth earned from running these models under real production load. With the NVIDIA Blackwell NVL72 platform, that rack-scale foundation is already built and proven, delivering the highest performance per watt to maximize revenues and the lowest token cost to maximize profit margins. It’s this foundation that the NVIDIA Vera Rubin platform builds upon next to further elevate rack-scale energy efficiency. Maximizing Performance per Watt for Frontier AI  Each new generation of frontier models brings architectural changes that unlock greater intelligence while demanding new optimizations to run efficiently at scale.  Across the newest generation of leading open models, NVIDIA GB300 NVL72 delivers up to 25x performance per watt compared with the NVIDIA Hopper generation. These numbers reflect where Blackwell stands today, a starting point that continues to improve.  Any single number only tells part of the story. Different workloads demand different operating points: some optimize for latency, others for throughput and cost — and most need to move between the two.  To best represent these operating points, NVIDIA showcases Pareto curves for each model rather than a single point and provides tools such as DynoSim to help teams find their optimal point on the Pareto frontier before spending a single GPU-hour on validation. NVIDIA GB300 NVL72 systems deliver up to 25x performance per watt over NVIDIA Hopper on DeepSeek V4 Pro. 1 NVIDIA GB300 NVL72 systems deliver up to 20x performance per watt over NVIDIA Hopper. 6, a model purpose-built for long-horizon agentic tasks. That codesign touches every layer of the stack.   For example, NVIDIA NVLink Switch, critical for rack-scale performance, is purpose-built for scale-up GPU domains, not adapted from general-purpose networking. Now in its sixth generation with the Vera Rubin platform, its capabilities are designed specifically for AI workloads such as SHARP, which performs in-network computing directly in the switch, offloading work from the GPUs themselves. NVIDIA’s inference software stack, including NVIDIA Dynamo and TensorRT LLM, as well as SGLang and vLLM, is built to run the full range of optimizations: NVFP4 quantization, disaggregated serving, large-scale expert parallelism, KV-aware routing, KV cache offloading and more. These stack together to multiply the performance each GPU delivers.