NVIDIA puts megawatt figures behind running AI clusters below full power

NVIDIA has published the operating figures behind an idea it has been promoting for months: that running more GPUs, each capped below full power, produces more useful output than running fewer at full power within the same electrical envelope.

The worked example is Lambda. Nineteen nodes held at 85 per cent power were compared with 16 nodes at full power. Cluster-wide throughput rose from 4 million to 5 million tokens per second, a gain of 24 per cent, with performance per watt up 23 per cent. NVIDIA projects up to 40 per cent additional GPU capacity for next-generation Vera Rubin NVL72 installations within an equivalent power budget — a projection rather than a measurement.

On grid response, the 15 September post says a demand signal from Silicon Valley Power cut consumption from 4 megawatts to 3 in under a minute with no jobs lost, and reports 200 demand-response signals handled without failure. NVIDIA separately projects a 3 to 5 per cent end-to-end efficiency gain from 800-volt DC distribution over the 54-volt approach used today, with hardware expected in 2027.

These are vendor figures on the vendor’s own hardware, and 200 signals without failure is a count, not an audited reliability record. But a megawatt shed inside a minute is a specific claim a utility can check.


Related