NVL72.dev

Chapter 06

Power & Cooling

~120 kWRack powerSources disagree: 120 kW nominal · 125–135 kW operating (Supermicro) · 132 kW fully loaded (Schneider Electric) Supermicro’s datasheet states an operating power of 125–135 kW and 132 kW of installed power-shelf capacity. Steven Carlini, writing for Schneider Electric: "When fully loaded into a rack, the latest NVIDIA-based GPU servers require 132 kW of power." The commonly quoted ~120 kW is the nominal design figure, not a measured ceiling. Supermicro · Steven Carlini, Schneider Electric — Forbes Technology Council · ServeTheHome in a 0.64 m² footprint. The worldwide mean rack density is 7.6 kWWorldwide mean rack density Up from 6.8 kW the prior year; 8.4 kW if racks above 30 kW are excluded. An NVL72 is roughly sixteen average racks in one footprint. Uptime Institute . Everything in this chapter follows from that ratio.

Sixteen racks in one

Rack power density, to scale

Uptime Institute 2025 mean vs NVL72
Worldwide mean rack 7.6 kW GB200 NVL72 ~120 kW Next generation 240 kW — announced, not measured (Schneider Electric)

This figure is interactive and needs JavaScript. The prose around it states every number it shows.

A normal computer rack in a data centre uses about as much electricity as a few homes. This one uses about as much as a small apartment block — in the same floor space.

Two things follow. The electricity has to arrive differently: instead of each machine having its own power supply, the whole rack is fed by a single thick copper bar and every tray taps into it. And the heat has to leave differently: there is far too much of it for fans and air, so cold water is piped directly onto the chips.

On the electrical side, per-server PSUs are gone. 8 × 1UPower shelves Arranged 4 + 4, each 33 kW from six 5.5 kW supplies — 132 kW of installed shelf capacity feeding the busbar. Supermicro power shelves — two banks of four, six 5.5 kW supplies each — convert facility input once and drive a shared busbar rated for 1,400 ABusbar current capacity NVIDIA Technical Blog , which runs the height of the rack; trays draw from it directly.Supermicro,NVIDIA Technical Blog Converting once at rack scale beats converting fifty-four times at server scale on efficiency, and it reclaims the volume that redundant supplies would have occupied — volume this design needs for cold plates.

On the thermal side there is no air-cooled fallback. Direct-to-chip cold plates cover every significant package, fed from a rear manifold through blind-mate quick disconnects, with a CDU isolating the rack loop from facility water. Inlet is deliberately warm — 32 – 45 °CCoolant inlet temperature NVIDIA ACS reference design. A 45 °C maximum inlet and 65 °C maximum return are attributed to QCT across several sources; that document could not be read directly, so treat the limits as second-hand. Warm water is the point either way: it is what enables free cooling. QCT per NVIDIA's ACS reference design, with QCT specifying 45 °C maximum inlet and 65 °C maximum return.QCT

Warm-water operation is the entire economic argument for liquid at this density. If the inlet can be 45 °C, then for most of the year in most climates the facility side needs only dry coolers and pumps — no compressor, no chiller plant. PUE approaches the pumping overhead rather than the refrigeration overhead. The same 45 °C return water is also hot enough to be worth something to a district-heating loop, which is why heat reuse appears in operator planning around these deployments rather than as an afterthought.

The counterweight is that thermal margin is now a facility property, not a rack property. A tray has no fans to spin up and no throttle-and-survive mode that removes 5 kW into still air. Loss of flow is measured in seconds to thermal event, which moves CDU and pump redundancy from a nice-to-have into the same tier as power redundancy.

The loop, and a number that does not add up

The research behind this site flagged coolant flow as a figure the sources disagree about: 2–3 L/min per module, 30–40 L/min per rack in NVIDIA's ACS reference, and up to about 130 L/min per rack from QCT. Those are not small differences. Rather than pick one, it is more instructive to check them against physics.

Coolant loop

Drag flow and inlet temperature
Facility water CDU heat exch. rack loop — 18 compute trays cold plates supply 32 °C return 45 °C
Heat removed
120 kW
Temperature rise
13.2 K
Return temperature
45.2 °C
Within 65 °C limit
yes

This figure is interactive and needs JavaScript. The prose around it states every number it shows.

Return temperature from ΔT = Q ÷ (ṁ · c_p), with Q = 120 kW, water at 4.18 kJ/kg·K. The point of the widget is the sanity check: at 30–40 L/min a 120 kW rack would need a ΔT of 40 K or more and would blow straight past the 65 °C return limit — so those published figures cannot be describing a full-load rack loop.

Power and thermal figures

Power and cooling
Rack power ~120 kW ± sources disagree: 120 kW nominal · 125–135 kW operating (Supermicro) · 132 kW fully loaded (Schneider Electric). Supermicro’s datasheet states an operating power of 125–135 kW and 132 kW of installed power-shelf capacity. Steven Carlini, writing for Schneider Electric: "When fully loaded into a rack, the latest NVIDIA-based GPU servers require 132 kW of power." The commonly quoted ~120 kW is the nominal design figure, not a measured ceiling.
Power per GPU ~1,200 W The commonly cited per-GPU board power for Blackwell in GB200. Rack power divided by 72 lands higher because the figure excludes the CPUs, NVSwitch trays, NICs and conversion losses.
Worldwide mean rack density 7.6 kW Up from 6.8 kW the prior year; 8.4 kW if racks above 30 kW are excluded. An NVL72 is roughly sixteen average racks in one footprint.
Coolant inlet temperature 32 – 45 °C NVIDIA ACS reference design. A 45 °C maximum inlet and 65 °C maximum return are attributed to QCT across several sources; that document could not be read directly, so treat the limits as second-hand. Warm water is the point either way: it is what enables free cooling.
Coolant flow ~2–3 L/min per module ± sources disagree: 2–3 L/min per module · 30–40 L/min per rack (NVIDIA ACS) · up to ~130 L/min per rack (QCT). Figures vary by whether they are quoted per cold plate or per rack, and by the assumed ΔT. Always state the basis.
Cost of liquid cooling per MW ~$2M retrofit ± sources disagree: ~$2M per MW to retrofit · upwards of $11M per MW for a new greenfield liquid-cooled build. STL Partners, May 2026. A widely repeated "$5–10M per MW" retrofit figure is often attributed to Schneider Electric; it does not appear in the Schneider article this site cites, and no primary source for it could be found — so it is not used here.
Rack weight ~1.36 t About 3,000 lb — a floor-loading problem for most existing halls. NVIDIA’s OCP contribution describes over 100 lb of added reinforcement steel in the frame alone.

The load-swing problem

A rack like this does not draw a steady 120 kW. Thousands of GPUs in a training run enter and leave collectives in lockstep, so the facility sees something closer to a square wave, and at sufficient scale that becomes a grid-interaction problem rather than a data-centre one.

NVIDIA's multi-node tuning guide names Power Smoothing as the implemented answer for bulk synchronous workloads.NVIDIA Docs The GB300 NVL72 version combines programmable power caps, energy-storage-enhanced power shelves carrying integrated electrolytic capacitors, and a hardware power burner, with software shaping the ramp-up, steady-state and ramp-down phases. NVIDIA reports a measured reduction in peak grid demand of up to 30 %Reduction in peak grid demand from power smoothing Measured on GB300 NVL72 training Megatron, using programmable power caps, energy-storage-enhanced power shelves with integrated electrolytic capacitors, and a hardware power burner across ramp-up, steady-state and ramp-down. NVIDIA states the feature is also coming to GB200 NVL72. NVIDIA Technical Blog while training Megatron, and states the feature is also coming to GB200 NVL72.NVIDIA Technical Blog

The power burner is the counter-intuitive part and worth stating plainly: part of the solution to a load that falls too fast is to deliberately waste power so that it falls more slowly. Grid equipment cares about the rate of change, not only the magnitude.

Why this section used to be one sentence

The research this site was built from flagged power smoothing as its thinnest-sourced topic and recommended finding a dedicated NVIDIA or OCP power paper before writing it up. That source has since been located — NVIDIA's own tuning guide and the GB300 power blog — so the section above now exists. It is a reasonable illustration of how the known-gaps list is meant to work: a gap is an open item to be closed, not a permanent disclaimer.