← Back to Intel
InfrastructureJune 18, 2026

An ASHRAE Case Study Cut Cooling Power 15% by Running the Spares Hot. It Ran at 8.8 kW a Rack.

Solomon Obadimu wants operators to stop parking redundant cooling equipment in the off position. Digital Realty's program manager for OT remediation made the case in DataCenterDynamics' June 18 opinion pages. Hot standby sparing should displace cold standby sparing as the industry default. He builds the section on Cho et al. (2024), published in the February 2024 ASHRAE Journal. The paper modeled a 30MW facility and reported that running the redundant cooling units at partial load instead of leaving them dark cut total cooling power consumption 15 percent and moved cooling PUE from 1.23 to 1.20. His thesis: "How redundancy is implemented is more impactful than how much redundancy is in place." Obadimu also puts cooling failures at 51 percent of data center downtime, the single largest cause of interruption, a figure he takes secondhand from a study Cho cites.

The Case Study Ran at 8.8 kW a Rack

Here is the number that never makes the headline. Cho's 30MW facility carried 8.8 kW of IT load per rack. That is air cooling. A GB300 cabinet draws past 100 kW, and the industry has already moved from 2 kW to 130 kW per cabinet in forty years. Obadimu titled the piece for the AI-driven data center, then sourced his cooling evidence from a hall running roughly fifteen times below AI density. The 15 percent holds inside Cho's assumptions. Whether it survives the jump to a direct-to-chip loop is a separate question, and the piece never tests it.

Liquid Loops Do Not Give You Minutes

The redundancy math changes when the working fluid changes. An 8.8 kW room has thermal mass to spend: room air, floor plenum, chilled water volume. All of it buys an operator time to start a cold spare and let it stabilize. A direct-to-chip loop carries almost none of that. The integrator Introl puts ride-through at under 10 seconds at full load once the pump stops, shorter than the start sequence on most standby mechanical equipment. That is the strongest argument available to Obadimu, and he never reaches for it. At AI density, hot standby is a reliability requirement. A CDU with 1+1 pumps backed by a cold spare chiller still cooks the die if that chiller needs sixty seconds to come up. Manifold flow balancing is the other half of it. A failed unit does not degrade evenly across a row.

Power and Servers Get the Same Argument

Obadimu runs the identical logic through two more pillars. On power he cites Takci et al. (2025) putting 10 to 50 percent of UPS capacity unused even during outages, and proposes that 2N and 2(N+1) sites sell that headroom back to the grid as a revenue line. On servers he cites Shaukat et al. (2022) in Sensors, where an energy-aware fault-tolerant scheduler cut unfinished processing tasks 80 percent during failures with servers simulated at 75 percent load, while requiring 27 percent fewer backup devices than GreenCloud, against a fleet where utilization sits between 12 and 30 percent and idle servers still burn roughly 66 percent of what an active one burns. He closes on aviation, recommending dissimilar redundancy across different hardware and software to kill common-mode failures.

Check the commercial angle and it runs against him. Obadimu works for an operator that sells redundancy tiers, and his conclusion argues for buying less of it and running it harder. The harder problem sits downstream. If NVIDIA's 45C warm-water architecture removes the mechanical chiller entirely, the cold-versus-hot standby debate loses its subject, and redundancy relocates to the CDU, the pumps, the manifolds, and the dry coolers. Somebody needs to run the Cho analysis again on a 130 kW rack with a 45C supply and a ten second failure window. Nobody has published that number.