Redundant on Paper: How to Tell Whether Your CDU's Spare Pump Will Protect the GPUs
A coolant distribution unit is normally built with a spare pump. Whether that spare protects the GPUs depends on how it is configured, how long the controls wait before using it, and how much load has been added since commissioning.

At 10:00 on a Monday, the coolant distribution unit (CDU) in a GPU rack stops its running pump and starts the spare. That schedule is the default in at least two in-rack CDU manuals, and one of them gives the changeover sequence as about a quarter of a second. The rack keeps running and no alarm appears.
That quiet is what a spare pump is for, but it proves less than it seems. A scheduled changeover is a healthy pump handing over to another healthy pump at a time someone chose. A real failure is a pump that weakens, trips or loses power at a time nobody chose, and the GPUs may have less time than the controls take to notice.
Whether N+1, one more pump than the load needs, protects your chips depends on four things you can check: which kind of spare you have, how long the controls wait to use it, which thermal mass stays connected meanwhile, and how much of the spare added racks have already taken.
New CDUs put conditions on the spare
Two CDU launches in late September put conditions on their redundancy. Schneider Electric's announcement of 23 September rates one unit at up to 3.5 MW at 1.5 liters per minute of coolant per kilowatt, or 2.5 MW at 2.0. N+1 pump and valve backup is an option for deployments up to 2.5 MW.

LiquidStack's page for its CDU 2.X, launched by Trane Technologies on 29 September, says single-pump mode "supports internal N+1 pump redundancy" and dual-pump mode "supports CDU-to-CDU redundancy". The mode decides where the spare sits.
NVIDIA's DSX reference design (Design Guide v2.0, 19 August 2026) puts CDUs in N+1 groups and asks for at least 1.5 liters per minute per kilowatt in the technology cooling system (TCS), the loop between the CDU and the cold plates. That is the ratio behind Schneider's 3.5 MW headline, while Schneider offers its N+1 option only up to 2.5 MW.
We found no public measurements of how quickly current GPUs throttle when that flow stops. In June 2026 an ASHRAE committee approved the recommended bidder for a research project on direct-to-chip failure modes and their throttling impact on IT equipment. No results are public yet.
Standby or sharing: know which spare you have
A CDU with two or three pumps can use them in two broad ways, and the difference decides what happens in the first seconds of a failure.
Run and standby
One pump, or one set of pumps, carries the flow while another waits idle. If the running pump stops, the controller must notice, start the standby and bring it up to speed. Until it does, flow falls.
Sharing the flow
All pumps run together at part speed. If one stops, no pump has to start from rest, but the others must speed up to cover the gap, which they can do only if they have speed to spare.
The same hardware can sometimes do either. Vertiv's guide specification for its CoolChip CDU 600 and 1350 describes pumps that are "configurable for either run/standby" or "simultaneous operation for maximum TCS flow". That second option is where a spare can vanish: if the design flow needs every pump running, none of them is spare.
In its 2024 resiliency bulletin, ASHRAE's technical committee for data centers (TC 9.9) recommends active redundancy on loop pumps so that cooling continues through a changeover. The bulletin does not define the term. Our reading is pumps already running and sharing the flow, rather than a pump that must start after another stops.
The bulletin also reports that IT equipment makers have seen TCS temperature jumps of as much as 27 degrees C when CDU pumps changed over without active redundancy. A changeover is not always as quiet as the Monday one. Whatever wording your supplier uses, check the mode actually configured in the CDU controller.
The alarm can be slower than the chip
Put the published timings on one ruler and the gap is plain: switching takes a fraction of a second, but noticing can take minutes.

In the default settings of two in-rack CDU manuals, Lenovo's RM100 and Vertiv's CoolChip CDU 121, a pump that keeps running but cannot reach 90 percent of its target is given 100 seconds before the controller declares a fault. The target is differential pressure, the supply-to-return pressure difference the pump holds, or in Lenovo's unit flow or differential pressure. In Lenovo's unit, the fault is also when the standby pump is started, and the manual describes no faster trip for a pump that stops outright.
In a 2017 Binghamton University laboratory study, server processors began to throttle, slowing themselves to stay within temperature limits, 23 seconds after all coolant flow stopped, at full load with 45 degrees C coolant.
The study used CPU servers of that era, not today's GPUs, so 23 seconds is not a forecast for your hall. The pattern is more useful than the seconds. Load mattered most: at idle, throttling began only after 305 seconds. Colder coolant helped less, stretching the full-load time from 23 to 34 seconds at 20 degrees C. ASHRAE's bulletin likewise notes that a lower supply temperature class can give more ride-through, so warm-water designs, chosen for efficiency, start with less headroom.
So compare two numbers for your own system: how long the controller waits before switching away from a pump that is not delivering, including one that has stopped outright, and the server maker's time-to-throttle for your busiest GPUs at your supply temperature after a complete loss of flow. If the first is longer, the spare pump starts after the chips have begun to slow down.
Ask the CDU supplier whether a faster signal, such as a pump drive fault, also triggers the changeover, and whether the delay can safely be shortened. ASHRAE adds that any plan to move IT load away from a cooling failure must work within that shortest time-to-throttle.
A big loop only helps while the coolant moves
Stored coolant is often described as a buffer. That is true in one failure and close to false in another.
If heat rejection is lost but the CDU pumps keep running, every liter in the loop keeps passing the cold plates. The whole loop warms together: in the teaching example below, 20 K (20 degrees C) of headroom lasts about 80 seconds. If the pumps stop, the coolant in the pipes is no longer connected to the chip. What is left is the cold plate metal and the small amount of coolant inside it, which warms about ten times faster and uses the same headroom in about 8 seconds.

ASHRAE's bulletin recommends uninterruptible power supplies (UPS) that keep the TCS operating when utility power is lost. NVIDIA's DSX reference design says its GPU-side cooling plant is not on UPS and relies on the thermal mass of the facility chilled-water loop until generators start. The page does not say whether the CDU pumps are on UPS, and that stored cold helps the chips only while pumps on both sides of the CDU keep running.
So find out which pumps and controls ride through a power dip, and how they come back. ASHRAE asks designers to allow for the rate of change during recovery, not only during the failure.
Vertiv's CDU 600/1350 specification has pumps restarting within 5 seconds after a system reboot, but it does not say how long the reboot itself takes. Lenovo's RM100, by default, resumes running when power returns and ignores high and low temperature alarms for the first 20 minutes after start-up, so those alarms will not flag a fast swing during the recovery.
Redundancy is used up, not switched off
N+1 holds only while the remaining pumps can carry the demand. Suppose three identical pumps serve a loop. While the demand is within what two of them can deliver, losing one is survivable. Once demand goes above that, every pump is needed and there is no spare, although the panel still shows three healthy pumps. The loop is redundant on paper only. Nor will the Monday changeover warn you as the margin shrinks: it swaps one healthy pump for another, and the same number keep running.

The same thing happens one level up, when several CDUs share a load as a group, and there the controls may not catch it. Vertiv's manual for its in-rack CoolChip CDU 121 says that if the load exceeds the capacity of the running units, the standby units "will not kick in automatically", and it tells the operator to increase the configured number of duty units when load has been added.
A working draft of an Open Compute Project CDU guide proposes delivering at least 95 percent of rated cooling capacity under N+1 redundancy. It is a draft, not a standard.
After every load addition, recheck two things at your design flow per kilowatt: can the CDU still deliver that flow with one pump out, and is the controller configured to use the spare it has?
Treat the weekly changeover as a free test
A weekly rotation gives you 52 small, planned disturbances a year, each at a known time. This only works where pumps run as duty and standby. Pumps that share the flow give no handover to record, so the acceptance tests in the next section matter more. To see the changeovers, log flow, differential pressure and GPU data every second or faster around that time. A trend averaged over a minute can hide the whole event. For each changeover, keep:
- Time, outgoing and incoming pump, IT load
- Lets you compare like with like across weeks
- Loop flow and differential pressure before, during and after
- Shows the size of the dip and how quickly it recovers
- CDU supply temperature
- Shows whether the changeover disturbed the coolant
- GPU clock-event reasons and thermal margin on the hottest racks
- Clock-event reasons show why a GPU lowered its clocks, for example a power cap or thermal slowdown. NVIDIA's Data Center GPU Manager (DCGM) also reports, where supported, each GPU's margin to its nearest slowdown threshold.
One changeover tells you little. If the dip is larger every time one particular pump takes over, or grows over several weeks at similar load and supply temperature, look first at the incoming pump, the check valve (non-return valve) on the pump that stopped, and the controls. Our earlier article on diagnosing hotter GPUs lists a recent pump changeover among the first work-log entries to check when chips run warm.
Choose the changeover time deliberately, through the manufacturer's settings. A busy hour gives more useful evidence, and higher stakes, than a quiet one.
Test the failover without creating the incident
A scheduled rotation only tests the easy case. Factory and site acceptance tests are the right place to prove the harder ones. The Open Compute Project draft lists redundancy switchover among the items a factory acceptance test should verify, but it sets no pass criteria in seconds or flow, so write your own.
ASHRAE recommends hydraulic and thermal commissioning with thermal test vehicles or the real IT equipment. A test vehicle is a heater: it shows how the loop responds, not when a GPU slows itself, so get the time-to-throttle figure from the server maker.
After handover, any live failover test belongs inside the site's change process and the manufacturer's procedure, at a known load, with someone watching the GPU telemetry and a clear way to stop.
Do not pull power from a live pump to see what happens. If the test cannot be done safely at load, that is a finding in its own right, and it belongs in the risk register.
Where Reliability Engine fits
Reliability Engine connects loop behavior, coolant condition and GPU thermal context, so that events such as pump changeovers are compared over time rather than forgotten. It does not replace the CDU's controls, its protection or the manufacturer's procedures. Our direct-to-chip maintenance page describes the approach.
To start recording your own pump changeovers against GPU behavior, talk to Reliability Engine.
Subscribe to updates
Get the latest engineering perspectives sent straight to your inbox.
References
- ASHRAE TC 9.9: Liquid Cooling: Resiliency Guidance for Cold Plate Deployments, technical bulletin, September 2024
- ASHRAE TC 9.9: 2026 Annual Meeting presentation, Austin, research update on project 1972, slides dated 28 and 29 June 2026
- Alkharabsheh, Ramakrishnan and Sammakia, Binghamton University: Failure Analysis of Direct Liquid Cooling System in Data Centers, ASME InterPACK 2017 (IPACK2017-74174)
- Lenovo: Neptune DWC RM100 In-Rack CDU Operation and Maintenance Guide, accessed 4 October 2026
- Vertiv: CoolChip CDU 600/1350 Guide Specifications, SL-80239, Rev A, September 2025
- Vertiv: CoolChip CDU 121 Operation and Maintenance Guide, SL-80278, Rev A, October 2025
- Open Compute Project community: Qualification and Performance Specification Guidance for CDUs in Modular Technology Cooling Systems, rough draft, accessed 4 October 2026
- NVIDIA: DSX Facilities Infrastructure Reference Design Overview, citing Design Guide v2.0 of 19 August 2026, accessed 4 October 2026
- Schneider Electric: Schneider Electric launches new Motivair coolant distribution unit integrating liquid and air cooling for more flexible data center deployments, press release via GlobeNewswire, 23 September 2026
- Trane Technologies: Trane Technologies Launches LiquidStack CDU 2.X for Faster, More Flexible AI Cooling Deployments, press release via GlobeNewswire, 29 September 2026, and LiquidStack: CDU 2.X product page, accessed 4 October 2026
- NVIDIA: DCGM Field Identifiers, accessed 4 October 2026
- Dow: DOWFROST LC heat transfer fluid operating guide, Form No. 176-01641-01-0623, 2023 (copy hosted by a distributor), used for Figure 3
- NIST Chemistry WebBook: Copper, condensed phase heat capacity, accessed 4 October 2026, used for Figure 3
- US Department of Energy: Improving Pumping System Performance: A Sourcebook for Industry, second edition, May 2006, and Optimize Parallel Pumping Systems, Pumping Systems Tip Sheet #8, October 2006, used for Figure 4