Liquid Cooling Reliability Is Now an AI Compute-Uptime Problem

Sep 17, 2026

A leaking fitting can be a small repair and a big interruption. Can you stop one rack to fix it, or does the repair mean pausing work across the room?

That's the question behind liquid cooling reliability. Keeping chips within temperature limits matters, of course. So does giving the team a way to handle a fault without taking more equipment offline than safety requires.

A sensor can raise the alarm. What happens next depends on the pipes, valves, controls, maintenance records, and people behind it.

How much has to stop?

Think about a leaking tap. With its own shutoff valve, you can work on it without turning off water to the whole building. Without that valve, a local repair becomes a much bigger interruption.

A cooling system is more demanding, but the same question comes first: can you separate the faulty section from the healthy ones? That separation is called isolation. It has to be built into both the pipework and the operating plan.

For a rack branch, that may mean closing both supply and return connections while other branches keep circulating. Whether that's safe depends on the installed valves, pump controls, minimum flow, and protection systems.

The size of the repair and the size of the shutdown are not necessarily the same.

Loading cooling isolation diagram...

A running rack isn't always a running job

There is a catch. If one AI job spans three racks, keeping two of them cool may still leave the job unable to continue when the third stops.

The facilities team needs to know which cooling paths can remain available. The compute team needs to know which jobs depend on the affected equipment and how those jobs will recover. Those decisions belong in the same incident plan.

Finding the fault matters just as much. A problem in the facility water system, a failed connection inside a rack, and an incorrect sensor reading don't call for the same response. A temperature alarm is a reason to investigate, not a diagnosis.

What real systems have taught us

Purdue: some jobs waited, others kept running

In March 2026, a cooling leak affected CPU nodes on Purdue's Gautschi cluster. Operators paused scheduling on those nodes, while the AI nodes continued operating.

During the repair, nodes using water cooling were temporarily shut down. Purdue later asked users to check affected jobs and restart them if necessary.

The notice doesn't identify the failed part. But it does show a useful distinction: the affected CPU nodes and unaffected AI nodes were handled differently. The whole cluster was not treated as one problem.

Blue Waters: one cooling unit, four cabinets

Blue Waters used a different arrangement from today's direct chip cooling. Its cooling system combined air and liquid, with external units that transferred heat from refrigerant to facility water. One unit served four compute cabinets.

If a unit failed, cabinet exhaust temperatures could rise enough to force shutdowns. Operators documented pump gaskets degrading and allowing refrigerant to leak. They also replaced pump gaskets and valve control arms proactively, addressing many problems before they interrupted computing.

These were refrigerant leaks in an older system, not leaks from GPU cold plates. Before you service a shared cooling component, you need to know everything that depends on it.

JUWELS and Blackwell: cooling is part of the computer

JUWELS Booster paired GPU computing with direct liquid cooling years before current Blackwell racks. Its system documentation describes heat exchangers linking internal and external cooling circuits.

In NVIDIA's DGX GB racks, manifolds distribute coolant to cold plates on CPUs and GPUs. Other components, including networking and storage, still use air cooling. Even inside one rack, there is more than one heat removal path to understand.

These designs make a practical point: the people planning compute capacity also need to understand how that capacity is cooled and serviced. Cooling cannot be left out of the availability plan.

Fixing the leak is only part of the job

Replacing the failed part can be the shortest step. The team still needs to check the repaired area for leaks, confirm pressure and flow behavior, verify the relevant alarms, and review fluid condition wherever the work could have changed it.

The exact checks, any required observation period, and the authority to restart belong in the site procedure and manufacturer requirements. A dry floor or an open valve is not enough to approve a return to service.

Give that verification time in the recovery plan. Otherwise, you've planned how to replace the part, but not how to restart the work.

Know what was fitted, not just where it failed

If a connector fails, someone will soon ask whether the same part is installed elsewhere. That's an easier question when you can look up its part revision, supplier batch, installation date, and test record.

Think of those records as a service history for the loop. They help you find related components without treating every connector in the building as equally suspect.

Keep repairs and supplier notices with that history. The next person investigating a problem should be able to find the answer without tracking down whoever remembers the installation.

A baseline needs to keep up with the loop

A higher coolant return temperature can be entirely reasonable if the racks are doing more work. The same rise at a similar load, supply temperature, and flow rate may deserve a closer look.

That is why a baseline should describe how the loop behaves under known conditions, not just preserve one clean sample from opening day. As equipment, workload, and coolant history change, the comparison needs that context.

Useful monitoring helps the team answer four questions:

  • Heat: What are supply and return temperatures doing as the workload changes?
  • Flow: Are branches receiving the flow they need, and what pump effort and pressure difference does that require?
  • Fluid condition: Has coolant chemistry changed? Do samples, treatment records, and added coolant help explain it?
  • Recent work: Was a valve moved, a filter replaced, or coolant added before the trend shifted?

Leak alarms and other equipment signals should be tied to the relevant location, too. A team investigating rack B should not have to guess which pipe or sensor a warning refers to.

The question to ask before the next alarm

Pick one rack and ask your team: if its cooling path had to be isolated today, could we say what stops, what keeps running, and which checks would allow a restart?

It's much easier to agree on those steps before there's a job waiting to finish.

At Reliability Engine, our focus is continuous fluid monitoring and predictive analysis for liquid cooling loops. We help operators understand changing coolant condition alongside the way their systems are running, so they have more to work with when investigating a problem or planning maintenance.

That work sits alongside equipment protection and required testing. Your site's team still decides when it is safe to restart.

Talk to us about monitoring your cooling loop.

Subscribe to updates

Get the latest engineering perspectives sent straight to your inbox.

References

  1. NVIDIA DGX GB Rack Scale Systems User Guide: Hardware and leak detection
  2. Purdue RCAC: Scheduling paused on Gautschi CPU cluster due to cooling leak, March 2026
  3. Blue Waters reliability analysis, CUG 2021: Liquid Cooling System Failures
  4. Forschungszentrum Jülich: JUWELS Booster direct liquid cooling announcement
  5. JUWELS system documentation: Rack heat exchangers and external cooling coupling
  6. Data Center Dynamics: Report on GB200 technical issues and shipment ramp, May 2025