Engineering guide 05 / 08
Liquid Cooling Reliability for GPU and HPC Clusters
A GPU or high performance computing (HPC) cluster can stay online while jobs finish more slowly because of cooling, workload, power or communication changes. Start with the lost throughput, then connect device behavior to the physical loop to find what deserves investigation.
On this page
Investigate slow or hot GPUs by comparing the same workload and power settings, then following each device to its cooling branch. Check coolant supply, flow, fluid condition and specific thermal events together; thermal slowdown identifies a temperature constraint, while corrosion or fouling requires separate evidence.
- For
- GPU infrastructure and HPC operators
- Scope
- Single-phase direct-to-chip GPU and HPC clusters with NVIDIA telemetry available for the installed platform. The investigation method also applies to other accelerators using their documented telemetry; field availability, units, temperature limits and operating actions remain platform-specific.
Key decisions
- GPU temperature and clock events describe device behavior, not the chemical condition of the coolant.
- A useful comparison matches workload, power limits, coolant supply and topology rather than GPU temperature alone.
- Assess reliability through completed work, recovery and operating margin as well as device availability.
- 01Compute
Identify the job, device, throughput and clock event.
- 02Topology
Map device to server, branch, manifold and CDU circuit.
- 03Context
Compare workload and power with temperature, flow and coolant history.
- 04Decision
Confirm the mechanism and use the authorized response and recovery test.
Correlation selects the next investigation. Physical checks and controlled comparisons support a cause.
Follow the slow device to its cooling branch#
For a graphics processing unit (GPU), start with its universally unique identifier (UUID). Map it to the server, rack and manifold branch, then to the server cooling loop and coolant distribution unit (CDU). Record the upstream facility water service and dependencies shared by other racks.
Keep the map current when servers move or CDUs are reassigned. A 'liquid cooled' label identifies a cooling method, but it cannot tell a responder which branch to inspect.
Keep device, circuit and job IDs separate. A coolant sample describes its circuit and location; joining it to GPU telemetry does not create an individual GPU coolant reading. Record the limits of a bulk sample when investigating local components.
Capture representative jobs under accepted operating conditions. Record hardware, software, job configuration, clocks, power settings, coolant supply, flow and control state so later comparisons describe comparable work.
References: Open Compute Project: Water-Based Transfer Fluid Guidance for Technology Cooling Systems
Use NVIDIA telemetry for the question it can answer#
NVIDIA Data Center GPU Manager (DCGM) reports device behavior. Its latest application programming interface (API) uses some names that differ from deployed dcgm-exporter metrics. The table pairs examples; verify your installed versions and collection configuration before writing queries or renaming a live metric.
Thermal violation and clock event reasons may need explicit collector configuration. They are not necessarily enabled in the default exporter counter list, so check collection before concluding that no event occurred.
Keep unavailable data distinct from zero. Check scrape timing, counter resets and whether a value is instantaneous, a bitmask of reasons or an accumulated duration. The latest DCGM reference gives thermal violation time in nanoseconds; other NVIDIA interfaces can use different units.
| Exporter or legacy name | Latest DCGM API field | Meaning and limit |
|---|---|---|
| DCGM_FI_DEV_GPU_TEMP | DCGM_FI_DEV_GPU_TEMP_CELSIUS | Device temperature in degrees C; not coolant temperature |
| DCGM_FI_DEV_MEMORY_TEMP | DCGM_FI_DEV_MEMORY_TEMP_CELSIUS | Memory temperature; availability depends on platform |
| DCGM_FI_DEV_POWER_USAGE | DCGM_FI_DEV_BOARD_POWER_WATTS | Device power in watts; not total rack heat |
| DCGM_FI_DEV_SM_CLOCK | DCGM_FI_DEV_SM_CLOCK | Streaming multiprocessor clock; interpret with load and limits |
| DCGM_FI_DEV_CLOCK_THROTTLE_REASONS | DCGM_FI_DEV_CLOCKS_EVENT_REASONS | Reason bitmask; the former API name is deprecated |
| DCGM_FI_DEV_THERMAL_VIOLATION | DCGM_FI_DEV_THERMAL_VIOLATION | Accumulated thermal violation time; verify collected units and reset behavior |
References: NVIDIA: DCGM Field Identifiers; NVIDIA: DCGM Exporter Default Counter Configuration
Use a thermal event to choose the next check#
Lower clocks do not always mean inadequate cooling. NVIDIA distinguishes power-related events from thermal behavior, and a general hardware slowdown indication may need further interpretation. Look for the specific event reason supported by the platform.
Hardware thermal slowdown means temperature is high enough to reduce clocks. It establishes a device thermal constraint, not why the heat path changed. Warmer supply, local flow, thermal contact, workload and equipment behavior remain possible contributors.
Coolant chemistry and debris can support a fouling or corrosion investigation. A GPU temperature alarm cannot measure those mechanisms. Use the device event to select the next checks, and retain the uncertainty until suitable analysis or inspection resolves it.
| Symptom | Possible explanations | Checks that separate them |
|---|---|---|
| Lower clocks during a job | Power cap, thermal event, operating setting or reduced demand | Specific event reasons, configured limits and job state |
| Higher temperatures across a CDU group | Supply change, shared flow or synchronized load increase | CDU supply, common controls, flow and workload timeline |
| One device is warmer than peers | Local cooling path, different power, sensor or hardware behavior | Matched device load, branch evidence and server inspection |
| Temperature difference across coolant rises | More heat, reduced mass flow or a measurement issue | Simultaneous flow, power context and sensor checks |
| Coolant condition drifts with stable GPU temperatures | Fluid change without an immediate thermal symptom | Fluid history, repeat sampling and targeted laboratory analysis |
References: NVIDIA: NVIDIA System Management Interface Documentation
Compare the same work under similar conditions#
A warmer GPU may simply be doing more work. Compare the same job stage, batch configuration, concurrency and device class where practical. Training, inference, communication and idle phases can have different power and temperature profiles.
Keep precision, software versions, power policy and input data with the record. A familiar job name does not guarantee comparable work if those settings changed.
Align the symptom and its lead-up with coolant temperatures, branch flow, CDU state and maintenance. Record sampling resolution and delays: slow coolant sampling cannot establish precise event order against much faster device telemetry.
For a permitted repeatable test, stay within the accepted operating range and keep other settings stable. After an intervention, check whether workload or power policy also changed. A faster benchmark is encouraging evidence, but by itself it cannot establish a cause or long-term reliability.
- Keep device, job and cooling topology versions with the comparison record.
- Annotate workload starts, control changes, maintenance and telemetry gaps.
- Use completed work and job duration alongside power, clocks and temperature.
Why a 10-second delay can matter to the whole job#
For a tightly synchronized high performance computing (HPC) job with a 40-second baseline step time, consider a participant taking 50 seconds. If every step waits for it and the work per step stays unchanged, each step now takes 50 seconds.
Each step takes 25 percent longer, while completed steps per unit time fall by 20 percent. Both percentages are correct because one describes elapsed time and the other describes output rate. Jobs with different synchronization behavior may be affected differently.
The operator first checks compute and communication records. Network behavior, host contention, power policy and application imbalance could explain the slower participant. If the same record includes a specific thermal event, follow the device to its cooling branch and check supply, local flow and hardware as well.
From step time to completed work
At 40 seconds per step, the rate is 1/40 step per second; at 50 seconds, it is 1/50. The rate reduction is 1 - 40/50 = 0.20. For 100 steps at those rates, elapsed time rises from 4,000 to 5,000 seconds. Use the application's measured step times to calculate actual job impact.
Step-time increase = (50 - 40) / 40 = 25%; step-rate reduction = 1 - 40 / 50 = 20%Coordinate workload protection with loop operations#
Decide whether the evidence calls for an advisory investigation or an immediate protection response. Follow platform limits and site-approved procedures rather than inventing an action from a trend.
The authorized IT operator may need to limit new scheduling, checkpoint a job, remove a node or take another specified action. The mechanical operator needs the approved authority and the affected load map before changing the loop. Coordinate the two so workload and cooling actions happen in the right sequence.
Automatic control actions require a reviewed design and protection sequence. Name the decision owner, applicable equipment limits and the fallback when required measurements are missing.
Preserve telemetry, fluid history and permitted samples before opening the loop. Coordinate the window with job owners and facilities, record any loss of redundancy and define recovery criteria before intervention.
Check that useful output recovered too#
Lower temperature is only part of the recovery check. Verify accepted coolant and hydraulic conditions, device behavior and completed work, with hardware and software configuration recorded. If output recovers only after reducing concurrency or power, record that operating constraint.
Measure repeat events, unavailable node time, failed jobs and checkpoint losses from scheduler and application records where relevant. Applying one node's utilization loss to the entire cluster can overstate impact. Report observed losses and remaining limits.
Record whether closure establishes a mechanical cause, a compute cause or an unresolved correlation. Follow-up should address the suspected mechanism; a short acceptance test cannot prove an intermittent problem has disappeared.
GPU cooling investigation packet
Job and stage: [ID]. Device and server: [UUID and ID]. Cooling path: [branch, circuit and CDU]. Symptom: [event, timing and impact]. Telemetry versions and quality: [record]. Matched workload: [configuration]. Coolant and hydraulic evidence: [measurements and samples]. Other explanations tested: [checks]. Approved action: [owner]. Recovery and follow-up: [criteria, results and remaining limits].Common questions
Does hardware thermal slowdown prove coolant fouling?
No. It identifies a device thermal constraint. The cause requires additional evidence about workload, power, supply conditions, local flow, thermal contact and fluid condition. Fouling or corrosion needs independent investigation.
Which NVIDIA metrics help investigate liquid cooling?
Device and memory temperature, device power, clocks, specific clock event reasons and thermal violation duration are useful context where supported. Join them with actual coolant and hydraulic measurements; none of these GPU fields is a coolant chemistry measurement.
Why do DCGM field names differ from dashboard metrics?
API names, aliases and exporter configuration vary by release. Check the installed versions and collector configuration. The latest documentation deprecates CLOCK_THROTTLE_REASONS in favor of CLOCKS_EVENT_REASONS, but a deployed dashboard may still expose an older metric name.
Can one slow GPU affect a whole HPC job?
It can when the application waits for that participant at synchronization points, but the impact depends on the job. Use application and scheduler evidence to measure the effect rather than assuming every job or every node shares the same loss.
Can coolant monitoring quantify recovered GPU hours by itself?
No. Quantifying compute impact requires workload, scheduler, availability and recovery records with a defensible comparison. Coolant condition can support the cause investigation but does not independently measure completed compute or prove avoided downtime.
Sources and further reading
Reliability Engine
Connect coolant condition to operating decisions
Reliability Engine provides coolant monitoring hardware and software that can support an investigation alongside CDU, loop and GPU context. Discuss available side-stream installation and the evidence required for your cluster; product scope and integrations should be confirmed for the deployed system.