Engineering guide 03 / 08
Coolant Monitoring, Maintenance and Troubleshooting
A coolant reading has changed, but the servers still look normal: the next decision is whether to check the reading, sample the fluid or plan maintenance. Follow the evidence from a trusted baseline to a specific next check, without turning every trend into a fluid replacement.
On this page
Investigate a changed coolant reading in context. First verify the measurement, then compare it with the approved fluid limits, previous results, recent service and current load; use the appropriate laboratory test or inspection to confirm a cause before changing chemistry or equipment.
- For
- Facilities and maintenance teams
- Scope
- Single-phase water-based or glycol-based direct-to-chip systems. Match limits, sampling and interventions to installed equipment, approved coolant and site procedures; immersion and refrigerant systems require different programs.
Key decisions
- Conductivity, pH and particle trends are investigation signals; no single trend establishes a complete diagnosis.
- Compare pressure and temperature measurements at a known flow and comparable heat load.
- Preserve pre-service samples and records, then establish a new baseline only after the accepted change is understood.
- 01Validate
Confirm sensor status, sample location, units and temperature.
- 02Compare
Check the baseline, operating load and recent service events.
- 03Investigate
Select laboratory tests or hydraulic checks for competing causes.
- 04Verify
Record the intervention and demonstrate stable recovery.
A trend narrows the investigation. Confirmation and approval determine the maintenance action.
Start with the right loop and a useful baseline#
Before interpreting a reading, locate the fluid it describes. The technology cooling system (TCS) carries coolant from the coolant distribution unit (CDU) to the server cold plates. Where a heat exchanger separates it from the facility water system (FWS), a facility-side sample does not describe the server-side fluid.
Give sample points, filters and sensors stable circuit identifiers. This small naming decision prevents an easy mistake: comparing two results from different loops as though one fluid changed over time.
Build the baseline after accepted fill and commissioning. Keep the coolant product, batch, mixture, wetted materials, laboratory results, instrument checks and approved limits together. Add temperature, flow, pump command and load so later comparisons describe similar operating conditions.
The Open Compute Project (OCP) publishes separate guidance for water-based and propylene glycol-based fluids. Use the relevant reference alongside supplier requirements. A numerical limit from one fluid program is not a universal alarm setting.
References: Open Compute Project: Guidelines for Using Water-Based Transfer Fluids in Single-Phase Cold Plate-Based Liquid-Cooled Racks; Open Compute Project: Cold Plate Project and Published Fluid Guidelines
Check the reading, then choose the next test#
A smooth chart can still contain bad data. Check for stale values, a moved sample point, recent sensor work, unit conversion errors and temperature compensation. Keep data quality visible: a flat line during a communications failure does not establish a stable loop.
Once the reading is credible, use the table to choose the next check. Each row offers competing explanations, because different faults can produce similar signals.
Set escalation limits and persistence rules from the approved fluid and equipment specifications. Name someone to review unusual trends that remain inside an absolute limit; passing a limit check does not explain a persistent change.
| Observed change | Possible explanations | Next evidence to obtain |
|---|---|---|
| Conductivity rises | Temperature effect, chemical addition, makeup fluid or dissolved contamination | Compensation settings, sample temperature, additions log and relevant laboratory chemistry |
| pH changes | Fluid chemistry movement, different sampling conditions or measurement error | Instrument check, repeat sample under the same method and supplier review |
| Particles or turbidity rise | Service debris, corrosion products, fluid instability or introduced contamination | Repeat sample, filter inspection and particle or deposit characterization |
| Filter differential pressure rises | Loading, increased flow or changed fluid properties | Pressure across that filter at comparable flow and temperature |
| Branch flow falls | Valve position, restriction, air, pump behavior or a faulty reading | Valve state, local pressure, other branches and independent measurement where available |
| Return temperature rises | More heat load, less flow, warmer supply or impaired heat transfer | Simultaneous supply temperature, flow and workload or power history |
Send the laboratory a question, not just a bottle#
Continuous measurements help show when a change occurred. Laboratory analysis can help establish what is in the fluid. Choose a test for the question: dissolved metals for a corrosion investigation, mixture concentration after a top-up, or particulate analysis when the filter contains unexplained material.
Other questions may need inhibitor reserve or microbiological testing. Ask the fluid supplier or laboratory to specify volume, containers, preservation and handling before collection; different tests may need different bottles.
Record the sample point, circuit, time, collector, operating state, fluid temperature and sample temperature. Include recent maintenance, the collection method and any temperature compensation or reference temperature. Follow the approved port preparation and label the bottle before it leaves the site.
When the conditions and procedure allow, preserve a sample before adding chemicals, draining or refilling. An intervention can remove the very evidence that explains the event. After an approved adjustment, keep the old and new results rather than rewriting history around the new reading.
- Include the fluid formulation and recent top-up or treatment quantities with the laboratory request.
- Agree who interprets the result and who may authorize a treatment or replacement.
- Record the collection method and any departure from it, including temperature and handling delays.
Find whether the problem is local or shared#
Begin with the measurement location. Pressure across a CDU, a filter and a server branch are different quantities. Higher pump speed might be compensating for restriction, or simply following changed demand. Check the controller state and setpoints before deciding which explanation fits.
The coolant's return-to-supply temperature difference also has more than one explanation. More heat or less mass flow can both raise it, so the difference alone does not prove fouling.
Compare similar workloads, supply conditions and coolant properties, with sensor placement and timing checked. Graphics processing units (GPUs) running different jobs can have different thermal behavior even when both loops are healthy.
| Pattern | Useful comparison | Limit of the conclusion |
|---|---|---|
| All racks on one CDU change together | CDU supply, common controls, FWS conditions and shared maintenance | Common timing suggests a shared contributor but does not identify it |
| One branch differs from its peers | Branch valves, hoses, flow and local service history | Peers must have comparable equipment and load |
| Filter pressure changes after a load increase | Filter differential pressure at the previous flow | Flow-dependent pressure cannot be interpreted as loading alone |
| GPU temperature rises while loop temperatures are steady | GPU power, event reasons, local flow and server thermal path | Bulk coolant readings can miss a local issue |
| Chemistry changes without a thermal symptom | Repeat chemistry, additions and laboratory results | Stable temperatures do not confirm that the fluid meets its specification |
A filter alert that needs a fair comparison#
If a filter's pressure difference rises from 18 to 31 kilopascals (kPa), but flow also rises from 90 to 120 liters per minute (L/min), the comparison needs a second look. Higher flow can increase pressure loss even without additional filter loading.
Verify pressure taps, instruments, temperature and controller history. Compare the filter at the earlier operating point during a permitted test window. If the pressure difference is still elevated, inspect the filter; the reading does not yet explain the material it has captured.
Keep the element and its identity, and select an appropriate analysis. Corrosion products, installation debris and material introduced during service call for different follow-up actions.
Before and after, with the conditions held constant
If inspection and analysis identify debris consistent with recent service work, carry out the approved cleaning and filter replacement. Then repeat the comparison at the same operating point. Close the work order only when recovery criteria are met, recording the cause, action, remaining uncertainty and follow-up interval.
Use the installed filter's limits and approved test conditions to decide whether service is needed.
Use the trend to improve the maintenance decision#
Keep scheduled maintenance grounded in supplier requirements, the service contract and operating evidence. Include instrument checks, leak inspections, filters, fluid sampling and standby equipment. Monitoring helps decide where to look between those scheduled visits.
Review the loop more closely after filling, component replacement, unexplained top-ups or prolonged idle periods. Those events can change what the old baseline means, even before a limit is crossed.
A work order needs the circuit, observed change, competing explanations, proposed checks, approved intervention and acceptance test. Coordinate with information technology (IT) teams before work affects cooling capacity, redundancy or access. Chemistry adjustments require the designated approval.
ASHRAE highlights the value of commissioning records for operating teams. Retain those records as evidence of accepted behavior and add the maintenance history alongside them. When a change is intentional, state what changed and why the acceptance criteria remain valid.
References: ASHRAE: Commissioning and Performance Validation
Prove recovery before closing the work order#
An alarm disappearing after a reset is an observation, not a recovery test. Show what improved under comparable conditions, and retain the trends, laboratory reports and component or fluid records before, during and after service.
State whether the cause is confirmed or still plausible. The next shift should be able to see the evidence behind that distinction, rather than inherit an unexplained 'resolved' status.
Assign the follow-up check and reopening criteria. Improved temperatures do not resolve chemistry outside its accepted range. If an instrument caused the alert, annotate the affected history as well as repairing the hardware.
Maintenance investigation record
Circuit and component: [ID]. Change first observed: [time and value]. Data quality: [checks]. Operating context: [load, flow, temperature, control state]. Recent service: [events]. Samples: [IDs and methods]. Confirmed cause: [evidence or unresolved]. Approved action: [owner and procedure]. Recovery criteria: [matched test]. Follow-up: [owner and date].Common questions
Does rising coolant conductivity prove corrosion?
No. Temperature, treatment additions, makeup fluid and dissolved contamination can change conductivity. Validate the measurement and history, then use suitable laboratory analysis and inspection to investigate corrosion when the evidence supports it.
Can continuous monitoring replace laboratory testing?
Continuous monitoring and laboratory testing answer different questions. Trends help establish when a change occurred; suitable laboratory methods can examine composition, contaminants or protective chemistry that the installed monitoring system does not measure.
How often should a direct-to-chip cooling loop be sampled?
Use the approved fluid supplier and equipment requirements, then adapt the plan to commissioning, service events and demonstrated stability. Document the interval and the events that require an additional sample rather than adopting a universal schedule.
Does a warmer GPU mean the coolant needs replacement?
No. Workload, supply temperature, local flow, thermal contact, power settings and other equipment behavior may explain the change. Fluid replacement needs evidence and approval against the applicable specification.
When should the coolant baseline be reset?
Create a documented new baseline after an accepted intentional change and a suitable stabilization and verification period. Preserve the previous baseline, the reason for the change and the evidence linking the two.
Sources and further reading
Reliability Engine
Connect coolant condition to operating decisions
Reliability Engine offers monitoring hardware and software for coolant condition and reliability investigations. Discuss the available side-stream monitoring arrangement, measurements and review workflow for your CDU or loop, alongside the laboratory and maintenance program already required by the site.