Engineering guide 04 / 08

Liquid Cooling Responsibilities for Colocation Providers

A tenant receives a coolant alarm, but the provider controls the valve needed for the response. A useful colocation agreement resolves that handoff in advance: who sees the evidence, who decides, who acts and who accepts recovery.

Reliability EngineUpdated 10 min read
On this page

Assign one decision owner and a responsible team to each cooling circuit and task. Agree who operates the facility and server-side loops, services the coolant distribution unit (CDU), approves fluid changes and responds to alarms, with authority and acceptance tests tied to the site's approved procedures.

For
Colocation providers and tenant teams
Scope
Colocation direct-to-chip deployments with provider facilities infrastructure and tenant IT equipment, including tenant-owned or shared CDUs. Assign responsibilities against the installed topology, service agreements and equipment requirements.

Key decisions

  • Physical fluid separation, equipment ownership and response authority are distinct decisions.
  • Each operating activity needs one accountable owner, an executing team and an accessible record.
  • Shared infrastructure requires coordinated response and change control across every affected tenant.
Define the physical boundary before assigning the work
  1. 01Facility

    Identify provider water supply, return and permitted operating envelope.

  2. 02Interface

    Name the CDU, heat exchanger, valves and isolation authority.

  3. 03Technology

    Map the fluid reaching manifolds, hoses and server cold plates.

  4. 04Operations

    Assign monitoring, response, maintenance and acceptance owners.

In a liquid-to-liquid arrangement, the heat exchanger separates fluid circuits. Equipment ownership and service responsibility still need agreement.

Draw where the fluid goes, then assign the work#

A rack alarm does not identify the owner of the coolant or the operator of an upstream valve. Start with the installed drawing. In a common liquid-to-liquid arrangement, the CDU heat exchanger separates the facility water system (FWS) from the technology cooling system (TCS) that feeds the server cold plates.

Record the fluid paths, equipment, isolation points and dependent loads. Then add ownership and contracted service. The provider might maintain a tenant's CDU or own a CDU serving several customers; either arrangement needs explicit operating responsibilities.

A shared cabinet does not necessarily mean shared coolant. Identify separate heat exchangers, pumps and circuits before deciding whether a fluid change can affect another tenant.

Finally, mark the contractual delivery point and its measurement method. Service might be defined at a flange, rack manifold or specified cooling condition. The Open Compute Project (OCP) distinguishes the TCS fluid program from the facility side; Vertiv describes several ways to divide facilities and information technology (IT) responsibilities.

References: Open Compute Project: Water-Based Transfer Fluid Guidance: TCS and FWS Scope; Vertiv: Liquid Cooling Services: IT and Facilities Responsibilities

Make the cooling promise measurable#

Terms such as 'liquid cooling ready' leave important decisions open. Agree the actual supply conditions, capacity, measurement points and normal control behavior. Describe what happens when the service moves outside its permitted operating range.

Include planned maintenance: what capacity or redundancy will be unavailable, which tenants depend on it and who approves the window. Assign approval for new server generations whose requirements differ from those originally accepted.

Acceptance needs evidence for both infrastructure and tenant connections. Keep fill and cleanliness records, equipment readiness, approved fluid, alarm tests, control state and the operating baseline. Make clear which tests cover the tenant's circuit and which cover only the upstream plant.

The decisions to settle at each connection
InterfaceDecision to documentAcceptance evidence
FWS deliveryPoint of delivery and permitted supply conditionsMeasurement locations, limits and recorded test results
CDUOwner, operator, service party and dependency groupAsset record, controls test and affected load map
TCS fluidRecipe, materials, fill approval and chemistry authorityApproved fluid specification, fill records and laboratory baseline
Rack connectionWho connects, disconnects and checks the interfaceApproved method, inspection and connection acceptance
MonitoringShared signals, access permissions and record ownershipAlarm routing tests, data quality and retention agreement
MaintenanceWindow approval, isolation authority and return to serviceWork order, operating authorization and recovery criteria

References: ASHRAE: Commissioning and Performance Validation

Put a decision owner against each activity#

A RACI matrix makes the division of work readable: Responsible, Accountable, Consulted and Informed. Use one accountable decision owner for each row, alongside the team that performs the work. The same party can hold both roles.

In the table, A approves or owns the decision, R executes, C is consulted beforehand and I receives the update. Writing 'shared' without naming a decision owner leaves the difficult handoff unresolved.

For provider-owned FWS and tenant-owned CDU and TCS, the matrix below assigns fluid work to a specialist and equipment advice to the original equipment manufacturer (OEM). Confirm the assignments against the service agreements. A provider-owned shared CDU needs revised rows and participation by dependent tenants in acceptance and outage approval.

Responsibilities for a tenant-owned CDU and TCS
ActivityProvider operationsTenant operationsService specialistEquipment OEM
Operate and maintain FWSA/RICC
Approve TCS coolant and material changesCARC
Collect and interpret scheduled TCS samplesIARC
Inspect and service tenant CDUCARC
Connect tenant server to approved manifoldIA/RCC
Acknowledge TCS alarm and start incident coordinationCA/RII
Authorize provider FWS isolationA/RCIC
Accept TCS return to serviceCARC

Connect each alarm to an authorized responder#

A tenant engineer may see a CDU alarm while only provider personnel can access an upstream isolation valve. That arrangement can work when the escalation route is clear and tested. It fails as an operating model when neither party knows who has authority to act.

Name who can acknowledge alarms, view trends, enter the rack area, change setpoints, operate valves and authorize fluid work. Receiving an alert should lead to a defined contact and action rather than a search for the equipment owner.

Share circuit and asset IDs, event time, state, trend quality, associated load and service history. Agree a time basis and the system retaining original records. Test responders' read access or export process during an incident exercise.

For shared service, register affected tenants, escalation contacts and permitted information sharing. The coordinator needs to identify dependent loads while respecting tenant access boundaries. Vertiv identifies shared visibility and escalation as important parts of the operating model.

  • Name primary and backup responders for mechanical, IT and fluid issues.
  • Record access restrictions and the method for emergency entry or remote support.
  • Test alarm delivery, acknowledgement, escalation and data access together.

References: Vertiv: Liquid Cooling Services: IT and Facilities Responsibilities

Agree the response while the system is healthy#

A chemistry advisory, loss of flow and a suspected leak call for different decisions. Define each alarm class, the evidence needed for escalation and its project-approved response before a live event.

Identify who may isolate equipment, protect a load or coordinate a controlled compute stop. The approved procedure must account for the actual circuit and protection sequence; a blanket shutdown instruction cannot resolve those differences.

Document protection steps, communication order, authority and evidence preservation. Coordinate authorized IT workload actions with mechanical valve and plant actions so an intervention does not unexpectedly remove cooling from an operating load.

Practice a scenario with unavailable contacts and incomplete telemetry. Check that a backup can make the required decision and that responders understand which readings are stale. The exercise should end with a return-to-service decision, not only acknowledgement of the first alarm.

Joint incident record

Incident and circuit IDs: [IDs]. Alarm class: [event]. Current evidence: [values, quality and time]. Affected loads: [dependency register]. Incident coordinator: [name]. Authorized mechanical action: [procedure and owner]. Authorized IT action: [procedure and owner]. Approvals and notifications: [record]. Recovery evidence: [tests]. Return-to-service decision: [accountable owner].

One sample, three tenants: scope the service window#

For a provider-owned CDU serving three tenant branches on one TCS circuit, an out-of-range coolant sample needs a coordinated review. Even if only one tenant reports a symptom, a common fluid intervention could affect all three branches.

Use the agreed incident coordinator to contact all three tenants and arrange confirmatory tests with the fluid specialist. The circuit drawing identifies who could be affected; the operating agreement identifies who may authorize the work.

Verify history and sampling before choosing an intervention. If confirmatory tests support a fluid correction, obtain approval and agree the permitted compute state with each tenant. Record the procedure, quantities, batches and resulting tests. The designated owner accepts return to service when recovery criteria are met.

What changes when circuits are separate?

If the same cabinet instead contains separate TCS circuits, the team scopes fluid work to the affected circuit after verifying the design. Common FWS supply, electrical power or controls can still create shared dependencies. The dependency register therefore distinguishes fluid sharing from other shared failure paths.

Keep the agreement useful after handover#

Warranty and service coverage follow the actual contracts and supplier terms. Record the required fluid, materials, service qualifications, tests and notifications for the deployed equipment. A guideline or monitoring installation does not itself establish coverage; proposed changes need review by the party authorized under those agreements.

Keep a handover pack that a new responder can use: the circuit register, responsibility matrix, approved procedures, fluid records, commissioning evidence, spares, alarm routes and contacts. Name who retains samples and laboratory reports, and how the relevant parties access them after a staff change or equipment refresh.

Review responsibilities after hardware, capacity, formulation, topology or contractor changes. Date and approve the revised document, update contacts and retain the associated training record.

Before accepting a liquid cooled tenant deployment

0 of 10

Common questions

Who owns the coolant in a colocation data center?

Ownership and maintenance responsibility depend on the contract and the actual circuit. Define the FWS, TCS and CDU separately, then assign fluid selection, sampling, treatment, disposal and acceptance to named parties.

Does a shared CDU mean tenants share coolant?

Not necessarily. A cabinet or service may include separate circuits or heat exchangers. Confirm fluid paths and other common dependencies in the installed design before deciding the scope of a change or incident.

Should the provider or tenant receive CDU alarms?

The authorized responders on both sides need the alarms and evidence relevant to their responsibilities. Assign acknowledgement, escalation and operating authority explicitly so shared visibility leads to a defined action.

Does a coolant alarm require shutting down the entire data hall?

The response depends on the alarm, verified condition and approved protection procedures. Define isolation and compute actions by circuit and dependency, with authorized mechanical and IT decision makers.

Does using an OCP fluid guideline preserve an equipment warranty?

A guideline alone does not establish warranty coverage. Check the actual supplier terms and contracts, and retain the approvals and service records required for the installed equipment.

Sources and further reading

  1. Water-Based Transfer Fluid Guidance: TCS and FWS ScopeOpen Compute Project
  2. Liquid Cooling Services: IT and Facilities ResponsibilitiesVertiv
  3. Commissioning and Performance ValidationASHRAE

Reliability Engine

Connect coolant condition to operating decisions

Reliability Engine's monitoring hardware and software can contribute coolant condition records to a provider and tenant investigation. Discuss the CDU or side-stream arrangement, information access and review ownership as part of the operating agreement; the site's approved teams retain authority over equipment and fluid interventions.