Engineering guide 08 / 08

Integrating Liquid Cooling Signals Into BMS and DCIM

A cooling value arrives in your building management system (BMS), but its units, age or equipment state are missing. Integrating it with data center infrastructure management (DCIM) starts by preserving that meaning so operators can choose the right response.

Reliability EngineUpdated 11 min read
On this page

Start with read-only monitoring and define every point's location, units, source time, quality and expected equipment state. Give each alarm a response owner, and distinguish operator acknowledgment from physical recovery. Test the complete workflow, including stale data and outages, before relying on it.

For
BMS, controls and DCIM engineers
Scope
Monitoring across CDUs, technology cooling loops, facility systems and available coolant instruments. Verify proposed protocols with suppliers. Control writes require separate authority and tested sequences.

Key decisions

  • Carry units, source time, quality and maximum usable age with every point.
  • Define expected startup, standby, isolation and maintenance behavior.
  • Separate monitoring from control authority, and acknowledgment from recovery.
Turn an observation into an accountable response
  1. 01Observe

    Acquire a value with identity, units, time and quality.

  2. 02Validate

    Check freshness and the equipment operating state.

  3. 03Evaluate

    Apply the approved condition and persistence rules.

  4. 04Respond

    Route the event to its owner and record acknowledgment.

  5. 05Close

    Verify recovery and retain the event evidence.

Data receipt, operator acknowledgment and physical recovery are different events and need separate records.

Give each alarm one clear response owner#

Begin with the person who will act when cooling conditions change. What do they need to see, where will they see it and who can operate the affected equipment? Draw the route from coolant distribution units (CDUs), monitors and facility controllers through gateways to BMS, DCIM and the historian, which stores the time series.

Choose a source of truth for each value and event. Two routes to the same signal can produce duplicate alarms with different times or thresholds. Assign owners for equipment protection, coolant investigation and maintenance dispatch. Shared visibility is useful when the responsibility for the next step is clear.

Default to read-only monitoring and retain the equipment's approved local protections. Any command that changes operation needs separate authority, interlocks and a tested recovery sequence. DMTF's Redfish standard can help describe cooling assets and their relationships; confirm the resources and permissions each supplier actually implements.

References: DMTF: Redfish Resource and Schema Guide, 2025.2

Define each point so the next engineer can use it#

A point is one item of information, such as loop flow, supply temperature or pump state. Give it a stable identifier tied to the asset and location. Record its units, scaling, range, resolution and update interval. These definitions form the point dictionary that both ends of the integration use.

Include data quality. Good, suspect, invalid, stale and maintenance states need agreed meanings and a tested mapping. Preserve source quality where available. An unavailable flow reading converted to zero could be mistaken for a physical loss of flow, so keep missing information distinguishable from a measured value.

Identify calculated values too. A temperature difference needs its input points and subtraction order. A coolant assessment needs its producer, revision and interpretation limits. Verify the actual available measurements before promising them in the interface. The table is a candidate schedule to adapt to the installation.

Candidate points for a liquid cooling monitoring schedule
PointUnits or typeDefinition to approve
Loop supply temperaturedeg CExact location, source sensor and valid equipment states
Loop return temperaturedeg CMatched boundary and source sensor for comparison
Loop flowL/min or approved alternativeMeasured versus estimated flow, direction and conversion
Supply and return pressurekPa or approved alternativeGauge versus absolute reference and sensing locations
Filter differential pressurekPaInlet minus outlet, associated filter and valid-flow conditions
Pump or CDU stateEnumerated stateComplete state list and expected transition behavior
Leak alarmBoolean or eventSource zone, native severity, latching and approved response
Coolant conductivity, if availableuS/cmSample location, temperature and compensation basis
Coolant pH, if availablepHProduct-compatible method, calibration and validity conditions
Monitor sample-flow stateEnumerated state or measured flowWhether a fresh representative sample is reaching the instrument
Data quality and ageEnumerated quality; secondsSource validity, maximum age and handling of uncertainty
Maintenance or isolation stateEnumerated stateOwner, start time, expiry and return-to-service approval

Required metadata for every accepted point

Point ID; asset ID; loop ID; physical location; source system and address; data type; measured or derived basis; units; scale and offset; precision; approved range; source timestamp in UTC; receipt timestamp in UTC; update cadence; maximum valid age; quality mapping; operating-state validity; alarm owner; read/write permission; calibration or derivation revision; retention and aggregation rule.

Show when the fluid was observed#

Keep two times: when the source observed the value and when the platform received it. Store them in Coordinated Universal Time (UTC), then convert for the local display. Document clock synchronization and uncertainty. If only the gateway time is available, label it as acquisition time so readers understand the limit.

Agree a maximum usable age for each point and purpose. A value suitable for a long-term trend may be too old for an immediate response. Evaluate age from the source observation time. Repeatedly receiving the same old value does not refresh the physical observation.

After an outage, retain delayed values for history with their original time and quality. Test ordering, duplicates and replay so historical events cannot be mistaken for new ones. Show gaps on the trend. An uninterrupted line through missing data can make an unknown interval look stable.

A new message can carry an old reading

For a flow point updating every 10 seconds with a maximum usable age of 30 seconds, a reading observed at 12:00:00 UTC and received at 12:00:45 UTC is stale. It is 45 seconds old. Another copy arriving at 12:00:50 keeps the original observation time. Save the value in history and apply the agreed data-availability alarm rule if the condition persists.

Observation age = 12:00:45 UTC - 12:00:00 UTC = 45 seconds

Set the update interval and maximum usable age for each point according to its source behavior and the response workflow.

Keep acknowledgment separate from recovery#

Pressing acknowledge tells the system an operator has seen the event. The pump may still be stopped, the temperature may still be high or the data may still be missing. Record acknowledgment and return-to-normal separately. BACnet, a building automation communications standard, provides distinct notification and operator-acknowledgment services.

Give each alarm a trigger, severity, persistence period, clear condition and response owner. Where appropriate, use hysteresis: a different clear threshold that prevents rapid toggling near the trigger. Specify whether the source equipment or supervisory software creates the event, and correlate duplicates deliberately.

When a persistent low-flow condition meets the trigger criteria, the event becomes active, an operator acknowledges it, and the team follows the approved response. Record recovery when the defined clear criteria are met. Final closure may also require a work record. The table keeps those steps visible without treating a button press as a repair.

Alarm states and evidence to retain
Lifecycle eventMeaningRequired record
Condition detectedThe approved trigger and persistence criteria were metEvent ID, source, condition, value, quality and source time
Active and unacknowledgedThe condition is active and no operator acknowledgment is recordedSeverity, response owner, route and escalation deadline
Active and acknowledgedAn operator acknowledged the event while the condition remains activeOperator identity, acknowledgment time and response note
Returned to normalThe approved clear criteria were met; acknowledgment may still be pendingRecovery time, clear evidence and any latching requirement
ClosedThe required recovery and response documentation are completeClosure owner, evidence, action taken and linked service record
Suppressed or under maintenanceA defined rule temporarily alters notification, with a named ownerReason, scope, expiry, visibility and audit history

References: BACnet: BACnet: A Standard Communication Infrastructure for Intelligent Buildings; Schneider Electric: About Alarms

Ask whether this state is expected right now#

Low flow during active cooling duty and low flow during approved standby need different interpretations. The same applies to a sample monitor intentionally isolated for service while the main loop remains operational. Define these states with the equipment owner and verify how they reach the receiving platform.

Give temporary alarm suppression a reason, scope, owner and expiry, and show it on the operating screen. Test that maintenance handling preserves events that remain relevant under the protection design. This lets a responder see why a notification is suppressed and when normal monitoring should resume.

  • Startup: define stabilization time and conditions requiring immediate attention.
  • Normal duty: apply limits to valid, fresh observations.
  • Standby: define expected flow and pump behavior.
  • Sample isolation: mark observations unavailable until circulation and validity return.
  • Maintenance: record suppression scope, expiry and return-to-service checks.
  • Communications loss: apply the availability rule and show the last observation's age.

Choose a connection that preserves the agreed meaning#

Compare the actual interfaces offered by the equipment and receiving software. BACnet, Modbus, Simple Network Management Protocol (SNMP) and Redfish may fit different parts of the architecture. Check supported objects or registers, units, state codes, update behavior and invalid-data handling. Sharing a protocol name alone does not settle those details.

If a gateway translates between systems, keep its mapping under revision control and test the conversions. Limit accounts and network access to the required monitoring scope. Document credential rotation and support access. Any approved control write needs its own permissions, conversion checks and failure tests.

Choose history settings for the intended investigation. A long average can conceal a short temperature excursion. Agree which raw observations, minimum and maximum values, events and service records to retain, with time and quality. Check that operators can retrieve them.

References: DMTF: Redfish Resource and Schema Guide, 2025.2; BACnet: BACnet: A Standard Communication Infrastructure for Intelligent Buildings

Test the response all the way to closure#

Follow each accepted point from its physical source to the operator display and history. Use approved simulations or a test environment for abnormal conditions. Verify the units, time, quality and state mapping, then follow an event through notification, acknowledgment, recovery and closure.

Test stale values, invalid sensors, outages, restart and delayed replay. Ask responders to identify what is known and missing. Check duplicate events and maintenance expiry. The record should explain what happened, who acted and which evidence allowed closure.

BMS and DCIM integration acceptance schedule

Architecture: [drawing and owners]. Interface: [protocol, version and mapping]. Points: [dictionary]. Time and quality: [UTC source/receipt times, clock uncertainty and age limits]. States: [startup, duty, standby, isolation and maintenance]. Alarms: [source, trigger, persistence, severity, acknowledgment, recovery and closure]. Access: [read-only scope and separate write authority]. History: [retention and aggregation]. Failure tests: [invalid, stale, outage, restart and replay]. Acceptance: [results, deviations and approvers].

References: Schneider Electric: EcoStruxure IT Data Center Expert User Manual, 9.2.0

BMS and DCIM integration acceptance checks

0 of 10

Common questions

What liquid cooling points should be integrated into a BMS or DCIM?

Start with temperatures, flow, pressures, equipment states, alarms and data availability needed for the response workflow. Add verified coolant observations. Every point needs a location, units, time and quality definition.

Is alarm acknowledgment the same as clearing an alarm?

No. Acknowledgment records that an operator has seen the event. Return-to-normal records that recovery criteria were met. Closure may also require service evidence. Keep these events separate.

How should stale coolant data appear?

Show its source time and age, mark it stale for the relevant use and apply the availability rule. Keep it for history where appropriate; never replace it with a misleading zero.

Should a monitoring integration control the CDU?

Default to read-only monitoring. Control needs separate authority, permissions, interlocks, failure behavior and tested sequences. Retain approved local equipment protection.

How do you choose a monitoring interface?

Check the interfaces available at both ends. Confirm the protocol, points, security model and permissions for the proposed Reliability Engine configuration before specifying the integration.

Sources and further reading

  1. Redfish Resource and Schema Guide, 2025.2DMTF
  2. BACnet: A Standard Communication Infrastructure for Intelligent BuildingsBACnet
  3. About AlarmsSchneider Electric
  4. EcoStruxure IT Data Center Expert User Manual, 9.2.0Schneider Electric

Reliability Engine

Connect coolant condition to operating decisions

Reliability Engine connects available coolant observations and operating context for investigation. Agree the point dictionary, interface and response owners before deployment. Verify protocol support and any control authority for the proposed configuration.