Reliability Engine Insights11 min read

When Cooling Fails: What 2026 Data Center Outages Actually Reveal

The cooling system failed before the software did. The software simply made the failure visible to everyone else.

On May 7, 2026, AWS reported rising temperatures inside one data center in a single US-EAST-1 Availability Zone.

EC2 instances and EBS volumes were impaired.

Coinbase later said multiple chillers had failed in the affected hall, racks entered thermal-safety shutdown, and nearly all trading stopped.

Azure and Google Cloud published different versions of the same physical lesson.

Electrical disturbances reduced or removed cooling capacity, temperatures rose, hardware protected itself,

and service recovery continued long after the initiating event had passed.

In Phoenix, the evidence was visible in the hardware itself.

NVMe temperatures climbed together toward 85 C, thousands of hosts became unavailable, and diverted work created queues in another region.

A data center does not run on electricity alone. It runs on electricity plus a continuous path for heat to leave.

First, name the failure correctly

None of the primary incident reports used here describes a coolant leak, a blocked cold plate, a chemistry excursion, or another failure inside a direct-to-chip liquid loop.

AWS, Azure, Google Cloud, and Blacksmith describe facility cooling or thermal events. Coinbase documents the service consequences of the AWS event.

Meta reports a temperature-related training-throughput variation, not an outage.

This is therefore not a list of liquid-cooling failures.

It is a set of documented cooling incidents that exposes dependencies every high-density data center, including one using direct-to-chip liquid cooling, must understand.

What the primary sources actually establish
RecordReported initiating eventDocumented consequenceNot established by the source
AWS and CoinbaseThermal event in one AWS data center; Coinbase reports multiple chiller failures in the affected hall.EC2 and EBS impairment, rack thermal-safety shutdowns, loss of matching-engine quorum, and a prolonged Coinbase outage.No direct-to-chip loop failure is reported. AWS does not identify the chillers as root cause in its public Health record.
Azure West US 2Thunderstorm-related voltage sag and swell events caused parts of mechanical cooling to enter protective lockout.Elevated temperatures, proactive infrastructure shutdowns, and a staged recovery across cooling, compute, storage, and telemetry.No rack-side coolant failure or fluid-quality event is reported.
Google Cloud europe-west4-aA 3 ms upstream voltage drop and utility breaker action disrupted power and cooling.Rising data-hall temperatures and controlled shutdowns of hosts, storage, and network equipment.The report does not attribute the event to a liquid loop, coolant, or cold plate.
Blacksmith PhoenixBack-to-back utility interruptions took all four facility chillers offline.NVMe temperatures crossed 85 C, storage lost quorum, and roughly 2,400 hosts were unavailable at peak.The postmortem does not report a direct-to-chip liquid-cooling failure.
Meta Llama 3 405BHigher mid-day temperatures affected GPU dynamic voltage and frequency scaling.A 1 to 2 percent diurnal training-throughput variation based on time of day.Meta does not publish the raw time series, temperature values, cooling topology, or a full-day fleet-capacity loss.

AWS and Coinbase

Reported initiating event
Thermal event in one AWS data center; Coinbase reports multiple chiller failures in the affected hall.
Documented consequence
EC2 and EBS impairment, rack thermal-safety shutdowns, loss of matching-engine quorum, and a prolonged Coinbase outage.
Not established by the source
No direct-to-chip loop failure is reported. AWS does not identify the chillers as root cause in its public Health record.

Azure West US 2

Reported initiating event
Thunderstorm-related voltage sag and swell events caused parts of mechanical cooling to enter protective lockout.
Documented consequence
Elevated temperatures, proactive infrastructure shutdowns, and a staged recovery across cooling, compute, storage, and telemetry.
Not established by the source
No rack-side coolant failure or fluid-quality event is reported.

Google Cloud europe-west4-a

Reported initiating event
A 3 ms upstream voltage drop and utility breaker action disrupted power and cooling.
Documented consequence
Rising data-hall temperatures and controlled shutdowns of hosts, storage, and network equipment.
Not established by the source
The report does not attribute the event to a liquid loop, coolant, or cold plate.

Blacksmith Phoenix

Reported initiating event
Back-to-back utility interruptions took all four facility chillers offline.
Documented consequence
NVMe temperatures crossed 85 C, storage lost quorum, and roughly 2,400 hosts were unavailable at peak.
Not established by the source
The postmortem does not report a direct-to-chip liquid-cooling failure.

Meta Llama 3 405B

Reported initiating event
Higher mid-day temperatures affected GPU dynamic voltage and frequency scaling.
Documented consequence
A 1 to 2 percent diurnal training-throughput variation based on time of day.
Not established by the source
Meta does not publish the raw time series, temperature values, cooling topology, or a full-day fleet-capacity loss.

The final column is deliberate. A technically useful postmortem records both what the evidence proves and what it does not prove.

AWS and Coinbase: the physical event was only the beginning

AWS reported increased temperatures in one data center inside use1-az4.

Its Health record says affected EC2 instances and EBS volumes experienced a loss of power during the thermal event.

It also says servers automatically shut down after temperatures exceeded operating thresholds.

Primary evidence: AWS Health recorded increased temperatures in one data center, a thermal event, loss of power, and impaired EC2 and EBS resources. Screenshot captured August 27, 2026.

Coinbase provides the facility detail that AWS does not publish in that record.

According to Coinbase, multiple chiller units failed simultaneously in the affected data hall at 7:20 PM ET.

Cooling loss triggered rack thermal-safety shutdowns.

The distinction between sources matters. AWS publicly calls it a thermal event and describes the infrastructure impact.

Coinbase attributes the initiating facility problem to chiller failure. Combining those statements without preserving their attribution would overstate the AWS record.

Cooling explains how hardware became unavailable. It does not, by itself, explain why Coinbase remained severely disrupted for roughly 8 hours.

Coinbase says AWS terminated three of its five matching-engine nodes at 9:29 PM ET, causing the cluster to lose quorum.

The exchange lacked automated cross-zone failover for that system.

A separate AWS Managed Streaming for Apache Kafka control-plane defect prevented automatic partition-leader reelection and extended recovery.

The cooling event lit the match. Application architecture determined how far the fire traveled.

Azure: protection can create a recovery problem

On May 29, a severe thunderstorm produced utility voltage sag and swell events across multiple facilities serving Azure West US 2.

Microsoft is explicit that this was not a complete loss of utility power. Power remained present, but it was unstable.

A subset of mechanical cooling components detected abnormal electrical conditions and entered a protective lockout state by design.

Some equipment did not return automatically after the disturbance. Manual diagnosis and several restart attempts were required before stable cooling capacity returned.

Cooling was restored at 05:55 UTC, about 91 minutes after customer impact began.

Approximately 95 percent of affected underlying virtual machines had recovered by 12:00 UTC.

Storage validation remained the bottleneck later in the day, and Application Insights and Log Analytics continued processing backlogs until 02:30 UTC on May 30.

A protective state can prevent equipment damage and still become an operational failure if people cannot identify, reset, and validate the affected devices quickly.

Google Cloud: 3 milliseconds did not mean 3 milliseconds of impact

Google reports a 3 ms voltage drop on the upstream utility feed, followed by utility breaker protection.

The event affected both utility feeds and disrupted power and cooling at a facility serving europe-west4-a.

As temperatures rose, hosts, storage clusters, and network switches shut down to protect equipment and data.

The environment then had to be stabilized before infrastructure could be restarted and verified in a controlled order.

The report gives an overall incident window of 14 hours and 55 minutes. Individual service impacts were shorter and different: 9 hours 24 minutes for Google Cloud VMware Engine,

8 hours 31 minutes for Google Cloud NetApp Volumes, and 12 hours 57 minutes for Bare Metal Solution.

That precision matters. The report does not say every affected service was unavailable for 14 hours and 55 minutes.

A millisecond-scale electrical disturbance can open an hours-long recovery sequence because power, cooling, network, storage, and compute do not all return at the same instant.

Phoenix: the hardware drew the incident timeline

Blacksmith reports that back-to-back utility interruptions took all four chillers offline at its Phoenix facility.

Temperatures began climbing across the fleet at about 06:00 UTC.

At 09:37 UTC, monitoring alerted as NVMe temperatures crossed 85 C, up from a normal range in the 50s and 60s C.

Primary evidence from Blacksmith: NVMe temperatures across the displayed top 20 traces rose together from the 50s and 60s C toward 85 C during the Phoenix cooling loss.

The synchronized rise is important. One hot drive can be a component problem. Many drives warming together point toward a shared environmental boundary.

Storage clusters failed first and lost quorum. Hosts responsible for adopting GitHub Actions jobs then became unavailable, peaking at roughly 2,400.

Primary evidence from Blacksmith: unavailable hosts peaked at roughly 2,400 and remained elevated during phased recovery.

Blacksmith could not power-cycle machines for much of the day because the facility remote-management network also went down.

Moving work to other regions was the main mitigation.

At peak, 62,300 jobs waited in Phoenix and 27,300 waited in Ashburn. Heat rose in one building. Queue time rose somewhere else.

Meta: performance moved before availability did

The Meta evidence is different and should not be presented as another outage.

In The Llama 3 Herd of Models, Meta reports a 1 to 2 percent diurnal throughput variation for Llama 3 405B training based on time of day.

The paper attributes the fluctuation to higher mid-day temperatures affecting GPU dynamic voltage and frequency scaling.

That does not mean Meta lost 1 to 2 percent of fleet capacity for an entire day.

It means measured training throughput varied by 1 to 2 percent over a daily cycle.

The paper does not provide the raw time series, temperature readings, cooling architecture,

or enough information to calculate a daily energy or capacity loss.

There is no reconstructed curve in this article because inventing one would turn a reported relationship into fictional telemetry.

The same paper notes that even one slow straggler can slow thousands of other GPUs during synchronized training.

That is why a small, repeatable thermal effect can matter operationally without becoming an outage.

What can legitimately be carried into liquid cooling

A direct-to-chip system changes the final part of the heat path. It does not remove the upstream dependencies.

In a common liquid-to-liquid architecture, the technology cooling system carries heat from cold plates to a coolant distribution unit.

A heat exchanger transfers that heat into the facility water system. The facility still needs enough flow and a low enough supply temperature to accept the load.

A heat exchanger is a bridge, not a landfill. Heat entering one side must leave through the other.

The steady-state sensible heat carried by a single-phase coolant is approximated by Qdot = m_dot * c_p * (T_return - T_supply).

Fluid density and specific heat must match the actual mixture and temperature.

Sensor uncertainty and unsteady operation also limit how literally a field calculation should be read.

Delta T alone is not a health score. A larger temperature rise can reflect more IT load, less flow, or both.

A smaller rise can reflect lower load, more flow, bypass, or a measurement problem.

Workload, flow, temperatures, and control state must be interpreted together.

Direct-to-chip systems also introduce rack-side failure modes that these public incidents do not document, including pump loss, valve misposition,

trapped gas, branch imbalance, fouling, contamination, leaks, and loss of fluid chemistry control.

Read relationships, not isolated numbers

No universal threshold can replace the commissioned design envelope, supplier limits, sensor accuracy, and the site operating plan.

Useful detection begins with a known-good baseline at comparable load and control state.

Signal combinations that deserve investigation
Observed relationshipPhysically plausible interpretationCheck before concluding
Facility supply temperature and TCS supply temperature rise together at similar IT load.The upstream heat sink or heat-exchanger approach may have degraded.Confirm sensor calibration, valve position, facility flow, control setpoints, and actual workload.
Branch flow falls and the pump-speed to differential-pressure relationship moves away from baseline.Hydraulic resistance, air, valve position, or pump performance may have changed.Verify where pressure is measured, valve commands, pump state, strainers, and instrumentation.
Rack return temperature and GPU temperature rise at comparable load and flow.Coolant supply, heat transfer, or package-to-coolant resistance may have changed.Exclude workload transients, firmware limits, sensor bias, and local branch effects.
Coolant chemistry changes after makeup or maintenance.Fluid composition or contamination history has changed.Use the approved chemistry method and limits. Chemistry alone does not prove a hydraulic restriction.
GPU clocks or throughput move with thermal conditions.DVFS or another thermal control may be affecting useful compute.Exclude scheduler behavior, power caps, firmware, network stragglers, and workload changes.

Facility supply temperature and TCS supply temperature rise together at similar IT load.

Physically plausible interpretation
The upstream heat sink or heat-exchanger approach may have degraded.
Check before concluding
Confirm sensor calibration, valve position, facility flow, control setpoints, and actual workload.

Branch flow falls and the pump-speed to differential-pressure relationship moves away from baseline.

Physically plausible interpretation
Hydraulic resistance, air, valve position, or pump performance may have changed.
Check before concluding
Verify where pressure is measured, valve commands, pump state, strainers, and instrumentation.

Rack return temperature and GPU temperature rise at comparable load and flow.

Physically plausible interpretation
Coolant supply, heat transfer, or package-to-coolant resistance may have changed.
Check before concluding
Exclude workload transients, firmware limits, sensor bias, and local branch effects.

Coolant chemistry changes after makeup or maintenance.

Physically plausible interpretation
Fluid composition or contamination history has changed.
Check before concluding
Use the approved chemistry method and limits. Chemistry alone does not prove a hydraulic restriction.

GPU clocks or throughput move with thermal conditions.

Physically plausible interpretation
DVFS or another thermal control may be affecting useful compute.
Check before concluding
Exclude scheduler behavior, power caps, firmware, network stragglers, and workload changes.

These are diagnostic relationships, not universal alarm limits. The sign and magnitude depend on system design, sensor location, controls, fluid, and workload.

Five questions before the next thermal event

  1. What is the real shared failure domain?
  2. Map power feeds, halls, chillers or dry coolers, facility-water paths, CDUs, network, and storage alongside software availability zones.
  3. How much thermal time is available? Measure the interval from degraded heat removal to derating, protective shutdown, and possible equipment risk for the actual load.
  4. Will the evidence survive the incident? Keep sensors, control networks, historians, and alert paths available when the primary environment is impaired.
  5. What must recover first? Document and exercise the order for cooling, controls, network, storage, compute, and workload return.
  6. Which claims can the telemetry truly support? Separate direct measurements, calculated values, model outputs, and operator inference.

Where Reliability Engine fits

Facility controls show what the building is doing. Server telemetry shows what the silicon is doing.

Liquid-cooled operations also need the evidence between them.

Reliability Engine combines continuous fluid monitoring with thermal, hydraulic, maintenance, and operating context.

Predictive analysis can then compare present behavior with a known-good state, identify connected changes, and help teams investigate before a trend becomes a compute event.

It does not replace the building management system, CDU controls, OEM interlocks, laboratory methods, or commissioning tests.

It makes the cooling-loop evidence easier to interpret together.

The lesson is not that liquid cooling failed

The documented lesson is more precise.

Heat removal is part of compute availability, and its failure can begin in the grid, facility controls, heat-rejection plant, rack loop, or workload response.

AWS, Coinbase, Azure, Google Cloud, Blacksmith, and Meta do not tell the same story. That is why their records are useful together.

They show different points where the physical system reached the digital service.

The first signal may be physical. The first customer symptom may be digital. Good operations must be able to read both.

Operating or commissioning liquid-cooled AI infrastructure? Reliability Engine helps teams turn cooling-loop telemetry and fluid evidence into decisions they can defend.

Get new Insights by email

Practical reads on coolant health, GPU thermal margin, and what to check next.

References

  1. AWS Health: EC2 thermal event in US-EAST-1, May 7, 2026
  2. Coinbase: A postmortem of the May 7, 2026 outage
  3. Microsoft Azure: Power and cooling issues in West US 2, May 29, 2026
  4. Google Cloud: Power and cooling failure in europe-west4-a, July 15, 2026
  5. Blacksmith: Phoenix facility cooling loss, August 13, 2026
  6. Meta: The Llama 3 Herd of Models
  7. Open Compute Project: Cold Plate Cooling Loop Requirements
  8. ASHRAE: AI Data Center Framework