# When the Cloud Gets Too Hot
Canonical: https://www.reliabilityengine.com/insights/too-hot-to-compute-2026-data-center-cooling-failures
Markdown: https://www.reliabilityengine.com/insights/too-hot-to-compute-2026-data-center-cooling-failures/markdown
Site: Reliability Engine
Author: Gaurav Dhir
Author profile: https://www.reliabilityengine.com/authors/gaurav-dhir
Published: 2026-08-26T10:00:00.000Z
Updated: 2026-08-27T13:47:06Z
Categories: Liquid Cooling, Data Center Operations, AI Infrastructure, Thermal Management
At 7:20 PM ET on May 7, a cooling problem inside an AWS data hall became a cloud problem. According to Coinbase, multiple chillers failed, racks entered thermal-safety shutdown, and nearly all trading on the exchange stopped.

The first broken link lived in the facility. The customer felt it in software.

Cooling is the part of cloud computing nobody notices until the cloud gets too hot to compute.

That same chain appeared in several forms during 2026. A thunderstorm disturbed Azure cooling equipment. A 3 millisecond voltage drop disrupted power and cooling at a Google Cloud facility. In Phoenix, NVMe temperatures climbed together as four chillers went offline. Meta documented a quieter effect: hotter parts of the day moved AI training throughput.

These were facility cooling and thermal incidents, not documented failures inside direct-to-chip coolant loops. They still matter to liquid-cooled data centers because a cold plate can only move heat as far as the rest of the cooling system allows.

## Five incidents, one physical dependency

Each event started differently. Every one eventually reached the workload.

### How heat reached the customer
_The initiating events differed, but the path was familiar: heat removal weakened, equipment responded, and software inherited the consequence._
| Incident | What happened first | What the customer experienced |
| --- | --- | --- |
| AWS and Coinbase | Coinbase reports multiple chiller failures in an AWS data hall. | Rack shutdowns, impaired EC2 and EBS resources, lost matching-engine quorum, and a severe trading disruption. |
| Azure West US 2 | Storm-related voltage instability put some cooling equipment into protective lockout. | Elevated temperatures, infrastructure shutdowns, and a long staged recovery. |
| Google Cloud europe-west4-a | A 3 ms utility disturbance disrupted power and cooling. | Hosts, storage, and network equipment shut down as the data hall warmed. |
| Blacksmith Phoenix | Two utility interruptions took all four facility chillers offline. | NVMe temperatures crossed 85 C and roughly 2,400 hosts were unavailable at peak. |
| Meta Llama 3 405B training | Higher mid-day temperatures affected GPU clock behavior. | Training throughput varied by 1 to 2 percent with time of day. |

## The night cooling became a trading outage

AWS recorded increased temperatures in one data center inside use1-az4. Its Health record says affected EC2 instances and EBS volumes lost power during the thermal event. Servers shut down automatically after temperatures exceeded their operating thresholds.

Coinbase adds the facility detail. Its postmortem says multiple chiller units failed simultaneously in the affected hall and cooling loss triggered thermal-safety shutdowns at the racks.

Those shutdowns protected the hardware, but protection is not the same as availability. A seat belt can save the passenger without keeping the journey on schedule.

At 9:29 PM ET, Coinbase says AWS terminated three of its five matching-engine nodes. The cluster lost quorum. The exchange had no automated cross-zone failover for that system, and an AWS Managed Streaming for Apache Kafka control-plane defect delayed partition-leader recovery.

The result was severe disruption for roughly 8 hours. Cooling removed the machines. System architecture decided how difficult they were to replace.

## When protection slows the restart

On May 29, a severe thunderstorm produced voltage sag and swell events across facilities serving Azure West US 2. Utility power did not disappear completely. It became unstable, which was enough to change the state of the cooling plant.

Some mechanical cooling components detected the abnormal conditions and entered protective lockout as designed. Several did not restart automatically. Operators had to diagnose the condition, attempt resets, and prove that cooling was stable before infrastructure could return.

Cooling came back at 05:55 UTC, about 91 minutes after customer impact began. By 12:00 UTC, about 95 percent of affected virtual machines had recovered. Storage validation took longer, while telemetry backlogs continued clearing until 02:30 UTC the next day.

The equipment protected itself correctly. The harder question was how quickly the entire system could be understood, restarted, and trusted again.

## Three milliseconds at the grid, hours at the service

Google reports a 3 ms voltage drop on an upstream utility feed, followed by breaker action. Both utility feeds were affected, and power and cooling were disrupted at a facility serving europe-west4-a.

As temperatures rose, hosts, storage clusters, and network switches shut down to protect equipment and data. The environment then had to be stabilized before infrastructure could be restarted and verified in a controlled order.

The overall incident window lasted 14 hours and 55 minutes, although individual services experienced different durations. Google Cloud VMware Engine was affected for 9 hours 24 minutes, NetApp Volumes for 8 hours 31 minutes, and Bare Metal Solution for 12 hours 57 minutes.

The voltage disturbance lasted less than a blink. Recovery took hours because power, cooling, network, storage, and compute could not all return at once.

## The graph every operator hopes never to see

Blacksmith reports that back-to-back utility interruptions took all four chillers offline at its Phoenix facility. Fleet temperatures began rising at about 06:00 UTC. At 09:37 UTC, monitoring alerted as NVMe temperatures crossed 85 C, well above their usual range in the 50s and 60s C.

![Blacksmith incident chart showing NVMe temperatures rising across multiple Phoenix hosts toward and above 85 C.](https://cdn.sanity.io/images/7c899jfp/production/818e0426cc67bc2022d3f602bc5a2e1f5a065e56-1938x772.png?w=1200&fit=max&auto=format)

One hot drive can be a component problem. Many drives warming together are the building speaking through the hardware.

Storage clusters failed first and lost quorum. Hosts responsible for accepting GitHub Actions jobs then became unavailable, peaking at roughly 2,400.

![Blacksmith incident chart showing unavailable Phoenix hosts peaking near 2,400 during recovery.](https://cdn.sanity.io/images/7c899jfp/production/74dcc278a7f96f4a4e51c7aad8bded9778f634fd-1956x670.png?w=1200&fit=max&auto=format)

The remote-management network was also down, so Blacksmith could not power-cycle machines for much of the day. Diverting work to other regions became the main mitigation.

At peak, 62,300 jobs waited in Phoenix and 27,300 waited in Ashburn. Heat rose in one building. Queue time rose somewhere else.

## The quieter cost of running hot

Cooling trouble does not always announce itself with an outage. Sometimes useful compute simply becomes slower.

In The Llama 3 Herd of Models, Meta reports a 1 to 2 percent diurnal variation in Llama 3 405B training throughput. Higher mid-day temperatures affected GPU dynamic voltage and frequency scaling, or DVFS. In plain English, the GPUs adjusted their clock behavior as thermal conditions changed.

This was a daily throughput swing, not evidence that Meta lost 1 to 2 percent of full-day fleet capacity. The paper does not publish the raw temperature series or cooling topology. For operators, the useful point is simple: ambient and facility conditions can move AI training performance before they cause an outage.

Synchronized training makes small slowdowns matter. Meta notes that one straggling GPU can slow thousands of its peers because the group must wait at coordination points. A modest thermal effect on one part of the system can therefore tax a much larger job.

## Why this matters when coolant touches the chip

Direct-to-chip cooling gives heat a short, efficient route out of the package. It does not give heat somewhere to disappear.

In a common liquid-to-liquid design, the technology cooling system carries heat from cold plates to a coolant distribution unit. Inside the CDU, a heat exchanger passes that heat into the facility water system. Pumps, valves, controls, and the facility heat-rejection plant all have to keep the route open.

A heat exchanger is a bridge. If the far side cannot accept traffic, the near side eventually backs up.

For a steady, single-phase coolant loop, the heat being carried can be approximated by Qdot = m_dot * c_p * (T_return - T_supply). In ordinary language, heat transfer depends on how much fluid is moving, how much heat that fluid can carry, and how much warmer it becomes across the equipment.

That is why delta T alone cannot tell you whether a loop is healthy. A larger temperature rise may mean more IT load, less flow, or both. A smaller rise may come from lower load, more flow, bypass, or a bad measurement. Temperatures only become useful when read beside workload, flow, pressure, and control state.

The rack side adds its own risks: pump loss, a misplaced valve, trapped gas, branch imbalance, fouling, contamination, leaks, and chemistry drift. None of the public incidents above documents those failures, but the same physical rule applies. Once heat loses its exit path, compute starts paying the bill.

## Read the system as one story

A useful baseline is not one sample taken on commissioning day. It is a living, load-aware picture of how the loop behaves when it is known to be healthy. That picture becomes more valuable as the system ages, maintenance changes it, and workloads move.

### When signals move together
_These patterns guide an investigation. Actual limits must come from the commissioned design, approved fluid specification, sensor accuracy, controls, and operating plan._
| What you see | What it may be telling you | What to check next |
| --- | --- | --- |
| Facility and rack supply temperatures rise together at similar load. | The upstream heat sink or heat-exchanger approach may be losing margin. | Facility flow, valve position, setpoints, sensor accuracy, and actual workload. |
| Pump command rises while branch flow falls. | Resistance, trapped air, valve position, a strainer, or pump performance may have changed. | Pressure locations, valve commands, strainers, pump state, and instrumentation. |
| GPU and rack return temperatures rise at comparable load and flow. | Coolant supply or local heat transfer may have changed. | Workload transients, sensor bias, branch balance, and package-level behavior. |
| Chemistry shifts after makeup or maintenance. | The fluid composition or contamination history has changed. | The approved fluid specification, maintenance record, and validated chemistry method. |
| GPU clocks or throughput follow thermal conditions. | Thermal controls may be reducing useful compute. | Power caps, firmware, scheduling, network stragglers, and workload changes. |

## Five questions to answer before the next hot day

1. Where is the real shared failure domain? Map power, heat rejection, facility-water paths, CDUs, network, and storage beside the software availability design.

1. How much thermal time do we have? Measure the interval from degraded heat removal to derating and protective shutdown at realistic load.

1. Will our evidence survive the incident? Keep sensors, control networks, historians, and alert paths available when the main environment is impaired.

1. What has to recover first? Exercise the sequence for cooling, controls, network, storage, compute, and workload return.

1. Can we see drift before an alarm? Trend connected changes in flow, pressure, temperature, chemistry, controls, and useful compute against the healthy operating picture.

## Where Reliability Engine fits

The building management system knows the plant. The CDU controller knows its equipment. Server telemetry knows the silicon. Liquid-cooled operations still need a continuous story between those layers.

Reliability Engine combines fluid monitoring with thermal, hydraulic, maintenance, and operating context. Its predictive analysis compares current behavior with the loop at known-good conditions, finds connected changes, and helps teams investigate while the evidence is still a trend rather than an outage.

The platform works alongside building controls, CDU controls, OEM interlocks, commissioning tests, and laboratory methods. The goal is simple: turn scattered cooling-loop signals into a decision an operator can act on.

## Cooling is compute infrastructure

The 2026 incidents began in different places, but each exposed the same dependency. A workload can run only while its heat has somewhere to go.

The first warning may appear in a chiller, a voltage trace, a pump curve, a fluid trend, an NVMe sensor, or a slowing GPU. The customer may see only a timeout or a longer queue. Strong operations connect those two views before the gap becomes an outage.

The first signal may be physical. The first customer symptom may be digital. The best response can read both.

Operating or commissioning liquid-cooled AI infrastructure? Reliability Engine helps teams turn continuous cooling-loop telemetry and fluid evidence into earlier, clearer action.

## References

1. AWS Health: EC2 thermal event in US-EAST-1, May 7, 2026

1. Coinbase: A postmortem of the May 7, 2026 outage

1. Microsoft Azure: Power and cooling issues in West US 2, May 29, 2026

1. Google Cloud: Power and cooling failure in europe-west4-a, July 15, 2026

1. Blacksmith: Phoenix facility cooling loss, August 13, 2026

1. Meta: The Llama 3 Herd of Models

1. Open Compute Project: Cold Plate Cooling Loop Requirements

1. ASHRAE: AI Data Center Framework