Confidence: high · Complexity 8/10
Microsoft Azure West US 2 power and cooling PIR
A severe thunderstorm caused power disturbances across multiple datacenters in West US 2, triggering cooling lockouts and protective shutdowns across two physical availability zones. The hard part of the investigation was proving why recovery remained slow after cooling returned: storage validation and telemetry backlog had become the new bottlenecks.
Problem statement
Customers in West US 2 experienced failures to access or manage many Azure services between 04:24 UTC on May 29 and 02:30 UTC on May 30, 2026.
What investigators first believed
- Availability zones should contain most datacenter-scale failures.
- Restoring cooling would likely end the majority of customer impact.
How the diagnosis unfolded
- 01
Detect utility sag/swell events and thermal alerts across datacenters.
Teams established that the initiating event was region-wide power instability rather than a single isolated datacenter fault.
- 02
Diagnose cooling-system protective lockouts and begin manual restoration.
Cooling was restored within roughly two hours, stabilizing temperatures.
- 03
Assess the health of networking, compute, and storage systems after thermal stabilization.
Compute recovered gradually, but storage and dependent systems lagged.
- 04
Identify the remaining recovery bottleneck.
Storage validation and related networking constraints became the dominant reason impact persisted.
- 05
Work down telemetry and service backlog after core infrastructure recovery.
Application Insights and Log Analytics required additional hours beyond primary infrastructure recovery.
What narrowed the fault domain
Utility sag/swell and cooling lockout observations
The incident began with widespread power instability that pushed cooling components into protective lockout.
Thermal alerts and staged recovery milestones
Cooling restored relatively early, but compute, storage, and telemetry recovered on very different timelines.
Storage validation and telemetry backlog
Sequential validation and backlog processing, not just cooling, dominated total outage duration.
Assuming the restored cooling milestone meant the hard part of the incident was over.
Key turning points
- Recognizing that two physical availability zones were affected broke the zone-containment assumption.
- Declaring storage the remaining recovery bottleneck changed the response from incident initiation to recovery engineering.
Root cause
A severe thunderstorm caused voltage instability across multiple West US 2 datacenters, which triggered cooling-system protective lockouts and consequent shutdowns of compute, network, and storage infrastructure across two physical availability zones.
Resolution
Azure restored cooling manually, then prioritized network, compute, storage validation, and finally telemetry backlog processing until remaining impacted services recovered.
Lessons from the response
- Recovery bottlenecks can outlast the triggering fault by many hours.
- Availability-zone assumptions need to be tested against shared regional dependencies and physical failure behavior.
- Manual recovery decision trees for physical systems must be explicit before the incident.
Troubleshooting principles
- 01
Recovery is a separate troubleshooting problem from failure initiation.
- 02
Physical incidents become software incidents once shared control planes and storage paths are involved.
The case crosses physical infrastructure, regional platform design, and long-tail recovery sequencing, with ambiguity shifting from cause discovery to restoration bottlenecks.
Questions answered
What happened in the Microsoft Azure West US 2 power and cooling PIR incident?
Customers in West US 2 experienced failures to access or manage many Azure services between 04:24 UTC on May 29 and 02:30 UTC on May 30, 2026.
What was the root cause?
A severe thunderstorm caused voltage instability across multiple West US 2 datacenters, which triggered cooling-system protective lockouts and consequent shutdowns of compute, network, and storage infrastructure across two physical availability zones.
How was the incident resolved?
Azure restored cooling manually, then prioritized network, compute, storage validation, and finally telemetry backlog processing until remaining impacted services recovered.
Original incident source
Root Cause separates reported facts from analyst synthesis. This public record was explicitly approved before export.