Weekly caseMay 29, 2026Cloud Infrastructure

Confidence: high · Complexity 8/10

Microsoft Azure West US 2 power and cooling PIR

A severe thunderstorm caused power disturbances across multiple datacenters in West US 2, triggering cooling lockouts and protective shutdowns across two physical availability zones. The hard part of the investigation was proving why recovery remained slow after cooling returned: storage validation and telemetry backlog had become the new bottlenecks.

01 / Observed failure

Problem statement

Customers in West US 2 experienced failures to access or manage many Azure services between 04:24 UTC on May 29 and 02:30 UTC on May 30, 2026.

02 / Starting hypotheses

What investigators first believed

  • Availability zones should contain most datacenter-scale failures.
  • Restoring cooling would likely end the majority of customer impact.
03 / Investigation path

How the diagnosis unfolded

  1. 01

    Detect utility sag/swell events and thermal alerts across datacenters.

    Teams established that the initiating event was region-wide power instability rather than a single isolated datacenter fault.

  2. 02

    Diagnose cooling-system protective lockouts and begin manual restoration.

    Cooling was restored within roughly two hours, stabilizing temperatures.

  3. 03

    Assess the health of networking, compute, and storage systems after thermal stabilization.

    Compute recovered gradually, but storage and dependent systems lagged.

  4. 04

    Identify the remaining recovery bottleneck.

    Storage validation and related networking constraints became the dominant reason impact persisted.

  5. 05

    Work down telemetry and service backlog after core infrastructure recovery.

    Application Insights and Log Analytics required additional hours beyond primary infrastructure recovery.

04 / Diagnostic evidence

What narrowed the fault domain

physical infrastructure signal

Utility sag/swell and cooling lockout observations

The incident began with widespread power instability that pushed cooling components into protective lockout.

time correlated telemetry

Thermal alerts and staged recovery milestones

Cooling restored relatively early, but compute, storage, and telemetry recovered on very different timelines.

dependency health signal

Storage validation and telemetry backlog

Sequential validation and backlog processing, not just cooling, dominated total outage duration.

Dead ends

Assuming the restored cooling milestone meant the hard part of the incident was over.

05 / Direction changes

Key turning points

  1. Recognizing that two physical availability zones were affected broke the zone-containment assumption.
  2. Declaring storage the remaining recovery bottleneck changed the response from incident initiation to recovery engineering.
06 / Mechanism

Root cause

A severe thunderstorm caused voltage instability across multiple West US 2 datacenters, which triggered cooling-system protective lockouts and consequent shutdowns of compute, network, and storage infrastructure across two physical availability zones.

07 / Restoration

Resolution

Azure restored cooling manually, then prioritized network, compute, storage validation, and finally telemetry backlog processing until remaining impacted services recovered.

Lessons from the response

  • Recovery bottlenecks can outlast the triggering fault by many hours.
  • Availability-zone assumptions need to be tested against shared regional dependencies and physical failure behavior.
  • Manual recovery decision trees for physical systems must be explicit before the incident.
08 / Reusable reasoning

Troubleshooting principles

  1. 01

    Recovery is a separate troubleshooting problem from failure initiation.

  2. 02

    Physical incidents become software incidents once shared control planes and storage paths are involved.

8/10
Diagnostic complexity

The case crosses physical infrastructure, regional platform design, and long-tail recovery sequencing, with ambiguity shifting from cause discovery to restoration bottlenecks.

09 / Direct answers

Questions answered

What happened in the Microsoft Azure West US 2 power and cooling PIR incident?

Customers in West US 2 experienced failures to access or manage many Azure services between 04:24 UTC on May 29 and 02:30 UTC on May 30, 2026.

What was the root cause?

A severe thunderstorm caused voltage instability across multiple West US 2 datacenters, which triggered cooling-system protective lockouts and consequent shutdowns of compute, network, and storage infrastructure across two physical availability zones.

How was the incident resolved?

Azure restored cooling manually, then prioritized network, compute, storage validation, and finally telemetry backlog processing until remaining impacted services recovered.

10 / Provenance

Original incident source

Vendor PIRPost Incident Review (PIR) – Multiple services – Power/cooling issues in West US 2 →

Root Cause separates reported facts from analyst synthesis. This public record was explicitly approved before export.

Continue investigating

Related cases