Microsoft 365 Outage July 2026: What Broke and Why
A maintenance bug in Microsoft's West US Azure region took down Teams, SharePoint, OneDrive and Copilot for hours on July 23, 2026. Here's the breakdown.
The productivity backbone of a large share of corporate America went intermittent for most of a workday afternoon. On July 23, 2026, Microsoft 365 suffered an outage that degraded access to Teams, SharePoint Online, OneDrive, Copilot Chat, Power Automate and other services for roughly three and a half hours, affecting customers whose traffic routed through the company’s West US Azure region. Microsoft tracked the incident as MO1437424 and, by its own account, traced the fault not to an attack or a hardware failure but to a bug in a routine maintenance procedure.
The disruption was narrow in cause and broad in reach — a familiar shape for cloud incidents, and a pointed reminder that the tools most enterprises treat as always-on ride on infrastructure that occasionally is not.
The timeline
By Microsoft’s status updates, the trouble began at roughly 10:44 a.m. ET on Thursday, July 23, when users started reporting failures across multiple Microsoft 365 workloads. The company opened an incident, identified a recent networking change as the likely cause, and began reverting it. Microsoft completed the reversion at about 2:26 p.m. ET, then confirmed through service telemetry and customer reports that impact had cleared for the vast majority of users.
That puts the active window at just under four hours — long enough to derail an afternoon of meetings, document collaboration, and automated workflows, but short of the multi-day sagas that dominate outage postmortems. The impact was also uneven: the incident primarily affected customers whose Microsoft 365 traffic was routed through the West US Azure region, so organizations served from other regions saw little or nothing while others lost core tools for hours.
What actually broke
The failure did not originate in the applications themselves. According to Microsoft’s account, the root cause sat in the network layer beneath them.
Engineers were performing routine device maintenance in the West US Azure region, a process that involves isolating specific network paths so equipment can be serviced without disrupting live traffic. During that work, a bug in the request-conversion system — the component that translates maintenance instructions into actions on the network fabric — incorrectly marked additional network devices as part of the maintenance event. Devices that should have kept carrying traffic were instead treated as if they were being taken offline, and the paths that Microsoft 365 services depended on degraded as a result.
In plain terms: a maintenance job meant to touch a narrow slice of the network fanned out to hit more of it than intended, because the system that scoped the work got the scope wrong. Once Microsoft recognized that the recent change was the trigger, the fix was to reverse it rather than push a new one — the safest path when a configuration change is the suspect. The incident is a textbook example of a change-induced outage, where nothing failed on its own; a routine, planned action had an unplanned blast radius.
Who felt it
Because Microsoft 365 bundles so many services behind one platform, a single network fault surfaced as a scattering of seemingly unrelated failures:
- Microsoft Teams — chat functionality degraded, with images failing to load in conversations.
- SharePoint Online — users hit “Something went wrong” errors when trying to reach sites and documents.
- OneDrive — intermittent access to stored files.
- Power Automate — automated flows failed to load, stalling the background automation many organizations rely on.
- Microsoft 365 Admin Center — loaded slowly or not at all, which is doubly painful because it is where administrators go to diagnose an incident in the first place.
- Copilot Chat and other AI-assisted features — delayed or unavailable, alongside collaboration tools such as Loop and analytics in Power BI.
The spread across chat, storage, automation, and administration from one underlying cause is the whole story. None of those workloads failed independently; they shared a network path, and when it degraded, they degraded together. For IT teams, the Admin Center slowdown was a particular sting — the dashboard meant to provide visibility into system health was itself caught in the same fault it was supposed to help diagnose.
A familiar pattern
This is not a new kind of failure. The most consequential cloud outages of recent years have repeatedly traced back to changes and configuration, not exotic attacks. The 2021 Facebook BGP outage that erased Facebook, Instagram, and WhatsApp from the internet began with a routine command that withdrew the routes to Facebook’s own data centers. The 2024 CrowdStrike outage that grounded flights and froze hospitals came from a faulty content update pushed to millions of machines. And just a week before this incident, the AWS CloudFront outage on July 16 knocked services offline worldwide after an internal limit stopped routing configuration from loading correctly.
Different providers, different mechanisms, same lesson: at hyperscale, the routine change is the dangerous one. The systems are too large and too interdependent for any single engineer to fully model the blast radius of a maintenance action, so the safety has to live in the tooling — in staged rollouts, automated scope-checking, and fast, reliable rollback. Microsoft’s ability to identify the change and reverse it within hours is the difference between a bad afternoon and a bad week.
What it means
The direct damage from MO1437424 was bounded — a few hours, one region, then recovery — but the incident fits a pattern that keeps recurring across every major cloud, and the takeaways are the same each time.
Concentration risk is structural, not accidental. The reason a single West US network fault could interrupt meetings, file access, and automated workflows for so many organizations at once is that those organizations have consolidated all of it onto one platform. That consolidation buys real reliability and integration most of the time; the cost is that a provider-side fault becomes your fault, with nothing you can do but wait. Microsoft 365’s dominance in enterprise productivity makes its outages unusually visible precisely because there is no diversity of exposure to blunt them.
The change process is the attack surface. No adversary was involved here — the damage came from Microsoft’s own maintenance tooling misjudging its scope. The defensive lesson for any team running infrastructure is that the most likely cause of your next outage is a change you made on purpose. Rigorous change scoping, blast-radius limits, and rehearsed rollback are worth more than any amount of after-the-fact heroics, and they are the same disciplines that turn a vague uptime promise into a measurable service-level objective.
What to watch next. Whether Microsoft publishes a detailed postmortem explaining why the request-conversion bug mis-scoped the maintenance event and what guardrail failed to catch it; whether enterprises respond by building genuine failover for their most critical Microsoft 365-dependent workflows, or simply absorb the risk as the price of the platform; and whether the steady cadence of change-induced cloud outages finally pushes more organizations toward the redundancy the pattern has been demanding for years.
Tagged
Keep reading
Chisato · · 7 min read Microsoft Maia 300: TSMC Order and Nvidia Challenge
Microsoft is in talks with TSMC to build 300,000+ Maia 300 AI chips, aiming for over 1 million units to cut its reliance on Nvidia. The plan and what it means.
Chisato · · 5 min read Multi-Cloud vs Hybrid Cloud: The Real Difference
Multi-cloud spreads workloads across public cloud providers; hybrid cloud connects private infrastructure to a public cloud. How they differ and why it matters.
Chisato · · 6 min read Google's $15B India Data Center Faces Water Protests
Google's $15B Visakhapatnam AI data center with Adani faces legal challenges and protests over water use and a nearby wildlife sanctuary. What's at stake.