Microsoft has revealed that a software bug in its automated network maintenance system caused the widespread Azure and Microsoft 365 outage that disrupted services for thousands of customers on Thursday.
The outage began at 10:44 a.m. ET on July 23 and mainly affected users accessing Microsoft 365 services through network infrastructure connected to the company’s West US Azure region. According to Downdetector, outage reports quickly surged to more than 2,400, with SharePoint accounting for most complaints, followed by Excel and the Microsoft 365 Admin Center.
Microsoft tracked the incident under ID MO1437424 and confirmed that multiple Microsoft 365 services were impacted. Users experienced intermittent access to OneDrive, “Something went wrong” errors in SharePoint Online, degraded Microsoft Teams chat, slow or unavailable Microsoft 365 Admin Center access, Power Automate loading failures, delays in Copilot Chat, and problems opening Microsoft Loop pages.
The disruption also affected several other Microsoft cloud services, including Fabric, Power BI, Power Apps, Copilot Studio, Windows 365, and Microsoft Defender. Some Microsoft Defender customers experienced delays in receiving responses from Defender Experts, while security investigations, automated workflows, and remediation actions through Threat Explorer and Advanced Hunting also encountered issues.
Microsoft initially attempted to reduce the impact by rerouting traffic through alternate network paths. While this restored access for some customers, many services continued experiencing disruptions. The company later identified a recent networking change as the source of the problem and began rolling it back. The rollback was completed at 2:26 p.m. ET, after which Microsoft confirmed that affected Microsoft 365 services had recovered based on system telemetry and customer reports.
In a preliminary Post Incident Review, Microsoft explained that the outage occurred during routine maintenance work in the West US Azure region. The maintenance process was designed to isolate specific network paths while ensuring that one of two redundant routes remained available. However, a bug in the system that converts maintenance requests into machine-readable instructions incorrectly included additional network devices in the maintenance operation.
Because of this error, IP routes were removed from more network devices than intended, disrupting traffic between Microsoft’s West US datacenters and its wide-area network. Microsoft said network traffic that remained entirely within the West US region continued operating normally, while traffic entering or leaving the region experienced connectivity failures and increased latency.
The networking issue also affected a wide range of Azure services, including Azure App Service, Application Gateway, Azure AD B2C, Azure AI Search, Azure API Management, Azure Cosmos DB, Azure Databricks, Azure Firewall, Azure Kubernetes Service, Azure Monitor, Azure Virtual Desktop, ExpressRoute, Log Analytics, Microsoft Graph, Microsoft Sentinel, Power BI Embedded, Virtual WAN, and VPN Gateway.
Engineers began investigating immediately after the outage started and initially observed widespread route instability across Microsoft’s wide-area network. They eventually traced the issue to a datacenter in the West US region and linked it to the ongoing maintenance activity. Microsoft began reversing the maintenance changes at 1:45 p.m. ET and completed the rollback 41 minutes later. Most Microsoft 365 services recovered shortly afterward, while 3:41 p.m. ET fully restored the remaining Azure services.
If this article helped you, please consider supporting our work. Every small contribution keeps Abijita.com independent and running.
Microsoft said it is now conducting a full internal review of its automated maintenance processes and safety mechanisms to understand why the bug bypassed existing safeguards. The company plans to publish a complete Post Incident Review after finishing its investigation, which it expects to complete within the next two weeks.
Microsoft Explains Cause of Azure and Microsoft 365 Outage That Disrupted Cloud Services





