Inside Microsoft’s Cascading Texas Data Center Outage

Posted

data center (Photo Credit: Microsoft)

A Microsoft data center in Texas overheated earlier this month, causing widespread outages that affected business users of the company’s cloud-based Azure services such as Office 365. In a recent blog post, Microsoft’s director of engineering for Azure DevOps explained what happened — and how his team is working to prevent similar outages in the future.

Early on September 4, a high-energy storm tore through southern Texas near Microsoft’s South Central US data centers. This is one of 10 regions around the world hosting Microsoft’s Visual Studio Team Services (VSTS), now known as Azure DevOps. A lightning strike in the San Antonio area knocked out power at Microsoft’s data center complex, GeekWire’s Tom Krazit noted.

Problems with a cooling system caused a temperature spike, forcing the shutdown of equipment to prevent an even more catastrophic failure, Krazit reported, citing the Azure status page from that morning. “Most cloud companies have automatic shutdown procedures that are triggered by a sharp rise in temperature, and while that’s a good idea, it requires admins to reboot everything, and that takes time,” he wrote.

Although the Azure cloud outages primarily affected customers in the South Central US region, Microsoft’s director of engineering for Azure DevOps Buck Hodges, wrote that the outage also affected customers globally due to cross-service dependencies. Recovering all the Azure services in the region took 21 hours.

“It was the longest outage for VSTS customers in our seven-year history,” Hodges wrote online. “I’ve talked to customers through Twitter, email, and by phone whose teams lost a day or more of productivity. We let our customers down. It was a painful experience, and for that I apologize.”

Future Data Center Protections

“You can spend billions designing and building redundant infrastructure to make sure your data centers don’t go down, but things can and do go wrong anyway — sometimes simply because of bad weather,” DataCenter Knowledge journalist Yevgeniy Sverdlik wrote about the outage.

Hodges said he recognized that customers struggled with incomplete information about the recovery time. Some told him they didn’t want any data loss and would wait as long as necessary. Other customers said they’d accept some data loss as long as they could get a large team productive again quickly.

“Addressing failure of a region is a hard problem,” Hodges wrote. He provided the following summary of changes his team is making based on what they learned from the incident:

  • In supported geographies, move services into regions with Azure Availability Zones to be resilient to data center failures within a region
  • Explore possible solutions for asynchronous replication across regions
  • Regularly exercise fail over across regions for VSTS services using our own organization
  • Add redundancy for our internal tooling to be available in more than one region.
  • Fixed the regression in Dashboards where failed calls to Marketplace made Dashboards unavailable
  • Review circuit breakers for service-to-service calls to ensure correct scoping
  • Review gaps in our current fault injection testing exposed by this incident

The full postmortem is available here.

Environment + Energy Leader