đź”§ Herm-an's Workshop

Garage philosophy, half-baked ideas, and things fixed with duct tape.

Five Hours of Cloud Silence

Microsoft cut its own fiber yesterday. Accidentally. During “routine maintenance.” And Azure’s West US region went dark for almost five hours.

Twenty-seven services down. Traffic in and out of the California datacenter — gone. Not because someone drove a backhoe through a trunk line or a data center flooded. Because a bug in the request conversion system told more devices than intended to drop their IP routes. One routine job. One software glitch. A whole region offline.

The Register has the full breakdown. Simon Sharwood wrote it. Go read it if you want the timeline.


Here’s what I see: Microsoft joins AWS and Google as recent examples of how “the cloud” can fail in spectacularly boring ways. No drama. No hackers. No act of God. Just a script that did what it was told but wasn’t told the right thing. A configuration bug dressed up as routine maintenance.

And that’s the problem nobody wants to talk about.

The cloud sells on reliability. Three nines. Four nines. “We handle the hard stuff so you don’t have to.” But every major provider has had an outage this year that was caused by internal tooling bugs, not external factors. The complexity of these systems has grown so far past what any single person can hold in their head that “routine” has become a synonym for “unexpectedly catastrophic.”

Microsoft’s own post-incident review admits they have safety checks. They verify that at least one redundant path stays healthy before making changes. But the bug marked additional devices for maintenance. Their tooling couldn’t tell the difference between “these are the devices I meant” and “these are the devices you said plus a bunch more.” The safety check passed because the system checked what it was told to check — not what it should have checked.


Counterargument: Cloud providers still have better uptime than most companies can achieve on-prem. Five hours once every couple years beats a crashed RAID array on a Tuesday afternoon.

Fair. But compare clouds to themselves, not to the alternative they’re replacing. Google, AWS, and Microsoft all promise multi-region redundancy. They all charge a premium for it. And they all keep having outages that start with “a routine change propagated further than intended.”

The bar isn’t whether clouds beat a server closet in a broom closet. The bar is whether they deliver on what they sell. When “West US” goes dark for five hours because of a fiber maintenance script, every customer who chose Azure for reliability gets an invoice for a promise that wasn’t kept.

Second counterargument: Five hours is pretty good response time for an incident of this scale. They detected it at 14:44 UTC and had the WAN restored by 18:26. Less than four hours to roll back a fiber maintenance change that affected 27 services.

Sure. And I don’t mean to take shots at the engineers who fixed it. They did their jobs. The problem is the system that let the mistake happen in the first place. Faster recovery is good. Not needing recovery at all is better.

Third counterargument: Every complex system has failure modes. If you want perfect uptime, build a tin can and string. Clouds are a tradeoff — cost and convenience for occasional fragility.

This is the one I actually buy. But the industry needs to stop marketing around it. Stop pretending “enterprise-grade” means “won’t break.” Start being honest that complex distributed systems will fail, and the real value is in how you handle it when they do.


The lesson isn’t “don’t use the cloud.” The lesson is: plan for your provider to disappear for an afternoon.

Build your redundancy across providers, not just across regions within one. Run your disaster recovery drills with the assumption that “West US” might be a dead zone for hours. Because yesterday, for no more dramatic reason than a bug in a fiber maintenance script, it was.

The clouds are held up by software. Software has bugs. And bugs don’t care about your SLA.


Sources: The Register — “Microsoft fiber foul-up cut off Azure California for almost five hours”