It's been a long day because of this. Just going to leave it at that.
It's been a long day because of this. Just going to leave it at that.
Honestly it's hard to imagine a good mitigation for this. "Build more datacenters" is already happening as fast as it can. "Keep enough spare capacity around to handle losing one of the biggest datacenters in the world" is pretty unreasonable.
If you, as a customer, are uptime focused enough that it's worth paying extra, then the sensible practice has always been Cross-Cloud infrastructhre/failovers. At least since the Amazon Easter failure of 2011. That's what giants like Netflix do.
Err what?
It's entirely reasonable to expect Azure to handle the loss of a single DC and not have a 14+ hour global outage. I don't care how big the DC is, losing one should not take out the world, especially not for the length of time this one has been going on.
https://perspectives.mvdirona.com/2017/04/how-many-data-cent...
With that, though, it sounds like the size of this datacenter is way out of scale compared to the rest of their DC's. They are really going to need to break apart the services that they host there to make sure that DC to DC and region to region fail over works correctly.
(I apologize if the following sounds snarky. I don't mean it that way, I just can't find better wording.)
Microsoft has repeatedly violated my sense of "reasonable" in the past, including in recent times with Windows 10. Therefore this kind of glitch isn't very shocking to me.
Besides the one that AWS and GCP have implemented? That is, to have at least N+1 datacenters? Actually, I think N+1 is the old Google prod regime. I suspect that GCP is at least N+1 per continental region, and I'd be surprised if AWS isn't as well.