Postmortem: Azure DevOps (VSTS) Outage of 4 Sep 2018
blogs.msdn.microsoft.com
blogs.msdn.microsoft.com
Anyone have any data around how frequently such a failure occurs?
This shouldn't matter, only to them and whoever operates the datacenter.
> and their systems are not designed for speedy restoration into another region.
This is what matters. You have a cloud service that is not really designed for the cloud.
They linked to the Azure service details in the Azure Status link if you wanted more information: https://azure.microsoft.com/en-us/status/history/. It sounds as though they're compiling a report on this problem which will be released in the coming weeks.
> The primary solution we are pursuing to improve handling datacenter failures is Availability Zones, and we are exploring the feasibility of asynchronous replication.
This I do not understand. I was also amazed when I saw that Azure AZs are not available on all regions. In AWS, the bare minimum is 2 AZs (except for one odd region). Same thing for Google Cloud.
For instance, S3 promises that it is replicated between 3 AZ's in a region. That guaranteed is available in regions that only have two publicly available AZ's.
It bothers me that increasing number of large companies dump their own data centers to jump into The Cloud. Thia means future outages (which will undoubtedly happen) will have wider and wider impact on end users.
For example, if your email is hosted on AWS and it goes down, you loose access to your email. No big deal. However, if your email, VOIP and IM/chat go down at the same time, you may loose all ability to communicate electronically. This can be a very big deal in certain situations.
This is not true. I've watched some full region outages going on while our systems were chugging along just fine. I have also seen temporary network issues in one region(that caused our monitoring to go haywire) where others were completely unaffected.
What tends to happen is that AWS locates a bunch of services in us-east-1 (Route53 being one example). This may knock out the control plane for a bunch of services, which shouldn't be a thing.
Yeah! It has a green dot!
Really makes you think that with all these big players pushing chaos engineering and multi-cloud deployments, the cloud suppliers should really be more prepared.
How about at least two "chaos regions" for every supplier, available at half cost, but with random outages, to really test our applications, and the suppliers infrastructure!
Experience kind of matters in this case and allows smaller IT departments to outsource resilience to more skilled companies/workers.
Nowhere near 99%.
Running a large service on top of Azure or AWS is not a walk in the park. You don't manage hardware and you deal less with provisioning, but you need a significant level of expertise in dealing with idiosyncrasies of the particular hosting you've chosen.
I've seen companies transition to AWS. Very often they spend months "figuring things out". And then it turns out that they've just traded low-level problems which everyone knows how to solve for high-level problems no one knows how to solve.
Also, just because AWS/Azure is running doesn't mean your service is up. You end up managing high-level company-specific infrastructure on top of low-level one-size-fits-all infrastructure. Both can fail.
A lot of companies jump onto AWS/Azure/other big clouds because of hype, bandwagon effect and the general feeling everything will become "world-class" by magic. In many cases there is no legitimate cost-benefits analysis that compares costs of AWS switch to costs of fixing whatever problems company's current infrastructure suffers from. Also, almost no one factors in the costs of vendor lock-in and "too big to care" effect.
There are definitely use cases where Big Cloud makes sense, but it's nowhere near 99%.
Source: https://www.internetsociety.org/internet/history-internet/br...
Separately, I think we have enough history of working with the cloud at this point to demonstrate that major providers' availability is on par with, or better than, the availability of the typical small entity. Sure, the impact is potentially wider spread (although this can be mitigated with a cellular architecture, which first-class providers do employ), but there's a perverse advantage that when outages occur, they tend to get fixed a lot faster because the complaint volume is much higher.
On the other hand, they can be much harder to fix, because the sheer scale of failures and complexity of the infrastructure. There is a higher probability of complex systemic issues, as demonstrated by this very outage.
There are plenty of smaller providers that beat Azure VMs in uptime. Plus, smaller websites/services can employ much simpler failure mitigation strategies.
The necessity of low-latency-but-decoupled-physical-plant AZs is well known in the art by now, and these issues will no doubt be addressed as Azure matures. Remember, they're 5 years behind AWS.
Availability zones are a mitigation. The issues is the sequence of events and dependencies described in the postmortem. The description has six paragraphs.
And how could your perfect model, whatever that is, survive a similar catastrophic DC failure without availability zones?
Of course, Azure has such a massive footprint that the bigger issue is, why wasn't that used as an advantage when South Central US went into a shutdown mode? This does play against the notion that with so many regions this sort of event should have been preventable.
Simply google AWS outage and see the results. Thankfully an entire datacenter outage is rare and Azure is taking steps to mitigate that risk.
Not again.
TBH I'd do the same.
Just check the github status pages, there are orange level issues almost weekly and you don't need to go back too far for full outages.
As this case tells, Azure, isn't spending a lot on failure testing. Also, having experience with being an big Azure customer, I can tell you that things look better than they are.
Azure looks shiny from the outside but we've had way, way more problems, from uptime to bad APIs to awful language SDKs to bad user interfaces to licensing hell, than I've ever had on AWS or GCP. It's so bad that I am currently weighing whether or not to advocate for a migration off of it, at nontrivial expense, because I cannot pretend to provide reliable services for our customers.