Microsoft Azure suffers outage after cooling issue
datacenterdynamics.com
datacenterdynamics.com
> A severe weather event, including lightning strikes, occurred near one of the South Central US datacenters. This resulted in a power voltage increase that impacted cooling systems. Automated datacenter procedures to ensure data and hardware integrity went into effect and critical hardware entered a structured power down process.
It's always quite fascinating whenever cloud platforms like this have "leaky abstractions." GCP had a very long storage service degradation today, as well. [1] I don't know if it's related.
[0] https://www.mysanantonio.com/news/weather/article/Several-re...
[1] https://status.cloud.google.com/incident/storage/18003
edited for formatting
> SAN ANTONIO - Heavy rain caused flash flooding Monday night in northern Bexar County and extending into Comal County. Some areas had up to 9 inches of rain, and water rose on some roadways, including Interstate 10.
> Tuesday morning, the National Weather Service said via social media that parts of San Antonio had up to 6.07 inches of rain, which "smashes" the daily rainfall record of 1.76 inches from 1889.
> Over 9 inches of rain was observed around Stone Oak Parkway, and more than 8 inches between Shavano Park and Camp Bullis, according to the NWS website.
https://www.mysanantonio.com/news/weather/article/mysananton...
If Microsoft didn't own GitHub, this may have prompted a move, but since they do it seems a little redundant given that Github will likely be on Azure too before long.
Unless it’s life critical (911, air traffic control), if it’s down its only going to hamper productivity, but it’ll be back eventually. Time to stretch and get a coffee, and if it’s all day, going home and we’ll start fresh tomorrow.
We’re not saving lives, we’re just building websites. Downtime isn’t shameful, it happens to all of us.
There's a reason why Azure has a SLA and Reddit doesn't ;)
Also if you start comparing the big companies with GitLab then we don't have to continue talking anymore. It's not ok to nuke your production db and that's why everyone in the tech scene laughs about GitLab and comparing them to Azure is like comparing a lego house to brick and mortar.
The one thing we are proud of is our transparency. The community really appreciates our openness and we are happy about it.
Here's the one example https://about.gitlab.com/2017/02/01/gitlab-dot-com-database-...
You can use git outside of GitHub and Microsoft though, I mean, you could always use bitbucket.
I incidentally think this is a completely terrible outcome.
I doubt they will move back to Azure or any cloud for that matter. It's the same story with Dropbox and similar companies. Once past a point in scale, and depending on the case, for example need to control the data security to have certain certifications, it's essential to have your own infrastructure.
It's been a long day because of this. Just going to leave it at that.
Honestly it's hard to imagine a good mitigation for this. "Build more datacenters" is already happening as fast as it can. "Keep enough spare capacity around to handle losing one of the biggest datacenters in the world" is pretty unreasonable.
If you, as a customer, are uptime focused enough that it's worth paying extra, then the sensible practice has always been Cross-Cloud infrastructhre/failovers. At least since the Amazon Easter failure of 2011. That's what giants like Netflix do.
Err what?
It's entirely reasonable to expect Azure to handle the loss of a single DC and not have a 14+ hour global outage. I don't care how big the DC is, losing one should not take out the world, especially not for the length of time this one has been going on.
https://perspectives.mvdirona.com/2017/04/how-many-data-cent...
With that, though, it sounds like the size of this datacenter is way out of scale compared to the rest of their DC's. They are really going to need to break apart the services that they host there to make sure that DC to DC and region to region fail over works correctly.
(I apologize if the following sounds snarky. I don't mean it that way, I just can't find better wording.)
Microsoft has repeatedly violated my sense of "reasonable" in the past, including in recent times with Windows 10. Therefore this kind of glitch isn't very shocking to me.
Besides the one that AWS and GCP have implemented? That is, to have at least N+1 datacenters? Actually, I think N+1 is the old Google prod regime. I suspect that GCP is at least N+1 per continental region, and I'd be surprised if AWS isn't as well.
Having worked for a cloud provider, the reason they are saying that is because they are actively working to understand and fix the problem but haven't come to a well resound solution and thus they cannot give you a decent time estimate because you will probably get even more mad if they under/over estimate the time it took to fix it.
It's akin to waiting for surgery and the doctor saying "we're working on it". I don't want/need the details for the surgery, but tell me everything is ok, what comes next and give me some estimates to set expectations.
When your customer's demands are so directly opposed, you're somewhat caught between a rock and a hard place.
Those who don't need to be bothered with the details can refrain from reading them.
All this would take to keep both sides happy is a little link with "more info" below it.
Why things like this are so difficult, I will never understand.
As a consumer, the lesson here is that Azure is one, big basket. It would probably be prudent to think of AWS and GCS as single baskets too.
https://www.reddit.com/r/AZURE/comments/9cvgn2/is_there_a_st...
Their status page is back up now, but my stuff's still broken. :\
And so does DynamoDB, and Aurora ...
(source: friend who's an engineer at T-Mobile)
But do AT+T engineers carry T-Mobile phones?
If yes then they should put themselves together a deal so that none of the on-call engineers have to worry about running up big bills using their phones. When there are freak weather events they are all in it together.
I suspect we are gonna have to wait at least one other day at best for this to resolve. Meanwhile my local code goes even more out of sync.
I’m probably just gonna spin up a git repo on my local machine and use that to share code with my team.
Anyway I have a good guess as to what most of the employees there are going to be doing for the next six months.
Is failing builds due to formatting issues really a sound setup?
They can also just format their code according to the rules by hand..
When you have a style guide - test for it, if the test fails, then fail the style linter job and don't allow the change to be accepted.
It's failed to meet your acceptable code criteria after all.
If you find you are making your code unreadable just to pass, then your style guide is wrong. That needs fixing, not the CI job.
If you find an urgent "this needs to merge, style rules be dammed" change, allow your senior team members to overrule the style CI job and merge it anyway.
Also unable to lodge a support ticket because the portal fails to identify me as having paid support (that API request appears to timeout).
- https://azure.microsoft.com/en-us/status/
- https://twitter.com/AzureSupport
"NEXT UPDATE: The next update will be provided by 07:00 UTC 05 Sep 2018 or as events warrant."
As I finished writing this they finally updated with essentially the same message except stalling for an additional two hours.
So, if you're thinking "Well, surprises will happen". Yeah, and Microsoft is not actually prepared for that at all, so, sucks to be their customer I guess?
[0] https://www.skyhighnetworks.com/cloud-security-blog/microsof...
https://cloud.google.com/customers/
You know....services consumers actually use.
I refreshed a couple times, and sure, I saw more (on average 1 or 2) that I recognized on each page. But I don't think your response is particularly persuasive. Are you suggesting that the services that I use that run on AWS are in fact, not services I actually use?
Or am I not a consumer? I'm confused.
Edit: Do you hold any Alphabet/Google stock? I've noticed your comment history trends toward dismissing criticism of Google, praising their products, and taking opportunities to speak about the flaws of their top competitors.
So the first page was supposed to be indicative of all of the popular consumer facing services they host? Here, let me help you out: Spotify, eBay, Twitter, Apple iCloud, Verizon, Vimeo, Netflix, etc
>I refreshed a couple times, and sure, I saw more (on average 1 or 2) that I recognized on each page. But I don't think your response is particularly persuasive. Are you suggesting that the services that I use that run on AWS are in fact, not services I actually use?
What popular consumer services were on AWS again?
>Edit: Do you hold any Alphabet/Google stock? I've noticed your comment history trends toward dismissing criticism of Google, praising their products, and taking opportunities to speak about the flaws of their top competitors.
Do you own Microsoft stock? Because quote a few of your posts seem to praise their products and services. Do you work for them?
Single-purpose accounts are not allowed here, especially not when pushing an agenda, and most of all not when pushing corporate propaganda. Of all the things that make HN users angry, that's at the top. And I agree with them.
Most of the time we tell HN users that they're not allowed to accuse each other of astroturfing. When we do find a clear-cut case of abuse that's been getting away with it for this long, I get pretty steamed.
You've also frequently broken the site guidelines by being uncivil, so much so that we've warned you at least half a dozen times. That's more than enough reason to ban you in its own right.
Google has ~3% market share compared to Microsoft's ~28% and Amazon's ~40%. Not even in the same league at the moment. Google is more on par with IBM and Rackspace, for now. Google will undoubtedly make strides in the space, but they haven't been tested.
[1] https://www.google.com.au/amp/s/www.cnbc.com/amp/2018/04/27/...
Where does this number come from?
If it is based on the revenue reported, be very careful with Microsoft's numbers. They report a lot of products as "Azure intelligent cloud", including Office suite subscriptions, on-premise server licences, and software (Windows, SQL Server) licensing revenue from other cloud providers in that number.
Pretty soon their claimed growth is going to flatten out, because they couldn't find any more revenue to report as "Azure intelligent cloud", like PC hardware ...
No other regions were affected, except for global APIs (e.g. create S3 bucket), which one shouldn't rely on on your critical path.
Many new customer features have been delivered to allow mitigation of this kind of failure (e.g. cross-region S3 replication).
AWS had a power failure at one data-center in us-east-1 earlier this year, which had very little impact (basically only customers who didn't have sufficient redundancy in other AZs were affected).
But, it doesn't make the headlines like a 2-hour S3 outage in a single region, which must mean something ...
The last incident I'd personally classify as major lasted 39 minutes and was widely reported: https://status.cloud.google.com/incident/cloud-networking/18...
Disclaimer: I work at GCP but am not speaking for them. I also wish a speedy recovery for our colleagues at Azure: an outage like this can only result from many things going sideways simultaneously, and both the cause and recovery can be complicated in ways that flippant "well why didn't you just N+1 it" commenters here on HN can only guess at.
Besides the one you listed:
* https://status.cloud.google.com/incident/compute/18005
"Google Compute Engine VM instances allocated with duplicate internal IP addresses, stopped instances networking are not coming up when started." - 22 hours
Newly-launched instances, or instances that were stopped and started, received duplicate IP addresses. 4.5 hours in a mitigation was provided, but it was only resolved after 22 hours, and customers may have had to still fix individual instances. As far as I remember, this was global, and there is nothing on the status page indicating it was limited to one region or a subset of regions. So, for 4 hours, if you needed to create a VM with working networking, you couldn't, anywhere on GCP, and no mitigation was available. Do you not consider this to be "global"?
* https://status.cloud.google.com/incident/compute/18009
"Instances using Local SSD might experience VM failures. This affects GCE VMs globally. No data corruption has been observed." - 5 hours
The original claim was:
> So AWS has had some big outages, as has Azure. Has GCP had any big outages yet?
I said:
> GCP has had multiple many-hour (6+) GLOBAL outages in the past year. I think it's at about 3 so far this year.
So, maybe it's only 2 major global outages, or maybe it's 3 5-hour+ global problems, but the only way anyone can claim that Google hasn't had any big outages is if they don't have enough market share for a global outage to affect many websites or end-users.
Sometimes I can't understand people.