MS Azure down: An emerging issue is being investigated
status2.azure.com
status2.azure.com
So, the outage appears to be that DNS for `azure.com` and maybe also `windows.net` (blob storage for us, but I'm not sure) is not resolving.
So, the OP's link here is broken. Tweets indicate that it might be intermittently resolving.
https://twitter.com/AzureSupport/status/1377737333307437059
> Warning sign We are aware of an issue affecting the Azure Portal and Azure services, please visit our alternate Status Page here https://status2.azure.com for more information and updates.
Which is the link above, and is also down for me & many others.
Edit: seems hit or miss. A coworker got a successful resolution of status2.azure.com of 104.84.77.137 , so manually, I can get there now. They're directing customers to an all green status page, except for the "An emerging issue is being investigate." bit at the top… (I know of at least two services that are not happy…)
It had to have been less than a month ago that AAD caused a cross-service global outage. Now it's DNS. It's always DNS.
Our current provider, with whom we have a number of dedicated servers, might be piss poor in some ways (takes ages to get changes made, lack of pricing transparency, kind of expensive for what they are), but I don't remember the last time they had an outage. There's definitely been one in the last year, but two in the space of a month is ridiculous.
Is this normal for Azure? Is AWS any better?
Sure, you can screw up on prem. But if you are only marginally competent spinning up services/servers/clusters takes but an instant, it's dirt cheap, and vastly more reliable.
Our only significant outages in the past 10 years have been our services dependent on big clouds. Including the Pre-Thanksgiving AWS one last year.
AWS outage history here https://en.wikipedia.org/wiki/Timeline_of_Amazon_Web_Service...
Do you mean on premise you can spin up virtual servers and virtual clusters instantly? Because I imagine you'd have to order, physically rack, and setup the bare metal if one doesn't have the capacity already.
There is no need to make blanket statements about on-prem vs. cloud. You have to weigh the pros and cons as any other business decision.
It’s about controlling you’re own destiny with either decision.
Now I work for an Global Enterprise and the benefits of cloud aren't necessarily about reliability--they're more about the speed and efficiency that pet projects can be spun up and spun down without international labor laws and multi-year leases.
I sure hope there aren't any emergencies in Canada until this is resolved...
Meanwhile Azure is the only cloud with two Canada regions (Canada "Central", Canada East)[0], AWS has one (Canada "Central")[1], GCP has one (Montreal same city as AWS)[2].. there's really only one player if you want to use managed cloud services.
[0]: https://azure.microsoft.com/en-us/global-infrastructure/geog... [1]: https://aws.amazon.com/about-aws/global-infrastructure/regio... [2]: https://cloud.google.com/about/locations/
[0]: https://www.publicsafety.gc.ca/cnt/mrgnc-mngmnt/mrgnc-prprdn...
Governments of the world: pay your IT people more money to prevent brain drain.
I’m sure there are some cases where the cost is worth the complexity, but I don’t think it’s cut and dried as written here. Most cloud providers are very reliable, and my guess is that you are more likely to have a self-inflicted outage due to the complexity of your infrastructure than the cloud provider having an outage.
It’s atypical motivation but one of the few verticals I know of where the cloud is largely out of the equation.
I think Azure is cheaper for certain workloads as well, and at one point had DCs in places where AWS didn't. But it's mostly the "we already buy X Microsoft product, and they cost about the same, so..."
But if I had to make an a priori prediction about which systems would have the better reliability over the long term - cloud-based or those run out of a corporate DC - my money would be on the cloud-based systems.
The cloud providers are going to have better management of power/network/hardware than the most mature government agency, simply because it's a core capability for them.
In contrast UK gov servers like hmrc.gov.uk (for example) are pretty stable, I can't remember the last time I heard about an outage.
Cloud providers are certainly better at some things (like staying up to date with latest tech), but I'd contend reliability is not one of those things based on their prominent and frequent outages (at least several a year).
Maybe not, but even just a "flick the switch to go back to a dumb system" option is worth maintaining
The right lesson is "Be a person that strives for excellence in all spheres of life".
I've been doing "Local IT" since '96 and I spun up an EC2 instance the day I saw it announced on slashdot and have been using both ever since. Both excel in some ways in the hands of good people.
Anyone thinking "I'll move to the cloud (or to on-prem) and all my problems will magically go away" is fooling themselves. If you suck at on-prem those underlying issues will carry into the cloud. If you have excellence in a well run on-prem install you'll experience great benefits leveraging the cloud.
For instance, Last month I was discussing a "move to the cloud" with a bank CTO. They had an unreliable on-prem network and moved to the cloud.. and they just discovered after suffering an outage in the cloud what an "availability zone" was. The same attitudes that made their on-prem unreliable, insecure, expensive will make the cloud the same way for them.
But wanting is one thing, having the money to implement and maintain such a solution, a quite another thing :-/
I don't think they could build infrastructure even if they wanted to.
[0] https://www.tbs-sct.gc.ca/agreements-conventions/view-visual... [1] https://www.levels.fyi/company/Microsoft/salaries/Software-E...
You can't make this stuff up.
Think about it - what kind of work does involve "reliable infra" today?
Instead of knowing one Linux system well and maintaining multiple servers in different data centers with different providers you now have the stupid overhead of multiple cloud infra providers with all their lockin-pitfalls and incompatible specialities.
The promises of cloud have not been delivered.
It is all fake.
Look at the comments on any entry and it’s clear people do stuff like report “outages” for Google because a random website won’t work in Chrome.
There’s a comment on the AWS page complaining that iTunes gift cards are being slow to arrive. The Google page has people who think the comments are the place to talk to Google support.
A major outage will make itself known in far clearer ways.
The website simply answers the question "is anyone else having trouble accessing this?" In practice that's all it does, and that is useful.
I wish those companies would actually have status pages.
So it's just a mess.
>As discussed, we found that your case could be part of a global scale incident wherein multiple Dynamics services were rendered inaccessible due to a suspected Azure DDoS attack. This issue started at approximately 21:30 UTC on 1st April 2021 and, along the symptoms, our customers may experience intermittent issues accessing Microsoft services, including Azure, Dynamics, and Power Automate. On this regard, the Microsoft teams involved have rerouted traffic to separate resilient DNS services and are seeing improvement in service availability. With this confirmation, as agreed, I will proceed to lower the severity of this case and transfer it for an agent working in your time zone to continue following up with you and confirm these services have returned to an expected operational state.
InformationAzure DNS - Investigating
We are currently investigating reports of an issue affecting Azure DNS. More information will be provided as it is known.
This message was last updated at 21:59 UTC on 01 April 2021
----------------- WarningDNS issues - Investigating
Engineering is investigating an issue with DNS that is impacting several downstream Azure services.
This message was last updated at 22:07 UTC on 01 April 2021
Then GitHub Actions and some services stopped working and went on holiday today: https://news.ycombinator.com/item?id=26666782
Can't sign into portal.azure.com, can't hit Azure File Shares, etc.
The last outage a few days ago was enough for my company to up and move most of our stuff to AWS. This new outage is enough for us to fully migrate away from Azure.
What a cluster.
Maybe Microsoft should invest less in their shiny new AI/ML platforms and more into stability of their core services.
It seems like most larger customers that use Azure do so because management got shiny presentations from Microsoft and now it's their "strategical partner".
A lot of overselling with huge discounts gets them their in. Already seen this at multiple companies. Azure is a nice platform for your Windows administrators to shift some load to the cloud. But to build large applications on?
Edit: And I kinda feel bad for saying this, since I assume that there are indeed pretty competent engineers working on Azure. But somewhere something isn't right.
But seriously, as inflexible, painful and tedious to support our on-prem infra last time we had critical outage was ~3 years ago, for like an hour. We've almost finished migration to azure. I hope this outage and AD outage earlier this month are outliers.
ex: hidden primary in cloudflare, with some sort of automated secondaries in az + aws (and how does replication get automated?)
At least use well formed English when posting your update...
To be clear: I’m not ridiculing anyone in particular (native or non-native). I’m pointing out that a multi-billion dollar company experiencing a major outage could probably proof read the statements they put out about that incident before publishing. Else why not just go “if ur dns is bad its probs us soz, were on it”