Azure Active Directory down
twitter.com
twitter.com
https://twitter.com/MSFT365Status/status/1371554704518352896
They actually acknowledged it was due to a "recent change to an authentication system" where Azure Support just vaguely mentions an outage
Took out Outlook, Teams and all meetings over the course of 5 minutes, then moved on to internal service alarms going off across multiple (10+) major customer facing digital products that got vaporized due (presumably) to various dependencies on AD. I think some of this was service accounts on backend that lost connection to various dbs, file systems, etc.
What a total shit show.
Company basically down from 2pm CST and barely coming back up 5 hours later.
IT Buddy who is 2 people away from CIO pinged me and said “DNS issue, data center power outage or Azure just crapped its pants.”
a. the status page update took as long as it did.
b. the status page still only claims that AAD is down. You're not "up" if a required dependency of yours is down, IMO, and most of the Azure services right now are severely degraded to the point of not being useful due to the AAD outage. Now, that would make your status page look bad, yes. (And I would clearly message on it that the other outages are fallout of AAD on that.)
c. I don't have any real expectation that we'll get a public PM about this. (And a PM must include sufficient information for me to understand what went wrong, and what's been done to prevent it going forward.)
I had been starting to wonder what sort of incident it would take for them to actually update the status page. I guess this is it. We've witnessed a number of outages in various services over the past few months, none of which made it to the status page. I pushed a support rep on the issue, and was told that they don't want to cause a panic.
I looked at their past status histories, and they have provided pretty detailed postmortems for much more minor incidents. I don't have any reason to believe that they won't do that this time. I'm new to Azure, so this kind of outage really looks bad in my mind, but it's par for the course with Cloud stuff. They'll fix it and there will be one less thing to break in the future.
Documentation that contains triple negatives, UI elements that have no consistency (clicking off a modal on one page does nothing while doing so on another closes it and clears your data), widgets and modals forgetting state when you navigate to different submenus.
The sort of unexciting stuff that by itself is ok but it slowly makes you descend into madness.
The best we could get was ~45ms using pgbench vs ~10ms local and ~18 using AWS Aurora.
Their solution, wait for the flexible server offering which is not available in AUS but when tested in another region showed ~30ms so still not fit for purpose.
We eventually had to switch to running PG in AKS which has its pros and cons.
The funniest thing was, running Azure VM to AWS PG was faster than running Azure VM to Azure PG.
We hit the 4tb storage limit that they didn't tell us about on it when the documentation said 16tb. They told us we were on the older storage tier and that we needed to rebuild a huge analytics database on our time because they couldn't migrate us to a backend storage tier. So we had a database server that was in read only mode until we could migrate off of it.
Then we had random restarts and outages with their proxy that handles all of the connections. Which is why you actually need the username@db-name, because without the @db-name it doesn't know where to route you, because the postgres connection string is actually just mapping to one IP that they manage for multiple databases!
Fun trick is you can make your url anything that maps to the IP they use for their proxy and it'll take you to whatever the @db-name is in your connection string.
Their solution to the random connections dropping for minutes at a time was to "implement retry logic" on our applications.
Our solution was to migrate away from it.
People always talk about cloud as an entity which can/shouldn't have problems at all. When it's more about convenience then anything else. So I wouldn't worry too much about this if I were you.
Messaging around transient issues helps. E.g., AWS would often email us to tell us "hey, sorry, such & such VM is being hosed by the underlying hardware, it is migrating but you'll see a forced reboot". That's all I need to know: I expect some interruption to that portion of the system, and I expect to see it heal within a certain time frame, and if it does, great, no support ticket required.
Azure sort of has that with "Resource health"; but, e.g., one of our recent support tickets is that we had a VM reboot unexpectedly due to a resource health event with its disks. And then, a few hours later, another VM, same resource health event, same issue, same reboot. And then again, a few hours later. And that pattern, with no further communication, requires me to write the (obvious) ticket of, "So… what's happening?"
(That ticket ended with such and such service was experiencing an incident. Never made the status page. It did make the internal "Service alerts" page.)
Part of issuing a PM is also to help show & convince me, the customer, that I won't be filing support issues for the rest of my life. But this last 6 months has just felt like support ticket after support ticket, and it's sort of depressing, since I'd rather be coding. I do think, last spring or so, it was no nearly this bad.
¹Every place I've worked at… we page on 5xx codes sent to a client. If Azure is internally doing that, it sure is hard to tell from where I sit.
I had to use the IBM cloud for a project once upon a time and the status monitoring was almost always wrong ...
I really hope that as customers we can succeed in making "excellent status reporting" an absolute minimum requirement for cloud services. Observing the quality of the status page reporting history is one of the first things I do now when evaluating a new service -- and no I'm not just looking for a page that shows all greens -- I'm looking for a page that connects to some conceivably meaningful metric as well as detailed reports about major and minor outages ... something to prove that they actually even know if their service is working at a given point in time ...
The IBM cloud page was a page that tended to update in response to my own issue reports (usually if I also pointed out that nothing about the outage appeared on the status page ...). I don't think that's at all an acceptable standard ...
It turned the whole afternoon into learning time at our company. Thankfully our Okta integration goes through our on prem AD servers and not purely AAD. Otherwise I wouldn't be able to get to learning resources which authenticate through AD!
Strange that HN was down at exactly the same time, although this is reported to be unrelated.
> CURRENT STATUS: Engineering teams are currently rolling out mitigation worldwide. Customers should begin seeing recovery at this time, with full mitigation expected within 60 minutes. This message was last updated at 21:12 UTC on 15 March 2021
DX10501:+Signature+validation+failed.+Unable+to+match+keys:+ kid:+'[PII+is+hidden]',+ token:+'[PII+is+hidden]'.
and sure enough, it seems it was some authentication thing going sideways
What is more crazy is that low level service-to-service authentication in Azure is primarily based on Azure AD. This includes Key Vault and Storage Accounts. Both were mentioned in the Microsoft notification page as affected by this outage.
How can anyone build a reliable service on top of Azure if what is essentially the failure of the "Office 365 login" also breaks your VM's access to its SSL certificates (or whatever)?
Trying to provide globally available, replicated service that meets every and all needs 24/7/365 is basically making Active Directory about 600x more complicated than it would be otherwise. It's basically impossible for a service that complicated to meet the uptime of... installing Windows Server 2019 on a VM and patching it monthly. (Bear in mind, if I know when my business is not affected by an outage, I can do it without being disruptive. Microsoft, by definition, cannot. Every moment is critical since their customers are everywhere.)
There are types of businesses to which the former might be a necessary solution, but most would be better off with the latter. It'll be really interesting when people start realizing the cloud is mostly just a scam to get people on subscription revenue streams, and not actually providing any greater reliability or less management overhead than what they had before.
Microsoft is always a shitshow this time of year. Infrastructure changes usually land ahead of spring releases.
Azure AD taking down pretty much the entirety of Azure / O365 / Teams is frankly inexcusable. Astounding incompetence, and an astonishing single point of failure that needs to be re-architected.
You have to make the service as resilient as possible, which clearly they failed at, but you can't very well fail over from your AD service to something else.
Huh? x509-type authentication works without a single point of failure. I can authenticate myself to any client via public-key cryptography without any external dependencies, after we've agreed on a root of trust.
Active Directory chose a solution that requires centralized availability. They didn't have to do it this way (though it does make certain administrative tasks, like revocation, simpler).
Note that even AD itself does not require a single point of failure though! You can define secondary (i.e. failover) DCs (usually via DNS). Then DNS becomes your point of failure -- and again most DNS implementations support failover, such that any single point of failure gets pushed lower to the stack (e.g. NICs), which also support failover or multipathing, etc, if you really need high availability.
Anyway, it's totally possible to make authentication not be a single point of failure. Of course, you have to make some tradeoff to do it (see: the CAP theorem).
> Azure AD taking down pretty much the entirety of Azure / O365 / Teams is frankly inexcusable
My objection was that the Microsoft Azure AD service is down, there's nothing that Microsoft can do to protect the services that rely on it.
Setting aside your valid x509 counter-argument, the remaining options you listed to increase availability, to me, have to do with resilience of the Azure AD service. They're not alternative services.
It depends on where you draw the boundary around the service. Again, clearly Microsoft has failed to provide sufficient resiliency, not arguing against that at all. But once the service is down, practically by definition Teams and Exchange and everything else is doomed.
Early signs are not pointing to this as a service resilience issue - this appears to be sheer incompetence - pushing an update that was likely not tested well enough that broke pretty much everything. More than that, why aren't updates to Azure AD being rolled out regionally? Why is Azure AD architected in such a way that the entire thing going down can do so much damage?
It's pretty hard to be understanding with a preventable screwup with this kind of global scale.
Some days later and after our complaints, Microsoft rolled back the change.
I would call this evidence that they do practice gradual deployments of changes.
I would guess that this outage was caused by something more complex than MS not slow-rolling their deployment.
Between Exchange 365 outages and Azure AD outages, I'm not sure how anyone thinks moving Windows infrastructure to the cloud is not an existential risk to business operations.