We wanted to use their hosted kubernetes solution(I forget the name) and pods would just randomly lose connection to eachother, everything network related was just ridiculously unstable. Host machines would report their status as healthy, but then be unable to start any pods, making scaling out the cluster unreliable. I also remember a colleague I regarded as a bit of a wizard being very frustrated with cosmosdb, but I cannot for the life of me remember what the specific issue was.
Our solution was actually quite well written, if I do say so myself, we had designed it to be cloud agnostic, just on the off chance that something like this would happen (there may have been rumours this acquisition would happen ahead of time).
But Azure was just utterly unable to deliver on anything they promised, thus the write-off on my part.
To see the positive side of it, it’s a Chaos Monkey test for free. Everything you deploy must be hardened with reconnections, something you should be doing anyway.
What’s frustrating is that it happens rarely enough to give the illusion of being stable, making it easy for PM to postpone hardening work, yet often enough to put a noticeable dent in your uptime if you don’t address it. Perfect degree of gaslighting.
Keep in mind that in Azure it's a must-have, whereas everywhere else it's either a nice-to-have or a sign your system is broken.
But my impression is that it's better now. In general my experience with azure is that the base services, those who see millions of hours in use, are stable. Think VMs, storage, queue etc. But the higher up you go in the stack, the fewer hours of use do they see, and the lower quality it gets.
Thank you for the insights!
Or put it in another way, if Mojong were to start with Azure but couldn't manage to migrate to AWS, which provider is the parent going to write off?
It was easier to go to AWS than to Azure; and I've done both in the past ~4 years. Migrating to AWS was just technical work. Migrating to azure was 'fight unexpected bugs and the fact it doesn't actually work in surprising situations'.
The only reason to go to azure was that Microsoft waved a big fat $$$ discount to use their services.
Migrating to AWS was a breeze.
We have a very long list of Azure features and services that we've banned people from using.
Just got off a call with someone at Azure today who told us to setup our own NAT gateway instead of using Azure's because of an outage where we made too many requests and then got our NAT Gateway quota taken away for the next 2 hours.
Maximum 55k concurrent connections. After that, they make you deploy NAT gateways in other availability zones. And a max throughput of 10 Gbps.
I imagine AWS would also tell you to deploy your own gateway if you were running into the 55k concurrent connection limit of managed NAT.
AWS tends to be flexible with their quota enforcement in my experience, though.
Care to point out a concrete example? I've worked with Azure a few years ago and I wouldn't describe it as buggy. At most, I accuse them of not "getting" cloud as well as AWS. For example, the whole concept of a function app is ass-backwards, vs just deploying a Lambda with a specific capacity. That is mainly a reflection of having more years working with AWS, though.
My experience with Azure is that it simply breaks in ridiculous and frankly unacceptable ways, all the time. It's like someone is unplugging network and power cables every couple hours just for fun.
I'm not sure how well founded the stability argument is. I still remember the infamous series of AWS outages that took place a couple of years ago.
The fact that AWS invests heavily in vendor lock-in is a problem created by AWS, not their competitors.
You decide if that’s more or less stable than AWS.
I’d say the evidence is pretty empirical, but hey, all I can say is my experience was utterly unambiguous.
You can argue a lot of things, but hundreds of azure fails in a big giant list is probably one of the tougher ones to go “no, this is fine compared to AWS!” about, imo.