400 reasons to not use Microsoft Azure
azsh.it
azsh.it
A couple of years back, I was working at Mojang (makers of Minecraft).
We got purchased by Microsoft, which of course meant we had to at least try to migrate away from AWS to Azure. On the surface, it made sense: our AWS bill was pretty steep, iirc into the 6 figures monthly, we could have Azure for free*.
Fast forward about a year, and an uncountable amount of hours spent by both my team, and Azure solutions specialists, kindly lent to us by the Azure org itself, we all agreed the six figure bill to one of corporate daddy's largest competitors would have to stay!
I've written off Azure as a viable cloud provider since then. I've always thought I would have to revaluate that stance sooner or later. Wouldn't be the first time I was wrong!
Or put it in another way, if Mojong were to start with Azure but couldn't manage to migrate to AWS, which provider is the parent going to write off?
It was easier to go to AWS than to Azure; and I've done both in the past ~4 years. Migrating to AWS was just technical work. Migrating to azure was 'fight unexpected bugs and the fact it doesn't actually work in surprising situations'.
The only reason to go to azure was that Microsoft waved a big fat $$$ discount to use their services.
Migrating to AWS was a breeze.
We have a very long list of Azure features and services that we've banned people from using.
Just got off a call with someone at Azure today who told us to setup our own NAT gateway instead of using Azure's because of an outage where we made too many requests and then got our NAT Gateway quota taken away for the next 2 hours.
Maximum 55k concurrent connections. After that, they make you deploy NAT gateways in other availability zones. And a max throughput of 10 Gbps.
I imagine AWS would also tell you to deploy your own gateway if you were running into the 55k concurrent connection limit of managed NAT.
AWS tends to be flexible with their quota enforcement in my experience, though.
Care to point out a concrete example? I've worked with Azure a few years ago and I wouldn't describe it as buggy. At most, I accuse them of not "getting" cloud as well as AWS. For example, the whole concept of a function app is ass-backwards, vs just deploying a Lambda with a specific capacity. That is mainly a reflection of having more years working with AWS, though.
My experience with Azure is that it simply breaks in ridiculous and frankly unacceptable ways, all the time. It's like someone is unplugging network and power cables every couple hours just for fun.
I'm not sure how well founded the stability argument is. I still remember the infamous series of AWS outages that took place a couple of years ago.
The fact that AWS invests heavily in vendor lock-in is a problem created by AWS, not their competitors.
You decide if that’s more or less stable than AWS.
I’d say the evidence is pretty empirical, but hey, all I can say is my experience was utterly unambiguous.
You can argue a lot of things, but hundreds of azure fails in a big giant list is probably one of the tougher ones to go “no, this is fine compared to AWS!” about, imo.
We wanted to use their hosted kubernetes solution(I forget the name) and pods would just randomly lose connection to eachother, everything network related was just ridiculously unstable. Host machines would report their status as healthy, but then be unable to start any pods, making scaling out the cluster unreliable. I also remember a colleague I regarded as a bit of a wizard being very frustrated with cosmosdb, but I cannot for the life of me remember what the specific issue was.
Our solution was actually quite well written, if I do say so myself, we had designed it to be cloud agnostic, just on the off chance that something like this would happen (there may have been rumours this acquisition would happen ahead of time).
But Azure was just utterly unable to deliver on anything they promised, thus the write-off on my part.
To see the positive side of it, it’s a Chaos Monkey test for free. Everything you deploy must be hardened with reconnections, something you should be doing anyway.
What’s frustrating is that it happens rarely enough to give the illusion of being stable, making it easy for PM to postpone hardening work, yet often enough to put a noticeable dent in your uptime if you don’t address it. Perfect degree of gaslighting.
Keep in mind that in Azure it's a must-have, whereas everywhere else it's either a nice-to-have or a sign your system is broken.
But my impression is that it's better now. In general my experience with azure is that the base services, those who see millions of hours in use, are stable. Think VMs, storage, queue etc. But the higher up you go in the stack, the fewer hours of use do they see, and the lower quality it gets.
Thank you for the insights!
> S3
Hasn't this recently been an issue where Amazon arbitrarily changes the S3 contract and all software following it as a spec has to play catch-up?
Although perhaps it’s about time we made a proper standard based on S3, yeah.
The S3 system is proprietary to Amazon, and it's your fault if you're not using Amazon but you're relying on Amazon to not change it anyway, because they have no obligation to you.
The concept of object storage is not proprietary. You should be able to change your code to use a different object storage provider.
I was trying to point out that S3's protocol is "proprietary" so if you're using it you're still (somewhat) using "proprietary bullshit" in your stack!
Not that most of our customers whose AWS environment we audit do much of this, at least not beyond some basics like creating a vlan ("virtual private cloud"), three layers of proxies to "load balance" a traffic volume that my old laptop could handle without moving the load average above 0.2, and some super complex authentication/IAM rules for a handful of roles and service accounts iirc
(The other half of the point is infinite scale, so that you can get infinite customers signing up and using the system at once (which hopefully pay before your infinite AWS bill is due), but you can still do that with VPSes and managed databases/storage.)
And one small note: apart from S3, virtually all AWS services are tied to a VPC, any kind of deployment starts with "ok, in what VPC do you want this resource?".
At Azure or GCP you pay a similar price but you don't even get the reliability so literally why would you use them? The only reason I see is that "cloud" means you can start instances at any time without a setup fee or contract duration. But with the amount of cost difference, you could have three times your baseline "cloud" load running all the time at a non-cloud hoster, and still save money!
> Isn't at least half the point of AWS to use their SaaS thingies
It is. (That’s how they lock you in!) I think it’s okay to use some AWS stuff once in a while, but I’d be wary of building your whole app architecture around AWS.
I’m in the self-hosters camp myself :-) I’m building a Docker dashboard to make it easier to build, ship, and monitor applications on any server: https://lunni.dev/
I’ll rethink the landing page animations a bit later! (I was thinking about redoing it from scratch again, anyway :^)
Yep yep, we're on the same page here, just that it sounded a bit like you recommend people to use EC2, managed Postgres, and S3, when that seems a bit like that defeats the purpose of using a big cloud vendor in the first place since those products can be gotten much cheaper elsewhere
Just clicked the link to your website: your hero statement I sorta have on a t-shirt (there is no cloud, just someone else's computer). The footer statement "To improve your browsing experience we don’t use cookies" I have nearly literally on one of my websites. We are definitely on the same page xD. The website seems to think I'm from Talinn btw, maybe a reverse proxy issue (that it looks up the IP address' location of a proxy instead of that of my German IP address)? I do second the remark posted in a sibling comment about fading in content, but (maybe you adjusted it already) it fades fast enough that it doesn't bother me as much as on some other websites. Could still be faster though imo
Good luck with the product!
Except Azure - AFAIK it's pretty much the only cloud provider that doesn't support S3 API, see e.g. https://learn.microsoft.com/en-us/answers/questions/1183760/...
And there’s also an issue with client libraries and their compatibility with various implementations. I recently discovered this issue:
https://github.com/boto/botocore/issues/3394
This is GCS implementing the S3 API incorrectly in a way that really ought not to break clients, but it’s still odd because the particular bug on GCS’s end seems like it took active effort to get wrong. But it’s also boto (the main library used in Python to access S3 and compatible services) doing something silly, tripping over GCS’s bug, and failing. And it’s AWS, who owns boto, freely admitting that they don’t intend to fix it, because boto isn’t actually intended to work with services that are merely compatible with S3.
As icing on the cake, to report this bug to Google, I apparently need to pay $29 plus a 3% surcharge on my GCS usage. Thanks.
Time to check out OpenDAL, I suppose.
That's the price of a support contract, not a "bug report". And it's not "plus", it's "or": support costs $29/month or 3% of your monthly billing, whichever is greater. It comes with SLA agreements for fixing or working around your reported problems. Though obviously in this case they'll probably just tell you to use their own python library and not boto.
And this situation is bad business. Google advertises that GCS has S3 interoperability support. And they have customers who use it in its interoperable mode. Presumably those customers could use GCS’s biggest competitor, too. Shouldn’t Google try to make the S3 interop work correctly?
Honestly it seems to me like you're excited to have found a bug and want to report it for glory; we've all been there. But No One Cares about that stuff in the world of commercial software. They fix bugs for real customers, not internet rock stars. If you aren't losing even $29 (one mid-tier meal!) from this bug, well... does it even rise to the level of "yell about it on HN?".
What’s the point of S3 compatibility then?
Again, this isn't an open source product. It's not like there's a universal agreement that coordination and compatibility are ideals to be respected, nor a public place for discussion and development. Google is trying to make GCS as S3 compatible as possible, but obviously they prioritize their paying customers (who seem not to have reported this issue). Amazon does not want GCS to be boto-compatible at all, really.
Now of course, that's a very sizable if...
We migrated from AWS to GCP in 2016/2017 (mostly VMs and related stuff, CloudFront, etc - no lambdas) and it was pretty painless and everything worked smoothly until the end of that company.
https://www.zdnet.com/article/ms-moving-hotmail-to-win2000-s...
...and then .NET and SQL Server started shipping for Linux.
Can't say much more, but I worked on a huge (internal) Sybase ASE on Linux based app (you've _all_ bought products administered on this app ;) ) way back (yes, pre-SSD, multi path fiber I/O to get things fast, failover etc.) and T-SQL is really nice, as is/was ASE and the replication server. Been about 20 years tho, so who knows.
https://www.microsoft.com/en-us/sql-server/blog/2016/12/16/s...
Can't say the same for Oracle...
Microsoft SQL Server has long stop being Sybase SQL Server, and works on Linux by making use of Drawbridge.
https://www.microsoft.com/en-us/sql-server/blog/2016/12/16/s...
I guess wording wise for my comment, the "Hah, they didn't actually write that themselves, they just bought the Sybase rights to everything license" got the better of me :)
To be fair again, from what I hear, SQL server at Microsoft did some nice things on top of Sybase but the base, T-SQL, is just nice overall and by itself. I really want to like Postgres (and I do) but some of the awesome things I had with actual Sybase ASE 20+ years ago, Postgres still does not have. And that was a piece of software that had those features I loved for 10+ years prior to when I started working with it. The app we're talking of here was 15 years old when I worked on it 20+ years ago and it's probably still around and very probably still uses Sybase ASE (tho the actual app was converted from Smalltalk to Java ;)
I also later on had to use Oracle and had the same "WTF? You can't do that?" experience :shrug:
https://learn.microsoft.com/en-us/iis/get-started/introducti...
Yeah, the advantages (RCE) were copied by modern web browsers. /s
unlike shit show that was windows 95/98/ME
I don’t hate windows 2019 but Linux is better, easier, faster and a relief after any futile attempts to use IIS or sql server in 2025.
The first generation of tabletised 8/Metro interfaces made me audibly groan every time I had to RDP into machines running 2012.
ninja proof: https://i.imgur.com/l29rDVo.jpeg
(My fever dream wish is for a "distribution" of NT that boots in text mode and has an updated Interix subsystem alongside Win32. Throw in ZFS and it would be awesome.)
I too wish for an NT that was CLI-only, striped of services as much as possible.
Starting Windows NT...
C:\>
It's too bad Microsoft has no interest as a business in on-prem software.I agree re: MSFT having no interest in on-prem software. It saddens me.
Like I said in my earlier post, text mode NT is my fever dream fantasy. Maybe you were saying the same thing.
Text-only mode would be wonderful even if all you could do is look at a blinking cursor.
I'm not as much against windows as I uses to be but I'm not budging off Ubuntu LTS even though they too try really hard to rock the boat.
Really? Oh, compared to other Windows versions...
Because it never came close to the stability of OS/400, Netware 3, AIX, Solaris or even OS/2 v2.
I saw years of uptime on those systems whereas Win2000 iirc needed a reboot for every single update of the OS, and even for applications like IIS or Exchange.
Compared to NT4 it was probably very stable, since I remember telling most clients to just shut it down Friday evening and boot it Monday morning cause the pre-SP4 NT4 could not stay up more than three weeks.
Compare that to AS/400, where we pushed updates all over the country, without warning clients, to system running in hospitals, and there never was even the slightest problem. It sounds irresponsible to do that today, but those updates just worked, all the time and all applications continued to work.
This just means security updates were never installed.
(Or you claim that all those operating systems never had kernel-level security issues which seems doubtful...)
Most were only locally connected (for example OS/2 had a Token Ring in one building). The WAN connection (for AS/400) was trusted.
That is easier to achieve when your operating system only runs on your own proprietary hardware. (No mess of millions of drivers to write for one).
It worked well for years without any sysadmin touching it.
Well my mom was trained to be the "sys admin", which meant rotating backup tapes.
Wrt Win95 & it's kind - all processes in that family essentially run in a single address space, and data "isolation" were "achieved" only through obscurity. If you knew some magic constants that were easily obtainable from disassembly, you could do anything there. So no wonder it was as bad as the worst program you've installed..
The tone and content of this document is shockingly candid and frank. I think it did a ton to make Windows Server a better product. I have a lot of respect for the people at MSFT who reviewed the company's own product in such a critical light.
The 90s were the dark ages of cloud computing. It was the age of system administrator, desktop apps, Usenet, and the start of the internet as a public service. At the time concepts such as infrastructure as code, cloud, and continuous deployment, were unheard of.
AWS, which today we take for granted, was launched on 2002, and back then it started as a way to monetize Amazon's existing shared IT platform.
Of course migrating anything back then was a world of pain, specially when it's servers running on different OSes. It's like the rewrite from hell, that can even cover the OS layer. Of course it takes years.
There existed different names and solutions for things like cloud. I worked with Grid Engine in 2000 after Sun acquired Gridware, but that project started in 1993. By 2000 we were experimenting with running Star Office on the grid and serving UI to thin clients (kind of what Google Docs or Office 365 do now, but on completely different stack).
Some more game-oriented technologies of course have helped in the years since though.
Edit: AWS -> Azure :)
Did you mean Azure?
Can I have my Mojang account back?
I literally cannot log in to it after it was forcefully migrated to Microsoft. Microsoft doesn't recognize my computer as not-a-bot. Something to do with being Linux, I imagine.
Or, can I get a refund?
In the end I gave up and he plays it on an old Ubuntu laptop instead.
I think I finally managed to make it work in a new/clean Chromium browser profile.
The first time it was a throwaway account only needed for one day so I abandoned it. The second time I had to give them my phone number. Very scummy.
This was during the "Microsoft <3 Linux" campaign, and I think I cited that and then told em Minecraft would not be able to move forward with Xbox account migration until they stopped such idiocy.
Since I was the dev tasked with migrating Mojang accounts to Xbox accounts, I felt I had at least SOME credibility to my claim that it was blocking me.
But honestly, modifying my user agent was easy, it just pissed me off.
They did fix that the same day tho, so I guess the believed me!
Reminder to whomever is reading: if you bought the game during alpha, you have the right to all future Minecraft games and a premium account forever. Microsoft barely tried to uphold this by giving a free Bedrock license to alpha buyers for a limited time several years ago. I suppose you'd have to sue them now if they break it, and the judge will wonder why you bothered to bring a $20 dispute to court.
Again, this is a while ago: I remember when I started we were just starting to replace the old yggdrasil servers with the new micronaut based system which I think is still in use today?
I still remember that application fondly as the best architectured piece of software I've ever worked on. I hope all is well!
Those services and any written since then have all been .net based. Still lots of older services hanging around though :)
For the most part it was just fine, until we started using CosmosDB (then called DocumentDB).
DocumentDB, in its first incarnation, was utterly terrible. The pricing was extremely hard to predict, so we would end up with ridiculous bills at the end of the week, the provided .NET SDK for it was buggy and horrible, but the very worst part was the WebUI appeared to be directly tied to your particular instance of CosmosDB.
Why is this bad? Because if you under-provisioned stuff for your database, it might start going slow, and it would actually lag the web interface you would use to increase the resources! We got into situations where we had to turn off the entire application, just to bump up the resources for Cosmos. It felt like it was a complete amateur hour from Microsoft.
My understanding is that Cosmos has gotten a lot better, but man that left a sour taste in my mouth. If I end up getting some free credits or something, maybe I'll give Azure another go but I would definitely not recommend it right now.
What the ever loving heck... seriously?! Why wouldn't this be a control plane API that reconfigures a data plane?!
I think they did fix it eventually, because I know that multiple people on my team complained to Microsoft about it. Very short-sighted decision on their end, but to be fair it was a brand new product.
What's so surprising about that? If your CRUD operations take longer to do, and you're doing those to drive a GUI, of course the GUI will lumber along.
The reverse also applies; by separating them you can have issues with your control plane but not have the database go down.
A couple of years ago I stumbled upon a Azure project which started off using the old timey Cosmos DB. Looking at the repository history from those days, I saw a bunch of Entity Framework configurations and navigations and arcane wizardry that would take an engineer months to really figure out.
Then there was an update to CosmosSDK, and all that EF noise was replaced by single CRUD operations that took the unserialized object, id and partition key as input. That's it.
Worlds of difference, and a simple key-value query takes ~10ms to do.
Yes, it's worlds of difference.
Unless that query goes over the Internet to another continent, that's a really long time isn't it?
It is an absolute eternity though. A KV lookup is fractions of a microsecond in a managed language like c#. A http request is in the microseconds range on localhost and a smidge more on a performant local network. A poorly behaved local network (busy wifi on an ISP router) is 1-2ms, and I can do a round trip to my nearest AWS region in 10-15ms from my home network.
It’s an absolute eternity, and when thinking about this stuff and scaling, think “is it worth slowing down the normal case by 1000x to introduce an external service”
ConcurrentDictionary<K, V> read latency is going to be around 7-15ns for the data hot in the cache, scaling with key length if it is string-based, and anywhere between 75ns and 150ns for reading that out of RAM. Alternate implementations like NonBlocking.ConcurrentDictionary can reach down to 3.5-5ns for integer-based keys given all data is in L1 and the branch predictor is fully primed, on modern x86_64 cores.
Redis is cool but people use it too willy nilly. It's expensive to run, and you have to pay a cost of (de)serialization, and network latency. Sometimes that's necessary, especially if you need to share stuff across multiple nodes and you're afraid of hammering the database too much, but a lot of the time, I'd say even most of the time, you can get away with just a big ol' global ConcurrentDictionary (or whatever the equivalent is in your language of choice).
Basically any thread-safe dictionary is fine, because no matter what kind of locking strategy they're using, it will almost certainly be faster than anything that has to hit the network, by several orders of magnitude, and you have one less thing to manage in your application. You can figure out the best way to invalidate old entries, or find a library to do it (e.g. something like a Guava cache), and you'll likely much better performance than arbitrarily farming to Redis.
That's fine if all you're doing is caching limited data independently on each node and you have no requirement for consistency or even durability.
Most of the usecases you stumble upon do not fit that scenario. That's why "a lot of the time" suggesting a dictionary is plain wrong.
You've assumed you have multiple nodes here. Our point is that you can scale _significantly_ more with a ConcurrentDictionary (or equivalent) than you think - far more than a single node backed by redis will go. You have a point about durability, but in my experience, most apps that lean on Redis don't actually solve that problem. If the app goes down mid request, it will leave broken state behind. The data itself in redis might be valid, but the app can send partial data. The solution to this is the same as handling durability in a non-redis case too.
Also, runnign Redis on the same machine for persistance is an option for durability.
> Most of the usecases you stumble upon do not fit that scenario. That's why "a lot of the time" suggesting a dictionary is plain wrong.
I disagree. Most of the usecases _do_ fit this scenario, and Redis is (over) used as a tool to allow for horizontal scaling caused by an unwillingness to scale vertically. At the level of Meta/X yes you absolutely need it. But for a web app with about 10k concurrent users, you can handle it on a single node.
I think you need to revise your assumptions. Even optimized DynamoDB setups shows average response times in the 15ms range and P99s at the 20ms range.
One of the mistakes you seem to be making is assuming that to do a query on a globally distributed database with multiple partitions, you'd only need to hit one single box and no data exchange would take place. That bears no resemblance with reality.
If you're hosting a service in a cloud provider and you implement your services so that your call to the cloud provider's database goes over the internet to another continent, you have serious problems but none of them are caused by the cloud provider.
Also, CosmosDB is globally distributed.
I don't think anyone making that sort of claim knows anything about cloud services. A single roundtrip of a no-op request within the same data center takes 0.5ms. Add querying across multiple partitions and data seeking, and you don't find cloud providers doing better.
To frame how oblivious that claim is, DynamoDB is lauded for it's response times being sub-20ms.
https://stackoverflow.com/questions/34552625/how-to-get-sub-...
It was interesting seeing the biweekly status updates, they basically all started with “This is how Jet.com broke Azure core services this week”.
As much as it sucks, this was a deliberate strategy all the way from Satya - every employee knew Azure was a joke, but the only want to actually fix shit was to get internet scale customers to break it daily and weekly.
I don't get it. There's lots of distributed systems theory that could provide a more robust, analytical approach to a scalable architecture. If a system is regularly breaking like this, it sounds like it should be a "back to the drawing board" moment.
The only thing that seems to get the fiefdoms to work on an issue is if a big enough customer calls enough of an attention to the issue so that no leader/fiefdom wants to be blamed in such a "high profile" issue that everyone got their work in gear for. This leads to organizational dysfunction since now everyone understands that there's no point in trying to fix anything until there is enough attention on it, and that is mainly achieved when a big customer complains.
I've even had projects that took months of engineering work that would resolve many user complaints about a lack of really basic functionality, and different elements of leadership would block the launch. I can only guess that some fiefdom did not like change and would not state their objection publicly, so I only got stonewalled, and I never got an explanation to throw away months of work. I am as certain as I can be that if a big customer complained loudly enough about the problem that such basic functionality was missing, the organization would demand that the fix be launched.
For some values of "better", I guess. Performance is still terrible, their data visualization/inspection tools are shameful, their SQL dialect is finicky and has no error reporting beyond "something is wrong with the input value", and their official Python SDK has a race condition bug that can silently clear out your documents when under heavy load.
I used to work at a Cosmos-heavy house and I would utter "fucking Cosmos" around 15 times a day.
I kicked up a stink and we migrated everything to AWS in under a week.
We raised support tickets, which were mostly closed to the tune of “seems fine to us”. They seemed to think 20-30ms for a basic Postgres query, which for the same schema, data and hardware was <7ms on RDS.
Being tied to AWS and being unable to shake off a huge bill is not a trait of its competitors. It's a trait of AWS, and stresses the importance of not picking them as your cloud provider.
Also, I think it's unbelievable that a monthly 6-figure invoice charged to a company already with cloud engineers in their payroll is not justification enough to shift their workload elsewhere.
...in the US.
In Europe, even the likes of Amazon pays it's SDEs 70k/year. In Sweden, for example, Microsoft pays it's SDEs south of 800k SEK, which is about 70k dollars/year.
Low 6 figures is an entire team of Microsoft SDEs working full time for a year.
Each case is a special case how everything gets configured, but between Azure, GCP, AWS and IBM clouds, the ones with smoother experience on my case, have been on Azure, based on Java and .NET technologies.
And we also have our share of support tickets across all of them.
Now Azure back in its early days, 2010 - 2016 was kind of ruff, maybe this is the timeframe you're referring to?
Azure is a mess designed by smart people with no time and little budget. Azure flat out lies about the AZs they have by claiming two halved of one data center is two AZs
it was around 2007 (I'm not that old!)
> Founded in 1996 by Sabeer Bhatia and Jack Smith as Hotmail, it was acquired by Microsoft in 1997 for an estimated $400 million
Wrong decade, they really did acquire Hotmail in the 90's.
Do you think it was the migration itself or the services on Azure?
Having worked with all three, there's certainly things that suck about all of them but I've found aws "most reliable" but also seems to have a large amount of disparate services needed to do things that were simpler on Azure.
GCP was pretty meh, but depends on what services you used.
Azure is a good choice for .net and sql server (azure sql or whatever it is now) but in but sure a service built for aws is going to "just work" on Azure (or vice versa).
But the documentation and every other reference to it will retain the old name.
"But it's easier!" ... yeah, we'll see...
Unless you need to switch providers, at which point it may take more time to adjust for differences in how those managed services operate.
Managed services are absolutely not the main reason for moving to the cloud. Companies do it for the flexibility that comes with renting the real estate/energy/hardware instead of owning it.
For example, we'll used Managed Postgres, but not Azure or AWS's home grown databases.
Makes migrating much easier.
I would love for you to briefly describe how and where this can be done. I wasted a significant amount of time searching for this exact capability for Azure SQL Database and only ran into dead-ends.
Not sure how to do the same for Azure Postgresql databases, but looks like standard pg_dump and pg_restore are supported.
[0]: https://learn.microsoft.com/en-us/azure/azure-sql/database/d...
i.e. in 2025 managed Kubernetes is not _that_ different between providers
People use things like RDS and EKS/GKE to avoid all the administrative overhead that comes with running these things in prod. The database or its underlying hardware has a problem at 1am? It’s Amazon engineers getting paged, not you (hopefully… assuming the fault hasnt materialised to operational impact yet)
It's different for fully featured SaaS. It's a matter of the abstracted complexity vs. interface complexity ratio that is so common for everything you do in software.
Before 2014 or so, AWS would periodically reduce prices on major services passing on falling technology costs.
Azure didn't like that, so they aligned their prices to AWS's, matching immediately the same discounts on the same service.
This is a form a predatory pricing, because the goal is to kill the incentive for competitors to reduce prices, by denying them market share gains when they do.
we had one customer that needed IPsec tunel to vpc where production servers were living, we didn't want to maintain such setup just for single customer so we check Aws offerings. and look at that they have managed IPsec solution, great.
until client called that tunel is down and solution wa that they need to restart it on their end to resume connection. why? you can enable some logging to S3 but according to them everything should work. what we should do next?
but even if you stick to just ec2 thing can go weird. our recent incident: ec2 instance stopped responding but ASG didn't replace it, any action on it throwers error that instance is not running but it was in running state.
I've also waffled several times on the Azure FaaS offering. I am now firmly and irrevocably at "Don't use it. Run away. Quickly.". The experience around Azure Functions is just too weird to get comfortable with. Getting at logs or other binary artifacts is a gigantic pain in the ass. Running a self-contained .NET build on a blank windows/linux VM is so easy it doesn't make sense to get in bed with all this extra complexity.
Also, things that break automation, like calling back to say your sql server is up and running when in fact it’s not ready for another 20 minutes. I am half sure the terraform time_sleep was written specifically to counter azure problems.
App Services takes some getting used to, but it's a locked down Win Server/IIS container with built in FTPS, self-healing healthcheck endpoints, deployment by pointing to a repository, auto-scaling options, and a 99.95 SLA.
A few years back, it was a bit of a dog performance-wise, but the modern CPUs have been no problem for a 2+ vCPU, Premium level SKU. Pricier than a VM, but dealing with security and updates for a webserver VM is a ton of work.
But the Windows containers have more features. I stuck with them for quite a number of websites.
Significantly cheaper than a VM as you noted just based on maintenance that would otherwise be required.
Migrated off hosted Postgres because performance was tragedy - now their India-based expert led us to use different volumes type and after instance restart database didn’t start up because of I/O latency. Expert don’t want to meet for 3 straight days now, because he is busy. RCA (half pager, written probably by some LLM) says it’s not their fault, but charts says different story.
The only thing they crash GCP and AWS with is dashboard that loads everything so quickly... sad you can’t run e.g. making 2 similar network operations in parallel because they will fail or take 10x the time they would take when run one after another.
Run, don’t use.
In general I think it’s sad that most buy in to consuming these ”weird” services and that there’s jobs to be had as cloud architects and specialists. It feeds bad design and loose threads as partners have to be kept relevant.
This is my take on the whole enterprise IT field though!
At my little shop of 30 so developers, we inherited an Azure mess, built abstractions for the services we need in a more ”industry standard” way in our dev tooling, and moved to Hetzner after a couple of years.
A developer here knows no different, basically - our tooling deals with our workflows and service abstractions, and these shouldn’t change just because new provider.
1/10-th of the monthly bill, and money partly spent on building the best DX one can imagine.
Great trade-off, IMO!
Only two cases come to mind for using big cloud:
- really small scale: mvp style
- massive global distribution with elasticity requirements.
Two outliers looking at the vast majority of companies out there.
Even as someone that had minimal exposure to other clouds, I could easily see how Azure user experience lags due to the lack of proper care.
The amount of pages with a filter bar that will not work properly until you remember to click the load more should clearly be zero at this point, this is an objectively bad pattern that existed for years and should be "easy" to fix. But the issue will probably never be prioritized.
The fact is that unless tackling those issues are part of the organization core values or that they are clearly hitting their revenue stream, they won't be fixed. Publicity and visibility of those issues will always be crucial for the community of users.
I love me some AWS, but my god every time I have to dive into an unfamiliar environment and try and reverse engineer how everything connects - I need a drink afterwards.
I wouldn't have believed it, but while testing out a server for a business, they deleted my account, and didn't reply when I emailed support about it.
In practise, even the good ideas are implemented poorly.
A great thing about mice is that they are fungible and don't change without the user's consent, unlike software, so you can keep buying and using the same mouse forever.
The average mouse of the time was blocky and uncomfortable.
"Here...don't miss this settings page! Seriously! Look!"
AWS UIs are generally snappy and smartly designed individually, but they are horrendously organized at the general level. AWS is built as if you are exploring a relational DB containing your resources, instead of a deployment tree.
Your VM doesn't have a NIC in AWS, it has a foreign key to your entire VPC's NIC table, which lives in the VPC service, not the EC2 service. And then your NIC doesn't have an associated subnet, it has a foreign key to the subnet. And then when you get to the subnet table, you look up the routing tables table, and finally in the routing tables table, you'll find the settings for the routing table. This all works through following links, but the constant context switching and tabs upon tabs that AWS UI requires are extremely unpleasant for me at least to use. I'll take Azure's sluggish UI that organizes all of this in one-two pages instead of four any day.
In all seriousness, even in the face of IAC, the one thing Azure can do that AWS can't [at the time this happened to me], is have a global view of everything that's running right now and costing me money. It was years back, it was a $5 bill, but the principle of it had me livid. I did my best to tear down everything after my evaluation, yet something was squirreled away costing money.
So yeah, absolutely, sluggish UI all the way (I also find the Amazon storefront profoundly ugly and disorganized).
If you click an instance and go to its networking tab you get a list of ENI IDs that are clickable links to the resource, same for vpc and subnet. If you click subnet you can just click the route table tab, so if you're on an instances networking tab the route table is 2 clicks away.
But rather than doing this you could use reachability analyzer that allows you to check routing tables and security groups for a source and destination IP/resource and port on same or different VPCs connected with peering or TGW and it will tell you if you're missing routes or SG rules in either direction. I created a slackbot that allowed our devs to input src/dst IP/domain and port an that used this API to do the check for them, saved a lot of time troubleshooting.
I had an absolutely horrendous time working in Azure a few years ago (as a network engineer), we did have quite a complex setup with custom route tables and Azure Firewall though and VPN connectivity between Azure and AWS, but stuff like their VPN gateway taking 40+ minutes to change instance size on, wtf? I've filed 2-3 bugs to AWS in the almost 10 years I've worked with it, all for newly created APIs/services, they were all fixed within a week or two. I filed 8+ bugs to Azure in the first month using them, none of them were fixed as they had workarounds instead. And their documentation is absolutely useless, I could never trust that I understood what I read correctly, I always had to verify that it worked that way by testing it.
Azure's APIs are atrociously slow. Azure's UI design is pretty nice. There's not much the UI designers can do about their API colleagues.
I don't understand how dumb you must be to design a web site that way. It's like a brewery that sells their beer in plastic shopping bags and thinks that's good.
and
> The one time I tried out azure for a few days the portal was absolutely painful.
conflict with each other. Here's what you sound like:
"I don't believe you because I have very little experience in something and it doesn't comport with that."
* WHEN the resource was created * nor WHO created the resource
IMHO this is unforgiveable, but on second though, it is probably intentional rather than any sort of oversight.
I am quite surprised to hear someone say this. MS documentation has a horrible reputation that is well-deserved and and in my direct experience the docs for Azure are much like any other MS docs (incomplete, out-of-date and poorly organized -- usually all three at once)
[1]: https://blog.cloudflare.com/container-platform-preview/
You need at the very least containers and persistent volumes to be interesting to me at least.
Rust is still a wasm target. Not everything easily works. It doesn’t have all the Cloudflare sdk features js/ts has either.
And yes, you can test your local Rust code. It works nicely on your machine, but breaks with a really nasty error on their platform.
The target is `wasm32-unknown-unknown`, which allows you to use `fetch` as your only source of IO. Ok, their workers has a hacky socket implementation nowadays. Non-standard of course. And most of the ecosystem won't work really without forking everything and fixing the bugs by yourself.
We pivoted to a native Rust project. We still have one worker running in Cloudflare. We isolated that code from the workspace so that renovate updates will not touch it. You know, a random version upgrade might break the service...
The reason aws is dominant is because it’s the default and a known quantity. But developers and cost conscious organizations will look at alternatives. Not saying it would be easy but the prize would be huge. Plus AWS seems in chaos with a lack of sensible leadership.
I'm guessing it's just because they're so extremely all-in on every single of their products being on the edge, and you can't make that work with VMs without becoming far more expensive than any of their customers would pay.
Or maybe I'm clueless. But it sure looks that way from the outside.
I am part of a team building an automation tool for cloud provider creation of Projects/Accounts/Subscriptions (depending on provider). Our primary provider is GCP, and implementing that was fairly easy. Some gotchas, but easily surmountable.
Now we have gone multicloud to Azure and we need to add support (We were historically on AWS but moved 95% off, we still have some teams on there but we rarely build tools for it outside of Terraform modules). And the Azure API, MS Graph API and Go SDKs for bith are the biggest piles of trash I have ever worked with. Everything is a pointer, even string literals need to be made pointers, but sometimes they aren't....
Documentation is in accurate. Some apis take just the ID, others take a full path. Some of it is documented, many apis have the wrong one documented.
None of the APIs return related resources IDs, you have to search for all of it. So many name based searches. I had to add a caching layer for IDs during creation so I didn't have to lookup the same resource over and over (we use a state machine for creation and it can be resumed midway and other fun things, so we need a lot of checks and resume based code.
Overall it is the worst designed and implemented cloud provider. I would never recommend or choose it if given the power.
The world is a big place
However, all the comments about the Portal are baffling to me. AWS portal is just all over the place, I feel like people are expecting AWS awfulness and when portal wants to be consistent, it's breaking people brains.
Oh yea, Day 313 with Public IP to put into DNS. Alias record that you noob. :P
Even doing classwork involving AWS was an exercise in frustration. I couldn’t actually believe the sort of button trails and on-hover menus I was told to use to access various functions.
I don’t have the experience to evaluate the technical functionality, or whether AWS’s interface is better for experienced users, but I can definitely say it was far less approachable as a novice.
Chiefly among them is their famously bad support. Just google it- I've never spoken with anyone or seen a single written word saying that Azure support is even decent.
It's a race to the bottom platform, and I'm starting to get to the point where I want to start selling AWS.
I cannot count the number of times I've found a Microsoft support forum question that's exactly my problem too and the official tagged Microsoft support person fully misunderstands the question and then doesn't even properly answer their own misunderstood question
My number one concern with Azure is availability of resources. Working within US regions, we've had to shift regions during production rollout because one or more of the resources we needed -- a current gen Azure SQL database or App Service Plan -- were simply not available. Rolling out an inexpensive VM (think equivalent of a t3/t4g.micro) is always a ride too, between unavailable SKUs or excessive quota gatekeeping.
Spending gotchas exist on any cloud, but we also know someone who got caught off guard in a completely new way recently. In late-December, the team needed to automate a database event once per day on an Azure SQL instance. Scheduled jobs aren't natively available inside Azure SQL, and so they reached for an elastic job agent. Everything went smoothly until someone dug in to a price increase on the January bill and asked why Sentinel had jumped from under $200 to over $3,000.
A colleague and I helped them dig in and quickly discovered that the controller for the elastic job agent is running dozens of batches per second in order to schedule that one job per day. With default security audit settings on Sentinel to meet compliance obligations, this generates over 600GB of BATCH_COMPLETE log messages per month at a cost of $5/GB for ingest!
Not sure if this was an Australia/Oceania limitation - or just an ongoing product limitation.
My requirements weren’t complex. I needed to manage my domains (not AD), spin up virtual machines, and associate the two.
I also found the UI, overall; tedious. Finding the right offering under their ambiguously named services was difficult. And this comes from an AWS user.
I wanted to like Azure, but for very least reasons above; it’s not the product for me.
Ive been using it for 3 years now to register names for my job as Azure bill just gets paid. Anything outside Azure is a bureaucratic mess.
I work at a startup that runs on Azure, and we're only here because of Microsoft's monopolistic behavior. We switched because Microsoft gives Office 365 discounts to our customers as long as all the the SaaS services they use are hosted on Azure, and so our customers demanded we use Azure. Part of the monopoly playbook: "using a monopoly in one area to create a monopoly in another".
I used to work at GCP, and I thought it was almost shameful that we were in 3rd place behind Azure. Now it just makes me mad (especially since I had to migrate our startup from GCP to Azure).
Also, I suspect part of the reason people are hesitant to use GCP because Google is perceived as a company that will gladly kill products off on a whim. Not great for something mission-critical.
Case in point at work: we need to set up Azure infrastructure per-customer. Hitting the Azure RM endpoint from outside the Azure network is not reliable; the API endpoint's DNS record points to one of two IP addresses in westus, and when the DNS record flips (presumably for blue/green deployments) the no-longer-referenced IP address immediately aborts the connection. The official Azure Terraform provider throws an error when this happens and it usually results in Terraform state losing track of something that it already created. Azure support just says "well all we see is 200 OK from our side".
The "solution" is to run the Terraform workload from within Azure. The SLA is only really guaranteed if you're connecting to the Azure RM API from within Azure. Cue the insanity.
Whereas with other providers, it can also be about money (because they offer a big discount or because the migrating partner is cheaper), but it can also be about wanting one service/feature that only X is providing; and once you're in, people tend to prefer to put everything there, instead of doing poly-cloud.
In my region, Azure salesmen are very active, providing huge discounts, so Azure is the most popular amongst big companies. Meanwhile smaller ones will go on AWS because its easier to find information and (actual) knowledgeable people.
I used to work in a company using AWS : everything was managed through Terraform and we were as cloud agnostic as possible (mostly containers). Then we were acquired by a bigger company with a Azure deal, so they told us to migrate from AWS to Azure. They provided us with their own experts to help us, but six months in, we were still unable to have anything remotely viable for UAT. The experts were starting to acknowledge that even with they years of experience, they still weren't convinced with this whole Azure stuff so they actually relied heavily on a legacy on-prem DC. That's when I left. Last time I heard about old coworkers, the product was still running perfectly fine on AWS, while there was still a team working on the migration. It's been more than two years now.
And I had other bad experiences with Azure. I know that cloud providers are not fun if you don't start with two weeks of training, so I try to stay open-minded, but no matter how many Azure experts I talked to, I never found one who was actually confident in using it.
- we use mainly GCP
- we do not want to use AWS because of random political issue (absurd, in my opinion but whatever)
- we are raped by GCP and would like an alternative to help keeping the price "acceptable"
Welcome azure ..But what about a solo founder running their Web site on a "full tower" computer they plugged together themselves?
So, why use a cloud server farm with its expense and complexity?
Or, get a mother board, a processor with 16 cores and a 4+ GHz clock, 128 GB of main memory, some rotating and/or solid state disks for a total of 20 TB or so, some external disks for backup, a recent copy of Windows Server, applications software from .NET, and a 1 Gbps Internet connection?? The computer -- tower case, power supply, motherboard, processor, disk, solid state disks, and Windows Server -- costs ~$3000?
So, a 1 Gbps Internet connection, ~$100 a month, would have capacity of, say, 100 MBps. If sending a Web page with 200 KB, then the peak capacity would be
100 MBps / 200 KB = 500 pages/second.
Then 500 pages a second with 5 ads per
page with revenue of, say, $2 per thousand
ads sent (CPM), that would be 500 * 5 * 2 / 1000 = $5/second
at peak capacity or maybe an average of
half that for $2.50/second.Then at 16 hours a day that would be revenue of
2.50 * 60 * 16 * 30 = $72,000/month
For the electric power, at 200 W and
$0.10/KWh, that would be 200 * 24 * 30 * 0.10 / 1000 = $14.40/month
How many users?With peak capacity of 500 pages a second and average of half that, 250 pages a second, for 16 hours a day, that would be
250 * 60 * 16 * 30 = 7,200,000
pages a month. If on average send 5 pages
per user, that would be 7,200,000 / 5 = 1,440,000
users per month.If users come on average 2 times a week that would be 8 times a month or
1,440,000 / 8 = 180,000 users
from one tower case and some Web page
software.So, why use a cloud server farm with its expense and complexity?
I advice CEOs of SMEs on this, and I can tell you that the main concern they have is availability of people to build and maintain the systems. Because cloud / k8s is more popular these days, that's what they go for. If we could reliably find smart system operators that will happily maintain a couple of racks of servers for years, it would be a more viable option.
By the time you spec out a real server (redundant power, higher quality components than Newegg stuff), rent some space in a rack, pay for bandwidth (50Mbps is going to be about it without paying a premium), you're going to be looking at $5000 + $300/m. All that effort, whereas you could spin up something in the cloud for a bit more per month.
This does flip quickly, however. Once you get into the high 5-figure monthly spend, running your own hardware makes sense again. DHH's blog posts on 'Leaving the Cloud' are a great read.
I looked at Hetzner's $150/m for dedicated Xeon server with 1Gbps.
As someone who ran many IIS boxes since IIS 4, I greatly prefer Azure websites over having to worry about all the “other stuff” that comes with onprem. Yes, the cloud needs an RP, WAF, etc, but they’re always HA and simple services, not another box to maintain.
Thanks! Yup, it will be "a hobby project" until, if ever, it gets some users and revenue. Then, have several servers, some load balancing and redundancy, uninterruptible power with a generator outdoors on a concrete slab, contact Cloudflair or some such and have them do what is needed, etc.
Just looked up SSL, reverse proxies, firewalls, etc. Okay.
I mean, your intuition is correct that "cloud" is mostly just a bunch of boring, standard computers running boring standard software and there isn't anything they do that you can't.
But at the same time, boring standard software is (by definition!) commoditized and if you're spinning up some new and interesting thing, it's only going to be differentiated by the parts that are not boring and standard. So put the boring standard stuff on a credit card and do the interesting stuff instead.
If you're virtualized on your host, 2x HAProxy on top of OpenBSD utilizing carp. It's great fun to set up and run -- and once you have it running, it's stupid stable. Very little maintenance required.
You can do this with Linux & keepalived, as well.
If you need an active component near your customers for low latency responses, the cloud makes it very cheap to deploy tiny VMs or small containers all over the place. It’s trivial to template this out, scale up and down for follow-the-sun or to account for local traffic spikes.
If you need it, you need it, and nothing else meets this need except perhaps some CDNs with “edge compute” capabilities — however those are quite limited.
Because the constant attacks you’ll be under from china/russia that start 5 minutes after your servers are live.
Because as a solo founder you should be hammering on new features not putting out sysadmin fires.
I am happy to give AWS $200 a month to look after all that crap.
The cloud part (VMs, k8s etc.) is something that I touch if I am being forced to. Even creating a VM is way more complicated than it should be.
Would you tell us how the monoculture you speak of (which, for some reason, uses AWS, K8S, Postgres, Mongo, etc.) hurt you?
Granted, the price of services is higher than on other platforms, but you would be mistaken to thing that's the price you are paying at scale, on a multi-year deals with reservations.
If I were to start a B2B startup I'll definitely go with Azure. For B2C or e-commerce, I'll probably look at others
Working on enterprise or higher level Microsoft is a way to get Grey hair fast. All the way back to Server 2003, we had the infuriating inconsistency of group policy, roaming profiles, DFS drives. Everything is full of errors, you will have a larger IT team as a result to deal with the headaches.
After using Google Workspace for IT, and AWS for infra, I always tell people to stay far away from Microsoft.
Even now, I have a friend who can't honestly deploy intune, because of inconsistencies of the "type" of enrollment, and it being able to execute a winget script as a result. Despite both machines enrolled in intune, the one that was enrolled during OOBE can run the scripts, but the machine enrolled in the OS cannot. Microsoft support has had that ticket for weeks.
I was tasked with a project where it made sense to use Azure Durable Functions. Again... BUGS ... I reported a couple of them and even went and spoke with the product team about those. One BUG was due to a misunderstanding of how the framework works (in my defense, the documentation was very unclear) and the rest of the bugs are still not fixed almost 2 years later.
I decided to fail the project and restart with a different approach and framework.
But I tried Azure for my most recent startup because I was offended by AWS, and GCP did not have enough adoption among my customers, and Azure worked - fine.
What do you really need out of a cloud?
I want them to rent me VMs, for them to not go down, and to make it easy to do standard stuff like an object store, run containers, run databases, etc.
Azure was as good or better than AWS
Sometimes in business the deal falls through
I was on the receiving end and didn’t appreciate it.
I will just say that Azure seems to want to do shit different for the sake of being different.
It is really annoying to write infrastructure as code for aws/gcp then go to do the same for azure and realize how dumb some of their stuff is.
Just my personal experience.
If you use standard tools you don’t have this problem.
Containers running on VMs is standard.
A mesh of microservices that depend on cloud queues and managed services is not.
One argument against standard containers is saving dev time. You can still save dev time by using standard open source software. How many different ways are there to implement a queue or a load balancer?
If you really need access to some proprietary technology then by all means use the cloud that offers it. Eg if your customer demands GPT4.5, then go with Azure.
But if you need something standard, don’t get caught in the trap.
Also, Azure APIs are incredibly slow.
Per the parent:
>That’s how clouds try to lock you in, by making you use a custom tool that is different for the sake of being different.
> If you use standard tools you don’t have this problem.
The overall UX of AWS is absolutely crazy. It's easy to "lose" a resource... in there... somewhere... in one of the many portals, in some region, costing you money! Meanwhile, Azure shows you a single pane of glass across all resource types in all regions. It's also fairly trivial to query across all subscriptions (equivalent to AWS accounts).
Similarly, AWS insists on peppering their UI with random-looking internal identifiers. These are meaningless and not sortable in any useful way.
Azure in comparison allows users to specify grouping by "english" resource group names and then resources within them also have user-specified names. The only random identifiers are the Subscription GUIDs, but even those have user-assignable display name aliases.
The unified Portal and scripting experience of Azure Resource Manager is a true "second generation" public cloud, and is much closer to the experience of using Kubernetes, which is also a "second gen" system developed out of Borg. E.g.: In K8s you also get a single-pane-of glass, human-named namespaces (=resource groups), human-named workloads, etc...
like region names...
That's Microsoft's MO in a nutshell in my experience, and I say this as a recent(~5yrs ago) convert to Linux who built a career on Windows endpoints, servers, ADDS, Exchange, SCCM, you name it. It's how they achieve lock-in to their ecosystem, and it's incredibly frustrating to see how they've just layered that method of operation over and over again, decade after decade, rather than fix anything.
I need a cloud to be reliable and secure. I've used Azure extensively and it's neither. I'll take GCP or AWS over Azure any day.
Happened twice with my kubernetes deportment, first something with node groups made them incompatible and had to recreate the cluster from scratch, then one of their scripts to rotate key access to volumes (that one has to run manually, go figure) stopped working and caused my volumes to detach from pods, and has to recreate the cluster again, and I just give up.
I was super happy as well, the first two years. By years six I was fully migrated out.
Azure ostensibly fails outright on points 1 and 3, and limps by on 2.
The products are confusing mess, there’s way too many ways to auth things, docs are garbage, tried multiple times to manage stuff via Terraform which broke far too much to be excusable, to say nothing of the dumpster fire that is their UX.
I’m sure some people have either beaten it into submission, or have stockholmed themselves into putting it up with it.
<screenshot of a python error>
This is a scenario that is all too common with az.
I have the privilege of working on azure stack hub. To do IAM management you have to install a cli version that is years old at this point. On recent versions you get a trace error like this.
I plan to drop IPv4 and go nodns/IPv6(/64 or /96 prefix) for self hosting.
Give me 1 integrated service built with the same stack as Azure anyday. Builtin service connections, managed identities, etc.
It works, but as with all Microsoft products it's in desperate need of some love and polish
Boards works for basic Kanban projects, but if you want to dwell into scrum stuff like sprints, burndown charts, etc, it's very bad and cumbersome. You'll have to do a lot of stuff manually that Jira does automatically.
Wiki... It's not good. It's extremely slow, and lacks a ton of features that Confluence has.
Pipelines is godawful, and has been suffering severely from a migration from their "Classic" pipelines to YAML. The funny part is that if you go in depth into YAML pipelines, you'll notice there's a very large amount of things that aren't configurable by YAML. Also has a ton of bugs, many of which have been open for over 5 years. To make matters worse, it's currently in an identity crisis with Github Actions (which has more features and is continuously getting them over Pipelines).
I don't know what's the future of Azure DevOps, honestly I feel like they'll eventually shutter it and move everyone to Github Enterprise.
Never used it, never will.
Yeah you can die of thousands cuts but I don’t believe you don’t get the same on GCP or AWS.
Except for OCI. Holy shit they are the worst.
That said, I only use Azure for redundancy. Hosting one app I have there on anything but Azure would be pretty much infinitely cheaper, especially when my IO quota goes above ~1GB a month which with azure causes a need to instantly change to a 20$ plan for the month.
I've seen the 0 vote questions throughout many product and non-product specific subreddits. Bots or just reddit vote fuzzing?
It's an obtusely confusing interface with opaque pricing that you can just magically sign up for with innocent sounding names. It's like the dark pattern people from Intuit came on over for a house party and got drunk one night.
To create a new deployment (which is basically a PAYG model that doesn’t really require this limitation), you’d have to switch to the old UI to find the button to create one.
I agree with comments about cloud providers, most applications would be better hosted on VPS or services like Digital Ocean. But hey, software developers like to look smart by complicating things