Stop Doing Cloud
grski.pl
grski.pl
Right now I run a SaaS platform entirely in the cloud. I am fully aware I am paying more for a managed Postgres cluster versus spinning one up myself and I'm paying more for a managed docker service than I would if I just spun up some EC2's myself or went for a VPS.
But right now I'm also working on the project entirely solo so I have to be careful about what I commit my time to. Sure I could save some $$$ by managing my own infra but the time it would take to spin that up and maintain it will cut into the time I have to build new features for my app. I could hire a contractor to help me maintain that infra but then it goes from being the cheapest option to the most expensive option because I'm paying someone to help with it.
The cloud isn't cheaper than running your own infra, but it's sure as hell cheaper than running your own infra + paying someone to help you run it. Down the road when (hopefully) my company has more scale and we have dedicated folks for infrastructure, then I can start looking into lowering costs on the cloud side because those infra folks are gonna be working for me regardless of where I run the app, but for now when it's between doing it myself and hiring someone to help me, the value prop of a managed service is undeniable.
At the very small scale when trying to start something, do what you're comfortable with, because whatever you're not is likely to cost you time.
Above that, it's down to what your margins are like, and how large a part the hosting is, but it's rare for cloud to be the cheaper option. Certainly not at list prices - while you're small enough to pay list prices you're being screwed over massively.
> how much of a commodity the skill set to set up and manage non-cloud setups is.
I've seen so many teams failing thisSo many little teams actually managing non-cloud setups
"Commodity" ? Not here
The same problem of teams lacking ops skills exists in companies doing cloud stuff too - often because they've bought the fiction that cloud doesn't need devops.
Cloud has a role many places, but not in setups that are cost sensitive at scale.
There's also liability. If you pay someone for services rendered, you can throw liability on them. If you run stuff yourself, you're liable.
Data breach at <cloud> for your services? Hang them out to dry. Data breach on your own server? You're fucked.
If you think you're covered in any meaningful way with a standard contract, you most likely didn't read their terms.
If liability is a concern, you need insurance.
Two different things: 1. infrastructure (computers) and 2. running the infrastructure (running the computers)
Obviously, purchasing computers is more expensive that not purchasing them, i.e., using someone else's computers. The later is known, in marketing terms, as "cloud computing".
The remaining question is whether (a) paying for running own computers is less or more expensive than (b) paying for running someone else's computers, i.e., cloud computing. Or whether the costs are the same.
For some, the fundamental problem posed by "cloud computing" is not the cost. It is the fact that "the cloud" is someone else's computers. Handling data that does not belong to them. In some cases having customers share the same computers with each other.
AI does seem different in that regard, I think it is providing tremendous value to people. Just probably not in the random places where people are throwing it in where it doesn't actually fit.
[1] https://web.archive.org/web/20200418170753/https://twitter.c...
But we're entering leaner times. Many younger people haven't been through this before. Others may not remember. Investors are going to demand more responsible spending, and better balance books, sooner.
If you can start on a shoestring and hold on until you need to scale, you have an advantage in this market that you didn't have 5 years ago.
That and a lot of the cloud stuff comes with an inherent complexity. Architecting for cloud hosted K8 and various cloud services comes not only with a $$ cost but a software architecture complexity & labour overhead.
A single server can do a lot these days. Many many businesses think they need to scale out of the box a lot more than they do.
You can spin up a single server in the cloud as simply (or more so) than in a colo. Then you can add a load balancer and auto scaling in a few hrs (or less depending on experience) if you reach the point where you need to scale.
Had we not gone with k8s and just gone with a vm -- even without docker compose -- but just a monolith we'd have been 1000% further along than when I left.
Even doing everything cloud native means learning how to package your stuff for that provider and go through their docs, learn their idosyncracies, etc.
BUT!!!!
Let's say you hire a kid whose been tinkering with servers in their room since they were teens and know their way around a Linux distro and can package python and run a proxy etc etc
There's a lot of compelling reason to simplify the tech stack and cloud/k8s/etc makes sense only when the scale requires it.
---
That being said -- Hetzner is not the panacea its pricing might suggest. We run our self hosted gitlab on it and we've had 2-3 incidents in the last month that required them to either power cycle a machine or fix some networking thing etc etc. I kinda expected more. But it's hella cheap: where else will you find so much hardware for 200 a month.
If you're running a whole fleet of VMs and need to manage them the pain of doing that is quickly going to exceed the "pain" of using k8s. Do you need something like k8s ingress? Do you need the ability to move workloads if a machine goes down? certs? external-dns? Are you going to deploying tooling that's available off the shelf for k8s? (lessay Prometheus/Grafana?). VMs have overhead and likely waste a lot of resources.
I've seen this "monolith" and I've seen containerized/k8s. For anything that needs to scale and needs similar abstractions you get out of k8s you definitely should not try to reproduce this using VMs.
A lot of applications can fit on a single machine. Single machines have gotten so big: 256+ cores, terrabytes of ram, huge amounts of fast storage, 100+Gbps NICs.
Yeah, you should have 2-4 for continuity, but that doesn't need orchestration.
You really don't need to start thinking about horizontal scaling unless you're getting close to filling a single big machine. Especially since the longer it takes for you to scale to big machines, the bigger the machines will be when you get there.
If your application is very latency sensitive, then maybe you need to run machines in more places, and maybe you want something to orchestrate that...
I'm sure plenty of things fit this model. But many do not. There's going to be a point where you're taking more pain if you chose the wrong model.
Go can serve about 500M users per month on Xeon E3 v2 CPUs.
Why, except for HA do you need more?
Do you know how much 5000 (and more in reality) concurrent requests are, and I'm just talking HTTP1.1 requests?
And with HTTP/2 and 48 cores 96 threads and even 24c/48t commodity hardware that's multiples of 5k concurrent rps.
Of course if you insist on using low performance php, node, ruby, elixir that's your own fault.
"Scale" my ass. Just another FUD buzzword to sell expensive proprietary services.
> Go can serve about 500M users per month on Xeon E3 v2 CPUs. > Why, except for HA do you need more?
Because not everything is a simple CRUD website. The last place I worked had a website that appeared to only be moderately more complex than such a site, and had well under 500M users per month. But parts of the application involved complicated processing of many dozens of TBs per day of newly collected data to optimize the content being served (some stream processing, some batch). We had multiple clusters of 1,000+ machines coming and going for an hour or two at a time (often overlapping).
Most of my experience is with high performance software. C++/C and yeah lots of Go. My observations aren't related to performance. There's more to scale and providing a reliable service than performance. I am totally onboard with the idea that people often go horizontally when they don't need to performance-wise because they choose slow tools. This is not what I'm talking about here.
That said, there's probably a class of problems that are fine with just one VM. So if that's what you have, go for it (which is what I said in my original comment).
My app is a stateless monolith hosted with Fargate and RDS. I don't have to think about K8s, or look at it, or know how it works.
The pricing is relatively garbage, but it doesn't matter: now I don't have to think about devops as a concept.
I don't get how people are screwing up this cloud thing: it's as simple as you make it. Being pennywise and pound foolish trying to engineer complex solutions on the "cheaper" parts of cloud providers makes absolutely zero sense to me.
9 times out of 10 you hear X startup paying stupidly expensive bills it's because they tried to build out something that the cloud provider would have provided with a 50% markup but with an architecture you can't screw up.
Because we were dumb and thought k8s meant we had to go microservices.
> if the app can fit on a single VM …
Well of course. If and when it out grows the vm hopefully you’ve found product market fit and are bootstrapped and making money and have hired proper platform and dev teams to do it right and the single VM PoC is redone.
Or so that is what we are sold. In reality things happen piece-meal and more chaotically.
Yes u need to get to Pmf first but you can do that without a kubernetes architecture or deploying a highly available app in 5 different zones.
The reason many start with cloud is that most businesses quickly grow, or at least hope to quickly grow, to the point where they'll need it. Standardized cloud tooling helps onboard new developers more quickly. It drastically speeds up audits like SOC2 or ISO. Much easier scale-up/scale-down. etc etc
You also pretty quickly get to the point where your cloud costs are negligible compared to the rest of your business expenses. If you're employing 20 people and paying ~$200k/mo in salaries, is reducing your cloud bill of $2000 to $200 going to materially impact your business, especially given the opportunity cost of not doing other revenue-generating work?
And most don't. That's the point of the article.
Adding the requirement of separate private/public VPCs linked by a NAT gateway when we're comparing to a single physical machine with everything running in a single host running Docker is nonsense.
Second, he starts with sudo. Where's that command line coming from? Who installed the OS? Who maintains it? What about availability zones and fast failover?
You can do self hosting if you want, but it's not as easy or as cheap as he's making it out to be.
Cloud is still better for early stage startups, that's why everyone's using it.
I've noticed most of these anti-cloud articles omit this exact point consistently. So we'll replace "the cloud" with docker compose! Because apparently docker compose just spins up resources ex nihilo.
There's also often the contrived assumption that "the cloud" means a bunch of huge bills because you automatically launch every service available and run them constantly. Also you run the largest instances because reasons.
You're making OP's main point. He's arguing that (to use his numbers) 90-95% of startups don't actually need to worry about that. Hell, I would argue that a lot of services at established companies don't actually need to worry about it (and the ones that do, need to worry about it a lot). Tons of stuff inside of $megacorps is running on crummy "non-scalable" architectures because somebody sat down and did the "engineering" part of "software engineering" and went "y'know, I think all we're really ever need is..."
>Cloud is still better for early stage startups
My preference is cloud these days (especially if I'm just quick and dirty clicking buttons in console rather than CDK-ing things together). However, there are still lots of cases where sinning up a $10 droplet and running `docker compose` completely satisfies "the architecture" part of the business.
Yeah, those system architectures are the ones that end up in all the other rants, the ones about how companies never prioritize fixing technical debt and legacy code :)
But even if you neglect these issues you still need to manage your own machines, which is more time consuming than spinning up a new VM with preinstalled OS in EC2.
Here’s a plan for building a food cart for serving your first 1000 customers.
Yes, this might be a viable option for some use cases. In fact in Thailand and other SEA countries many restaurants never graduate past the food cart by choice.
But there are environments and some business plans that require a certain level of maturity, security posture from the get go.
At the same time I can deploy an application on AWS using IaC (AWS CDk) methods in minutes and use mostly serverless services without paying $200/mo.
Cloud or not, I think depends on your unique circumstances. Just becasue you can apt install everything once doesn't necessarily mean you should.
The bright shiny object / resume building in software decisions is in some ways fascinating
Handrolling: "wrote a few systemd services"
Cloud: put the EC2 in the VPC with an ENI and attached the EBS to connect to the RDS
Same security industry convinces you to upgrade every 15 seconds and then sells you solutions for when those upgrades fuck you over.
Additionally, not every update needs to be applied, you need to understand your threat model and only apply updates when they actually patch something that would affect you - this cuts down on the actual number of updates that you need.
But if you're going to manage your own servers, you need to do it right. This article is a perfect example of what I've seen in the past with not doing it right. It's inconsistent, it doesn't take backup seriously, and it doesn't take availability seriously.
Postgres is installed from their repo. Caddy is installed via & managed via ansible (instead of using something like caddy-docker-proxy). Actual backends are managed via docker compose. Backups are via a hacky script instead of something maintained & that another developer will be able to find docs on like borgmatic. Nothing is thought about for making sure the database state is consistent during backups (eg barman). Monitoring is an afterthought. Firewalling isn't even mentioned.
This setup is one hardware failure away from spending days or weeks trying to recover.
But;
A) You have to engineer your system to scale to zero, engineering time isn't free
B) The overhead cost of service (5x for Linux compute, 11x for Windows compute, 7x for managed DB -- based on my last reads) can easily swallow any savings, especially if you're not scaled to zero almost all the time.
It only takes about 8hrs of being "non-zero" and it would have been cheaper just to have a whole machine for a day somewhere that wasn't a hyperscaler cloud provider.
Someone that says this has never even reflected about availability requirements.
HA shouldn't mean "we deploy a few times a day which drops our uptime to three minutes so we have to spend a lot". One redundant server and a load balancer is hardly HA and increases costs by tens of dollars a month.
And frankly, a minute of downtime is the optimistic case. It all goes to hell when you have a failed software deploy that needs to be rolled back. Or when the application server has a bug that causes crashes and needs to restart every few hours.
A single-server setup is only useful for companies where the technology matters far less than the service being provided. If you can tolerate an uptime of only 2-3 nines, great, but that's absolutely not reasonable for companies where the technology is the service. It's not rocket science that paying customers are going to be pissed if your site is unreliable.
Top comment reflects on how "getting to market fast" is the most important thing, much more important than costs.
I agree.
Second comment reflects on reliability. Reliability is only a concern when you actually have a product that people want to use.
Cold hard truth:
If you are trying to get to market as fast as possible, most of the cloud abstractions are going to slow you down in most situations. It can be worth it, but your MVP should be as fast to develop as possible and there's nothing faster than writing code and running it.
Second, the reliability of a single machine is much more than you think. Until recently AWS didn't even cover their own reliability of a single instance, if they needed to boot you off a node and shut down your EC2 instance: you might get a bit of notice.
Ironically, you don't even need five-nines in most cases, and if you do, cloud only gives you a toolbox to get there, it's not going to get you there without serious human investment.
You will get 2-3 nines of availability with a bog standard PHP+MYSQL on a single machine.
You know how I know? Because we all did this 15 years ago and truthfully: servers and operating systems are significantly more robust now than they were then, so it's only gotten better for those who do things "wrong".
Cloud is a scam in multiple dimensions: - It's overpriced most of the time (I'm looking at you aws) - It wastes you're brain cycle about their stupid proprietary stuff instead of focusing on real problems - Makes you dependent on their bullshit - There's probably more to add to the list.
Overall, I'm a strong believer of "Refrain from doing cloud) :)
I opted of K3S which give me two benefits. 1. Predictable stable containers running and also restarting all the deployment automatically and 2. If I do need to migrate to bigger cluster / AWS cloud, I got everything ready in my K8S yml files.
- A lot of companies and startups can get by with a few modest sized VPSs for their applications
- Cloud providers and other infrastructure managed services can provide a lot value that justifies paying for them.
Edit: also why the obsession with dedicated hardware and so anti vCPU? For the vast majority of people this is an implementation detail that doesn't matter at all. A lot of that dedicated server is sitting idle.
Because it is the reason an application that appears fast and snappy on your mid-range laptop slows down to a crawl when deployed. It’s the difference between a database query taking minutes (and blocking everything else) vs seconds.
vCPUs and the lack of persistent direct-attach storage is the reason I personally don’t use the cloud for anything I build.
There is, however, one big disadvantage one should be aware of, and that is downtime. There will be downtime when you
* push a new version of your project
* install package updates for docker, postgres, ...
* install a new kernel and need to reboot
* need to upgrade the base OS because you can't stay on Ubuntu 16.04 for the next 10 years (this one can take hours)
Of course there are projects which can live with that. If your target audience is 9-5 office workers, some downtime at night might be acceptable. At least if you don't have customers in other timezones.
IMHO, the sweet spot would be to have two identical servers with automatic failover, but I did not find a good solution for that yet. Would be very interested if someone here has a solution for that and is willing to share.
I prefer Caddy myself, for it's simplicity. On load balancing, it can detect when an upstream server isn't working, and remove it from the list of upstream until it's working again. Sometimes you want control of this, for example for OS updates - Caddy has an admin API you can use to remove a server from upstreams and wait until all connections are drained, then you can perform the updates and add the server back again. Then do the same for other upstream servers.
(I haven't actually used the Caddy admin API in production yet, but plan to soon).
But the most undersold value of the cloud, and where I see so far unbeatable value, is automation (single, unified API) and the ecosystem of tools it allows to flourish from it.
Yes, Ansible could replace some Terraform, CloudFormation or ARM/Bicep but it doesn't manage out of the box a component's lifecycle. That is, knowing how and when should a resource be created, storing that last applied state somewhere and, when there's a new change, knowing if a resource should be recreated or just updated.
But when there's a demand there's (eventually) tools.
I think that with modern tools running your own / or rented metal (collocated) is doable and could make sense, but the could solves so many things that you don’t need to be responsible for, you really have to have a clear (businesses)case why you should not start in the cloud or migrate away.
I’m rooting for 37signals, but I’m not sure their move away from the cloud (which I root for in particular) will be really that effective in the end, but I hope it will.
I am happy for you though.
Or sorry that happened.
I kid, but the cloud may be expensive, but I too have used the cloud serving hundreds of thousands of users and have nowhere near the huge costs (around $100-130 a month), making well over $100K MRR.
Sometimes it depends on your business model and if you value your time.
Running a bunch of dev scripts on a VPS for me just wastes my time when I can just heroku push, point click and done, back enjoying my vacation.
But is it really a necessity that all those services are globally scaleable and have five nines?
What is the endgame, architecturally? How healthy is that the internet, all served bandwith increasingly is hosted by a few vendors?
Isn't Hetzner Robot still technically "cloud"? You would have to hire rack space in a data center and put your computer in there if you want to really self-host.
That's a lot of startups.
From reading comments in this thread it seems that many are completely oblivious to the complexity that the cloud can bring to an architecture. It's as though it doesn't cost any time. AWS documentation for everything that you opted in? Downloaded at night in your brain? What about Kubernetes? Just a pill away?
When I hear someone say that they're choosing the cloud for their new project because "managing" their own infra is too much, as someone who is familiar with what the author is pointing at, I wonder "what exactly is there to manage?" Are we talking about the few Python or PHP scripts that speak to Postgres and Redis? Scaling up and down? What are we talking about?
Getting to PMF first as a criteria to not consider the simpler alternatives suggested here is also a premature optimization. The author of this post is boasting a bit by saying that he can show you how to do this in less than 15 minutes and for less than 200$/m. I think it'd still be worth it if it took an entire day to set up. What kind of product are you building that can't spare a day to address your infrastructure?
If the cloud is so seamless and your project becomes so successful that it actually ends up needing all its toys, then you have nothing to worry about. When the time comes, you'll simply turn those magic keys and enter paradise.
RDS takes backups, out of the box with zero configuration, and further customization is very easy.
by draining cognitive resources you are able to manipulate people into making poor quality decisions. ex) mainstream news, aisle of grocery items, FANG-led development mythos and over-engineering
the last one is especially susceptible to manipulation.
"but we have to scale", "i read HN article from an ivy league founder", "someone with lot of followers on twitter said so"
and boom you are no longer in control of your mind or aware of losing control.
marketing is just like herding cattles, DDOS their cognitive resources so they can't think for themselves and gaslight each other, make them use your npm/pip package with a subscription/metering from day 1, get em while they are young.
The most simple setup is a some git repo software (forgejo, gitea, goes, etc...) with hooks and client side deploy script. Can't get more simple. You have systemd watching. You don't need docker at all.
Even 100Mbit/s is enough when you have nothing to show for.