Admins wonder if the cloud was such a good idea after all
theregister.com
theregister.com
I agree with others that the cloud vendors make it complex just to setup a simple service and push you to have these complex architectures. While they are beneficial for availability reasons, I question how often it is really needed.
It’s also designed to encourage and then monetize a certain style of complex, over-architected, high resource load development: microservices, use of single-cloud-proprietary services, tons of loosely coupled systems, etc. Cloud native is the J2EE of the present era.
This is an anti pattern that keeps happening over and over in our industry. It’s partly a result of the middlebrow meme— that mid developers love complexity— and partly something that is pushed because someone (consultants, cloud providers) are profiting.
Edit: I don’t mean to blast mid developers too much. One must often journey through the middlebrow meme to make it to the hooded dude on the right. The meme describes a learning trajectory. But we should try not to let the complexity explosion of mid developers become industry norm.
Modern businesses are full of outsourcing, temporary contracts and overdone management structures. This the shape of the software they produce.
For one of the projects where the demand appeared during working hours in my timezone, I did the math, and found this to be the case.
Pretty un-fixable without better stewardship over the software industry. That won't come for decades, so you may as well join the dark side and make the money and practice the way you want with hobbies or start your own company to enforce it.
I guess the stewardship can happen on a team level, but someone who's being restrained can always go find another job where they'll build the complexity they desire that ultimately gets maintained by someone else.
I keep trying and failing to make this argument at work. The majority of the team are primarily PHP developers and if you want their best and quickest work they can write you a perfectly good internal service in PHP.
But someone several levels above me drank the kool-aid and now everything takes four times as long and turns out crappy because half the team is spinning wheels learning new proprietary cloud tech and reimplementing the base-level tooling we already had. I am yet to hear any explanation why this cost is worth it beyond vague mutterings of “scalability” and “enterprise”.
The original system was blazingly fast — it was actually a set of composable TCP workers and coordinators, but repeatedly called dirty names like monolith, legacy, synchronous by ambitious managers trying to demonstrate the need to migrate. Proceed to bring on legions of contractors and offshore teams to move everything to HTTP and AMQP, fully coupled to the cloud vendor’s mechanisms for these with all of the goodies.
Today they are paying a fortune for a slower system, full vendor lock-in, real customers dismayed and prospective customers aghast at the prices.
Recently they came back to the original engineering teams to “make it a real-time system” and panicked at the realization that this capability was unceremoniously sacrificed in the name of microservices, the cloud for the sake of hot-shot aspiring upper management to make their name in “modernizing” the tech. Yikes
They need to balloon the amount of people they manage and be involved in as many projects as possible for bullet points on their CV when they're applying for the next position.
The mistake most companies and developers make is mapping their organization structure to their deployment architecture. And then over provisioning everything. Kubernetes is the goto solution for companies like that. We spend less on our cloud bills than it costs to spin up even a simple Kubernetes cluster that does basically nothing. It doesn't solve a problem we have because we have a monolith. All I need is two cheap vms and a loadbalancer. If it gets busy, I'll add vms.
There's some other stuff we need, of course. The most expensive thing is actually managed search and databases clusters. They are nice but also pricey.
But from a cost perspective the whole point for me is actually minimizing my time investment. A managed database means I don't need to obsess about uptime, backups, and other things you would normally pay some team to obsess about 24x7. I don't have that and I can actually go on vacation and it's fine.
Of course all this stuff is way overpriced. You get the equivalent of a raspberry pi worth of compute for a monthly price that would allow you to buy several of those. Cloud providers actually put multiple customers on their hardware for the cheaper vms. So, these things are basically paying for their own cost pretty much instantly. When you use a managed database, it runs on shared infrastructure that minimizes resource usage. So why is that more expensive than a non managed one that has more idle resource use? Convenience. It doesn't actually cost them more to manage it. But it's more convenient for you. So they just charge you more for less.
Cloud pricing bears no relationship to actual cost. Which is weird. There should be way more competition on price. But it seems that there just isn't. Of course, people just mindlessly buy AWS because they heard it's good. That's why Amazon is so rich. The retail business is a side hustle for them at this point. And you can get some amazing deals from smaller providers. Hetzner gets mentioned a lot and they are pretty aggressive with their pricing. But it's less convenient.
Like, you wouldn't rent a dozen ZipCars every day. If your business required a fleet of cars, you would quickly look for a better arrangement than pay-by-the-minute "convenient" app services.
In front of it are two Caddy load balancers, also running on VMs (Hetzner offers load balancers, but we wanted to support custom domains [2]).
The databases are running on root servers though, to get the maximum out of them.
When we started, we used Google Cloud. I think we would be paying at least $2,500/month for the same setup, compared to ~$400 now. And I have happily managed this all by myself. I think a proper sysadmin for a larger corp is worth the cost!
[1] https://pirsch.io/blog/techstack/
[2] https://pirsch.io/blog/how-we-use-caddy-to-provide-custom-do...
In the future, we will need a more robust and scalable solution of course.
https://www.civo.com/kubernetes
So I decided to try it to see it with my own eyes (EKS takes often half an hour or so...) and lo and behold, they use k3s! I mean, there is nothing wrong with that, but it's not mentioned anywhere on their product page so that was a bit misleading.
In any case they have a verification process that I haven't managed to pass. I was surprised because I used the same credit card for Gogle and AWS, but for some reason they closed my account.
Not dealing with procurement cycle for h/w purchase is priceless.
Where it is cost-effective, it can still be incredibly strategically limiting.
For instance, the IT department manager claims to need more headcount and money to fund it, because they believe that having more budget and employees shows how important their function is. Engineering requests a new VM, and the request sits and sits and sits, so long that you start believing that they are overloaded. Engineering complains that it is taking too long to provision machines, and IT (for the 100th time) says they need more headcount to keep up with internal demand. To fix the problem, management grants one or two hires, things get marginally better for a while, until the cycle repeats.
In this scenario, Engineering is trying to get their product completed, and looks bad because the schedule is slipping. If you are the Engineering manager, do you care more about fixing the organization, or sidestepping the issue by swiping a credit card?
Fixing things gets political fast, spending money less so.
Personally, I tend to push for PostgreSQL primarily with near self-service containerized services for the applications and mostly will just leverage the cloud's queues and blob/s3 storage, which are easy enough to replace in a software stack. I'm not big on lock-in.
As somebody who worked at a mid-tier company trying to run databases, I can attest that RDS was a godsend. Trying to hire both a DBA and an Ops team that knew how to write the chef cookbooks for a proper multi-node cluster postgres was a nightmare. Like, we never succeeded, and something that's ootb with RDS
You must not forget that also, many (most?) companies that run things themselves do not do it right. Like, with proper off-site backups that you're regularly testing and know you have options to easily spin up replicas or restore point-in-time backups
Jeff Atwood's been saying this from the initial SO podcasts from 2008. If you have the right people who are motivated, and provide the right equipment and resources, you've always had the opportunity to have lower TCO doing it yourself
I have since moved on to a small top-tier company and still prefer to "outsource" my DBA work by way of using Aurora. Yes Aurora is more expensive. No, I don't have the mental or monetary budget to hire up a proper ops team. I know my limits
First it took 3 weeks to decide what we needed, then 3 weeks to get a quote for new machines, then 3 weeks of negotiating and penny pinching because this was a 3 year lease. Then it took 3 months for them to get the machines in there.
In the cloud it's 3 clicks and you have the servers, and even better, scale on demand once you have good metrics in place.
Flexibility has a cost. But not being able to do anything also has a cost.
It's not like on-prem costs have magically stayed the same while cloud costs increase.
Not to say it's a slam dunk either way, you gotta compare, but on-prem is not without its own unique requirements.
Instead of having to hire a full sysadmin and buying infrastructure, you just click a server, enable snapshoting and you are already 10x better of of what you had before.
The biggest issue with these 'cloud' discussions and the pricing is simply solved: EITHER you have a good team who understands it and can and should determine if you are better of onprem or on cloud, or you do not have this team and you are not wasting money you are just paying for a lot better system which you would otherwise never had.
I have seen plenty of tremendes shitty setups in small companies from people who should know better. Databases reachable on the internet, slow ticket systems for getting hardware, costly upgrade prices (and partially internal cost centers were an server upgrade costs a few k).
You use too many slow lambdas? your Database costs tons of money? You can't control your micro services anymore? And you think your company would have been able to setup a better thing on prem without the help of cloud? Never
You need the same sysadmins (called something else to be trendy, but basically same role) either way.
Why? Because it was a lot less effort to just keep a few servers running (no hardware, etc.) and its not like a dev can't do anything system related.
I'm glad to read this! Unfortunately the industry has become so fragmented that a very significant percentage of younger developers no longer have any knowledge nor any interest in understanding how systems work. There is this idea that a developer only writes code and has no need to understand threads or processes or userspace vs. kernel or networking or anything at all other than the language syntax. It is very sad.
That being said, a design philosophy I’ve imposed is KISS with largely off the shelf OSS solutions, that way we have the ability to move the software elsewhere in the future (I’d love to run in house if we get big enough). Of course, nothing is that simple but it’s much harder if not effectively impossible with something built significantly out of vendor-specific libraries and platforms. I don’t mind managed solutions as long they are completely swappable.
Unfortunately, the people before me used DynamoDB as our persistence layer, but I digress.
Getting into things like functions as a service has always felt a little too hot for me. I don't mind integrating with cloud services though (e.g. S3-compatible APIs).
Hybrid seems to be most compelling if you can keep the monster in the box. Sign up for something like Azure or AWS to get at their IdP, DNS and CDN but keep all your actual machines in a cheaper provider (hetzner, et. al.).
What’s something that’s most like heroku today? You just upload your code and they handle everything?
I don’t want to be worried about applying OS updates, etc.
…otherwise you can try Render, Fly.io, Google Cloud Run, Railway, etc.
All that will not fail, because there are strong forces that explicitly want this to happen.
But it's always nice to hear that there are at least problems...
Historically, the main economic impetus for centralized government, is to coordinate the construction of irrigation infrastructure. The more dependent a society is on one big river (Ancient China/Egypt), the more centralized it becomes.
So unless your personal homelab needs to be hyperscalar, you don't need a public cloud; your needs will be served perfectly by a $5/mo VPS/PaaS, and/or a LACK rack.
Do companies need sovereignty? No, companies need to make money, that's their primary and main reason to exist. The choice of public cloud vs dedicated servers vs colocation vs on-prem vs etc amounts to whatever best fits your operating model. The public cloud was overhyped (like everything in tech), but smart companies chose the right tool to do the job.
I don't pay someone to knit my clothes, I buy shirts at the store.
I run all my shit on a VPS (which could be called a cloud) or a dedicated rented server but that is so easy to setup and I can run all projects on the same server. Easy, simple and if I need to scale I just rent a bigger server.
Scaling vertically is easy, scaling horizontally is hard. Most people never need to scale horizontally but does so anyway because they think they do.
You also get like 10x the perf for the same money. Using SQLite makes it easy to have backups and even time-specific testing databases.
Now in some miniscule amount of cases this is true and probably did help some people who's business 100x'd overnight, but in the vast majority of cases your business just will never get to the point it needs to be "cloud scale" in the first place. Nevermind accidentally shooting yourself in the foot with a recursive lambda here and there in certain instances or a misconfuguration causing a huge bill.
Edit: Another is because lots of companies who do actually end up succeeding negotiate a shit-load of credits with cloud providers so they can basically grow their business for free for a while. That is until those credits run out and they get hit with the actual costs.
Most of the value generated by startups is highly concentrated in the few that succeed. So naturally the industry should optimize itself to go big or go home, not penny pinch. That's also, why they hire expensive engineers rather than offshored developers, because speed matters more than cost.
As for non-tech companies. Their demand is more stable in the long run, but they are not tolerant of outages. Amazon cannot have its servers go down during a big sale, too many physical ongoing costs that gets wasted for every second the central nervous system is down. So the cloud is good for its reliability.
I got to see close up that a team of devs ran their whole solution (with a bunch of paying customers and everything) in the cloud, because cloud automation was good enough that they didn't need dedicated ops people.
Now I work for a cloud provider. I can't say that if I was running a business that I'd build it cloud-first instead of OnPrem. Certain use cases, sure. If I didn't need a lot of horsepower, I might build it on a cluster of VM's with some segmentation of duties - not quite microservices, not quite a monolith. Most likely if I was hosting on the cloud, I'd use the provider I work for, just because I know the system and how to get things done and how to talk to support.
I will say though - learning the ins and outs of cloud computing has made for a great career. Challenging, but lucrative.
FTA:
> Microsoft and Google decided not to officially comment on the survey's findings. However, a representative for one of the hyperscalers retorted that the figures seemed cherry-picked and pointed out that, as an example, customers using reserved instances could realize significant savings.
Reserved instances are a thing for sure. There's lots of other ways you can control cloud spend (enterprise agreements, dev/test subscriptions, spot instances, automated shut down / scale down, etc.) - it's enough complexity by itself that big companies hire entire teams of people to just work on tracking, projecting and controlling cloud costs.
"Back in the olden days", if your product was slow but the number of CPUs was fixed (or could not be increased instantly), the solution was to go and fix your code.
Basic system level skills are now no longer taught or practiced at the appropriate levels, so teams end up without engineers who actually know how to profile and optimize.
The cloud providers are the big winners here.
I've lost track of the times I've heard "compute is cheap! engineers are expensive!" Except... that compute cost will live forever. The time it takes someone to debug a bad loop or poor query is at worst a one time cost. Longer term, it may even make other stuff faster in the future.
And then you'll get responses like "pfft.... that's hardly the cost of one part time FAANG person who makes $680k/year - what's the point?"
And around we go...
It also means less engineers are needed for most companies.
It's not cheaper, it's just more opaque.
Back when your service was deployed on that 2 CPU box and it was too slow for obvious reasons, you optimized it and then it was good.
Today you just shrug and increase that kubernetes cluster from 16 to 48 nodes and forget about it. Costs a lot but the bill shows up somewhere else, in most groups the engineer doesn't even know what it is.
More often I think that's more of an overall engineering department time budgeting / culture issue.
Sometimes management is correct in that decision, sometimes it is "penny wise, pound foolish".
Fast forward and what we now have is a terribly complex beast with nets of dependencies that got developed partly by marketing and product teams, partly by demands of larger customers. And it's more or less clear that if you are small, you will be much better off using VPS (that's why Amazon decided to offer Lightsail), and if you are very big, you will save a lot of money moving at least part of your infra away from the public cloud.
But what remains is a large part of the market: medium sized businesses and large organizations that depend on the public cloud for many reasons. But they are not stupid: once a project becomes expensive, someone starts asking questions. And after you've exhausted the path of reserved instances, spot etc., and still burn a lot of money with not-always-stellar performance, you'll find a way of moving these workloads where it makes business sense.
But it has become so complex that instead of an OPS team you now have a Cloud team. With a huge wallet.
Best of both worlds is to setup your own Cloud on multiple VPS which is relatively easy nowadays: HAproxy, Rancher, Kubernetes, , Keycloack, Openwhisk, Gitlab, Harbor, Opentelemtry + Prometheus + Grafana and your devops will feel right at home.
When AWS was first getting big though they solved genuine, really hard problems for a lot of organisations that were large or growing quickly. NVME drives didn't exist, SSDs were expensive and a lot of servers still had spinning SAS drives – A little box with some ram and some NVME drives didn't scale as ridiculously far as they do now.
I do think as computers keep getting faster and smaller the number of use cases that need a 'cloud' shrink very quickly though.
writing a blank cheque to begin with is just an easy invitation for 'surprise' billing swindlers
sure cloud benefit might not look like a lot if you are an enterprise customer, but for startup, along with open source software, it changed the game and is still an enabler of what you see today in term of supporting startups from bootstrapping all the way to unicorn.
Terramark - your $10K per month per server (of very modest specs) for a NIST 800-53 High, convinced me of the fact that much cloud is a giant scam.