Moving from AWS to Bare-Metal saved us $230k per year
blog.oneuptime.com
blog.oneuptime.com
> In the context of AWS, the expenses associated with employing AWS administrators often exceed those of Linux on-premises server administrators. This represents an additional cost-saving benefit when shifting to bare metal. With today’s servers being both efficient and reliable, the need for “management” has significantly decreased.
I also never seen an eng org where substantial part of it didn’t do useless projects that never amount to anything
A team does not use AWS because it provides compute. AWS, even when using barebonea EC2 instances, actually means on-demand provisioning of computational resources with the help of infrastructure-as-code services. A random developer logs into his AWS console, clicks a few buttons, and he's already running a fully instrumented service with logging and metrics a click away. He can click another button and delete/shut down everything. He can click on a button again and deploy the same application in multiple continents with static files provided through a global CDN, deployed with a dedicated pipeline. He clicks on another button again and everything is shut down again.
How do you pull that off with "Linux on-premises server administrators"? You don't.
At most, you can get your Linux server administrators to manage their hardware with something like OpenStack, but they would be playing the role of the AWS engineers that your "AWS administrators" don't even know exist. However, anyone who works with AWS only works on the abstraction layers above that which a "Linux on premises administrator" works on.
You don't click to start and stop. You start with someone negotiating credits and reserved instance costs with AWS. Then you have to keep up with spending commitments. Sometimes clicking stop will cost you more than leaving shit running.
It gets to the point where $50k a month is indistinguishable from the noise floor of spending.
I worked on a web application that provided by a FANG-like global corporation that is a household name and used by millions of users every day, and which can and did made the news rounds if it experiences issues. It is a high-availability multi-region deployment spread about a dozen independent AWS accounts and managed around the clock by multiple teams.
Please tell me more how I "never actually ended up with a big AWS estate."
I love how people like you try to shoot down arguments with appeals to authority when you are this clueless about the topic and are this oblivious regarding everyone else's experience.
1) not as reliable as you think you are 2) probably wasting gobs of money somewhere
I’ve set up an “RnD” account where developers can go wild and click ops away. I also set up a separate “development” account where they can test thier IAC manually and then commit it and it gets tested through a CI/CD pipeline. Then after that it goes through the standard pull request/review process.
In a dream. In the real world of medium-to-large enterprise, a developer opens a ticket or uses some custom-built tool to bootstrap a new service, after writing a design doc and maybe going through a security review. They wait for the necessary approvals while they prepare the internal observability tools, and find out that there is an ongoing migration and their stack is not fully supported yet. In the meantime, he needs permissions to edit the Terraform files to update routing rules and actually send traffic to their service. At no point he does, or ever will, have direct access to the AWS console. The tools mentioned are the full-time job of dozens of other engineers (and PMs, EMs and managers). This process takes days to weeks to complete.
This only works that way for very small spend orgs that haven’t implemented soc 2 or the like. If that’s what you’re doing then probably should stay away from datacenter, sure
No, not really. That's how basically all services deployed to AWS work once you get the relevant CloudFormation/CDK bits lined up. I've worked on applications designed with high-availability in mind, which included multi-region deployments, which I could deploy as sandboxed applications on personal AWS accounts in a matter of a couple of minutes.
What exactly are you doing horribly wrong to think that architecting services the right way is something that only "small spend orgs" would know how to do?
Not everything is warehouse scale. You can serve tens of millions of customers from a single machine.
That doesn't preclude continuing to use AWS and other cloud service as a click-ops driven platform for experimentation, and requiring that anything that is targeting production to refactored to run in the bare-metal environment. At least two shops I worked at previously have used that as a recurring model (one focusing on AWS, the other on GCP) for stuff that was in prototyping or development.
That's part of the apples-and-oranges problem I mentioned.
It's perfectly fine if a company decides to save up massive amounts of cash by running stable core services on-premises instead of paying small fortunes to a cloud provider for the equivalent service.
Except that that's not the value proposition of a cloud provider.
A team managing on premises hardware barely covers a fraction of the value or flexibility provided by a cloud service. That team of Linux sysadmins does not nor will it ever provide the level of flexibility nor cover the range of services that a single person with access to a AWS/GCP/Azure account provides. It's like claiming that buying your own screwdriver is far better than renting a whole workshop. Sure, you have a point if all you plan on doing is tightening that screw. Except you don't pay for a workshop to tighten up screws, and instead you use it to iterate over designs for your screws before you even know how much load it's expected to take.
Anyone who says that hasn’t done it at scale.
“Infrastructure has weight”. Dependencies always creep in and any large scale migration involves regression testing, security, dealing with the PMO, compliance, dealing with outside vendors who may have white listed certain IP addresses, training, vendor negotiations, data migrations etc.
And then even though you use MySQL for instance, someone somewhere decided to do a “load data into S3” AWS MySQL extension and now they are going to have to write an ETL job. Someone else decided to store and serve static web assets to S3.
> Our choice was to run a Microk8s cluster in a colocation facility
they go on to describe they use helm as well. there's no reason to assume that "a a fully instrumented service with logging and metrics" still isnt a click and keypress away.
your points dont make a whole lot of sense in the context of what they actually migrated too.
Bootstrapped companies generally don't do this btw. This is a symptom of venture backed companies.
Your description applies to a substantial number of business units in that company. They also had a "research institute" whose best result in the last decade was an inaccurate linear regression (not a euphemism for ML).
> I also never seen an eng org where substantial part of it didn’t do useless projects that never amount to anything
Name one of business (tech or non tech) where this is ok/accepted and competitive in capitalism.
How long will we keep making these inflated salaries while being known for being wasteful, globally speaking?
If you're using less than a dozen servers manual configuration is simpler. Depending on what you're doing that could mean serving a hundred million customers. Which is plenty for most business.
I've worked at companies with their own data centers and manual configuration. Every system was a pet.
One that's not immediately obvious is to keep on staff experienced infra engineers that bring their expertise for designing future projects.
Another is the option to tackle project in ways that would be to costly if they were still on AWS (e.g. ML training, stuff with long and heavy CPU load).
Perhaps there are some good reasons to not choose such a provider once you reach a certain scale, but they now have their own versions of a lot of different AWS services, and they're more than sufficient for my own relatively small scale.
The total revenue so far for one cpu is 100x18x12x7 = $150k If used as a spot instance it’s 144/month, so about 200k
A standard i9-14700k gen has 32 threads, but it can run 12 of these instances (max 192mem). This CPU will cost you $800. Memory is cheap, so for about 1-2k you’re all set, and have a machine that’s way faster and cheaper.
Basically, buy a bunch of NUCs and you’re saving yourself around $1500 per month per NUC. It pays itself back in 1 month
Cloud hosting is —insane—
Not even touching memory ballooning for mostly idle applications.
Lastly, don’t give me reliability s an argument. These were all ephemeral instances that have no storage, so you’ll have to pay for that slow non-nvme storage platform.
Not sure how to factor that $ into the equation.
2x FTEs to manage the AWS support tickets
3x FTE to understand the differences between the AWS bundled products and open source stuff which you can't get close enough to the config for so that you can actually use it as intended.
3x Security folk to work out how to manage the tangle of multiple accounts, networks, WAF and compliance overheads
3x FTEs to write HCL and YAML to support the cloud.
2x Solution architects to try and rebuild everything cloud native and get stuck in some technicality inside step functions for 9 months and achieve nothing.
1x extra manager to sit in meetings with AWS once a week and bitch about the crap support, the broken OSS bundled stuff and work out weird network issues.
1x cloud janitor to clean up all the dirt left around the cluster burning cash.
---
Footnote: Was this to free us or enslave us?
And no one cares about AWS certifications. They are proof of nothing and disregarded by anyone with a modicum of a clue.
I’m speaking as someone who once had nine active certifications and I believe I still have six active ones. I only got them as a guided learning path. I knew going in they were meaningless.
I assume whichever provides more margin to Jeff Bezos.
Also what happens at hardware end-of-life?
Also what happens if they encounter an explosive growth or burst usage event?
And did their current staffing include enough headcount to maintain the physical machines or did they have to hire for that?
Etc etc. Cloud is not cheap but if you are honest about TCO then the savings likely are WAY less than they imply in the article.
Your math is incorrect. The savings are per year. The job gets done once.
> Also what happens at hardware end-of-life?
You buy more hardware. A drive should last a few years on average at least.
> Also what happens if they encounter an explosive growth or burst usage event?
Short term, clouds are always available to handle extra compute. It's not a bad idea to use a cloud load-balancing system anyway to handle spam or caching.
But also, you can buy hardware from amazon and get it the next day with Prime.
> And did their current staffing include enough headcount to maintain the physical machines or did they have to hire for that?
I'm sure any team capable of building complex software at scale is capable of running a few servers on prem. I'm sure there's more than a few programmers on most teams that have homelabs they muck around with.
> Etc etc.
I'd love to hear more arguments.
TFA states that they maintain their AWS account, and can spin up additional compute in ~10 minutes.
Not that we'd need them as we wouldn't have to write as much HCL.
Would've saved another ~30% for minimal difference in performance.
For me this doesn't look like a sensible move especially since with AWS EKS you have a managed, highly-available, multi-AZ control plane.
Unless their product is pretty static and not seeing much development, they're probably in the negative.
What's capex vs opex now? Thats 150k of depreciable assets, probably ones that will be available for use long after all the current staff depart.
Everyone forgets what WhatsApp did with few engineers and less hardware, there's probably more than enough room for them to grow, and they have space to increase capacity.
The cloud has a place, but candidly so does a Datacenter and ownership.
Now you're maintaining two tiers. That's more work, not less.
I think the real question is what's the cost to buy an equal number of training hours so you can pretend your resources are that competent.
Imagine a military that never fought, never did significant exercises and probably doesn't even have cleaning exercises any more on half its inventory.. That's basically how I view a company that had organic IT growth over a few years and hasn't done a major transition in anyone's recent memory.
The move does cost money, once. Then the savings over years add up to a lot. We made this change more than 10 years ago and it was one of the best decisions we ever made.
Recently I've been moving most projects to Hetzner Cloud, it's a pleasure to work with and pleasantly inexpensive. It's a pity they didn't start it 10 years earlier.
Why would you spread FUD? They have several datacenters in different locations, and even if they were as incompetent as OVH (they are not)[0], the destruction of one datacenter doesn't mean you will lose data stored in the remaining ones.
[0] I bet OVH is also way smarter than they were before the fire.
During that time, one of my servers was hacked once (I was stupid enough to start digging Monero on the same system I had some other services installed) and another time one of my users had a weak password and his account was sending spam. In both cases they notified me and gave me the time to fix the problem. I also appreciate human contact and quick replies.
Also, there are companies that manage spot price allocation for you, so you should essentially always pay spot+small_x% and never actually get terminated.
Let's say you're really pushing the connection and your p95 is 900megabit up. That is $200 at colo vs ~$8200 for amazon.
Wait, is this accurate?
If so I need to sign our company up for a savings plan... now. We use RI's but I thought savings plan only applied to instance cost and not bandwidth (and definitely not S3)
You're not saving anything doing it yourself.
And you've just given yourself the massive inconvenience of running a HA Kubernetes control plane.
Keeping a few racks of servers happily humming along isn't the massive undertaking that most people here seem to think it is. I think lots of "cloud native" engineers are just intimidated by having to learn lower levels to keep things running.
Keeping them humming along redundantly, with adequate power and cooling, and protection against cooling- and power failures is more of an undertaking, though. Now you are maintaining generators, UPSs and multiple HVAC systems in addition to your 'few racks of servers'.
You also need to maintain full network redundancy (including ingress/egress) and all the cost that entails.
All the above hardware needs maintenance and replacement when it becomes obsolete.
Now you are good in one DC, but not protected against tornadoes, fire and flood like you would be if you used AWS with multiple availability zones.
So, you have to build another DC far enough away, staff it, and buy tons of replication software, plus several FTEs to manage cross-site backups and deal with sync issues.
But yeah, in the long run, colocation becomes significantly cheaper than cloud. You use AWS, and you'll find yourself paying $200/month for hardware you could buy once for $2,000.
Sometimes I think people forgot colocation is an option that exists.
For $600 I can get hardware that outperforms a $1200/mo ec2. Easy.
I don't mean to call you out specifically, but I believe your comment is a perfect example of how the majority of developers have no clue what being on bare metal actually involves. You literally need to do none of these things. Cloud vendors make it sound overly complex and the myth just kinda self perpetuates because nobody knows better.
If anyone in the Bay Area is considering the move out of the cloud and wants to see in person what is really involved, I might consider putting together a group tour of one of my rack locations.
What happens when your own HVAC dies and your DC has about 4 hours until it overheats?
(I'm a software engineer who previously built and maintained racks of bare metal. Never again).
These datacenters are fed by multiple power substations, have onsite battery and generators, and contracts for delivery of fuel in the event of a disaster. But none of these things are any more your problem than if a power plant explodes knocking out an AWS region.
> > Now you are maintaining generators, UPSs and multiple HVAC systems in addition to your 'few racks of servers'.
> I don't mean to call you out specifically, but I believe your comment is a perfect example of how the majority of developers have no clue what being on bare metal actually involves. You literally need to do none of these things.
By the way, I worked for a double digit billion dollar company that built its own datacenters as well as placing resources in colos. They started out purely in colos, and put several colos out of business over HVAC and power costs (back when rack space was billed by area, not cooling). Even after that, they stayed in colos, and when I worked there, we constantly had to deal with the unreliability of colos- not just that they were smaller, with less cooling, and inadequate power, but also because they often didn't actually fulfill their contractual requirements. ATL was a great example.
If colos work for you, that's great. I just don't thinnk they are prepared to handle disasters nearly as well as the megascale cloud providers.
No, you seem to misunderstand how leased space works. If you rent a floor of an office building, the toilets flush without you having to own a water plant or redundant water pipes.
> when I worked there, we constantly had to deal with the unreliability of colos- not just that they were smaller,
It sounds like whomever was in charge of picking datacenters was shit at their job. The colocation market isn't what it used to be and it isn't just a dude with some warehouse space and swamp coolers. Colos are publicly traded companies or REITs and have good SLAs.
> I just don't thinnk they are prepared to handle disasters nearly as well as the megascale cloud providers.
I worked for a megascale cloud provider. I'm intimately familiar with the nuts and bolts of a few others. Some of it is the big owned and operated campuses you see in the glossy brochure photos, but a substantial part is also in the same colos you can lease yourself. They don't bring in any additional cooling or power over and above what the datacenter provides.
Seriously, a facility with multiple Internet connections, adequate power and cooling, passive cooling in case of outage, protection against power failures, adequate UPSes, backup generator power, and so on could be my Mom's house. There's nothing special about any of those things that makes them somehow frightening.
"buy tons of replication software"? What industry do you work in? Seriously, nobody who isn't in some clueless "enterprise" would pay good money for things that're widely available in open source.
I'm dismissive because these things aren't difficult if you've actually done them, so I can only assume you've never done them.
Rightly so, because they're cloud native engineers, not system administrators. They're intimidated by the things they don't know. It'll be a very individual calculation whether it's worth it for your enterprise to organize and maintain hardware yourself, or isn't.
And there are plenty of us who've spent time managing hardware and physical networks, transitioned to cloud, and are very happy to not be looking back.
I made the correct career choice. Downturn? What downturn?
It isn't until hardware failures happen, and that requires different skills to deal with effectively. Like the time a core networking switch presented with a dead PSU and no backups, so you Frankenstein it back to temporarily working with another switch to pull the config off of it.
With bare metal you may have to have more generalized, Jack of all trades staff to account for the unaccountable.
Always keep an eye on your business goals and not on the hype. For oneuptime obviously downtime is a huge problem but you'd be surprised for how many businesses it's much cheaper to be down for a few minutes here and there than engineering a complex HA mechanism. The aforementioned spiking problem often can be solved cheaply by degrading the hot pages to static and serving them from CDN (if you have a mechanism for doing this of course). And so forth.
Remember KISS.
When we were utilizing AWS, our setup consisted of a 28-node managed Kubernetes cluster.
Each of these nodes was an m7a EC2 instance. With block storage and network fees included,
our monthly bills amounted to $38,000+
The hell were you doing with 28 nodes to run an uptime tracking app? Did you try just running it on like, 3 nodes, without K8s? When compared to our previous AWS costs, we’re saving over $230,000 roughly per year
if you amortize the cap-ex costs of the server over 5 years.
Compared to a 5-year AWS savings plan? Probably not.On top of this, they somehow advertise using K8s as a simplification? Let's reign in our spend, not only by abandoning the convenience of VMs and having to do more maintenance, but let's require customers use a minimum of 3 nodes and a dozen services to run a dinky uptime tracking app.
This meme must be repeating itself due to ignorance. The CIOs/CTOs have no clue how to control spend in the cloud, so they rake up huge bills and ignore it "because we're trying to grow quickly!" Then maybe they hire someone who knows the Cloud, but they tell them to ignore the cost too. Finally they run out of cash because they weren't watching the billing, so they do the only thing they are technically competent enough to do: set up some computers and install Linux, and write off the cost as cap-ex. Finally they write a blog post in order to try to gain political cover for why they burned through several headcount worth of funding on nothing.
I've seen this before.
It happens when the people who manage the infrastructure are completely decoupled from the developers. The developers determine how much CPU / memory is "needed" for a given workload, and the infra team just adds it. Developers aren't responsible for infra costs. Infra team has limited ability to control costs.
Add a few years and teams with their own container-based applications, Kafka, Elasticsearch, multi-instance Postgres, etc., and soon enough, it's a 28 node cluster and costs are out of control. Infra team can only do so much, and devs aren't incentivized to help, either. Everyone's shrugging their shoulders because now it'll take significant refactoring and cross-silo work to actually fix.
If they told us what was running on those 28 nodes, we'd point that out immediately. But it's not just happening at this company. I've seen this pattern many times.
All deployed using capistrano.
To be fair, considering the pocket-calculator-grade performance you get from AWS (along with terrible IO performance compared to direct-attach NVME) I can totally understand they’d need 28 nodes to run something that would run on a handful of real, uncontended bare-metal hosts.
looks to me like they did this in Europe previously, and they are looking to do the same in the US now.
And your own computers require expertise so expensive and frightening that no sane company would host their own computers.
How Amazon created this alternate reality should be studied in business schools for the next 50 years. Amazon made the IT industry doubt its own technical capabilities so much that the entire industry essential gave up on the idea that it can run computer systems, and instead bought into the fabulously complex and expensive and technically challenging cloud systems, whilst still believing they were doing the simplest and cheapest thing.
Colocation is rarely worth it unless you have non-standard requirements. If you just need a general-purpose machine, any of the aforementioned providers will sort you out just fine.
Almost everything that the AWS specialist needs to know comes in after that and has some equivalent in bare metal world, so those costs don't disappear either.
In practice there are extra costs which may or may not make sense in each case. And there are companies that don't reassess their spending as well as they should. But there's no alternative realty really. (As in, the usually discussed complications of bare metal are not extremely overplayed)
Different solutions work best for different companies.
Each of these statements is utter BS
PS. Oopsy I just read their third paragraph ;)
https://github.com/OneUptime/interview/blob/master/software-...
Side note: I'm in slight disbelief at how high that salary range is compared to how minimal the job requirements are.
Been there, done that, don't care to return to that life.
Half that and half it again and I'd still be looking at a decent raise lmao
A lot of companies (like Amazon) will gleefully slash your salary if you try to move somewhere cheaper, because why should we pay you more if you don't just need that money to fork over to a landlord every month?
There's also all the things Americans go without, like socialized healthcare. Even with their lauded insurance plans they still pay significantly more for worse health outcomes than any other country
Edit: even then, TC is tied to how the stock market is doing, and not paid out by the company directly, so it only makes sense to compare with base wage plus benefits.
Fargate, S3, Aurora etc. These are managed services and are incredibly reliable.
Lot of people here seem to think these cloud providers are just a bunch of managed servers. It's far more than that.
And then there's the question of whether you're going to use Terraform, Ansible, CloudFormation, etc or click through the GUI to manage things.
My point is, nothing in AWS is 100% turnkey like a lot of folks pretend it is. Most of the time, it's leadership that thinks since AWS is "Cloud" that it's as simple as put in your credit card and you're done.
A bit in jest, but places I've worked where we've moved to the cloud ended up with more people managing k8s and building a platform and tooling, than when we had a simple inhouse scp upload to some servers.
For examples:
- How much data are they working with? What's the traffic shape? Using NFS makes me think that they don't have a lot of data.
- What happened when their customers accidentally sent too much events? Will they simply drop the payload? In bare-metal they lose the ability to auto-scale quickly.
- Are they using S3 or not, if they are, did they move that as well to their own Ceph cluster?
- What's the RDBMS setup? Are they running their own DB proxy that can handle live switch-over and seamless upgrade?
- What's the details on the bare metal setup? Is everything redundant? How quickly can they add several racks in one go? What's included as a service from their co-lo provider?
I also would love to see a comparison done by a financial planning analyst to ensure no cost centres are missed. On prem is cheaper but only by 30 to 50%. That is the premium you pay for flexibility, which you can partly mitigate by purchasing reserved instance for multiple years.
Depending on use case.
If you have traffic which isn't consistent 24/7 then AWS Spot instances with Gravitron CPUs will be cheaper than on-premise.
Because you have the ability to in real-time scale your infrastructure up/down.
fluctuation in traffic is handled by auto scaling
Saving money on stateless (or short start times) services is done with spots
They have a decent business case, but I don't feel like they executed well to meet the real objective. They don't want to be in AWS since they're an uptime monitor and they want to alert on downtime on AWS. But they have a single rack of servers in a single location. A cold standby in AWS doesn't mean a ton unless they're testing their failovers there... which comes at quite the cost.
I've worked on-prem before. Now I work somewhere that's 100% AWS and cloud native. You'd have to pay me quite a bit more to go back to on-prem. You'd have to pay me quite a bit more to go somewhere not using one of the 3 major clouds with all their vendor specific technologies.
The speed is invaluable to a business. It's better to have elevated spend while trying to find good product market fit. I didn't understand this until I worked somewhere with a product the market wanted. > 50% growth for half a decade wouldn't have been possible on-prem. > 25% growth for a full decade wouldn't have been possible on-prem.
I've been with my current company from $80MM ARR to $550MM ARR. We've never breached more than 1.5% of income on total cloud spend. We've been told that's the lowest they've seen by everyone from AWS TAMs to VC/PE people. It's because we're cloud native, we've always been cloud native, and we're always going to be cloud native.
You've gotta get over "vendor lock-in". With our agility it's not really a thing. We could move to another cloud or on-prem if we really wanted. Wouldn't be a huge problem moving things service by service over... though we've had a few more recent changes that would be troublesome, we'd be able to work around them because we're architected well.
On the other hand maintaining $100K a year of spend on AWS is unlikely to be worth the effort of optimizing and maintaining $1M+ on AWS probably means the usage patterns are such that the cloud is cheaper and easier to maintain.
AWS is... incentivizing scope creep, to put it mildly. In ye olde days, you had your ESXi blades, and if you were lucky some decent storage attached to it, and you gotta made do with what you had - if you needed more resources, you'd have to go through the entire usual corporate bullshit. Get quotes from at least three comparable vendors, line up contract details, POs, get approval from multiple levels...
Now? Who cares if you spin up entire servers worth of instances for feature branch environments, and look, isn't that new AI chatbot something we could use... you get the idea. The reason why cloud (not just AWS) is so popular in corporate hellscapes is because it eliminates a lot of the busybody impeders. Shadow IT as a Service.
As someone who has worked on an AI chat bot I can assure you it does not come from engineers.
It's coming from the CFO who is salivating at the thought of downsizing their customer support team.
I think the reality is no discipline, be it Engineering, Product or Finance, is immune to flights of fancy.
The dude nearly got fired, but your comment hit the spot. You made my night, thank you.
Sure, but that seems orthogonal to the pros and cons of having more layers of oversight (busybodies, to use your term) on infra spend. Badly run companies are badly run, and I don't think having the increased flexibility that comes from cloud providers changes that.
It's not entirely without downsides though and I think many shops are willing to pay more for a different set of them. It is incredibly rewarding work though. You get to do magic.
* You do need more experienced people, there's no way around it and the skills are hard to come by sometimes. We spent probably 3 years looking to hire a senior dba before we found one. Networking people are also unicorns.
* Having to deal with the full full stack is a lot more work and needing manage IRL hardware is a PITA. I hated driving 50 miles to swap some hard drives. Rather than using those nice cloud APIs you are also on the other side implementing them. And all the VM management software sucks in their own unique ways.
* Storage will make you lose sleep. Ceph is a wonder of the technological world but it will also follow you in a dark alleyway and ruin your sleep.
* Building true redundancy is harder than you think it should be. "What if your ceph cluster dies?" "What if your ESXi shits the bed?" "What if Consul?" Setting things up so that you don't accidentally have single points of failure is tedious work.
* You have to constantly be looking at your horizons. We made a stupid little doomsday clock web app that we put all the "in the next x days/weeks/months we have to do x or we'll have an outage." Because it will take more time than you think it should to buy equipment.
I think it's very useful for batch processing, especially owning a GPU cluster could be great for ML startups.
Hybrid cloud + bare metal is probably the way to go (though that does incur the complexity of dealing with both, which is also hard).
I guess before Amazon invented "the cloud" there wasn't any software companies...
So it's a fact that for most use cases it will be significantly easier to manage than bare metal.
Because much of it is being managed for you e.g. object store, databases etc.
Setting up Garage for obj store: 1 hour.
Setting up Longhorn for storage: .25 hour.
Setting up db: 30 minutes.
Setting up Cilium with a pool of ips to use as a lb: 45 mins.
All in: ~5 hours and I'm ready to deploy and spending 300 bucks a month, just renting bare metal servers.
AWS, for far less compute and same capabilities: approximately 800-1000 bucks a month, and takes about 3 hours -- we aren't even counting egress costs yet.
So, for two extra hours on your initial setup, you can save a ridiculous amount of money. Maintenance is actually less work than AWS too.
(source: I'm working on a youtube video)
Because there is a world of difference between installing some software and making it robust enough to support a multi-million dollar business. I would be surprised if you can setup and test a proper highly-available database with automated backup in < 30 mins.
Also, with EKS, you get literally all of this (except Cilium and Longhorn, if you need that, which you don't if you use vpc-cni and eks-csi), in ~8 minutes, and it comes with node autoscaling, tie-ins into IAM and a bunch of other stuff for free. This is perfect for a typical lean engineering team that doesn't really do platform stuff but need to out of necessity and/or a platform team that's just getting ramped up on k8s.
You also don't need to test your automation for k8s upgrades or maintain etcd with EKS, which can be big time-savers.
(FWIW I love Kubernetes and have made courses/workshops of exactly this work)
I'm not suggesting that this isn't a viable solution, but I would prefer to be well-prepared and have a team of experts in their respective fields who are willing to have an on-call duty. This team would include a specialist for Longhorn or Ceph, since storage is extremely important, one to set up and maintain a high-availability PostgreSQL database with an operator and automated, thoroughly tested backups, and another for eBPF/Cilium networking complexities, which is also crucial because if your cluster network fails, it results in an immediate major outage.
Certainly, you can claim that you have sufficient experience to manage all these systems independently, but when do you plan to sleep if you're on call 24/7/365? Therefore, you either need a highly competent team of domain experts, which also incurs a significant cost, or you opt for cloud services where all of this management is taken care of for you. Of course, this service is already included in the price, hence it's more expensive than bare-metal.
You know bare metal can be an automated fleet of cattle too, right?
> Server Admins: When planning a transition to bare metal, many believe that hiring server administrators is a necessity. While their role is undeniably important, it’s worth noting that a substantial part of hardware maintenance is actually managed by the colocation facility. In the context of AWS, the expenses associated with employing AWS administrators often exceed those of Linux on-premises server administrators. This represents an additional cost-saving benefit when shifting to bare metal. With today’s servers being both efficient and reliable, the need for “management” has significantly decreased.
This feels like a "famous last words" moment. Next year there'll be 400k in "emergency Server Admin hire" budget allocated.
Compute, storage, database, networking. You would be better off using Digital Ocean, Linode Vultr etc. so much cheaper than AWS, lots of bandwidth included rather than the extortionate $0.08 GB egress.
Compute is the same story. 2 VCPU, 4GB VPS is ~$24 using a VPS. The equivalent instances (after navigating the obscured pricing and naming scheme), is the c6g.large is double the price at $50.
This is the happy middle ground between bare metal and AWS.
> 5 years would be a very good lifetime for a heavy use server, especially HDs.
Read Backblaze report - a lot of their HDs are over 8yo and the afr is less than 2%. SSDs will actually fail faster under heavy write load around 5 years yes
It's been a big "duh" for 20+ years. Large-scale, consistent loads aren't suited to cloud infrastructure. Mostly its shops that don't care about costs, don't know any better, or lack technical capabilities outsource most of their infrastructure.
The use-cases for *aaS are:
- Early startups
- Beginning projects
- Prototyping
- Peaky loads on-demand or one-of
- Evade corporate IT department
AWS and VPSes can also be ill-suited for personal use if you live in a major city with bottom tier, cheap datacenters that can rent a 1 GbE uplink, a PDU plug, and 4U. For anything substantial, it's not hard to lease some dark fiber and run (E)BGP.
https://www.ripe.net/manage-ips-and-asns/as-numbers/request-...
Lift and shift is brutal and doesn't make a lot of sense.
And at this point you are completely locked in.
At "Storage and LoadBalancers" the NFS link point to https://microk8s.io/docs/nfs Should be https://microk8s.io/docs/addon-nfs
Rolling a droplet, load balancer and database costs like $30.
>single rack configuration at our co-location partner
I've got symetrical gigabit static ipv4 at home...so can murder commercial offerings out there on bang/buck for many things. Right up until you factor in reliability and redundancy.
Most colocation facilities include dumb hands for free. Push a button, tell me what lights are on, plug in a monitor and tell me what it says, move the network cable from port 14 to 15, replace the drive with the spare sitting in the rack, etc.
Smart hands are billed around $100-$300/hr or a flat rate per task from a menu. Write an image to a USB stick and reinstall the OS on a server. Unrack and replace a switch. Figure out which drive has failed in the server and replace it. etc.
I've ran computers in datacenters for 20+ years and maybe used smart hands 1 or 2 times.
I’ve definitely seen a lot of over engineered solutions in the chase of some ideals or promotions.
It seems crazy to me that the two options are AWS vs bare metal to save that much money. Why not a moderate solution?
I came to the conclusion that when you factor everything in—the time it takes to maintain such a massive infrastructure operating smoothly across the globe, investment in the further development of cloud services, employing security experts, paying developers competitive salaries, and of course aiming for a profit margin for the company—you end up with the prices that are evident among the major cloud providers. It simply isn't feasible to offer these services for much less (there are a few exceptions people rightly complain about, such as stupidly high egress costs). It has reached a point where Google, for example, has attempted to undercut prices to such a degree that their cloud operations have been running at a loss in the past[^1].
[^1]: https://www.ciodive.com/news/google-cloud-revenue-Q2-2022/62... (read last paragraph)
It's probably easier to optimize your stack TBH. I can't wait until I get a chance to use ARM on AWS.
[0] https://www.servethehome.com/falling-from-the-sky-2020-self-...
Last fcos release was a week ago - https://fedoraproject.org/coreos/release-notes?arch=x86_64&s...
I was under the impression that bare-metal means "no OS".
edit: when you selectively choose your data points, and ignore human and migration costs.
They're selling an observability solution...
AWS as a whole has never been down.
It's Cloud 101 to architect your platform to operate across multiple availability zones (data centres). Not only to insulate against data centre specific issues e.g. fire, power. But also AWS backplane software update issues or cascading faults.
If you read what they did it's actually worse than AWS because their Kubernetes control plane isn't highly-available.
Sounds like they already have their bases covered.
HA architectures exist for a reason because that last step is a massive headache.
A huge multi billion dollar company with "cloud" in its name recently had a big downtime because they did not follow "cloud 101".
Keep rubbing salt lol I live in a low cost area though. It's even pleasant some times of the year.
like, talk about decades removed! no, nobody has their servers in the coworking space anymore sir.
nice to see people attempting a holistic solution to hosting though. with containerization redeploying anywhere on anything shouldn't be hard.