Glory is only 11MB/sec away (2023)
thmsmlr.com
thmsmlr.com
When we grew up in the early 2000s, our bigger sales usually featured complicated stacks. They had redundant load-balancing, redundant firewalls, way more than many customer never needed. (But they did ask for it). The failover often cost more in management complexity than it saved when a box died in the "right" way to trigger the planned failover event.
We sold ourselves on cleverness when people asked for that. It worked to grow our business to 30 staff, a data centre in our home city, life was good! We responded to AWS with an API-based cloud hosting platform. But sales still peaked in 2012.
Customers wanted even more complex solutions than the ones we were selling - partially or wholly based on AWS. But - we figured - the hardware we were buying was hugely powerful compared to 10 years previously, and sites weren't that much more complicated. The bigger customers would (surely!) want fewer, less complicated boxes as a result. Unfortunately that is not selling on cleverness, that is selling on price. We never understood the financial ambition needed for that pviot. Nobody trusted a single cheap server, and even if they bought two, where was the scalability? It worked enough to keep revenue flat, but we obviously couldn't compete on building managed service stacks and software ecosystems quicker than Amazon.
When the new technical challenges had long dried-up, we sold in 2018.
My thinking (and so most of the company's product design) came from being bootstrapped where the possibility of an uncapped hosting bill seemed like an insane risk to take. Who would take it? (wait - what - why was everyone taking it??!)
AWS are embedded not just because VC makes their high-priced products feasible, but because their particular brand of cleverness is embedded in a generation of software developers. It obviously works! But the knowledge of when you might not need their cloud (or what the alternatives could ever be) feels quite a niche thing now.
Unfortunately AWS's complexity and cleverness is catnip for developers, and its support for résumé-driven development is second to none.
1. Traffic is not evenly spread. The figures from the article (400M page loads per month) are subject to recursive 80/20 rule. 80% of the requests (320M) are served within 20% of the time (6 days). Within those high-volume days, 80% of the requests (256M) are served in 20% of the time (~29h). And if you're serving particularly spiky traffic patterns, then 80% of the those requests (~205M) come in during just 20% of the time window (5.8h) -- 5.8h is a little shy of 21000 seconds. 205M / 21k is about 9.7k requests per second.
That's still doable on a single system, but it's no longer trivial. Especially if you want to run off a single DB with no read replicas. And while the total amount of traffic served remains the same, the necessary bandwidth cap for peak loads gets far above the optimistically averaged 11MB/sec.
2. Unidirectional end-to-end latency only applies to streaming data. A cold start in the real world (no HTTP/3) requires first to establish the underlying TCP connection which is three trips, then the TLS connection which is a minimum of two more trips, and then you get to send the actual HTTP request... which still needs to send the response back.
If you want to serve real humans, everything observable has to happen in less than one second[0]. After that there's a steep drop-off as users assume your system is broken and just close the tab.
Disclosure: in previous life I helped to run a betting exchange. The traffic patterns are extremely spiky, latency requirements are demanding and trading volume is highly concentrated in just a tiny fraction of the overall event window. For any activity involving live trades, we had to get the results on their screen within 100ms from the moment they initiated the action. That means their network roundtrip latency ate into our event processing budget.
0: https://www.nngroup.com/articles/response-times-3-important-...
And you absolutely can deal with that sort of load in single large server configurations, but now we're not just building a webserver, we're building a pretty hardcore frontend load balancer (that happens to have an embedded webserver).
They may well charge through the nose for it, but a hell of a lot of engineering has gone into AWS' load balancers and network infrastructure, so that the rest of us don't have to become experts in that whole segment of the stack.
I don't know if the business folks would rather sacrifice those in a load spike, or just load shed at a request level - for an business defined by engagement metrics and ad revenue, it's not clear the former is better than the latter.
Most of it hard earned, after having to deal with thousands of outlier events.
An old friend used to run the Finnish election results services' public-facing backend, and was an early AWS customer. He broke their load-balancers - twice.
Both times he got in touch with AWS support well in advance, asking to pre-scale and pre-warm the load balancers because he was expecting a major traffic surge once the results started to come in. First time around, he was confidently told that he shouldn't worry, since AWS can take on any amount of traffic.
On the election night, AWS load balancers failed to scale in time. The traffic estimates my friend had provided were indeed accurate within an order of magnitude, but AWS hadn't believed his numbers. NOBODY could have a service that legitimately required scaling from a few hundred requests per second to 2M requests per second in about one minute. Apparently their on-call engineers got burned badly that night.
Couple of years later, he approached AWS support again, with the same request. That time he was taken more seriously, and was assured that the autoscaling algorithms were now much smarter and their relevant engineering teams had prepared the ground for more rapid autoscaling needs. They were confident they didn't need to pre-scale the fleet.
The scaling wasn't still fast enough and their on-call team had to _again_ manually force their balancers to stay just ahead of the demand with the massive surges coming in waves. So much for the fine-tuned scaling algorithms.
Fast-forward to today, and AWS have a specialist "hot launch" service offering where they work with the customers up front to make sure their launches have sufficient capacity available and their perimeter load balancers properly pre-warmed to absorb these kinds of supposedly one-off situations.
Agreed that his "across the world" example is a bit silly. Because he doesn't take into account the connection construction.
His primary point is still reasonable. How many services need world wide reach? Did you build it for multiple languages also?
If you're in the US. Or you're in the EU. A nice centralized server will have <=30 ms of latency to the entire region you are serving.
Edge is over valued unless you do have true global needs and then you have to also manage global database (s).
Looks like the author is getting more than 11MB/s traffic. Here's an archived version: https://archive.is/UVpg0
I've never seen that error before. Various websites suggest it's caused by using a proxy, a VPN, or DNS-over-HTTPS. None of these apply to me.
A better way IMO is: don't scale prematurely.
Build things as you need them. In the vast majority of cases, even CDNs are an unnecessary cost (presuming you're not paying the exorbitant cloud provider bandwidth tax). If you start to see performance issues, then deal with it as needed.
And if your workhorse suddenly grows that coveted single horn?
That's a problem you want to have!
But hey if you think you need to start out with webscale, have two dozen micro services and an exploding number of failure states between them and blow through your VC money with an AWS bill before the first customer signs up, more power to you.
>Mark hosted Facebook (it was theFacebook.com back then) from his own computer in his Harvard dorm room. You can buy a domain name and point it to any static IP address. That's how we did things back then because we have fewer options for hosting services that didn't cost an absurd amount for a college student.
>Eventually, Facebook was hosted in a shared datacenter in California,..(off Quora)
No you aren't running tests because you will be locked in to aws and your cloud computing bills will eventually bankrupt you.
And not just you, it's a core part of cloud computing's monetization model.
Additionally, if anything goes wrong with your Digital Ocean account, Digital Ocean will not help you. So you are running a high risk.
How does that jive with your interpretation of cloud as a scapegoat?
Now if you're willing to go DB-less, keep your whole global state in literal memory on one big server (no round-tripping to Redis or whatever, actual in-process objects), just occasionally snapshotting that memory to disk (this part's tricky), and use a compiled, multi-threaded language -- then you can saturate a Gbit or bigger NIC and literally serve the world from one box. I kind of wish I had a real use case for that architecture.
simonw (datasette) has built troves and troves of tools and writings about SQLite's production use for content-heavy and/or data rich websites: https://simonwillison.net/2021/Jul/28/baked-data/
From the benchmarks I’ve seen, because it’s significantly faster, specifically round trip times. This would make sense since SQLite is in-process and doesn’t need serialization[1]. Which in turn offers a second, optional advantage within reach - serial processing of operations, which is significantly easier to test, reason about, and build supporting cache layers around.
[1]: If you don’t use a Unix socket you also have networking overhead - but you said same server so I’ll leave this as a side-note since it’s extremely common to put postgres on a different machine for isolation. In fact, it’s one of the main advantages with networked dbs.
The author means sqlite makes sense if you're running one machine, which is what makes most sense for most use cases.
I would only consider postgres/mysql if I outgrew vertically scaling a single box.
My takeaway is not a debate against the merits of cloud in the face of trade offs, but against the necessity of—the now ubiquitous—cloud architecture pattern (and related lock-in). The "this-vs-that" is a rhetorical device to introduce an alternative. Of course which solutions are right for different use cases will depend on so many different things, and it's those things that keep engineers employed! :)
That said, we can put our engineering hats on to solve for SRE concerns within the patterned proposed by the author; I'm thinking the "what about availability when your one server goes down?!" is a straw-man in the sense that we have different ways of solving the availability story than our Ubiquitous System, and of course the solution must depend on what's actually relevant.
I, for one, appreciated the thought experiment and I took it for that: a theoretical alternative that would otherwise seem untenable being held up against a known pattern. I take it as theoretical because the author didn't actually build the system they're proposing for BusinessInsider, and without putting it into practice, I can only see it as that.
In practice, articulating and defining the cost of making choices is still the responsibility of the engineer and I don't believe the author was ignoring availability to suggest that we should too. In fact, I'm left with: "I find this proposal interesting, of course there are availability concerns, how might I solve those concerns from here?" The difference between that thought and the "straw-man" I pointed out is that the straw-man argument reads: "throw this all out because that system is only going to fail."
The trick is that SQLite can run single-threaded a massive amount of qps. So you can pump large amounts of ops in serial which is trivial to reason about and allows for in-memory caching on the API side with trivial invalidation. And by still keeping your web serving separate, you can still enjoy edge-performance of handshakes, and skip db altogether for static pages. (The article underestimates the RTT issue - it’s very real - real world apps need more round trips than you think)
Most of cpu outside of the db comes from parsing, deserializing, copying data, tls (both number of conns and encrypting data). By offloading the big chunks of these you can get easily 10s of thousands of writes/s on an entry level machine. Reads are faster.
That said, I think it’s always worth benchmarking the typical bottlenecks especially io. Providers are lying and misleading a lot, so just run your own on their free tier. Just make sure to have some integration test/bench ready in case you need to switch up.
You need to be on the edge, they say.
Be close to your users. Minimize latency.
How much of an issue is latency in practice?Here is an example. My book recommendation project Gnooks, which I run on a server in Germany:
Does it feel too slow for anybody?
Over the last years, I have gotten many thousands of suggestions from users on this project. Yet, as far as I can remember, nobody ever touched the topic of latency. While the largest group of users is from the USA.
The cool part is, you can test this reliably, and see if it matters. That'll get you a concrete answer.
Would like to hear how it feels from Australia compared to other websites. Like Hacker News, Goodreads and Wikipedia for example.
Fun aside related to latency: I've sshed from France to Australia, modal editors are so much easier over those latencies, as you can effectively queue up commands and round-trip time leaves you with enough time to plan you next key strokes.
From Europe that would give you 2x the latency an Australian user would see which should help illustrate their worst experience.
Yes, DNS is a different issue. To bring that down, I would have to use a distributed DNS I guess.
> No matter how you design your site, SPA, SSR, some hybrid in between you can’t get around that if there is at least one database query involved in rendering your page, you have to go back to your database in us-east-1.
https://aws.amazon.com/rds/aurora/global-database/
https://aws.amazon.com/dynamodb/global-tables/
https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Conce...
https://docs.aws.amazon.com/AmazonElastiCache/latest/red-ug/...
And so on.. Have fun kludging together the equivalents yourself.
Ten or twenty years ago I would have agreed with you: the internet was new, and when the site went down, people blamed the site, and things got nasty fast. Nowadays, though?
People are much more likely to blame their internet provider first, or just plain try again later. They're used to the unreliable nature of the internet now.
Unless you're Google, but how many businesses are Google?
As someone else on this thread said, simpler systems are less likely to fail anyway. Keep backups, replicate your DB to a cold DR site if you're truly worried. That will handle the majority of non-FAANG businesses in the majority of non-FAANG cases.
And if that failure occurs? So long as you don't lose data, it will be a short-lived blip for most businesses. Even being down for a few days can be weathered if you have good PR, provided that it's a once in a blue moon sort of event. For that matter, data losses can be weathered in many cases.
What was that saying? Each nine of reliability doubles the cost or some such?
That needs to be accounted for when determining the ROI. How reliable does your business actually need to be?
That depends on your target market. The default answer from IT, however, is an automatic "100%!". This is wrong, IMO.
We’re all on redundant internet connections now.
- if the site is down for more than a few seconds
- if you attempt to access the site during its downtime
- if you pay an ongoing subscription (this could very well be a non-profit site, or one that doesn't primarily drive its revenue from subscriptions)
- if you feel like unsubscribing
How much does this *really* affect a non-moonshot business?
The article explicitly restricts the problem space to top 1000 websites.
I've taken 30 seconds to browse through top 1000 websites and outside of the usual suspects up top, most of them look like https://www.pro-football-reference.com or https://www.tempo.co which to me are not far off blogs and local businesses.
If people stop using your site after it's been down once it can't be that good.
Now that's what I call moving the goalposts :) suddenly we're at an hour a day.
Github wasn't the first OSS hosting platform, yet they managed to steal the show, and let me tell you a secret, it wasn't because they had less downtime than others. But if you want to keep telling yourself five nines is of utmost importance to your startup, more power to you.
You are assuming that there is a single thing called downtime (not a scale of slightly bad to outrageous) and a single set of considerations as to whether to move.
Also it may matter less to HN which is a cult following. For a FAANGy type company of course it matters.
But keeping the site up also cost money. So context is really important here. How much money do you spend on the last % of availability vs how much money so you loose on downtime. In my experience people have a more complicated setup then needed because of availability, but at the same time have not tested a real DR scenario.
The article advocates storing data in local sqlite on single server. That means that when the server goes poof you will lose data. How much data, depends on how frequently you run backups. But it will be non-zero amount.
And backups bring the other point. It will never be just one server like the author proclaims. You need another server for the backups. Now you have two servers, and a dependency between them. The dreaded complexity is creeping in again, almost if it was inevitable
Sure you'll still loose some data but that's always the case whatever solution you choose.
I don't know but if there's a domain where the five nines are a pipe dream it's the cloud: we've seen a shitload of major outages where this or that cloud provider had services down for hours. This blows the five nines for years, if not decades (6 minutes max downtime per year: good luck with your five nines when you've got a 300 minutes cloud downtime).
The cloud is so unreliable people are now simply used to the Web not working: "Oh shit, it's down again, I'll try again later today or tomorrow".
That's how bad the "x nines" expectation became because of the cloud.
Customers ignore downtime when you're a service with big lock-in or no real alternative (GitHub, Facebook, etc). Because then they can't do anything about it.
When you're a small/medium service, and you're down for 3 hours during a time where 4 employees of a company need you for productive work, they ask what's up, and if they must expect that this will happen again. Generally, shorter, one-off downtimes are forgiven, but longer or recurring ones are not.
We run on Hetzner dedicated, and I think the best way to run most sensible software is to run on 3 or 5 beefy machines. If you only have 1 machine, the following will take you down, all of which have happened in my production systems (except the last, which I only read about in the news for 3 hosters):
* You need to reboot for kernel security updates. A reboot causes 3 minutes downtime because that's how long a reboot takes on server hardware, including Hetzner's (multiple POST style screens, netboot timeouts and so on, most of which make sense). Kernel updates come in approximately every 3 weeks. You may skip some of them that are not security-relevant for you, but this requires careful analysis, which is more effort than just installing.
* The mainboard dies. It needs to be replaced. Immediately causes 3 hours of downtime.
* Some disk or RAM dies in a way that prevents the mainboard from booting. This should not happen, but it does. Also 3 hours of downtime.
* One of these happens at 1 am on Saturday. You are in Europe or Asia. If you are not set up to be paged out of bed, your US customer will see 8 hours of downtime.
* The top-of-rack or datacenter building core router has some trouble, and it takes the hoster 3 hours to fix it. 3 hours of downtime.
* There's a fire or sprinkler flooding in your data center. It takes the hoster 2 weeks to clean up the physical mess. 2 weeks of downtime, unless you manage to re-build the infra somewhere else and restore from backups, then maybe 48 hours of downtime.
Note most of this also applies if you run a single VM in AWS. If you're lucky, the restore times will be faster by being a VM. But that is not guaranteed; I've had instances not do their work for some hours and afterwards got the email "We noticed a problem with the physical host of your VM, we've automatically migrated it somewhere else now".
Most of these problems go away if you have 3 servers (of the type the original post describes) instead of one, which is only slightly more effort.
Concern about availability is overblown, not irrelevant!
There are shades of gray, though. Most sites don't need to engineer their system to hide catastrophic database failures from the user. Not being able to tolerate a single web server failure, though... that's just careless, and will end up giving you a bad reputation.
You probably don't actually need five nines -- but three probably won't go amiss...
Very low internal latency from an on server database or directly adjacent database allows your software to run database queries orders of magnitudes faster and spend less resources per client served, resulting in a speedier, snappy experience for your end users compared to what any of these managed database services can provide in practice.
Users are not noticing the difference between 5ms lookups and 20ms lookups. They are noticing outages, data loss, and inconsistency.
I am not proposing writing any additional cross-region consistentcy code, many CRUD applications have added zero cross-region consistentcy code and just left this entirely on Postgres to handle since bi-directional replicas are now a core feature.
The more complicated you make your stack the more likely you will end up with outages, data loss and inconsistency. Using plain jane battle hardened, well developed tools that can serve the internet at scale from a single modern server gives you a very simple architecture where most of that architecture is getting significant improvements every quarter as the libre software community improves Linux, Nginx, Postgres, etc versus whatever random person Amazon is letting work on RDS or DynamoDB for the next year until they PIP them (as they almost always do).
So will your service running on AWS or equivalent. Complexity brings its own pitfalls and I don't think there's a web service in existence that never screwed this up.
That is if it's not AWS itself having a screw-up.
VPS is not worse off here.
But the article completely mises out on the aspect of availability, which would be quite risky with a single server.
Why not do mix & match smartly? A CDN is already "the cloud". Having your static content hosted (and pre-generated fully) in the cloud and only go on dynamic generated / rendered responses if you really have to. Then you have good latency, high availability and still low costs.
You can solve one box availability with box 2 (hot backup) - all within the same architecture and price structure.
The equatorial African nations aren't that much farther away than Great Britain.
Of course that does still leave out Southern Africa, India and Australia.
I realized that I didn't actually know, either.
https://mybroadband.co.za/news/wp-content/uploads/2021/08/Te...
I don't know how credible this map is, but it appears to match all the others I was able to find in a quick image search. I picked this one because it's larger and shows more detail than some of the others.
It looks like equatorial Africa is reasonably well connected to the Americas and Europe.
I'm sure connectivity is crap once you get out of the large coastal cities, but on the other hand it's pretty crap in Alaska and the Canadian interior, too (at least before Starlink came along).
Google grew up (early days) using cheap white box consumer PCs while "best practice" was expensive server boxes.
It's a tried and true method of budget hosting.
The mini PC world is exploding and they make for a solid low cost, low power server platform.
It also makes having hot and cold spares cheap and easy.
You can hear all about this if you turn up to Harry's at 7pm on Thursdays, or at many of the other tech events in Seattle. Lots of current and ex-Amazonians out there with stories!
Regardless, that wasn’t my point — the outages you mention do drive customer churn and cause sales problems. Ie AWS outages do impact revenue.
These are called Fight Club issues, and are restricted access issues that don't show up for most team members. Gotta cordon off the real hairy bugs!
I expect AWS outages affect revenue. But on the score of the number of companies… very few are in a similar position of AWS.
In this situation, blowing up your system complexity to maybe get another 9 makes no sense. Then the revenue change is pretty irrelevant for modest downtime.
People under estimate single server uptime. If availability is really that important, buy a hot backup. Put it in another region. Done.
Jk. But kinda. I have a little hobby website. Mostly for fun, I keep rewriting it in different languages, stacks, and deployment methods. At one point it was on a $5 AWS server. It continuously ran out of memory and crashed. Then it crashed for some other reason, I think because it ran out of disc space from log files. Then the postgres database got encrypted for ransom because I stupidly left it open without a password.
So now I use things like Fly.io and Firebase. And they work great for hobby stuff. I'd like for my projects to grow, and this article makes a good case why I should be competent enough to run them on a server myself.
At work I help run a much larger website that we run with K8s and a managed db. The idea of directly running that on servers seems equally daunting.
But I know it shouldn't be that way. Thanks for the reminder.
Not just for the obvious reason to grow your abilities, because you can say exactly the same about everything and you can't become expert in everything.
But simply because someone makes a lot of money off of you if you feel like you need them, and they are big enough to do many indirect things and change the entire environment to make you feel you need them and never even question it, and suffer essentially ostricisation if you ever do question it.
Making those baby sysadmin mistakes is perfectly fine. Everyone must make them. Are you ever going to forget to secure the access to any db after that? That is not only one specific conig but an entire class of problem that you are alert to now.
Not only is it ok because it's low stakes and then you know better for high stakes at work after that, but really even at work it should be normal to suffer a breakage once in a while, because at work it's even more important that you know how to recover when it inevitably happens anyway even without making mistakes. I don't mean break things on purpose, I just mean if you never suffer a problem, you never become prepared to deal with a problem. That is not good.
Besides, no matter what you still suffer equivalent problems, cloud or no cloud. What's the difference between your db going down from a hardware fault or misconfig, or a cloud account getting killed because of a billing or tos error?
Also, k8s actually makes an otherwise manageable system into a daunting one.
All in all, more and better reasons to be brave than to be afraid.
Of course, not wanting to go through all of this is a valid choice too, especially if you have no interest in running your own services.
Never start a fist fight with an Australian. Let alone 26 million of us at the same time.
The classic Australian peel assures a continuous pipelined fighting response.
And I'm not trying to suggest that we haven't learned a lot since then, but:
There was a time when the corpo web server/Internet computer existed as a pizza-box Sun Microsystems machine on someone's desk.
And sure, bandwidth was a lot less than 11MB/sec back then, but that's not a stretch at all for today's modern equivalent to that expensive pizza box.
But I'm old, and man do I sure as fuck remember the Web being a generally-unreliable turd back then. It was common for a website to not work today, or for an ISP's solitary email server to be down for a week or more.
It sure felt more real (and I even built a couple of those email servers), but it was also obviously very broken some of the time in ways that people don't generally accept today -- especially with 400 million visits per month, which was largely unfathomable at the time.
Or, working oppositely: Knowing the IP of the user (from IRC, say), it was often possible to discern the accounts's username. And since usernames were often factually-descriptive back then -- sometimes for billing purposes (firstnamelastinitial, say): It was easy to give them details about themself or their family that they would not have guessed were possible.
The greater Internet was a fun time back then for me as a teenager, and I never did anything particularly damaging with it even though I absolutely fucked around with things from time to time.
It is certainly fun to remember, but there are aspects of how things worked back then that I'm not itching to bring back to life here in 2024.
(And that includes the solitary pizza-box server on someone's desk, as well as the wooden racks of modems and thinkpads.)
absolutely do not miss those days
I'm over here remembering that /etc/passwd was world-readable (and shadowless) on a given system's [singular] shell box, but you were more like the man behind the curtain.
Most businesses are not running a SPA which only makes a database call. Most infrastructure I’ve seen is not web-facing. Managing this from a consistent place reduces operations overhead.
The idea that you’re SSRing everything instead of making a few API calls is strange to me. Most businesses in the top 1000 will have optimized for caching when at all possible. They’re also not paying sticker price to CDNs.
Offloading works to CDNs comes with inherent operational benefit.
This all seems great until you have your infra fail. Stories of unrecoverable outages abound. Even if you run in this architecture, you should always always always be using an external party for backup and log retention.
More of the above… this doesn’t interact well with the real world, but everything old is new again and many new CTOs /founding engineers haven’t seen how infra breaks and how operations impact throughput.
Even if the math is a bit off on the 11 MB/sec due to spikes he's got a point...
The world's population grows at a much slower pace than server processors. Heck, we're even talking about population degrowth now in many paces. Meanwhile AMD shall keep coming up with beefier and beefier servers.
For many non-FAANGs what couldn't be done 15 years ago on a single box is now totally doable on one. And this trend shall only continue.
If you're comfortable with a single point of failure.
> Why do we need Docker, serverless, horizontal scalability again?
The above and I don't ever have to run security updates or patch an OS.
You would have to update the base image to benefit from those updates though.
This has to be satire?
Have you been involved in the xz project in the last 2 years? Curious.
I don't know your specific use case though.
When your "stack" and deployment process involves multiple tools and services, then yes you need high availability. Simple things with careful design and implementation tend to break less often.
Use tested software bits, test your own bits and the odds of a failure drop quickly.
There was a story recently about running the workload of Twitter from a single machine: https://news.ycombinator.com/item?id=34291191
Yes - possible but no redundancy, difficult maintenance
If you want to handle 10x spikes over 11MB/s, then you need a single gigabit port. That's not hard to get, nor hard to feed.
I haven’t done that, though. It doesn’t seem widely available yet, or at least not for hobbyist programs like me.
Businesses need something predictable and repeatable(great if scalable too) that can be offered to customers at higher value than cost, and to make it repeatable businesses demand infinitely distributable architechtures. Otherwise you end up with just various forms of losses, missed opportunities, trust issues, foxhole problems, technical debts and such.
The idea of a single mainframe in a shed maintained by a UNIX wizard generating boatloads of cash would be awesome. The last part has proven difficult.
And maintaining a linux server is common knowledge. People have done it for decades. No wizardry needed.
https://stackexchange.com/performance
9 IIS web servers, 1 SQL server (plus one hot standby SQL) with 1.5 TB of ram and 2 redis serving 1.3 billion views a month. And if anything they've been scaling down their architeture over the years as hardware gets better and their software improves: in 2016 they ran 11 web servers for example. Notice how little CPU usage they've got, even in peak usage, it's really about I/O ie lots of RAM and fast storage.
So it depends on the use cases, but I do not think "I need the cloud" is the automatic answer to "I am making something for the web". In fact the vast majority of people will never even get close to the level of traffic SO has.
Of course, even with low traffic, some type of software would highly benefit from a distributed cloud architecture. Can't really make a search engine that indexes the whole internet on so little. But are you making the next google?
Their stack is certainly a lot, lot, lot cheaper to run in the long term than going with Amazon's blood sucking costs.
- very unreliable power
- the internet was shut down completely in the entire country during the last election cycle
Just those two problems make cloud hosting very attractive. No amount of cheap prices can make me switch to a physical server.
I 100% agree with the article.
Hetzner is the way to go.
Is AWS not widely known as one of the most expensive cloud providers all round?
And the cloud is not cheap in general, it's just convenient. If you do the math, most cloud providers only need a few months to make back the cost of the hardware, after that it's all money printing.
If you need something running 24/7 for a year+ it's always cheaper to buy... unless you're a large firm who doesn't trust their high turnover employees with their own infrastructure, which is where AWS makes their money.
Big companies don't use cloud just for fun.
And yes it's expensive but a ton of people working on that are good but not security experts or ops experts.
Managed shit takes the complexity out of this.
For everything non corporate yeah go with SQL lite if you like but you should have enough money anyway that those 'optimizations' don't matter.
One expert is not cheap and just coming up with that basic blog post requires a little tof expertise too.