Do most of your customers expect close to 100% uptime? Yes.
Does one machine provide the uptime required by your customers? Most definitely not.
Do most of your customers expect close to 100% uptime? Yes.
Does one machine provide the uptime required by your customers? Most definitely not.
But I also think that many applications are far more tolerant of small outages than you imply. And a more complex setup also adds more points that can lead to an outage even though it reduces the chance that hardware will cause one.
If you are in one of those two categories then setting up something more robust than a single server isn’t much more complicated than a single server.
That is to say, I think it very much does change the equation around the original statement.
I kinda want to know where you shop, because commodity hardware has a reputation as just utter crap. Before cloud went mainstream, I worked somewhere that had some racked servers in a colo literally catch fire. (The remote hands put them out, disconnected, and refused to touch them ever again.)
The correct high-availability solution should take business requirements into account and there is no silver bullet. Running everything on a $5 VPS is no silver bullet, but neither is your typical "cloud-native" "best practice" stack that everyone keeps cargo-culting which often leads to unnecessary cost while leaving many hard questions (such as replicating CAP-bound stateful databases) unanswered.
If we're talking about reliability/outage recovery, we're considering the application as one single unit visible from the external client's perspective - so everything including the DB (or equivalent stateful component) must be redundant.
Sadly this is also where a lot of cloud-native tooling and best practices fall short. There are endless ways to run stateless workloads redundantly, but stateful/CAP-bound workloads seem to be ignored/handwaved away.
I've seen my fair share of stacks that are doing the right thing when it comes to the easy/stateless parts (redundancy, infinite horizontal scalability), but everyone kinda ignores the elephant in the room which is the CAP-bound primary datastore that everything else depends on, which isn't horizontally scalable and its failover/replication behavior is ignored/misunderstood and untested, and they only get away with it because modern HW is reliable enough that its outage/failover windows are rare enough that the temporary misunderstood/unexpected/undefined behavior during those flies under the radar.
And so no, most teams don’t need to worry about the hard problems you bring up.
And spinning up a new VM typically takes longer than a minute. Even on cloud providers.
Our environments are easily repeatable, they're very maintainable, but they're not quick to start up
I think this is mostly a self imposed requirement. Banks regularly have overnight technical breaks. My country national rail has 30 minutes of downtime every night (!).
Unless your product is already global, most people won't have a problem with occasional overnight scheduled downtime.
Most people never even test number 2. If your a novel or niche product, you can afford a “lot” of downtime
For number 3, it really all depend on your load. I’ve run most early stage startups on an incredibly simple setup with a proxy placed in front of an app server.
Occasionally, our monitoring will start reporting high p99’s so we go in and bump the servers to the next tier or scale horizontally just a bit more. Eventually, that breaks down but many startups will be well into Series A or Series B point. At that point, you know your customer’s needs and can hire dedicated engineers to solve reliability and uptime.
Also, I sure hope your loadbalancers don't suck. I've worked with some that had worse uptime than my servers, and worse capacity too. Went back to putting the two host addresses in DNS, which is mostly fine, but means a lot of waiting when you want to take a host out of rotation for disruptive maintenance.
(Regular maintenance meaning OS updates needing a reboot, which happens like what? Once per month for a minute?)