But people tend to want to distribute their app with 5k monthly users to 10 really weak instances instead of just 2 medium ones or 1 big one. And in that, they think it makes it more reliable but fail to account for the added complexity of distributed architectures and start to fail when doing deployments and changes instead.
Another part people underestimate is the performance of dedicated instances vs "virtual" instances. Many times I've seen people shocked at how last applications work when running on dedicated instances, and how far you can scale with them as you won't have to upgrade as often as you think. Cloud is mainly meant (for me at least) for software you need to be able to scale really fast up AND down, not just up as you get more users. But most businesses I've worked for, with and started myself, never really had that need.
To give a bit of an estimate on how far basic infrastructure can carry you; the entire RetroArch project[1], which I'd say has about the same userbase as any small tech company is entirely hosted, buildbot (so CI) included, on a single 40$ DigitalOcean instance.
Most software is suprisingly efficient/resilient if you just configure it properly. The only place where I've seen horizontal scaling as an actual necessity vs. just properly configuring the tools you're given in most cases is Rails. There's no real efficiency fix for Rails tooling from what I can tell. It's just slow as molasses and you're gonna foot an ever expanding server bill as your userbase grows.
[1]: No endorsement.
Services falling over themselves for no reason is really uncommon in my experience, and usually are because A) introduced changes, B) hitting performance bottlenecks that are really bugs in the code or C) downtime of 3rd parties that someone failed to account could actually be down so that brought down "our" application too.
The DB is the hard part (not with automation but that's more to learn), and storing the files to be seen by all nodes (if app needs to create files during its normal work) is, but in most cases delegating that part to cloudy cloud isn't a problem.
> In reality, complexity of deploys and introducing new changes are way more likely to shut down your service than anything else, unless you're dealing with very large scale. And for those problems, it doesn't matter if you have 1 or 10 instances, shit would go down anyways as you push out your change to all of them.
Right, but updating OS is one of those changes, can't do that hitless if there is only one node in the system. And you do update your OS (or whatever FROM you got your container), right ?
Sure, most downtimes will be more "developer fucked up" than "something happened with hardware". But those fuckups are usually on smaller scale than "well, server needs to be rebuilt from scratch/backup/CM manifest" and "just" need revert.
> And for those problems, it doesn't matter if you have 1 or 10 instances, shit would go down anyways as you push out your change to all of them.
If you have tens of instances you can be responsible developer and do some kind of staged deploy to cut those fuckups by order of magnitude or two.
But honestly if you want to not do ops, outsource it instead of half-ass it, there are plenty of options
> And you loose
*lose
Running on multiple instance ties you to reliability of the hardware and reliability of the whatever software you use for cluster, which is also heavily dependant on your skill with that.
So it will entirely depend on what you run.
If you say run a cache server (Varnish or whatever else) behind some LB instead of single Varnish instance behind same LB, skill required to set it up is zero, extra software is zero so you end up with more reliability. We had exact zero problem with those kinds of setup for decade+ and it never caused any failure.
On other hand, if you use something like Pacemaker (our usual config was pair of DRBD + DB nodes managed by Pacemaker) it is absolute nightmare to "get it right" and every mistake on your part in configuring it will make it less reliable, and there is plenty of edge cases to think of and handle, some might not even exist when you start configuring it.
For example we had a hiccup at one point where initial config had too short timeouts to shut down database and it assumed it hanged, then killed it.
Other times DB occasionally failed to start because pacemaker was (back in SysV days, no systemd) running /etc/init.d/db start & status in quick succession, and Java didn't managed to write PID yet which made status return "db is down", confusing the resource manager.
We now have a bunch of setups like that also running rock solid but it took a bunch of time and setup to get it right. No extra cost to set it up really coz we have that in CM but for a time it definitely was lower actual uptime than "just a node with DB"
And there are elasticsearch clusters that are just more stable with 3 nodes than one just because sometimes index could shit itself on crash...
For low traffic stuff, might seem OK, but no redundancy for a service that users pay for seems irresponsible, however Gitea might just be for them to track internal stuff, so maybe a little less so.
But again, when I've seen distributed infrastructure, it's very uncommon to see issues hitting one instance while others are not being hit by the very same issue.
When you introduce change, you introduce it to all running instances, all at once or over a period of time. You might be in luck if you're doing blue/green deployment, but you can achieve the same thing with a staging + production environment mostly.
Same with OS upgrades, I can't remember the last time a OS upgrade completely borked anything I have deployed. At worst, the kernel didn't boot but recover from a backup was fast enough.