> Is your reasonably small bank known for its reliability and fault tolerance?
I don’t know if they have a specific reputation for it. But their services were very reliable and resilient.
> The main reason banks don't go offline is because the core critical infrastructure is running on 50 year old mainframes that no one is allowed to touch because all of the greybeards with actual talent who made them are pushing up daisies.
This isn’t really true at all. The banks central ledger will likely be running on DB2, and that will almost never fail. But most of their infrastructure runs on more “modern” systems (DB2 is still modern if we’re being honest, it’s just more specialized). When you swipe your card somewhere, nearly all of the systems the transaction passes through are ordinary web services, running on ordinary architecture. That’s where the greatest risk of service disruption is. It’s just the final side effect is updating a record in a DB2 database somewhere.
Most banks also outsource a lot of the core system maintenance directly to IBM. A typical software or infrastructure engineer in a bank will have absolutely no contact with those systems.
> Nowhere did I say you should be on one incredibly large server
No, you just seem to think that reducing the number of components in your infrastructure is the key to reliability. If you want fault tolerance, then you generally want redundancy, and you want your services to fail gracefully (ie, not take down other non-dependant services when they do). Both of those things require deploying more servers.
If points of failure are a concern, then you also need to account for the fact that every single time you make a manual change you are creating a new point of human failure (with humans generally being the least reliable part of any well designed system). If you automate a change, you can spend more time scrutinizing it for mistakes, and be reasonably sure it’s deployed exactly the same everywhere. Combine that with blue/green, rolling or canary deployments, and automated testing, and you end up with something much more reliable than shelling your way across the entire infrastructure to deploy a single change. To suggest that deploying less infrastructure is a superior alternative really just comes across as ludditism.