Seriously though, payment infrastructure actually is known for its reliability and fault tolerance. Banking systems don’t go offline terribly often.
Seriously though, payment infrastructure actually is known for its reliability and fault tolerance. Banking systems don’t go offline terribly often.
Nowhere did I say you should be on one incredibly large server, nor that you should have a single point of failure. That wouldn't be simple, either, because it would fail to support the prime directive, or would require a great deal of gymnastics to. It's about balance. You don't need thousands of servers to make a reliable system.
I don’t know if they have a specific reputation for it. But their services were very reliable and resilient.
> The main reason banks don't go offline is because the core critical infrastructure is running on 50 year old mainframes that no one is allowed to touch because all of the greybeards with actual talent who made them are pushing up daisies.
This isn’t really true at all. The banks central ledger will likely be running on DB2, and that will almost never fail. But most of their infrastructure runs on more “modern” systems (DB2 is still modern if we’re being honest, it’s just more specialized). When you swipe your card somewhere, nearly all of the systems the transaction passes through are ordinary web services, running on ordinary architecture. That’s where the greatest risk of service disruption is. It’s just the final side effect is updating a record in a DB2 database somewhere.
Most banks also outsource a lot of the core system maintenance directly to IBM. A typical software or infrastructure engineer in a bank will have absolutely no contact with those systems.
> Nowhere did I say you should be on one incredibly large server
No, you just seem to think that reducing the number of components in your infrastructure is the key to reliability. If you want fault tolerance, then you generally want redundancy, and you want your services to fail gracefully (ie, not take down other non-dependant services when they do). Both of those things require deploying more servers.
If points of failure are a concern, then you also need to account for the fact that every single time you make a manual change you are creating a new point of human failure (with humans generally being the least reliable part of any well designed system). If you automate a change, you can spend more time scrutinizing it for mistakes, and be reasonably sure it’s deployed exactly the same everywhere. Combine that with blue/green, rolling or canary deployments, and automated testing, and you end up with something much more reliable than shelling your way across the entire infrastructure to deploy a single change. To suggest that deploying less infrastructure is a superior alternative really just comes across as ludditism.
What makes banking systems reliable is that there's so much built in latency in every operation that nobody really notices if something is down for six hours.
This is true for settlement, but not authorization. When you spend your money it could be hours or days until the transaction is correctly reflected in your balance. But in order to spend it in the first place the transaction has to be authorized, and this happens in real time. Depending on where you live, who you bank with, what kind of card you’re using, how you’re using it... this authorization may rely heavily on your bank being available for online transaction processing, or not very much. But unless you’re doing a transaction with one of those imprint devices, your transaction must be authorized in real time by somebody.
I carry some cash too. But the times I’ve needed it were for things like an outage of the internet connection at a restaurant, or an issue with a payment terminal provider (or for the very uncommon cash-only merchant).
> Nowhere did I say you should be on one incredibly large server, nor that you should have a single point of failure. That wouldn't be simple, either, because it would fail to support the prime directive, or would require a great deal of gymnastics to. It's about balance. You don't need thousands of servers to make a reliable system.
Heh those things go offline every night for 2 hour maintenance.
Also fun when those nice single points of failure crash (they do)
Source: work at a shop trying to _get out of_ greybeard mainframe to get more reliablity