Also HN: "Database X can fail in Jepsen tests during partitions, it's unusable"
I see HA as a mandatory part of what makes a system production ready...
Also HN: "Database X can fail in Jepsen tests during partitions, it's unusable"
I see HA as a mandatory part of what makes a system production ready...
Most websites are not Google or Amazon, and gain very little from more 9's past the initial 2~3, which is totally manageable on a single server. Why spend thousands setting up and managing distributed servers if going from 99.0% to 99.999% gains you hundreds? If you need more nines and stronger guarantees than that, absolutely, use distributed solutions, nobody is telling you otherwise.
Jepsen tests are valuable once you have decided to use a distributed solution. They showcase problems that would be extremely hard to track down in production settings but can cause serious problems there. If encountered in practice these problems can cause data corruption or other very strange bugs that would take days if not weeks to track down.
I wouldn't say that a system with failures in Jepsen tests is unusable, every system has bugs, but if a system has these problems on purpose (e.g. to make it look better on benchmarks) and there is no movement to fix these problems then I would definitely steer clear of that system.
"How much money?" Is probably the most import question, as it actually helps in evaluating if there's a business case for investing into the financial commitment, engineering effort and maintenance burden of an ha solution.
On a different note, I've seen mysql-based multi-master setups and it's really a joy to be able to treat a mysql crashing as a minor annoyance instead of a call to fire-fighting.
Edit: there is also more to be said about this. I'd research crawl-budget if you're interested.
It's likely the additional overhead worrying about HA is not worth it, not to mention the real possibility of HA just not working properly in actual failure scenarios
It seems like his business could at least afford an HA setup on a cloud provider. Moving to any hosted db with backups, updates and redundancy could be worth it.
As for HA failing, it's still less likely than not-HA failing.
Also an HA setup allows for maintenance and upgrades without downtime, which is way better than the very common "we don't upgrade if it's working, because it could break".
Also there is a third option between HA (with its increased cost and complexity) and "we don't upgrade if it's working, because it could break", which is "take the site down for a few minutes, do the upgrade, bring it back up". It's not for every site, but there's a range of sites for which that is fine.