You've completely elided the quorum thing to make it sound like TFA is nuts.
You've completely elided the quorum thing to make it sound like TFA is nuts.
I don't think TFA is nuts. I do think the premise that engineers using cloud can ignore CAP theorem is wrong, though. It's a decision you must consider when you're designing a system. As long as everything is running well, it doesn't matter. You've got to make sure that the behavior when something breaks is what you want it to be. And you can't outsource that decision. It can't be worked around.
You don't because their clients won't also be isolated from the other servers. That's TFA's point, that in a cloud environment network partitions that affect clients generally deny them access to all servers (something you can't do anything about), and network partitions between servers only isolate some servers and none of the clients.
You can engineer cloud services to be resilient. AWS has done a great job of that in my personal experience. But worst cases can still happen.
If you have an algorithm that runs in n*log(n) most of the time but 2^n sometimes, it can still be super useful. You have to prepare for the bad case though. If you say it’s been running great so far and we don’t worry about it, fine, but that doesn’t mean it can’t still blow up when it’s most inconvenient. It only means it hasn’t yet.
My goodness. So many commenters here today are completely ignoring TFA and giving lessons on CAP -- lessons that are correct, but which completely ignore that TFA is really arguing that clouds move the needle hard in one direction when considering the CAP trade-offs, to the point that "the CAP theorem is irrelevant for cloud systems" (which was this post's original title, though not TFA's).
What does this mean? What is the cloud being down? The cloud isn't a single thing; it's lots of computers in lots of data centres.
You've misunderstood the problem.
You cannot be consistent if you cannot resolve conflicts, taking writes after a network partition means you are operating on stale data.
Honestly the closest we've ever gotten to solving CAP theorum is spanner[0], which trades off availability by being hugely less performant for writes. I'm aware spanner is used a lot in google cloud (under the hood), but you won't solve CAP by having hundreds of PGSQL replicas, because those aren't using Spanner.
In fact, you can test it out, ElasticSearch orchestrates a huge number of lucene databases, a small install will have a dozen or so replicas and partitions, but you're welcome to crank it to as many nodes as you want then split off a zone, and see what happens.
I'm becoming annoyed at the level of ignorance on display in this thread, so I'm sorry for the curt tone, you can't abstract your way to solving it, it's a fundamental limitation of data.
[0]: https://static.googleusercontent.com/media/research.google.c...
Cloud is always available? I have a bridge to sell you.
The cloud isn't what's being described as available. The service is available or not. The service is hosted on multiple computers that talk to each other. If they stop talking to each other, the service needs to either become unavailable (choosing consistency) or risk having not up to date data (choosing availability).
In the same spirit, some businesses are fine with a single db, making backups every night and losing some data when an issue happens. It doesn't mean that people maintaining these systems get to tell others that distributed systems are essentially a non-problem.
You have to consider the system as a whole including external clients since the definition of the availability of the entire system ultimately depends on the point of view of the end users of the system.
The CrowdStrike incident has demonstrated that no matter what you do it is possible for any distributed system to be partitioned in such a way that you have to choose between consistency and availability and if the system is important then that choice is important.
Partition intolerance means you can only proceed when all nodes have been reached. This means partition intolerance not only deals with availability in the yes or no sense, but also in terms of latency.
The author of the article doesn't actually understand the CAP theorem.