Call me maybe: Redis redux
aphyr.com
aphyr.com
What I proposed is a toy system as described here: https://gist.github.com/antirez/7901666
It is a toy as it has a super strong coordinator that can partition away instances, that is never partitioned, that can reconfigure clients magically and so forth. Under the above assumptions the theoretical system is trivially capable of reaching linearizability I believe.
Aphyr tested a different model, with an implementation that is not capable to even guarantee some weak assumption about the model (for example, the slave reset the replication offset to 0 when restarted), so I'm not sure what the result means.
I could test the actual model I proposed even with Redis, but manually following the steps that I outlined in the Redis mailing list thread. The point was, if you can guarantee certain properties in the system, there is always the "transfer" of data and higher offsets to a majority of replicas, and the system becomes strongly consistent.
The properties are hard to achieve in practice once you try to move the features from the mythical super strong coordinator into the actual system, and this is why, for example, Raft uses epochs and other mechanisms to guarantee both safety and liveness.
Unfortunately the focus is in showing other people are wrong without even caring where the discussion is headed.
--- EDIT ---
Btw now that I read the full post carefully, Aphyr also cherry-picked parts of the thread to construct a story that does not exist, like if I was going to implement strong consistency into Redis based on the proposed toy system that was only useful to show that WAIT per se was not a system, but just a building block. Note that yesterday I wrote the opposite in my blog, that there is no interest in strong consistency in Redis Cluster.
Very unfair IMHO... I read only the analysis part at first, and was thinking this was just a "let's check anyway this model with the current implementation".
* Our main business is low latency, so very few will use synchronous replication.
* Redis cluster is composed of multiple master-slave systems that hold a subset of the key space each. This means that if you need a majority of replicas to promote a new master in order to achieve consistency, after a partition you end with different hash slots having the majority in different sides of the partitions perhaps.
* With the current weaker consistency guarantees Redis Cluster can elect a replica that is isolated by the others but is in the right side of the partition, where the majority of other masters are. Similarly you can have just two total nodes for every hash slot, a setup that I believe will be very used, and still get some availability if the master fails.
So the consistency premises and the tradeoff of Redis Cluster are not exactly compatible with strong consistency.
Is WAIT still a useful tool? I believe yes, if documented as such:
WAIT in the context of Redis Cluster / Sentinel is not able to provide strong consistency, but lowers the percentage of probabilities that a failure mode that results in data loss happens.
A trivial example, what happens if a client is partitioned away with just a master? Without WAIT there is a window of NODE_TIMEOUT to lose data, with WAIT there is not this problem in this specific failure scenario. There are still a number of failure modes for WAIT, but are a smaller number compared to the full set of failure modes that there are without WAIT, so practically it does not feature strong consistency but provides the user with a smaller probability of data loss.
Another example: manual failover for very important data of some kind. You may have 5 nodes and use WAIT 4 to always write everywhere and improve durability.
And so forth.
That is kind of my feel. Redis is an outstanding product with a beautiful code base. This replication feature has been tough though. It is kind of due to external factors as I've mentioned in the previous post. Everyone and their cousin are talking about distributed databases, everyone likes CAP, CRDTs, Vector Clocks, Raft, Zookeeper and so on. It is hard to come up and say "Here I have made this custom replication protocol". Everyone stares and asks, "Hey where is your whitepaper or your partition tolerance tests?". 5-7 years ago, there would be only nods and approvals. The other aspect is this is about a database, so it is potentially toying and touching user's valuable data. If that gets lost either by a bug, mis-communication in docs, bad default, anything, it will not be taken lightly.
In the end I think it is fine to have it as what it is, with the warnings and disclaimers that data could be lost and avoiding papering over or hiding issues.
As an extra side note, simply put partition tolerance is hard. Net-splits are the devil of the distributed world. Some claim it doesn't exist or doesn't happens often. Others fear and tremble its name is mentioned. When it does happen it means having to resolve conflicts, throwing away user data, stopping killing your availability to stop some from accepting writes in order to provide consistency. This is a tough test (that Aphyr runs) and not very many databases fair well in it. But it is good these things are discussed.
[2] http://www.infoq.com/presentations/partitioning-comparison
Overall it is very good for the industry, better for Aphyr to lose one of his integer key value pairs, make a blog post about it, than for your health insurance to lose your health records.
People use Redis, in large part, for its time and space guarantees on data structure operations. (Without those, you may as well be using a serialized object store.) Strong consistency requires rollbacks; and the book-keeping necessary to do rollbacks throws away the time and space guarantees. So either you have Strong Consistency, or you have Redis, but you don't get both.
But Redis Cluster is a compromise: something which is roughly good enough for most cases people actually use Redis for, while failing horribly at things Redis isn't used for anyway, and still providing Redis's time and space guarantees.
Theorists balk, because there are obvious places where Redis Cluster falls down, and they can demonstrate this. Engineers shrug, because Redis isn't being used in their companies in such a way that those demonstrations are relevant to their problems.
Most people who need Redis Cluster have already Greenspunned a Redis Cluster themselves, and they're already happily living with the compromise it entails. They'll gladly hand the support burden of writing cluster-management code upstream to antirez; it won't change any of the facts about the compromise.
When that option is available it tends to beat complex software solutions in every way.
> I wholeheartedly encourage antirez, myself, and every other distributed systems engineer: keep writing code, building features, solving problems–but please, please, use existing algorithms, or learn how to write a proof.
That person should be sure to note these experimental results:
> These results are catastrophic. In a partition which lasted for roughly 45% of the test, 45% of acknowledged writes were thrown away. To add insult to injury, Redis preserved all the failed writes in place of the successful ones.
If you're unhappy with the formal interpretation of your gist, publishing the formal interpretation you intended (or, better yet, code) would allow others to build the system you actually intended.
2) There was never any plan for it to go inside the Redis implementation. This was just an argument in a mailing list to show that synchronous replication as implemented by WAIT is dependent on the rest of the system.
What bug report we are talking about?
If your argument is #1 (that Aphyr tested the wrong thing), then a reasonable reply would be to provide the model you did intend. If you tested it as well, that would be great, but it is reasonable to require the "bug submitter" to retest.
If your argument is #2 (that WAIT is just best-effort replication, and that it does not provide any guarantees) then that's fine, just say so clearly. But you should then stop disputing the model that Aphyr tested, because to do so implies the existence of a model which does provide guarantees.
It's amazing how far you can get with a reliable leader election protocol in a distributed system. We're getting to the point where all of the other stuff is just picking specific algorithms with specific tradeoffs, much like you do today in a non-distributed setting.
Has anyone attempted to use Pacemaker to wrangle Reddis instances?