Raft Consensus Animated (2014)
thesecretlivesofdata.com
thesecretlivesofdata.com
I've also been trying on-and-off again some different techniques for doing the visualization as I'd like to do more of these. I'm currently looking at trying to make it work with Remotion[1]. The JavaScript version I did for Raft was time intensive and I ended up having to write an entire (albeit terrible) implementation of Raft to even get it to work. lol.
It appears you actually did implement it!
Thanks!
This webpage cleared some long-standing doubts about what distributed computing means, what a consensus algorithm is and what his Raft thing is.
Kudos to the developer. You got a newbie interested in the field!
- Is Raft alone in this space, or are there other popular algorithms/libraries that fill the same space?
- What happens when the node count gets larger than a handful? What happens when you hit hundreds or even thousands of nodes, that are trying to achieve consensus? In particular, the part where all of the nodes respond (semi) simultaneously to a broadcast node. In a radio spectrum world, that would be a disaster. N:1 communication slots are choke points for timely communication.
If you just need eventual consistency, CRDTs are also possible.
Going in the other direction, if you don't mind the latency full consensus with global locking, you could just do that.
- You should not have too many nodes to make a decision, this is usually reserved for leaders; if you have a large distributed system you may clusterize them or forward decisions to leaders, whom decide for consensus. If you clusterize, the leaders for each node can also be selected by consensus. If you can't do any of those then having a consensus protocol might not even be a good idea; you'd end up with a sort of merkle tree (or some sort of blockchain) to make sure all the data is registered, or maybe audit transactions. In any case this[1] might be interesting.
[0] https://en.wikipedia.org/wiki/Paxos_(computer_science) [1] https://doi.org/10.1016/j.neucom.2016.10.011
to give an obvious answer, blockchains are one of the methods used for trustless consensus (imagine how things could go wrong in RAFT's case if the leader was malicious)
The 'right answer' is going to be Paxos and it's many flavors; a few references to Lamport, Google's Chubby, "Paxos Made Real". If the right people are on hn today you're going to get a few references ZAB (Zookeeper Atomic Broadcast) and View-Stampted Replication. Oh, and go watch all of Heidi Howard's talks.
The problem spaces of byzantine and non-byzantine consensus are far enough apart ... well kinda like Atomicity in postgres and atomics is C++. You decide if that's a lot.
But like I said, I also think that as a class of problems — byzantine consensus is just super different from the normal consensus problems paxos and raft are addressing. Way more than they might seem at first. Enough to be an off-topic to the point of just being naive or misleading.
Is there a better way to proceed with tentative consensus, until a majority cluster can be realized, and then have a conflict resolution strategy? People operate this way.
I've played around with Hashicorp Consul on "edge boxes" - long-haul, wirelessly-connected embedded computers, with unreliable power supplies. Allowing edge boxes be Consul leaders results in all kinds of mayhem: split brain situations, corrupted state, stale DNS resolutions (Consul handles DNS as well), cats and dogs living together, mass hysteria. A much better topology is to have 3 server nodes on a LAN as the "head cluster" and letting all the edge boxes be clients of the head.
I haven't used it but Consul has a multi-datacenter mode, which I believe is designed to better handle such a situation, which I believe has a dedicated raft cluster per datacenter.
https://learn.hashicorp.com/tutorials/consul/federation-goss...
It's been a while since I read "Designing Data-Intensive Applications" though, but there's a chapter on that.
The same hunt also surprised me that there is no common way to do leader election among pods in Kubernetes.
https://kubernetes.io/docs/reference/kubernetes-api/cluster-...
Also, I personally think the current blockchain literature is much more intuitive and easier to follow, for learning about consensus. The Byzantine case isn't really that different than the crash case if we assume cryptography. On the other hand, Raft is a spiderweb of a protocol, very easy to get wrong.
Raft Visualization - https://news.ycombinator.com/item?id=25326645 - Dec 2020 (35 comments)
Raft: Understandable Distributed Consensus - https://news.ycombinator.com/item?id=8271957 - Sept 2014 (79 comments)
A couple of questions:
1) In the case of a network partition, the client that is currently connected to the leader, do they get notified that there's a partition, or that the cluster is not in a healthy situation?
2) If a client writes to the partition that will get rolled back, and all their transactions get rolled back after the partition heals, do they get notified that their data was rolled back?
The cluster - or any server of the cluster - finds out about network partition only when the timeout passes. At this point the leader - which becomes the former leader - can notify the client, or the client can see for itself that the timeout has passed.
> 2) If a client writes to the partition that will get rolled back, and all their transactions get rolled back after the partition heals, do they get notified that their data was rolled back?
Note that the client was never notified that their data was committed in the first place. So the client can assume that if the timeout passed without notification that the data wasn't set in the cluster.
Surely there could be problems between the client and the leader. Idempotent messages could be useful.
If implement raft on top of S3 I can kinda see this working. Is there a sensible "file system on top of S3 like storage" out there already?
That's fascinating. Got more information on that?
Thank you!
ETH has do deal with at least one thing that Raft doesn't have to deal with... bad actors trying to inject bad data into the system, also known as the Byzantine generals problem [1].
But the slow nature of the introduction to the elements on each incremental click is a bit irritating.
I'd recommend static image(s) with legend/highlights for each node and message, etc. And animations for each relevant scenario illustrated.
I really like this, but not being able to go slightly faster with arrow keys was aggravating.
Cool explanation though!
Granted, this animation hails from 2014, so the above might not be possible without significant effort.