The Raft Consensus Algorithm
raftconsensus.github.io
raftconsensus.github.io
Raft is a consensus algorithm designed to be easier to understand than Paxos. To measure Raft's understandability, we conducted an experimental study using CS students at two universities. We recorded a video lecture of Raft and another of Paxos, and created corresponding quizzes. This page makes our materials available for anyone interested. We think these are valuable resources for anyone learning consensus (whether Raft or Paxos or both).
[0] - https://ramcloud.stanford.edu/~ongaro/userstudy/
Raft Video - https://www.youtube.com/watch?v=YbZ3zDzDnrw
Paxos Video - http://www.youtube.com/watch?v=JEpsBg0AO6o&feature=youtu.be
I'm not saying that it isn't genuinely simpler. Just that in discussions about raft, I see most people blindly assuming that it is. But I don't believe that it has been proven yet. I think this is the nuclear power plant part of the bikeshed discussion. [1]
If you scrub to the 50 minute mark of this talk[2] by Leslie Lamport (creator of Paxos), my coworker David Varvel asks him this question and he gives more or less the same answer.
[1] http://en.wikipedia.org/wiki/Parkinson's_law_of_triviality [2] http://channel9.msdn.com/Events/Build/2014/3-642
I believe most developers just want to consume a working library. It seems likely that the plethora of Raft libraries and dearth of Paxos libraries is because it is easier to write a Raft library, but I also don't think it matters _why_, it only matters that there are good Raft libraries.
A quick google search brings up libpaxos and other open paxos implementations, so it appears that they exist.
I'm curious: what are these edge cases you're talking about? I've worked on a Raft implementation, and I don't remember any particular edge cases (though it's been a while).
There's definitely some interesting choices for when to replace nodes in the face of failure, but I don't think that's a Raft algorithm shortcoming.
The core raft state machine comes in under 600 LOC last time I checked and the docs are here: https://godoc.org/github.com/coreos/etcd/raft. There are a few non-etcd projects that have started working with it, most notably the cockroachdb folks have started to use it to implement the "multi-raft" that they need for their DB. You can find the discussion and work on github.com/cockroachdb.
It turns out getting a raft implementation that is easy to audit with a copy of the paper in hand and easy to test in a fast and consistent way requires a significant engineering effort. I hope that based on the lessons learned from goraft that this implementation is something that a number of Go projects can use and leverage. Particularly since we carefully separated the WAL, RPC and state machine from each other.
Description:
https://github.com/basho/riak_ensemble/blob/develop/doc/Read...
Code:
https://github.com/basho/riak_ensemble
Ensemble module, the way I understand it, provides ability to have "consistent" CP (in CAP theorem sense) updates in the otherwise default AP database setup in Riak.
Disclaimer: I am not affiliated in any way with Basho, just follow riak_core module development because I think it is a piece of very cool engineering.
Try Corosync/Pacemaker.
I believe most developers just want to consume a working library.
Right. Maybe they'd even prefer not to have to bother, they just want the availability thing solved for parallel service instances. Perhaps the more interesting question is, "how do you neatly fit raft/paxos like functionality at deployment time in to the average SOA service development process?" (subtext: without making devs learn either algorithm)
I think the answer there is an improved devops process that discourages any form of assumption around infrastructure type and can easily run failure tests at various levels of the deployment topology.
Some of my thinking in the area is documented at http://stani.sh/walter/pfcts
Isn't it formally proven correct? And "simpler" would be subjective right? So what do you mean by "proven to truly be simpler than Paxos"? Thanks, and also that talk looks really great.
https://ramcloud.stanford.edu/~ongaro/thesis.pdf (Chapter 8)
If you can read the raft paper and come away also knowing how group membership changes (wow) and log compaction works, in additional to the "basic" log replication consensus, that's huge.
edit: I'm sure it's still 'hard' to implement, like bringing a brand new replacement machine up from scratch (copying the full data set over, to warm him up, and then starting replication efficiently), i.e. there's still a lot of coding to do
I saw a good talk about RAFT given by a Basho engineer, video and slides are here: http://www.meetup.com/Erlang-NYC/events/131394712/
The speakers Erlang implementation: https://github.com/andrewjstone/rafter
"Raft is a consensus algorithm that is designed to be easy to understand. It's equivalent to Paxos in fault-tolerance and performance. The difference is that it's decomposed into relatively independent subproblems, and it cleanly addresses all major pieces needed for practical systems."
"It's equivalent to Paxos in fault-tolerance and performance. The difference is that it's decomposed into relatively independent subproblems, and it cleanly addresses all major pieces needed for practical systems."
One of the authors of Raft is Prof. John Ousterhout(TCL creator, among other things)