How we implemented replication in Memgraph DB – Our approach and challenges.
memgraph.com
memgraph.com
There is some basic knowledge on how to achieve the replication, but, in our case, a lot of it was improvised and we're not sure what is the best approach.
We tried to minimize the overhead (keep as little information as possible) and that was more or less the only guidance we used.
I know that Paxos is considered hard to implement correct, but is there a reason to believe that in the general case, the more strict guarantees we need the more complex the algorithm has to be?
In this case, replication consists of communication between 2 instances, and when one them fails you need to pick consistency or the availability.
Paxos and Raft introduce consensus using multiple instances communicating with each other to mitigate those problems. Implementing correctly communication with 2 instances is hard enough, implementing it correctly for n instances is for sure harder.
In Memgraph, we actually tried to implement Raft and we did manage to some degree, but the effort was too much, especially when you look at the results we got (protocol that didn't work for many edge cases).