http://blog.foundationdb.com/call-me-maybe-foundationdb-vs-j...
Sadly, fdb has been bought by Apple[1], and you can't download it anymore. I sincerely hope foundationdb gets opensourced or something.
[1] http://techcrunch.com/2015/03/24/apple-acquires-durable-data...
Jepsen tests cannot be used to prove the system is safe, but to prove it isn't.
It looks like very often he looks into source code in order to figure out how the system operates, that way he can find weaknesses and write test to his testing framework to demonstrate the issue.
I wouldn't trust any company that uses Jepsen to show that their product is safe.
Deleted comment
[1] https://github.com/rethinkdb/rethinkdb/issues/1493#issuecomm...
> Even though the Postgres server is always consistent, the distributed system composed of the server and client together may not be consistent. It’s possible for the client and server to disagree about whether or not a transaction took place.
If you want to avoid that kind of problem, use 2PC.
Do you (or the author) see this as a bug, or just something that might surprise people who haven't thought through the guarantees?
Quoting the article:
> The 2PC protocol says that we must wait for the acknowledgement message to arrive in order to decide the outcome. If it doesn’t arrive, 2PC deadlocks. It’s not a partition-tolerant protocol. Waiting forever isn’t realistic for real systems, so at some point the client will time out and declare an error occurred. The commit protocol is now in an indeterminate state.
It would be foolish for a client to issue a COMMIT and then assume the transaction aborted because of a connection drop. The client should wait until the connection can be reestablished and determine the real transaction state before making a decision based on it.
It's the same issue as with a power failure during fsync. The durability of that transaction is indeterminate, but it doesn't matter because the system is down. Before the system comes back, it will go through recovery, and either find the commit record or not, thus getting back to a determinate state.
"During the network partition, no requests are successful" is not the best result for a CP system, IMO."
HBase should provide partial availability in the face partitions.
Also, if you follow the rest of the twitter conversation you may realize, as they did, that only requests to the minority partition are unsuccessful - which is exactly what you want from CP.