Redis persistence demystified
antirez.com
antirez.com
Thanks for writing these!
Is this not true or at least no longer true?
> Status code reply: always OK since SET can't fail.
So, what kind of situation are you envisioning?
redis> multi
OK
redis> set hi 'hello'
QUEUED
redis> rpop hi
QUEUED
redis> set okay 'done'
QUEUED
redis> exec
1. OK
2. (error) ERR Operation against a key holding the wrong kind of value
3. OK
redis> get okay
"'done'"
redis> get hi
"'hello'"
redis>
Note that it keeps on processing, even though the second command fails. This also has nothing to do with persistence, as the exhibited behavior will be the same in a fully-alive system and from disk and anything else, because redis maintains a perfect total ordering of operations in its log format and sync behavior.Thus ideally, the (horribly slow) disk doesn't even come into play, especially for in-memory DBs. You buffer the data in memory before they are sent out of the machine/datacenter, but you make sure to mirror this buffer at multiple separate physical machines (which your database cluster should support), in case one goes down or over. Once the data are committed into a replicated store, you can clear that buffer. Fast and reliable.
This is not to say that there are zillion cases where harddrive is still the ideal persistence device. After all, it's very hard to destroy a harddrive in a way to make the data unrecoverable (of course, I'm talking about cases where RAID failed or wasn't present). However in reality, data from broken harddrives are seldom attempted to be recovered, mainly I guess due to the price and relatively long service waiting times.
If there was the time though, I'd love to see LevelDB[1] merge the gap between in-memory and on-disk. Inspired by Google's BigTable, all keys are kept sorted on disk (see: SSTable[2]). Keys, lists, sets, sorted sets and hashtables could be encoded in sorted key-value form and would be reasonably (though not tremendously) efficient to retrieve, especially if the query can be converted to a range query (i.e. list retrieval or set intersection). Keep hot data in memory, the rest ends up securely on disk.
Yes, Cassandra is based on the same lineage but it's not simple or clean to operate. Redis is a pain free setup, features a simple API but has no transparent way to overflow to disk for cold data. LevelDB is an optional backend for Riak but I must admit I've not explored Riak heavily... Have I missed a contender from another database crowd?
[1]: http://code.google.com/p/leveldb/ [2]: http://www.igvita.com/2012/02/06/sstable-and-log-structured-...
[1]: https://github.com/inaka/edis/issues/2 [2]: https://github.com/srinikom/leveldb-server
http://sqlite.org/atomiccommit.html
Check out in particular the sections "Hardware Assumptions" and "Things That Can Go Wrong"
The upside of the design is that it makes things simple (e.g., transactions, append only file).
(This is my understanding, please correct me if I'm wrong!)
However there are people that turn Redis async replication into synchronous replication with a trick: they perform:
MULTI
SET foo bar
PUBLISH foo:ack 1
EXEC
Because PUBLISH is propagated to the slave if they are listening with another connection to the right channel they'll get the ACK from the slave once the write reached the slave. Not always practical but it's an interesting trick.Can you expand on this? Specifically:
-Do you mean 'bound' as in 'limited by' or 'known'?
-Why are RDB snapshots I/O bound when other systems are not?
-Why is this an advantage?