"I trust Redis-on-disk every day less." --Salvatore
groups.google.com
groups.google.com
For Redis on disk to work well, you need:
1) Very biased data access.
2) Mostly reads.
3) Dataset consisting of key->value data where values are small.
4) A dataset that is big enough to really pose memory/cost problems on the ever growing RAM you find in a entry level server.
What is left of Redis semantics and advantages here? The intersection of 1+2+3+4 is small and fits exactly in the case where Redis for metadata, or as a cache, plus another datastore designed to work on disk is the right pick. So why should not we focus, instead, into doing what we already do (the in-memory but persistent data structure server) better? It would be already an huge success to enhance what we already have (not to say that Redis is so important, just that Redis can do his small part in the big picture providing something simple that works well, instead of trying to do everything and save the world).
See ma, no ass-covering, no excuses, no cooking benchmarks.
Actually in the persistence side we plan to do more work to make Redis better, for instance Redis 2.4 that is entering release candidate can save/load most databases on disk ten times faster.
In the future we plan to explore a new AOF format that is more compact and faster to process, and the ability to rewrite the log without a background process (BGREWRITEAOF).
What's more reliable? Data persisted to disk on a single server in a RAID 1 array, or data stored in memory of 10 machines which are completely isolated from each other? What about 100 machines?
It's wrong to think of disks as reliable/persistent and memory as not. You should think of both as stores where neither is 100% reliable. Memory is far less reliable than disk, much in the same way that a single disk is less reliable than RAID-6. However, just like you can make disks more reliable by replicating across other disks (and then even more reliable by replicating across multiple servers...in multiple locations), so too can you achieve reliability with memory.
EDIT (clarify what they are talking about):
They aren't talking about persistence, though I can see why you think that from the OP title. They are talking about loading the data set into virtual memory when it doesn't fit into memory.
This has nothing to do with persistence. Antirez isn't saying he doesn't trust Redis' persistence implementation. He's specifically talking about how the Redis VM handles more data than available memory.
He's saying: always have enough memory for all your data.
So the short answer is "no", you don't lose everything. You just lose the data that was written to memory since the last sync to disk.
Bad idea or no, it's what they do - performance gains are large. There is an ersatz logical log in the form of application logs, and these have been used to piece together transactional information before.
My use-case is processing XML files that are pushed by a third-party (see [2] for a blog post describing the setup).
One drawback is that the AOF file will keep growing (at least on my version) and could reach the maximum filesize of your system, if any - there's the BGREWRITEAOF command available to work-around that issue (not tested).
[1] http://redis.io/topics/persistence
[2] http://blog.logeek.fr/2010/8/2/on-jruby-resque-and-windows
In addition you can make Redis asynchronously save to the disk every so often.
If you want the cost/performance characteristics of disk storage, use MySQL or Postgres. If you want them and the speed and flexibility of the Redis data structure server, use Redis + SQL as a one-two punch.
We use Redis for "live" data, and SQL for archival/report/versioned data.
We have plenty disk-based k/v stores already. If that's what I need then tokyo, riak, etc. are my friends.
Redis plays in a slightly different and equally important game. The cluster-/sharding-features seem like a more natural path to scale out the use-cases where it shines.
This has nothing to do with persistence. Antirez isn't saying he doesn't trust Redis' persistence implementation. He's specifically talking about how the Redis VM handles more data than available memory.
He's saying: always have enough memory for all your data.