memcached had a design similar to redis when I first got involved with it a few years ago. It was fine for a lot of systems, but didn't scale to the larger ones. It took a lot of work for us to get it to the point where it could saturate all of the cores you could give it.
If anyone can ever produce a benchmark where memcached does not perform better than redis, please file a bug in memcached, or at the very least, start a conversation.
The closest I've seen were tests that degraded to client tests. In one case, someone took a very poorly performing memcached client and compared it against a very well-tuned redis client and a couple of relational databases and showed that memcached was the slowest of all. sigh
In most of these cases, the test environments perform at the very least an order of magnitude slower than what I get on my development laptop (I have no trouble sustaining 90k writes per second with either of a simple c++ thing I threw together yesterday or my high level java client).
On real production-quality machines with good networks, 300-400k ops per second isn't hard to achieve. Even between two old "large" (2 core, 7GB) EC2 nodes, I've been doing 90k ops per second all day (which is somewhere around the point where they start to throttle my bandwidth).
I agree with you that if Redis performance is better than memcached there is to fill a bug or start a conversation, as there is something odd. Both are written in C and use multiplexing internally, against the same OSes. The performances should be very similar, or the memcached performances should be better as it supports less features.
Edit: p.s. can you please explain why running multiple memcached processes, one per core, was not good enough?
Having more machines reduces the performance gained from multigets. This is often a pretty big performance boost. This is especially useful with synchronous clients that serialize requests across servers (which are very common in many popular languages because popular languages don't seem to encourage proper concurrency). Four times the number of servers will quite often make requests take four times longer.
Some of the memcached clients will do additional work to hash the keys using an independent node location key just to ensure locality of related data so the values can be retrieved more efficiently.
About multi-gets, this is how Redis handle this, we have a concept called key tags, that I think it is very similar to the one you described. Basically a key in the form foo{bar} gets hashed in a special way, only the part inside {} is hashed ('bar' in my example).
So a client willing to optimize for multi get will use {user0}:name {user0}:surname and so forth. Well, not the perfect example as with Redis in this specific case you could use an Hash that is much more space efficient and will provide locality of related keys automatically.
The process is pretty straightforward, but as you walk down the path, you end up with something that looks a bit different from where you started.
Lock-free hash tables without GC are kind of hard, but not possible. I'd certainly welcome a lock-free engine if you're working on one. :)
To be clear: in memcached you can effectively set-and-forget keys and never worry about running out of memory. It uses an LRU to figure out what to throw away when memory gets full. In Redis if you want to achieve that behavior you need to expire your keys, which is another round trip every time. That means expire alone is also not enough, because you could set the key successfully but fail to set the expire. That means you need to also keep track of all your keys and periodically do cleanup.
Redis is awesome, but I would definitely not use it to replace memcached except in the cases where I absolutely need the data to survive a reboot and am willing to go through the hassle of manually managing the cleanup of every key.
What I think about Redis as a cache is that in many contexts the rich data types and operations supported allow better caching.
An example: oknotizie.virgilio.it is a large social news site, for this site I built a Redis cache where the "latest" news IDs are pushed into a Redis list, but they are also added to the MySQL DB that was formerly here.
So to paginate the "latest news" page I only use fast LRANGE calls against Redis, but if it will return a short read (the cache is empty since the Redis server was restarted) I'll use instead the MySQL server. This actually never happens and everything goes on Redis usually, but since there is no persistent storage needed in Redis side, it is actually a cache.
When a news is deleted I use LREM instead. And so forth.
Consider a website where 10% of the cache is hot: you'd discard a single hot entry only once every thousand cache evictions. I think for a lot of sites that is likely to be Good Enough, depending on the impact of a cache miss.
Bonus: you can trivially tweak how sensitive that is by increasing the number of items to randomly sample.
For plain old caching I haven't (yet) had a need to use redis, which is younger and by nature more complex. RAM is so cheap nowadays, we just tend to stuff two memcached machines with 64G, which goes a ridiculously long way.
On the other hand, as antirez pointed out, redis shines when you need "a little more than caching". We have rolled a few custom queueing solutions on top of redis (similar to resque) and working with lists and the SUBSTR operation is an extremely pleasant experience.
Redis seems to be the optimal "roll-your-own-queue" toolkit at this point. We had some strong delivery- and persistence-guarantee requirements in our project but those were not a problem to meet with a simple WAL on the clients and a redis-pair. So far our solution wins over the previous rabbitmq setup big time in all areas (performance, reliability, complexity, transparency). It could have been done with memcached, but less efficiently due to the lack of the aforementioned list-operations.
Also there is no way to mount a precise LRU schema that is O(1). O(log(N)) complexity is needed.
Trying to achieve something similar using Redis by for example updating the expire on every read would be a bit of a pain and I think put extra load on the client and server. This isn't a complaint about Redis though, it has its own use cases which are wonderful and I have no objection to using both for their strengths.