Redis on Acid: 0.5M ops/sec, 1ms latency and ACID compliance
redislabs.com
redislabs.com
This where the durability argument has a problem as well.
At a first glance, this looks like an un-replicated system, which means that the loss of an instance is an availability nightmare.
The worrying quote from the article for me is this one
>> With Redis Enterprise, it is easy to create a master-only cluster and have all shards run on the same node
Another node has to be brought in and attached to the same network disk to restore access to that key range?
500k+ ops/sec is nothing to laugh about on a single node with 1:1 read-write ratios, however the fragility of this system is concerning.
Half a decade ago, I was working with row-level atomic ops in ZBase/Membase (set-with-cas[1]), which gets away with using replication instead of an ssd backing the durability of operations + an fsync - the 99% latencies were at 3-4ms, but the scalability and availability were baked in.
[1] - https://github.com/zbase/documentation/wiki/Data-Integrity
No changes are written to the AOF so I guess the operation is atomic.
I still wonder what the details are around the "almost" in "almost always" and stand by the conclusion that "almost atomic" is not the same as "atomic".
- How does oss redis compare to enterprise redis ?
- What are the differences when running the same benchmarks with both oss redis and enterprise redis ?
- what is the marginal utility of an additional cpu core/thread ? that is, what happens if I run those benchmarks on an AMD ThreadRipper ?
Re the second and third questions - I'm afraid I cannot answer. I hope the author of the post will reply tomorrow morning (it's the middle of the night here).
- Redis Enterprise adds some enhancements to the Redis storage layer (details are in the blog).
- This benchmark only tested the Redis Enterprise. The idea was to show how fast Redis (Enterprise) can run on a single node with ACID transactions and still keep sub-millisecond latency. Note the hidden point - at the moment you cannot achieve sub-millisecond latency over the cloud (any cloud) storage, I mean persistent storage that is attached to an instance and not the local storage, which is ephemeral by design. So in-order to see how far we can go we decided to test it over Dell-EMC's VMAX that doesn't have these limitations
- In theory adding CPUs/cores can of course help, as you can add more shards to the Redis cluster and increase the parallelism when accessing the storage. That said, we haven't tested it over AMD ThreadRipper.
[1] RedisConf17 - Building High Performance DB with Redis using Flash Memory - Cihan B. & Frank O. https://youtu.be/Tf8JRsE6w2U?list=PL83Wfqi-zYZF1MDKLr5djmLYU...
write coalescing - caching individual writes for later bulk inserts into a db to help insert rates.
Simple counts and counters which can be incremented and then purged/reset on interval.
Simple Pub/sub for medium throughout systems.
Many more applications than simple key/value.
In general, an extra layer on top of Redis is a must to make it even simple searchable database. Especially indices can't be ad-hoc implemented with separate Redis commands. Only transactions or Lua snippets that update both data and related indices atomically can avoid data corruption. For my own use, I wrote: https://github.com/ilkkao/rigidDB It supports indices and strict schema for the data.
I think somebody has described Redis at some point as a database-SDK. Makes sense to me.
Redis is great because of its multiple data structures. Depending on their "kind", these JSON objects are either `APPEND`ed onto Redis Strings (e.g. for time&sales or order history) or `HSET` (e.g. opening/closing trade) or ZSET (e.g. open order book).
Sometimes an object transitions from a SortedSet to a String. We used to handle this with `MULTI` but now we use custom modules to do this with much better performance (e.g. one command to `ZREM`, `APPEND`, `PUBLISH`).
We run these Redis/feed-processor pairs in containers pinned to cores and sharing NUMA nodes using kernel bypass technology (OpenOnload) so they talk over shared-memory queues. This setup can sustain very high throughput (>100k of these multi-ops per second) with low, consistent latency. [If you search HN, you'll see that I've approached 1M insert ops/sec using this kind of setup.]
We have a hybrid between this high-performance ingestion and long-term storage. To reduce memory pressure (and since we don't have 20 TB of memory), we harvest these Redis Strings into object storage (both NAS and S3 endpoints) with Postgres storing the metadata to facilitate querying this.
We also do mundane things like auto-complete, ticker database, caching, etc.
I love this tech! It's extremely easy to hack Redis itself and now with modules you don't even need to do that anymore.
Some of it was done with raw Redis commands, and for some we wrote a data-store engine with automatic indexing and optional mapping to models - https://github.com/EverythingMe/meduza
We also used Redis for geo tagging, queuing, caching, and other stuff I forgot probably. It is very flexible, but requires some effort when not used as just a cache.
Right now I'm developing a module for Redis 4.0 that can do full-text and general purpose secondary indexing in Redis. https://github.com/RedisLabsModules/RediSearch
[1] http://hyperdex.org/ [2] https://www.cs.cornell.edu/people/egs/papers/hyperdex-sigcom... [3] http://rescrv.net/papers/warp-tech-report.pdf
Hyperdex's problem all along was that the author — a very talented developer from what I can tell — seems more invested in his projects from the perspective academic research (he's at Cornell) than in delivering a practical, living open source project. He tried to form a company around Hyperdex (the transactional "Warp" add-on thing was commercial) even though nobody seemed to be using it; and he was the sole developer. Unfortunately, as interesting as Consus is, history seems to be repeating itself there.
But yeah, Hyperdex seemed to have real potential at one point. It was the only NoSQL K/V store (at the time) that had transactions.
What about FoundationDB?
Almost always achieved? That actually means it is "never really achieved", fixed for you.
It is just sad that cheap marketing materials like this one are keep being pushed to the front page of NH.
So yeah, I guess you could use redis as main datastore.
* you can tolerate some data loss in case of system crashes/power failures/etc, because redis only flushes data to disk periodically [1]
* your dataset can fit in RAM, since redis is an in-memory datastore
Also since Redis is memory backed you need more RAM then data. This can get very costly.
Another annoyance is that Redis is single process single threaded so you really have to avoid running long running queries unless you do extensive manual sharding.
(Disclaimer: it's been a few years since I had to think about these constraints so maybe some are removed in more recent versions of Redis)
This allows long running queries to do primitive cooperative multi-tasking, releasing the GIL and letting other queries have a chance; But there is no real parallel data access. You will only gain real parallelism if you have actual work to do that does not touch the data directly when a thread is not touching the GIL. There aren't many cases that this applies to - usually copying the data aside to do work on it is not worth the gain of parallelism.