HNHacker News
TopNewBestAskShowJobs

pushkarg

13 karma · joined February 23, 2024

submissionscomments
pushkarg··on IKV: embedded key-value store, 100x faster than Redis
The benchmark highlights the value prop to the end user who can now use a "local program" ie an embedded database and get the perf wins over Redis - without worrying about data management (backups/replication/what-not).

Read path don't invoke any RPCs (even on startup or a cache miss). Writes need RPCs - since they have to be propagated to the (many) readers.

pushkarg··on IKV: embedded key-value store, 100x faster than Redis
I just explained how implementing the Redis protocol will inherit inefficiencies of Redis in IKV - so I don't understand how the comparison will be fair or honest.

Hope this addresses your original question about why we wrote a custom benchmarking client.

https://github.com/inlinedio/ikv-java-client/blob/master/src...

It is quite simple and is available here. There is nothing malicious in there to make IKV appear faster. Although ff you do see a bug, I am happy to fix and republish results.

pushkarg··on IKV: embedded key-value store, 100x faster than Redis
Point taken. Not opposed to self-hosting, its just that we haven't seen interest in that so far. And the way things are right now, it can be very complex to setup.
pushkarg··on IKV: embedded key-value store, 100x faster than Redis
Exactly! Accessing data in RAM will be orders of magnitude faster than over the network (maybe this is a good sanity check of our benchmark numbers). The core principle behind IKV is that it allows you to access (large) data-sets without network calls.

Most DB tooling out there only works for a client-server model.

And implementing the Redis protocol would imply changing our architecture significantly and negatively affect performance (ex. a producer-consumer queue to serve requests, ser-deserialization costs).

pushkarg··on IKV: embedded key-value store, 100x faster than Redis
Correct, we are super early so there is no self-serve yet.

The primary usecase for this is serving features for ML inference (since eventual consistency is ok and sacrificing write latency for reads is a fair tradeoff). Right now, this is done by using a traditional client-server DB at the moment (Redis/DynamoDB/etc) - or if you're a big tech company that cares about latency you can implement this on your own (https://doordash.engineering/2022/05/03/how-we-applied-clien...).

As far as self-hosting goes - yes writes will be definitely faster. IKV is fully open source so we're not opposed to it, just haven't figured out the details yet (since self hosting will mostly be useful to very large usecase)

At the core, we use in-memory hashmaps that reference memory-mapped files. So, when a dataset doesn't fit in RAM - it spills to disk automatically.

Cold start - the database is seeded with a "base image", that is built periodically by the backend. That's how a user can add new nodes to their cluster, and still avoid any RPCs.

That being said, if you don't have 10TB of disk, you have to partition IKV (and by extension your application). We support partitioning by allowing documents (the data) to declare partitioning keys. If one shard/partition cannot fit on disk - the store won't startup.

pushkarg··on IKV: embedded key-value store, 100x faster than Redis
Its a managed embedded-store, ie someone can write data, forget about it, come back in a month with new hardware and still access all their data. You can't do that with a traditional embedded store (ex. rocksdb or a local redis instance)

There are data pipelines behind the scenes to distribute writes to the embedded store which needs some provisioning time. And yes we are super early.

pushkarg··on IKV: embedded key-value store, 100x faster than Redis
All benchmark details are linked in the Github readme. https://docs.google.com/document/d/1aDsS0V-AybpvXEwblBlahGLp...

For Redis - The benchmark client and remote server (ie AWS ElastiCache) were in the same AWS availability zone (us-west 2a) to minimize network latency as much as possible.

pushkarg··on IKV: embedded key-value store, 100x faster than Redis
IKV is a managed DB solution ex. a user does not need to build data replication or backup pipelines. A "local" Redis instance is not fully managed - the instance goes down - your data is unavailable (unless you build a full distributed system around it).

The benchmark is done considering what an end user will use directly in their application. The performance gain comes from avoiding remote network calls and minimizing serde.

The landing page is under construction :)

pushkarg··on IKV: embedded key-value store, 100x faster than Redis
Fully-managed; Eventually consistent; Embedded (no RPC for read-path); In-memory with option to spill to disk.

Detailed benchmarks linked in the Github Readme. Single digit microsecond read latencies.

Written in Rust - available for use in Java and Go.