What's the big deal about embedded key-value databases?
notes.eatonphil.com
notes.eatonphil.com
[0] https://en.wikipedia.org/wiki/DBM_(computing)
This codebase shows how SSTables, WAL, memtables, recordio, skiplists, segment files, and other storage engine components work in a digestible way. Includes a demo database showing how it all comes together to make a RocksDB / LevelDB competitor (not really).
[0] https://pragprog.com/titles/tjgo/distributed-services-with-g...
Dgraph wrote a great explanation article about why they wrote Badger and the tradeoffs and designs reasons: http://web.archive.org/web/20181116033431/https://blog.dgrap...
[0] https://github.com/dgraph-io/dgraph/graphs/contributors
I don't work with a lot of data, but typically my decisions base on basic factors and purpose:
PostgreSQL - SQL, structured data, cannot scale horizontally
MongoDB - NoSQL, unstructured data
Redis - key-value, distributed cache
I get it that you can replace storage engine and you can theoretically get more performance, but in practice compatibility and standardization is more important, because a lot of products (including third-party) will already use PostgreSQL/MongoDB/Redis, so it's no-brainer to use it as well for your solution.
However for me to pick RocksDB or some other, new, shining database/storage engine, there would have to be more compelling reasons.
What's also interesting is the trend of newer distributed "database systems" like Vitess[0] or SpiceDB[1] that forego embedded KV stores and instead reuse existing SQL databases as their "embedded database". Vitess leverages MySQL and SpiceDB leverages MySQL, PostgreSQL, CockroachDB, or Spanner. Systems built this way get to leverage many high-level features from existing databases systems such that they can focus on innovating in even higher-level functionality. In the case of Vitess, it's scaling, distributing, and schema management of MySQL. In the case of SpiceDB, it's building a database specifically optimized for querying access control data in a way that can coordinate with causality across multiple services.
Storage engines are different levels of complexity based on the query requirements. Simple K/V stores can run circles around Postgres/MySQL as long as you don't need the extra features.
Think of it as a high performance sports car like a Ferrari. It's not good at taking the kids to school or buying groceries. But if you need to prioritise performance at the expense of all other considerations then it's exactly what you need.
* DynamoDB and the Dynamo KV store
* LMDB (embedded kv)
* Dgraph (distributed graph db) and its embedded kv store BadgerDBApplication
|
MySQL
|
RocksDB
|
Filesystem
In general the MySQL layer is doing all the convenient stuff for application developers like supporting different queries and datatypes. The RocksDB layer is optimizing for performance metrics like throughput and reliability and just treats data as sequences of bytes.
https://minervadb.com/index.php/2018/06/06/a-friendly-compar...
Here's another example of a realtime database which uses RocksDB under the hood: https://rockset.com/blog/how-we-use-rocksdb-at-rockset/
And this is very cool, distributed SQLite with FDB: https://univalence.me/posts/mvsqlite
https://rockset.com/blog/converged-indexing-the-secret-sauce...
Rockset's converged indexes + denormalisation means you can have fast querying.
The list of examples are odd. For instance MongoRocks is cited for using RocksDB, but actual stock MongoDB uses Wired Tiger, which isn't mentioned.
Disclosure: I played a part in the late-beginning of this space when Netscape funded Sleepycat to develop BerkeleyDB. dbm and ndbm existed beforehand, but BerkeleyDB used in LDAP servers is I think the genesis point for this pattern as it exists today.
Right, I'm waiting for standard for a level above relational databases which is Object-databases. I know there are several ones already and there are Object-Relational mapping layers.
I think the key point there is that Object databases are a level ABOVE relational databases. They are not "better" but they deal with the higher level of objects rather than "tables", just like relational databases can be seen to be are a level above key-value -stores.
I would like Object databases to become better and easier to use and more standardized.
I think there is value in being able to see both level, the objects, and the relational data that makes up the objects.
1. Yugabyte's relational query layer sits on top of a document store (DocDB): https://www.yugabyte.com/blog/how-we-built-a-high-performanc....
2. You can put documents in a PostgreSQL JSON(B) column.
I think that is what object-to-relational mappers like Hibernate do.
I think it would seem quite natural to implement objects on top of, with the help of an RDBMS. But not sure if the opposite is true.
But you're also welcome to write your own post. :)
The need for parallelism killed the first approach and the cost of increasingly complex reduce steps killed the second. Now we're back to "how much can we fit in RAM on a local machine" and it turns out, if you can still bang bits for smart key formats, a hell of a lot.
I immediately thought of Kafka's streaming query stuff when I read the headline (ksqlDB). I'm not sure if that's the origin story of RocksDB, but it's the storage engine underlying that streaming query tooling in Kafka's ecosystem.
[0] https://engineering.fb.com/2021/08/06/core-data/zippydb/
Edit: I've added Redis Enterprise Flash to the list now. Thanks!
This field is pretty interesting when you're talking about performance vs space amp vs write amp vs read amp.
My whole container started to look like a sparsely populated relational table where every row/column intersection could have multiple values (e.g. a photo could have a tag for every person in the picture attached). I started experimenting with using the KV stores as columns to form regular relational tables.
It turns out that it was relatively easy and was extremely fast. I started building tables with 50+ million rows and many columns and performing queries against them. Benchmarking the system against other databases revealed that it was very fast (and didn't need separate indexes to accomplish this).
Here is a video showing how it does a bunch of queries 10x faster than the same data stored in a highly indexed table in Postgres: https://www.youtube.com/watch?v=OVICKCkWMZE
Also - no mention of LMDB? RocksDB and LMDB feel like the ones that stand out in that field - levelDB definitely had a reputation for corrupting data.
The built-in Data Explorer is a must-have for me and idk of any other database that has something similar.
I used RethinkDB back in the days because it was the first DB that had pretty good replication and sharding - it was zero effort. I felt the functional programming model to be strange, some stuff got executed locally, other parts remotely and it was not very straight forward when things didn't go as planned.
By the time the RethinkDB company folded, CockroachDB emerged and has been my go-to distributed DB since.
They don't use a key-value store library.
I know it's a bit of a fine line. But I'm talking about standalone libraries people embed across different applications/databases. That's what RocksDB/LevelDB/Pebble are.
[0] https://github.com/rethinkdb/rethinkdb/tree/v2.4.x/src/btree
> Yeah, there's no workaround that I can find for 3.4 (duplicate effects), 3.5 (read skew), 3.6 (cyclic information flow), or 3.7 (read own future writes). I've arranged those in "increasingly worrying order"--duplicating writes doesn't feel as bad as allowing transactions to mutually observe each other's effects, for example. The fact that you can't even rely on a single transactions' operations taking place (or, more precisely, appearing to take place) in the order they're written is especially worrying. All of these behaviors occurred with read and write concerns set to snapshot/majority.
While foundationdb uses SQLite I didn't otherwise think of SQLite as being relevant here. :)
Kafka Streams is the first kind; the source-of-truth storage is HA (as HA as the Kafka topics it's backed with at least) but can only be queried with high consistency when the consumer is active, and it goes down for rebalances when you scale out or fail over (and in many operational setups also when you upgrade).
For an example of the second kind, see Fly.io's Litestream explanation - https://fly.io/blog/all-in-on-sqlite-litestream/.
That being said, I think the etcd etc. examples are just meant to be in contrast to stock Redis or Memcache, which offer very little HA support, generally just failover with minimal consistency guarantee.