Show HN: Badger – an embeddable, persistent key-value store written in Go
github.com
github.com
Not sure how I feel about that philosophy. Transaction support is extremely complex and should be left to a storage engine if ever possible.
What is the benefit of this over something like Berkeley DB?
One can take Badger and build a transactional layer above it, and it would still be below the application layer.
Transaction support comes with a set of performance compromises, and different companies may want to make different compromises.
Then comparing it to Rocks DB is unfair (infact title of this post is misleading). Stuff like transactions and MVCC is way complicated than it sounds.
TransactionDB* txn_db; Status s = TransactionDB::Open(options, path, &txn_db);
Transaction* txn = txn_db->BeginTransaction(write_options, txn_options); s = txn->Put(“key”, “value”); s = txn->Delete(“key2”); s = txn->Merge(“key3”, “value”); s = txn->Commit(); delete txn;
Have you tried benchmarking on baremetal though? I've been shopping around for a KV store and came across this benchmark: http://www.lmdb.tech/bench/ondisk/#sec7
Interestingly, for 768-byte values, LevelDB is the clear winner when it comes to read scaling on a VM yet on bare metal it's drastically outperformed by LMDB.
anyway, not strictly an answer to the question you asked, but b/c I've spent so much time doing cgo I figured I'd mention it as a possibility.
I haven't had a chance to benchmark on baremetal. We chose to run the benchmarks on Amazon, because that's what most people would use to run the applications/DBs built on top of the KV store. This ensures that the hardware specs used for benchmarking are easily available to everyone.
It's really about as simple, and as bare bones as you can make a traditional KV storage while still preserving MVCC, transactions, and maintaining ACID.
Furthermore it's design is prefectly suited to exploit Optane/FRAM non-volatile RAM devices that appear to becoming.
LMDB does mmap while LevelDB does plain file I/O.
Every `malloc` under XEN (that can't be satisfied with unused memory already mapped into that processes userspace) requires a hyper call.
This VMM approach was _largely_ depreciated about 6 years ago when Intel created page table bits that let the bare metal OS give the VM a _range_ of memory it can virtually re-map, but AWS hasn't adopted this.
http://fallabs.com/kyotocabinet/
http://fallabs.com/tokyocabinet/
https://en.wikipedia.org/wiki/Tokyo_Cabinet_and_Kyoto_Cabine...
(The names contain Cabinet but they aren't related to city/national governments.)
I once played around with Kyoto in Lua but can't say anything good or bad about it's performance, scalability, etc. They have documentation and some presentation slides in English on their website and provide quite a lot of language bindings if you are interested.
[1] http://samza.apache.org/learn/documentation/latest/container...
RocksDB does go to some lengths to reduce amplification on SSDs, though it seems this design is taking that into consideration from the very beginning.
You can follow the discussion here: http://bit.ly/2tb19eX
Can you have multiple OS threads created at the beginning so whenever an OS thread is blocked the OS scheduler simply schedules another OS thread? That's how most databases handle blocking disk IO, right?
Doing blocking disk IO has to be a common task. How do other golang applications/databases handle this issue?
Go does that even now. If a thread is blocked, it would create another thread, but it does it with a small delay. I saw a visible increase in IOPS when starting Go with more OS threads (using GoMAXPROCS), because Go doesn't need to wait before having access to these threads, but on a longer running job with enough initiation for OS threads, that benefit would be nullified.
Most LSM based KV stores are designed to avoid random lookups -- and Go DBs tend to use RocksDB -- so it probably hasn't been such a big issue for them.