HSE: Heterogeneous-memory storage engine designed for SSDs
github.com
github.com
I am curious about the durability and how well tested all of that is though. On the one hand, filesystems put a lot of work towards ensuring that bytes written to disk and synced are most likely durable, but OTOH Micron is a native SSD vendor so they've probably thought of that.
I'm also curious whether RAIDing multiple SSDs together at the block layer and then running HSE on top of that will be faster or whether running multiple HSE instances (not the right word, it's a library, but you get what I mean) with one per drive and then executing redundantly across instances would be faster. Argument for the former is that each instance would have to redo the management work, argument for the latter is that there's probably synchronization overhead within the library so running more in parallel should allow for concurrency and parallelism gains.
That's why all the hyperscalers all have their own custom SKUs.
All of the SSDs that this software might be deployed to have power loss protection capacitors to ensure the drive can flush its write caches when necessary. So this software only needs to make sure that the OS actually sends data to the drive instead of holding it back in an IO scheduler queue (as you point out, they're already bypassing the FS layer). Since this software should be pretty good at structuring its writes in an SSD-friendly pattern, the operating system's IO scheduler should probably just be disabled.
The default on RHEL/Fedora has long been the noop scheduler. I'd be surprised if other distributions haven't followed suit, given the prevalence of SSDs.
Their benchmarks show significant gains compared to RocksDB.
> https://github.com/spdk/rocksdb
But what I'd really like to see is a comparison against RocksDB using SPDK
> https://dqtibwqq6s6ux.cloudfront.net/download/papers/Hitachi... Based on these results, SPDK performs significantly better than the kernel requiring only 1-2 cores to saturate IOPS on an NVMe SSD (compared to the kernel requiring 16)
[0] https://news.ycombinator.com/item?id=10511960
[1] https://news.ycombinator.com/item?id=22266503
However, there's still lots of reasons to use SPDK. Performance is still significantly better[0], and you can directly access all the of the NVMe features on the device without going through any abstraction layers.
The potential of using io_uring with eBPF in Linux will makes any HPC enthusiast drooling :-)
First, the IO threads in RocksDB are a thread pool that assume they perform blocking operations. That doesn't jive with SPDK's async model (nor io_uring's). We're having to message pass to an async thread and block on semaphores in the IO threads.
Second, RocksDB was heavily reliant on the kernel page cache to make it fast.
Both of these things could have changed since we last worked on the integration. I haven't kept up with RocksDB development recently.
source: am SPDK maintainer
Is there a maintained version of SPDK RocksDB? Or SPDK any DB?
I'm a long time Aerospike user with no connection to Aerospike.
Edit: I would love to see a benchmark with Aerospike.
Also, I have to wonder how narrowly "open-source storage engine for SSDs" is being defined here such that it excludes so many earlier storage engines in claiming the title of "first".
And of course RocksDB has had similar graphs to showing perf against other systems.
Every system manages to find a benchmark that fits their narrative :)
Reality is both RocksDB and WiredTiger are high performance storage engines, and they are both optimized for SSDs too. These type of benchmarks rarely tell real story.
> HSE optimizes performance and endurance by orchestrating data placement across DRAM and multiple classes of SSDs or other solid-state storage.
Orchestrating data placement? Isn't that what all storage engines do?
What am I missing here? Is this a block level rather than file-system level driver?
https://en.wikipedia.org/wiki/Hierarchical_storage_managemen...
https://en.wikipedia.org/wiki/ZFS#Caching_mechanisms:_ARC,_L...
PR is insane hot air referring to another hot air product (can you even buy their 3D Xpoint devices yet?)
Nope. The only product they've announced so far using 3D XPoint is the Micron X100 SSD, which they're only selling to a limited number of major customers; you won't find it for sale on CDW. Intel's Optane products do use 3D XPoint memory, and at the moment I believe that's all manufactured in a Micron-owned fab. (Intel used to co-own it, and I don't think Intel will have their own production line up and running until next year.)
Have you considered just reading the code? It's all available. Best way to learn is to look what the masters are doing.
And even complex software is typically just a collection of simple things together.
IMHO one of the biggest failing of many CS courses is that they never get above toy software. Would be much better to at least once dive into some production software and try to figure something out in a real code base.
Aside from the two above issues, everything looks correct and relevant, and I can't think of any missing details that deserve to be added to an introduction of that length.
I see a maintenance page instead:
"Sorry! The URL you requested was not found on our server.
Is it Sunday between 4 and 8 PM (CST)? If so, the server may be undergoing regularly scheduled downtime. Otherwise, please contact the maintainer of the referring page and ask them to fix their link. Thanks!"
Here's a Jan, 2020 snapshot: https://web.archive.org/web/20200122013800/https://pages.cs....
I got really excited thinking it was an open source nvme fpga core.
1. This is a KV store optimized (or claimed to be) for a combo of SSD and PMEM. Not a packaged, supported, appliance storage system.
2. People who pay for enterprise storage do so for reasons beyond being too bone-headed to appreciate the joys of cobbling together production systems from the white box low-bidder and open source software.