LevelDB: SSTable and Log-Structured Storage (2012)
igvita.com
igvita.com
How simple it is to setup and get going is still extremely appealing though.
But with LSM style databases, you are going to have a really bad time if you don't have a team member who has dedicated serious time to understanding the internal workings of LSM itself, and the details of your chosen implementation. That's a real mark against it, IMO.
tldr; LSM databases are like a colander in the world of leaky abstractions.
Think citusdb without consideration you're gonna expect good performance in sharded environment?
But one shouldn't be designing for implementation details. Usually, we start technologies out with leaky abstractions, and gradually get better at it. A good example is game development, where it used to be that you always used the drawing method of the display and the clock speed of the cpu to your advantage. Nowadays, we've moved past that, because it was working on horrible abstractions, and because the technology underneath improved.
I'm not saying I have a solution, and I agree that this problem rears it's head the most when you start bringing in distributed storage. But my point stands: these databases run on a highly leaky abstraction, and that's a big problem going forward.
See the difference in 1 box of `1 process per core` `scylladb`,`voltdb`,`redis` compared to all other dbs `1-process-for-all-cores`.
IMO, while still cool, LSM trees don't feel particularly novel at this point. It seems like every database and their dog has adopted/made available some kind of LSM tree storage engine, down to traditional relational databases (e.g. MySQL with MyRocks).
It's also nice to recognize the work of people who have had a significant impact on the industry. Jeff Dean & Sanjay Ghemawat are an incredible force.