Building a Distributed Log from Scratch, Part 1: Storage Mechanics
bravenewgeek.com
bravenewgeek.com
I like the OP article though, cause I learned about NATS streaming, which I hadn't heard of before - just Kafka. Will have to check it out.
The solution to this problem is simple and elegant with a Write-Ahead Log. Every page write is appended to the log and only merged back into the tree file when it's sure that the log is safely written to storage.
SQLite has an extensive documentation of its WAL file format, which is great for learning.
[1] https://github.com/NicolasLM/bplustree
[2] https://www.sqlite.org/fileformat.html#the_write_ahead_log
- couchdb uses a single file for each db, which means the write ahead log _is_ the storage. The atomicity is guaranteed by saying that the latest root is the valid root. If writes are interrupted, everything since the last root is invalid and discarded upon restart. A simple design that just works, although it tends to be wasteful and requires frequent compactions
- lmdb (https://en.m.wikipedia.org/wiki/Lightning_Memory-Mapped_Data...) uses copy on write to make sure space is properly used, and atomicity is provided by sharing only the strict minimal pieces of information, so small in fact atomicity is guaranteed by the os. Follow the different links in the wikipedia page, there's a lot of interesting stuff