LMDB uses a regular bucket for the freelist whereas Bolt simply saved the list as an array. It simplified the logic quite a bit and generally didn't cause a problem for most use cases. It only became an issue when someone wrote a ton of data and then deleted it and never used it again. Roblox reported having 4GB of free pages which translated into a giant array of 4-byte page numbers.
Afaiu, Bolt is a personal OSS project, github repo is archived with last commit 4 years ago, and the first thing you see in the readme is the "author no longer has time nor energy to continue".
Commercial cash cows like Roblox (a) shouldn't expect free labor and (b) should be wise enough to recognize tech debt or immaturity in their dependencies. Heck, even as a solo dev I review every direct dependency I take on, at least to a minimal level.
I can't speak to the incident response as I'm not an sre, but as a dev this screams of fragile "ship fast" culture, despite all the back patting in the post. I'm all for blameless postmortems, but a culture of rigor is a collective property worthy of attention and criticism.
The developers at Hashicorp are top-tier, and this doesn't substantially change their reputation in my eyes. Hindsight is always 20/20.
Let's end this thread; blaming doesn't help anyone.
As for HashiCorp, they're an awesome group of folks. There are few developers I esteem higher than their CTO, Armond Dadger. Wicked smart guy. That all being said, there's a lot of moving parts and sometimes bugs get through. ¯\_(ツ)_/¯
How does this happen so often? It's awesome to get the authors take on things. Also thank you for explaining and owning it. Where you part of this incident response?
:)
It's also interesting how much a tiny detail can derail a huge organization. My former employer lost all services worldwide because of a single incorrect routing in a DNS server.
OSS contributors are rarely noticed or appreciated. Did HashiCorp ever sponsor you or share any revenue with you? The OSS ecosystem is broken.
That being said, I gave away the software for free so I don’t have any expectation of payment. I agree the ecosystem is broken but I don’t know how to fix it or even if it can be fixed.
Is there an issue/bug for this somewhere?
I haven’t heard of anyone else with the same issue since then so I assume it’s probably fixed.
So BoltDB and LMDB affected users may switch to libmdbx as the Erigon (Ethereum implementation) does year ago https://github.com/ledgerwatch/erigon/wiki/Criteria-for-tran...
For now this is (relatively) easy since bindings for GoLang, Rust NodeJS/Deno, etc are available and the API is mostly the same in general.
---
The ideas that MDBX uses to solve these issues are simple: zero-cost micro-defragmentation, coalescing short GC/freelist records, chunking too long GC/freelist records, LIFO for GC/freelist reclaiming, etc.
Many of the ideas mentioned seems simple to implement in BoldDB. However the complete solution is not documented and too complicated (in accordance with the traditions inherited from LMDB ;)