How Dolt Stores Table Data (2020)
dolthub.com
dolthub.com
I think the Noms doc linked from this article [2] is clearer than the article itself. That said I sill cannot turn my head around to grasp how this entire thing work tbh. I hope they wrote a peer reviewed paper to serve the audience better.
[1] https://github.com/attic-labs/
[2] https://github.com/attic-labs/noms/blob/master/doc/intro.md#...
You just sent me into an small crisis wondering when that 2023 thing happened :D
However once you have chunks of fixed size (the maximum) instead of content-defined sizes, changes to the content become non-local: a change in one part of the data might change an arbitrary number of following chunks/nodes.
Did they find a solution to that? Did they ignore it?
However, from what I've seen, these methods generally come at the cost of deduplication and/or speed. The most reliable method to avoid pathological cases seems to just be setting the min/max chunk size to a low/high enough value respectively.
If you're talking in a purely theoretical sense, I would assume that the possibility of changes affecting non-local chunks is inherent to CDC. With well-chosen parameters the likelihood of any but the closest chunks being affected just becomes low enough to be negligible.
Here's a more recent write-up on chunking that details some of them, but it mostly comes down to using a better hash function, and having a fail-safe max chunk size for truly pathological cases.
Using hashes as links isn't cheap especially with sha-512 wide hashes (I think they use 20bytes in reality?). I'd estimate their fanout at between 50-200, which isn't that much either.
So my feeling is expensive nodes, combined with low-ish fanout => high cost of storage?