My use case is extremely redundant data (specific website dumps + logs) that I want decently quick random access into, and I was unhappy with either the access speed, quality/usability or even existence of libraries for several formats.
Glancing over the code this seems to use the following setup:
- Aggregate files
- Chunk into blocks
- Compress blocks of fixed size
- Store file to chunk and chunk to block associations
What I did not see is a deduplication step for the chunks, or an attempt to group files (and by extend, blocks) by similarity in an attempt improve compression.
But I might have just missed that due to lack of familiarity with Pascal.
For anyone interested in this strategy, take a look at ZPAQ [1] by Matt Mahoney, you might know him from the Hutter Prize competition [2] / Large Text Compression Benchmark. It takes 14th place with tuned parameters.
There's also a maintained fork called zpaqfranz, but I ran into some issues like incorrect disk size estimates with it. For me the code was also sometimes hard to read due to being a mix of English and Italian. So your mileage may vary.
[1]: http://mattmahoney.net/dc/zpaq.html [2]: http://prize.hutter1.net [3]: https://github.com/fcorbelli/zpaqfranz