BtrBlocks: Efficient Columnar Compression for Data Lakes [pdf]
cs.cit.tum.de
cs.cit.tum.de
I wish we could specify what we intend to do to data, craft a cost model, and let a 4th gen system optimize around that. Something that would pick and choose between the different compression techniques in this paper and also the ones from arrow and parquet.
This implies that any lower level storage formats we pick can morph to better optimized forms in a way that does not alter query semantics. It's not possible to solve this without adding refinements to the framing, such as whether data is mutable/immutable, and adding infrastucture to the solution, like metadata management and alternate projections of data. Vertica/C-Store introduced the latter a couple of decades ago. [1]
Plus ça change...
I have in this decade encountered systems that still have to run overnight. Meanwhile I’m spending multiple developer salaries maintaining a system that might be asked to answer a question in ten seconds, or might not be asked any questions for days at a time.
Somebody save me.
I recently had to look into various TPC benchmarks and some of them are very non-trivial to cost-estimate. I found at several queries in TPC-DS that join the same table to itself four (4) times. Even triangles (join with itself three times) are hard, squares like these in TPC-DS are even harder.
Hopefully the gaping performance gaps have closed.
> As someone who work with C++ parquet readers and writers I say that different configurations can easily result in 10x differences in size -- the default behavior is usually very poor (Tony Wang / @marsupialtail_2) https://twitter.com/marsupialtail_2/status/17021850155038883...
> TUM have written a great summary about performance niches around parquet file IO performance, @DatabendLabs adopted a lot of optimizations mentioned in the summary and result is great in practice. https://dl.gi.de/server/api/core/bitstreams/9c8435ee-d478-4b... (zhihanz / @zhihanz1205) https://twitter.com/zhihanz1205/status/1702196118472536166
> We introduced BtrBlocks, an open columnar compression format for data lakes. By analyzing a collection of real-world datasets, we selected a pool of fast encoding schemes for this use case. Additionally, we introduced Pseudodecimal Encoding, a novel compression scheme for floating-point numbers. Using our sample-based compression scheme selection algorithm and our generic framework for cascading compression, we showed that, compared to existing data lake formats, BtrBlocks achieves a high compression factor, competitive compression speed and superior decompression performance. BtrBlocks is open source and available at https://github.com/maxi-k/btrblocks.
I doubt I'd ever used columnar compression again as I felt it too difficult to fight DBAs on keeping the original sorting and schema preserved in an optimal way. I do find it really interesting though.
However, the data science/big data world has been bogged down by inefficient, clunky formats like Parquet for quite a while and so I applaud any steps towards bridging the gap between the two worlds.
Generally with “data stuff” you’re chained to whatever the dominant language and tools do (for the longest time, Python and Spark) and you’d have to accomodate them and their idiosyncrasies directly. Nowadays though, basically everything has a pq lib, there’s tools at all scales: DuckDB, ClickHouse, Databend, etc. The Arrow internal representation of parquet means we’ve escaped the dominance of single massive packages (e.g. pandas and spark dataframes) as the only way to deal with certain datasets.
I think it’s great what we’ve managed to get via Parquet, even if it’s a little bit clunky.
What I’m very much hoping is things stay in the current style of “mostly” language independent formats and tools. I’ve written more than enough Python and Scala/Spark and I’ve moved on that I don’t ever want to go back to the abysmal landscape that was “Python or bust”.