Float Compression 3: Filters
aras-p.info
aras-p.info
zfp performance is... "underwhelming" to say the least :( (fpzip is pretty good though, if only a bit slow at decompression). Maybe zfp is much better on 3D or 4D data, and/or in lossy mode. Which might be a topic for a future post.
The first byte of a IEEE float has 7/8 of the exponent. If these are successive measurements, chances are they have lots of series of identical values, and the delta-encoded stream will have lots of zeroes.
So "block" is not a lossless transformation/compression.
https://www.timescale.com/blog/time-series-compression-algor...
Otherwise, the compression is lossy.
My question is: In the case of component vectors like that, given that many operations wind up in operations like dot products, do the advantages still apply?
I don't have any experience programming with simd instructions specifically, so the memory access patterns get a bit murky to me.
The same goes for cuda, for example, which has a native float4 type. It's making row ordering so convenient, and going for SoA seems like you need to walk away from that.
1. Reordering to SOA helps a lot - this is the whole point of column-oriented databases.
2. Specialized codecs like Gorilla[3], DoubleDelta[4], and FPC[5] lose to simply using ZSTD[6] compression in most cases, both in compression ratio and in performance.
3. Specialized time-series DBMS like InfluxDB or TimescaleDB lose to general-purpose relational OLAP DBMS like ClickHouse [7][8][9].
[1] https://clickhouse.com/blog/optimize-clickhouse-codecs-compr...
[2] https://github.com/ClickHouse/ClickHouse
[3] https://clickhouse.com/docs/en/sql-reference/statements/crea...
[4] https://clickhouse.com/docs/en/sql-reference/statements/crea...
[5] https://clickhouse.com/docs/en/sql-reference/statements/crea...
[6] https://github.com/facebook/zstd/
[7] https://arxiv.org/pdf/2204.09795.pdf "SciTS: A Benchmark for Time-Series Databases in Scientific Experiments and Industrial Internet of Things" (2022)
[8] https://gitlab.com/gitlab-org/incubation-engineering/apm/apm... https://gitlab.com/gitlab-org/incubation-engineering/apm/apm...
[9] https://www.sciencedirect.com/science/article/pii/S187705091...
Comparing OLAP and OLTP is usually just a waste of time. They solve different problems.
This blog post linked within your gitlab issue is quite impartial, relatively speaking, and disputes most of these claims: https://www.timescale.com/blog/what-is-clickhouse-how-does-i...
This also similarly matches my experience operating properly tuned instances.
Lastly, regarding the following
> Specialized codecs like Gorilla[3], DoubleDelta[4], and FPC[5] lose to simply using ZSTD[6] compression in most cases, both in compression ratio and in performance.
I'd love to see some data because this is a very broad generalization that has never remotely been true in my experience. At comparable block sizes, with proper tuning, double delta can absolutely smoke zstd. In fact, while zstd can achieve quite impressive compression and performance, I almost never reach for it as it's rarely optimal when I benchmark, compared to lz4 or xz or snappy or even sometimes gzip. It depends on many factors -- what data, how much data, block size, disk or storage medium, network.
Your whole post puts a rather sour taste in my mouth about ClickHouse.
Another experiment is described in the linked blog post: https://clickhouse.com/blog/optimize-clickhouse-codecs-compr... It compares codes on the weather dataset (NOAA).
The comment about "DoubleDelta" codec from the ClickHouse codebase:
/** NOTE DoubleDelta is surprisingly bad name. The only excuse is that it comes from an academic paper.
* Most people will think that "double delta" is just applying delta transform twice.
* But in fact it is something more than applying delta transform twice.
*/