HNHacker News
TopNewBestAskShowJobs

azimafroozeh

9 karma · joined August 27, 2021

submissionscomments
azimafroozeh··on The FastLanes File Format [pdf]
We’re building a community around FastLanes, and we’d love any contributions or feedback! Please join our Discord so we can discuss things in more detail.
azimafroozeh··on The FastLanes File Format [pdf]
As we discuss in the FastLanes paper, the way BtrBlocks implements cascaded encodings (Vortex now) is essentially a return to block-based compression such as Zstd — which we're trying to avoid as much as possible. This design doesn't work well with modern vectorized execution engines or GPUs: the decompression granularity is too large to fit in CPU caches or GPU shared memory. So Vortex ends up being yet another Parquet-like file format, repeating the same mistakes. And if it still underperforms compared to Parquet... what’s the point?

We just released FastLanes v0.1, and more results — including ClickBench — are coming soon. Please do benchmark FastLanes — and keep us posted!

azimafroozeh··on The FastLanes File Format [pdf]
CUDA and GPU support are next on our list — but we’re definitely interested in WebAssembly as well!
azimafroozeh··on The FastLanes File Format [pdf]
That’s a very valid question. We’ve done zero optimization on the encoding side so far, and improving that is definitely on our roadmap. Technically, once we learn the best expressions, they can be reused — data is often very similar across row groups — which opens the door to caching and amortizing the cost.

For very wide tables, expression detection only needs to happen once. Beyond that, we’re also exploring techniques like grouping columns into smaller sets or applying more aggressive heuristics to prune irrelevant columns. These are areas we’re actively investigating, and we plan to support them in future versions of FastLanes.

azimafroozeh··on The FastLanes File Format [pdf]
Vortex borrows a few ideas from the FastLanes project, such as bit-packing and ALP. However, it’s unclear how well these are implemented — their performance on ClickBench appears worse than Parquet in both storage size and decompression speed, which is counterintuitive.

Technically, Vortex is documented more like a BtrBlocks-style format, which we’ve benchmarked and compared against in depth.

azimafroozeh··on The FastLanes File Format [pdf]
I totally agree — the C++ implementations of Parquet aren’t the best they could be. To be fair, though, it’s a tough problem: Parquet supports a broad range of encodings and compression schemes, and to do that, it pulls in a lot of dependencies and requires complex dependency management.

That’s one of the main reasons we built FastLanes from scratch instead of trying to integrate with Parquet.

With FastLanes, we’ve taken a different approach: zero dependencies, no SIMD intrinsics, and a design that’s fully auto-vectorizable. The result is simpler code that still delivers high performance.

azimafroozeh··on The FastLanes File Format [pdf]
We’d love to bring FastLanes into DuckDB! We're currently working on a DuckDB extension to read and write FastLanes file formats directly from DuckDB.

FastLanes is columnar by design, so partial downloading via HTTP Range requests is definitely possible — though not yet implemented. It’s on the roadmap.

azimafroozeh··on The FastLanes File Format [pdf]
We haven’t benchmarked FastLanes directly against LanceDB yet, but here’s a quick look at the compression side:

LanceDB supports:

FSST

Bit-packing

Delta encoding

Opaque block codecs: GZIP, LZ4, Snappy, ZLIB

So in that regard, it’s quite similar to Parquet — a mix of lightweight codecs and general-purpose block compression.

FastLanes, on the other hand, introduces Expression Encoding — a unified compression model that allows combining lightweight encodings to achieve better compression ratios. It also integrates multiple research efforts from CWI into a single file format:

The FastLanes Compression Layout: Decoding >100 Billion Integers per Second with Scalar Code (VLDB '23) PDF: https://dl.acm.org/doi/pdf/10.14778/3598581.3598587

ALP (Adaptive Lossless Floating-Point Compression) — SIGMOD '24 https://ir.cwi.nl/pub/33334/33334.pdf

G‑ALP (GPU-parallel variant of ALP) — DaMoN '25 https://azimafroozeh.org/assets/papers/g-alp.pdf

White-box Compression (self-describing, function-based) — CIDR '20 https://www.cidrdb.org/cidr2020/papers/p4-ghita-cidr20.pdf

CCC (Exploiting Column Correlations for Compression) — MSc Thesis '23 https://homepages.cwi.nl/~boncz/msc/2023-ThomasGlas.pdf

azimafroozeh··on The FastLanes File Format [pdf]
Arrow is primarily an in-memory format, while Parquet is commonly used for on-disk storage. Typically, data is stored in Parquet and then read into memory as Arrow. FastLanes is a new on-disk file format, comparable to Parquet, but offers around 40% better compression and faster decoding thanks to its data-parallel encoding design.

Disclaimer: I'm the first author of the paper.