HNHacker News
TopNewBestAskShowJobs

alamb

38 karma · joined August 20, 2021

submissionscomments
alamb··on Embedding a Tantivy Index in Parquet
This demo extends a Parquet file by embedding a Tantivy full-text search index inside it. A custom DataFusion TableProvider implementation uses the embedded full-text index to optimize wildcard LIKE predicates.
alamb··on Embedding user-defined indexes in Apache Parquet
> Note that the readers of Parquet need to be aware of any metadata to exploit it. But if not, nothing changes

The one downside of this approach, which is likely obvious, but I haven't seen mentioned is that the resulting parquet files are larger than they would be otherwise, and the increased size only benefits engines that know how to interpret the new index

(I am an author)

alamb··on Embedding user-defined indexes in Apache Parquet
> That is, start with Wild West and define specs as needed

Yes this is my personal hope as well -- if there are new index types that are widespread, they can be incorporated formally into the spec

However, changing the spec is a non trivial process and requires significant consensus and engineering

Thus the methods used in the blog can be used to use indexes prior to any spec change and potentially as a way to prototype / prove out new potential indexes

(note I am an author)

alamb··on Embedding user-defined indexes in Apache Parquet
We are actively working on supporting extension types. The mechanism is likely to be using the Arrow extension type mechanism (a logical annotation on top of existing Arrow types https://arrow.apache.org/docs/format/Columnar.html#format-me...)

I expect this to be used to support Variant https://github.com/apache/datafusion/issues/16116 and geometry types

(note I am an author)

alamb··on Tpchgen-rs: TPC-H benchmark data generation in pure Rust
See also related blog: https://datafusion.apache.org/blog/2025/04/10/fastest-tpch-g...
alamb··on Apache DataFusion
Specifically, DataFusion is faster when querying parquet directly.

Most of the leaderboard of ClickBench is for database specific file formats (that you first have to load the data into)

alamb··on Apache DataFusion
I think you would pick DataFusion over DuckDB if you want to customize it substantially. Not just with user defined functions (which are quite easy to write in DataFusion and are very fast), but things like * custom file formats (e.g. Spiral or Lance) * custom query languages / sql dialects * custom catalogs (e.g. other than a local file or prebuilt duckdb connectors) * custom indexes (read only parts of parquet files based on extra information you store) * etc.

If you are looking for the nicest "run SQL on local files" experience, DuckDB is pretty hard to beat

Disclaimer: I am the PMC chair of DataFusion

There are some other interesting FAQs here too: https://datafusion.apache.org/user-guide/faq.html

alamb··on Building Databases over a Weekend
BTW here is a fun exercise that takes this idea to the extreme. Who can build a custom file format that gets the best ClickHouse performance (on DataFusion):

https://github.com/apache/datafusion/issues/13448

Disclaimer I am on the PMC of Apache DataFusion, so am totally a fan boy.

alamb··on Using Parquet's Bloom Filters
In general, if you can partition your datasets on your predicate column, sorting is likely the best option

For example when you have a predicate like, `where id = 'fdhah-4311-ddsdd-222aa'` sorting on the `id` column will help

However, if you have predicates on multiple different sets of columns, such as another query on `state = 'MA'`, you can't pick an ideal sort order for all of them.

People often partition (sort) on the low cardinality columns first as that tends to improve compression signficantly

alamb··on Bringing GPU acceleration to Polars DataFrames in the near future
It would be amazing if the code for working with arrow on GPUs could be made open source -- I think that would drive a significant amount of adoption
alamb··on Show HN: Spice.ai – materialize, accelerate, and query SQL data from any source
So great to see another project built on DataFusion @!
alamb··on Apache Arrow DataFusion Comet
The Apache Arrow PMC is pleased to announce the donation of the Comet project, a native Spark SQL Accelerator built on Apache Arrow DataFusion.
alamb··on What I talk about when I talk about query optimizer (part 1): IR design
CMU's database courses are online and excellent:

https://15445.courses.cs.cmu.edu/spring2024/

https://15721.courses.cs.cmu.edu/spring2023/

alamb··on What I talk about when I talk about query optimizer (part 1): IR design
BTW you can see a version of what an industrial strength query optimizer / execution engine looks like in Rust https://arrow.apache.org/datafusion/

(can also use it in your own projects)

It is quite similar to what is described in this post

alamb··on Updates to the H2O.ai db-benchmark
The following paper describes some of the tradeoffs between different formats

Deep Dive into Common Open Formats for Analytical DBMSs https://www.vldb.org/pvldb/vol16/p3044-liu.pdf

alamb··on Updates to the H2O.ai db-benchmark
I do think it was important for duckdb to put out a new version of the results as the earlier version of that benchmark [1] went dormant with a very old version of duckdb with very bad performance, especially against polars.

[1] https://h2oai.github.io/db-benchmark/

alamb··on DuckDB 0.8
DuckDB is a great piece of software if you are

If you are looking for a query engine implemented in a safe language (Rust) I definitely suggest checking out DataFusion. It is comparable to DuckDB in performance, has all the standard built in SQL functionality, and is extensible in pretty much all areas (query language, data formats, catalogs, user defined functions, etc)

https://arrow.apache.org/datafusion/

Disclaimer I am a maintainer of DataFusion

alamb··on Demystifying Apache Arrow (2020)
Here is another blog post that offers some perspective on the growth of Arrow over the intervening years and future directions: https://www.datawill.io/posts/apache-arrow-2022-reflection/
alamb··on Database drivers: Naughty or nice?
For completeness, FlightSQL[1] (as mentioned elsewhere in this thread) aims to provide such an HTTP based protocol

https://arrow.apache.org/blog/2019/10/13/introducing-arrow-f...

alamb··on IOx: InfluxData’s New Storage Engine
Ah -- got it! This is the beauty of aligning ourselves with technologies like Arrow, Parquet and DataFusion. We can share as well as benefit from the efforts of the broader community
alamb··on IOx: InfluxData’s New Storage Engine
Author here -- it is "free" in the sense that all the effort we put into DataFusion flows directly into IOx. But we do put a lot of effort into DataFusion
alamb··on Blaze: A Rust-based vectorized accelerator to speed up your Spark jobs
Rust for the win!
alamb··on Apache Arrow DataFusion 7.0 released
Here is an announcement of the 5.0 release: https://news.ycombinator.com/item?id=28290616
alamb··on Apache Arrow Rust Implementation Highlights
Security, performance and async parquet reader for the rustlang implementation of Apache Arrow
alamb··on Apache Arrow Datafusion 5.0.0 release
<DataFusion committer here>

I do think that is the best current view of a RoadMap and Vision -- it would be great to flesh it out a bit more.

In fact, I'll make a note to try and add some more higher level context into the project on our goals.