38 karma · joined August 20, 2021
The one downside of this approach, which is likely obvious, but I haven't seen mentioned is that the resulting parquet files are larger than they would be otherwise, and the increased size only benefits engines that know how to interpret the new index
(I am an author)
Yes this is my personal hope as well -- if there are new index types that are widespread, they can be incorporated formally into the spec
However, changing the spec is a non trivial process and requires significant consensus and engineering
Thus the methods used in the blog can be used to use indexes prior to any spec change and potentially as a way to prototype / prove out new potential indexes
(note I am an author)
I expect this to be used to support Variant https://github.com/apache/datafusion/issues/16116 and geometry types
(note I am an author)
Most of the leaderboard of ClickBench is for database specific file formats (that you first have to load the data into)
If you are looking for the nicest "run SQL on local files" experience, DuckDB is pretty hard to beat
Disclaimer: I am the PMC chair of DataFusion
There are some other interesting FAQs here too: https://datafusion.apache.org/user-guide/faq.html
https://github.com/apache/datafusion/issues/13448
Disclaimer I am on the PMC of Apache DataFusion, so am totally a fan boy.
For example when you have a predicate like, `where id = 'fdhah-4311-ddsdd-222aa'` sorting on the `id` column will help
However, if you have predicates on multiple different sets of columns, such as another query on `state = 'MA'`, you can't pick an ideal sort order for all of them.
People often partition (sort) on the low cardinality columns first as that tends to improve compression signficantly
(can also use it in your own projects)
It is quite similar to what is described in this post
Deep Dive into Common Open Formats for Analytical DBMSs https://www.vldb.org/pvldb/vol16/p3044-liu.pdf
If you are looking for a query engine implemented in a safe language (Rust) I definitely suggest checking out DataFusion. It is comparable to DuckDB in performance, has all the standard built in SQL functionality, and is extensible in pretty much all areas (query language, data formats, catalogs, user defined functions, etc)
https://arrow.apache.org/datafusion/
Disclaimer I am a maintainer of DataFusion
https://arrow.apache.org/blog/2019/10/13/introducing-arrow-f...
I do think that is the best current view of a RoadMap and Vision -- it would be great to flesh it out a bit more.
In fact, I'll make a note to try and add some more higher level context into the project on our goals.