DuckDB Ducklake
github.com
github.com
There's a cool alternate rust/datafusion ecosystem initiative going on at https://github.com/datafusion-contrib/datafusion-ducklake, and think the Quack protocol opens up a lot of cool possibilities too.
If you need an idea for what do do with ducklake: recommend throwing all of your agent traces in it.
For context, we previously built custom catalogs optimized for specific use cases. But they were hard to maintain, especially as requirements changed, and Apache Iceberg was too heavy for our specific low-latency work.
Since Ducklake is only a spec, we implemented datafusion-ducklake, and it performs as well as any custom or specialized catalog we built. We use Postgres as the catalog store, and it does not get much simpler than that: a transactional database for transactional data.
Plus, it gives us a clear spec for implementing complex parts like time travel, snapshots, etc.
It's been a godsend.
We welcome and encourage contributors!
Here is our write up on the Ducklake blog: https://ducklake.select/2026/07/29/bringing-ducklake-to-data...
I always thought the catalogue was a duckdb file. E.g, data lives in partitioned parquet files, but which parquet files are current or soft deleted, etc, etc, is managed in a duckdb data file.
However, looking at https://ducklake.select/, it seems the catalogue lives in PostgresSQL - so it is not really a ducklake, but a postgresslake.
The more you know.
When it works well it’s really nice. And it beats handrolling a multi level parquet store.
For example I think a lakehouse is a bad choice for standard enterprise BI type analytics - you've got no column or row access controls, and no column masking. I don't see how this could ever be bolted on to the bucket and catalog.
https://www.tomwphillips.co.uk/2026/08/the-benefits-of-data-...
https://duckdb.org/2025/05/19/the-lost-decade-of-small-data....
Much faster, but adds a dependency... that you would have added anyway with database based catalogs (that are not the only kind of catalogs)