Directly running DuckDB queries on data stored in SQLite files
twitter.com
twitter.com
I’ve been using duckdb instead of Pandas these days because it’s much faster on larger datasets plus I can write more compact and complex SQL than I can Pandas constructs. The fact that duckdb can query in memory Pandas dataframes faster than Pandas itself is a plus.
It’s like having a local performant database engine that can query and join across Parquet, CSV, Pandas and now SQLite.
Kinda like a quick efficient local Spark.
It is a wrapper on Sqlite, but I'm looking to switch to duckdb because its dialect is more comprehensive.
Previously, I was using ruby, sometimes python, and sometimes postgresql to process CSV. But it wasn't convenient enough. Most of the times I just tried to use Excel, but formula is so hard to use.
Or are you talking about something else?
Just curious, what format did you use?
If it's Parquet, it does column compression via Snappy by default (which more optimized for speed than size).
You can change the compression to ZSTD or GZIP for higher compression ratios.
DuckDB's Node bindings leave a lot to be desired, as well. WASM bindings have a different API surface, too. From my reading of the code and docs, it looks like duckdb 3p bindings for anything other than python is in dire need fit and polish.
But I'm happy to discuss this further if you have any suggestions.
DuckDB, similar to SQLite, is an in-process query engine for local data.
Is this adapter copying the contents of the SQLite db into DuckDBs own memory layout in order to optimise the queries or is it just “proxying” to the SQLite library?