I've been using clickhouse-local for quite some time, instead of DuckDB.
There is also chDB.
After using pandas for 10 years, I favor SQL now, for some reason.
After using pandas for 10 years, I favor SQL now, for some reason.
I use SQL in data pipelines and processing that is going to require interoperability.
But for data exploration, I usually prefer Polars (imo it is easier to work with text, semi-structured data, etc.)
Regular CH also support external data sources, so I can read 500GB of JSON from S3 and group by it on production server very fast and in memory.