Alternatively, you can run DuckDB as part of a plpython function:
CREATE FUNCTION pyduckdb ()
RETURNS integer
AS $$
import duckdb
con = duckdb.connect()
return con.execute('select 42').fetchall()[0][0]
$$ LANGUAGE plpythonu;
SELECT pyduckdb();- Postgres is a row store optimized for transactional workloads, whereas Parquet is a column-oriented format optimized for analytical workloads.
- Database storage is typically very expensive SSD storage optimized for fast IO and high availability. Parquet files, on the other hand, can be stored in inexpensive object storage such as S3.
- Loading is an additional and possibly unnecessary step
But that's why I'm asking, I don't know.
Be kind. Don't be snarky. Have curious conversation; don't cross-examine. Please don't fulminate. Please don't sneer, including at the rest of the community.