This might end up being the best way to etl postgres tables to parquet. From everything else that I tried, doing a copy to CSV and then converting to parquet was the fastest but can be a pain when dealing with type conversions.
SELECT ... FROM postgresql(...) FORMAT Parquet
And you can run this query without installing ClickHouse, using the clickhouse-local command-line tool.
It can be downloaded simply as:
curl https://clickhouse.com/ | sh
In my case, I had parquet to begin with because I accidentally deleted some production data (oopsies) and when you export a snapshot from RDS to S3, it is in Parquet. Thankfully, I now have a few tricks up my sleeve to quickly restore data, but that was stressful for a bit haha
If you don't use arrays and composites, Spark should be able to do it, right?
Does that help or do you have any other questions?