In the case of AWS, repartitioning Parquet files in S3 via Athena CTAS statements in limited to 100 active partitions, which is a bummer to work around. Therefore, I’m using DuckDB with repartitioning queries, because it doesn’t have the 100 partition limit.
I wrote a blog post about it at https://tobilg.com/casual-data-engineering-or-a-poor-mans-da... Additionally, to get started with using DuckDB serverlessly in Lambda functions, you can have a look at https://tobilg.com/using-duckdb-in-aws-lambda