[1] https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tab...
Hopefully this will come at some point. Product looks very cool otherwise.
Then spin up duckdb and do some performance tests. I’m not sure this will work, there is some overheard with reading parquet, which is why it is discouraged to have small files and row groups.
Do you ever go back and reaggregate older data into bigger, sorted files? That is, maybe you originally partitioned by hour, but stale data is so infrequently accessed, you could roll up into partitions per week/month/whatever. Depending on the specifics, you might save some space from less file overhead and better compression statistics.
No more files. You might be able to avoid per usage pricing just by hosting this on a regular vps.