Also very interested in the parquet tuning. I have been building my data lake and most optimization I do is just with efficient partitioning.
Regarding reading materials, I found this DuckDB post to be especially helpful in realizing how parquet could be better leveraged for efficiency: https://duckdb.org/2024/03/26/42-parquet-a-zip-bomb-for-the-...
Tends to be that an optimal file size for Parquet is about 1GiB, once again, the "many small files" problem of Hadoop remains.
Then it's things like, can you organise your data in such a way to take advantage of RLE etc.?