Um, my data lake is measured in Pb, not Tb. How is that going to work exactly?
Here's a post about applying this same trick to SQLite from few years ago: https://phiresky.github.io/blog/2021/hosting-sqlite-database...
Once you know what you are reading, many parquet/arrow libraries will support streaming reads/aggregations, so the client doesn’t need to load the whole working set in memory.
For all others you'll need to download all columns you are filtering or selecting.