Also can the same thing work for files hosted on s3?
Also can the same thing work for files hosted on s3?
The trick it's using is the HTTP Range header, which lets a client request eg bytes 4500-4900 of an HTTP file.
Most static file hosting platforms - S3, GCS, nginx, Apache etc - support Range headers. They're most commonly used for streaming video and audio.
The Parquet file format is designed with this in mind. You can read metadata at the start of the file and use it to figure out which ranges to fetch. Columns are grouped together, so sum() against a column can be handled by fetching a subset of the file.
The offsets of said metadata are well-defined (i.e. in the footer) so for S3 / blob storage so long as you can efficiently request a range of bytes you can pull the metadata without having to read all the data.