Agreed on DuckDB, fantastic for working with most major data formats
```
def _get_duck_db_arrow_results(s3_key):
con = duckdb.connect(config={'threads': 1, 'memory_limit': '1GB'})
con.install_extension("aws")
con.install_extension("httpfs")
con.load_extension("aws")
con.load_extension("httpfs")
con.sql("CALL load_aws_credentials('hadrius-dev', set_region=true);")
con.sql("CREATE SECRET (TYPE S3,PROVIDER CREDENTIAL_CHAIN);")
results = con \
.execute(f"SELECT * FROM read_parquet('{s3_key}');") \
.fetch_record_batch(1024)
for index, result in enumerate(results):
print(index)
return results
```I ran the above on a 1.4gb parquet file and 15 min later, all of the results were printed at once. This suggests to me that the whole file was loaded loaded into memory at once.
To stream, fetch more batches.
What ddb does to get the batches depends on hand wavey magic around available ram, and also the structure of the parquet.
When I'm writing to postgres though I'm doing into entirely inside DuckDB with a `INSERT INTO ... SELECT ...` and that seems to stream it over.