Have you given DuckDB a try? I'm using it to shuttle some hefty data between postgres and some parquet files on S3 and it's only a couple lines. Haven't noted any memory issues so far
```
def _get_duck_db_arrow_results(s3_key):
con = duckdb.connect(config={'threads': 1, 'memory_limit': '1GB'})
con.install_extension("aws")
con.install_extension("httpfs")
con.load_extension("aws")
con.load_extension("httpfs")
con.sql("CALL load_aws_credentials('hadrius-dev', set_region=true);")
con.sql("CREATE SECRET (TYPE S3,PROVIDER CREDENTIAL_CHAIN);")
results = con \
.execute(f"SELECT * FROM read_parquet('{s3_key}');") \
.fetch_record_batch(1024)
for index, result in enumerate(results):
print(index)
return results
```I ran the above on a 1.4gb parquet file and 15 min later, all of the results were printed at once. This suggests to me that the whole file was loaded loaded into memory at once.
To stream, fetch more batches.
What ddb does to get the batches depends on hand wavey magic around available ram, and also the structure of the parquet.
When I'm writing to postgres though I'm doing into entirely inside DuckDB with a `INSERT INTO ... SELECT ...` and that seems to stream it over.