HNHacker News
TopNewBestAskShowJobs

ianmcook

43 karma · joined February 3, 2021

https://github.com/ianmcook https://twitter.com/ianmcook
submissionscomments
ianmcook··on What if Jev spoke Arrow?
Agreed. With the API as it is today, the bottleneck is the thousands of round trips (one per state). The post is a bit of a "what if" where we're imagining (and hoping!) that TypeSafe adds a batch endpoint, which would take the round trips out of the picture. Then serialization would start to matter, and that's where Arrow would come in

It would be straightforward to wrap this as an Arrow Flight (Arrow over gRPC) endpoint. But it would make it heavier and less browser-compatible.

ianmcook··on Adding concurrent read/write to DuckDB with Arrow Flight
@1egg0myegg0 that's great to hear. I'll check to see if it applies to Arrow.

Another performance issue with DuckDB/Arrow integration that we've been working to solve is that Arrow lacked a canonical way to pass statistics along with a stream of data. So for example if you're reading Parquet files and passing them to DuckDB, you would lose the ability to pass the Parquet column statistics to DuckDB for things like join order optimization. We recently added an API to Arrow to enable passing statistics, and the DuckDB devs are working to implement this. Discussion at https://github.com/apache/arrow/issues/38837.

ianmcook··on Adding concurrent read/write to DuckDB with Arrow Flight
Arrow developer here, we've invested a lot in seamless DuckDB interop, great to see it getting traction.

Recent blog post here that breaks down why the Arrow format (which underlies Arrow Flight) is so fast in applications like this: https://arrow.apache.org/blog/2025/01/10/arrow-result-transf...

ianmcook··on Microsoft is bringing Python to Excel
Anyone know what format they are serializing the data in to move it between Excel and Python? Are they using Apache Arrow?
ianmcook··on Apache Arrow 3.0
Thanks for the heads up. The post is intended to be up but there's an intermittent error happening. It's been reported to the Apache infrastructure team.
ianmcook··on Apache Arrow 3.0
Parquet is not based on Arrow. The Parquet libraries are built into Arrow, but the two projects are separate and Arrow is not a dependency of Parquet.
ianmcook··on Apache Arrow 3.0
From https://arrow.apache.org/faq/: "Parquet files cannot be directly operated on but must be decoded in large chunks... Arrow is an in-memory format meant for direct and efficient use for computational purposes. Arrow data is... laid out in natural format for the CPU, so that data can be accessed at arbitrary places at full speed."
ianmcook··on Apache Arrow 3.0
The Arrow Feather format is an on-disk representation of Arrow memory. To read a Feather file, Arrow just copies it byte for byte from disk into memory. Or Arrow can memory-map a Feather file so you can operate on it without reading the whole file into memory.
ianmcook··on Apache Arrow 3.0
Re this second point: Arrow opens up a great deal of language and framework flexibility for data engineering-type tasks. Pre-Arrow, common kinds of data warehouse ETL tasks like writing Parquet files with explicit control over column types, compression, etc. often meant you needed to use Python, probably with PySpark, or maybe one of the other Spark API languages. With Arrow now there are a bunch more languages where you can code up tasks like this, with consistent results. Less code switching, lower complexity, less cognitive overhead.