If you use pyspark and Spark for Parquet, you get:
* Easy type inference, even for nested maps and structs
* lz4 compression support
* SQL and directory partitioning out of the box
If you use pyarrow, you get: * Write support for nested types, but read support is broken / incomplete (it throws a TODO error)
* no lz4 support and a load of Jira politics blocking it
* You can query using Pandas (not SQL), but that querying can be much slower versus Spark (Spark is naturally parallelized)
pyarrow might need to hit 2.0.0 to be really viable. It’s definitely easier to use than parquet-mr though.