Some years ago I had huge performance increase by simply converting a csv dataset to parquet before processing it with Apache Spark.
One thing I miss though is how easy it is to inspect .csv and .xlsx. I kinda solved it using [1], but it only works on Windows. More portable recommendations welcome!
Its a vi(m) inspired tool.
It also handles xls(x), sqlitedb and a bunch of other random things, and it appears to support parquet via pandas: