Things I Wish I'd Known About Spark When I Started
enigma.com
enigma.com
I think data engineers 'should' add Spark to their toolset, even if just for ETL. I found myself having to go through a few hundred CSV files and import them into Oracle, Spark made it fast and easy.
Many of the pain points that you mention are universal whether using Scala or Python. I wish I'd have known about .par years ago.
I also find checkpointing to help a lot, especially if using JDBC data sources. If in reading in data that doesn't change, I read it into parquet, and then change to reading from parquet.
Arrow has made pyspark more pleasant to work with, especially if you primarily work on notebooks.
The one thing which I struggled with was latency on time-critical jobs. I'd sometimes get a few minutes' pauses on jobs that should take under a minute to run. I haven't checked how that's improved since Spark 2.0 though
I don't understand this part, it's not clear which major benefit you're giving up, or what you should do instead. Is it saying not to convert these formats to parquet? Or that you should create multiple parquet files to get the full benefits?
See https://stackoverflow.com/questions/27194333/how-to-split-pa..., https://parquet.apache.org/documentation/latest/, etc.
Whether it's better to have multiple Parquet files or a single parallelizable Parquet file is dependent on your environment and application. At my company, we've tended to have a single row group per file (and one HDFS block per file), in part due to historical reasons.
The way "normal" (non-HDFS) tools write data is by creating a file with an extension. Let's say you have a 2GB parquet file, if it's broken down into 4 smaller positions, your read rates will likely be quicker, especially if the file is distributed.