I don't understand this part, it's not clear which major benefit you're giving up, or what you should do instead. Is it saying not to convert these formats to parquet? Or that you should create multiple parquet files to get the full benefits?
I don't understand this part, it's not clear which major benefit you're giving up, or what you should do instead. Is it saying not to convert these formats to parquet? Or that you should create multiple parquet files to get the full benefits?
The way "normal" (non-HDFS) tools write data is by creating a file with an extension. Let's say you have a 2GB parquet file, if it's broken down into 4 smaller positions, your read rates will likely be quicker, especially if the file is distributed.
See https://stackoverflow.com/questions/27194333/how-to-split-pa..., https://parquet.apache.org/documentation/latest/, etc.
Whether it's better to have multiple Parquet files or a single parallelizable Parquet file is dependent on your environment and application. At my company, we've tended to have a single row group per file (and one HDFS block per file), in part due to historical reasons.