I wish it supported large csv to partitioned parquet. THAT is something I need a good solution for.
Part of the problem is the scale, but another part is that writing partitioned parquets seems poorly documented (I would love corrections, I spent a decent amount of time last week looking for good information)
It's literally one line of code. See http://labs.vistarmedia.com/2016/12/27/indexing-json-logs-wi... for an example (except you write to a local file system rather than HDFS).