Could this same problem be solved by using Apache Arrow, to convert to/from pandas, and cut down the complexity in the process?
This does use Apache arrow, that's how the data is transferred from spark to pandas and back.
Great, glad this was confirmed as I am about to solve a similar problem and my expectation is that Apache Arrow will be needed. So pyarrow = Apache Arrow.
This approach makes sense for predicting the data. Obviously, one could split the data to run distributed prediction. But, how does this work for training the linear model mentioned here with scikit-learn in a distributed fashion?