This approach makes sense for predicting the data. Obviously, one could split the data to run distributed prediction. But, how does this work for training the linear model mentioned here with scikit-learn in a distributed fashion?
However, for the example in this post, I would recommended using the logistic regression provided by MLlib to scale up.