A bit of both. I wanted an excuse to test out Spark to find the kinks which were ommited from the documentation (and boy did I find kinks), and also provide a practical demo.
Out of curiosity, what kinks did you find?
Then there are the massive shuffle read/writes that result in 50GB i/o which are not great for SSDs.