Thanks for writing this! I'm thinking about using Spark for a little 2M-data-point project that I'm working on, just for the learning experience.
Out of curiosity, what kinks did you find?
Out of curiosity, what kinks did you find?
Then there are the massive shuffle read/writes that result in 50GB i/o which are not great for SSDs.