Announcing Spark 1.6
databricks.com
databricks.com
That said, as part of Project Tungsten, we have some ideas about a batch columnar format that can be shared by Python, R, Scala and Java, and that should be able to eliminate most of the inefficiency in serialization across process boundaries.
Btw, I was critical about the issue above, but I do love Spark, using it on a daily basis. :)
Thanks for the reminder!
uses groupBy
I'm pretty sure based on previous comments you've made that groupBy was one of the things you'd rather eliminate from the RDD api, because of the performance impact compared to reduceByKey (which is almost always what people should be using instead).
Are you at all worried about confusion if groupBy now performs ok on datasets, but not on rdds?
I don't know if the OutOfMemory exception can still occur in recent versions of Spark, but the performance impact of groupByKey is very real.
In the section titled "Future Directions for Spark Streaming" there is a paragraph about Event time and out-of-order data and Backpressure. This would blow my mind to be able to use; this is a real pain currently.
That said, during Spark Summit Databricks guys themselves were most excited about Dataset API. Looking forward to giving it a try.