I'm one of the authors of the blog post as well as this new API. Feel free to ask me anything.
For some context - In our case, loading a reasonable set of data from HDFS can take upto 10-30 mins so keeping a cached copy of the most recent data with certain columns projected is important.
I guess this means DataFrames should be used all the time in the future, or will there still be a reason to use plain RDDs in the future?
You guys are doing great work !