Announcing Spark 1.3
databricks.com
databricks.com
in any case, great product, nice usage of Akka and Scala, and a very intuitive API. I feel lucky to be working with it on a daily basis.
We have a few important improvements and changes to GraphX planned for 1.4, including Java API, and possibly a Python API.
DataFrames impose just a bit more structure: we assume that you have a tabular schema, named fields with types, etc. Given this assumption, Spark can optimize a lot of internal execution details, and also provide slicker API's to users. It turns out that a huge fraction of Spark workloads fall into this model, especially since we support complex types and nested structures.
Is the core RDD API going anywhere? Nope - not any time soon. Sometimes it really is necessary to drop into that lower level API. But I do anticipate that within a year or two most Spark applications will let DataFrames do the heavy lifting.
In fact, DataFrames and RDDs are completely inter-operable, either can be converted to the other. This means that even if you don't want to use DataFrames you can benefit from all of the cool input/output capabilities they have, even just to create regular old RDDs.
The first step of all my Spark tasks is "turn this RDD[String] into an RDD of parsed JSON", or turning CSV into case classes.
What JSON parser will dataframes be using? I presume Jackson?
You can find some public ones here:
http://spark-summit.org/east/2015/agenda
https://cwiki.apache.org/confluence/display/SPARK/Powered+By...
(Also, you probably know this already, but many people don't really have data big enough to necessitate a distributed framework. If your datasets are counted in gigabytes, you can do everything more simply on one machine and/or with a traditional database)
Spark uses a lot of Hadoop under the covers, so you still benefit from that ecosystem.
What I really like about it is that it can be easily unit and integration tested.
The only thing special we do here is explicitly test that any reduce function which is required to be associative is actually associative, due to a dumb bug I wrote once.
With the integration test, we have structured our code so that the processing occurs in a function that takes an RDD and returns an RDD. We then start a SparkContext in local mode[1], create an RDD with test data and can easily test that our processing produces the correct results.
We're just using JUnit for this, as we're largely a Java shop, so while we're coding in Scala for the cleaner API, we haven't jumped into the Scala ecosystem fully.
We also run end to end tests on our cluster using a subset of production data stored on S3 (Spark workers have to read/write from a distributed file system, S3 is one that the Hadoop ecosystem supports), and just verify the output against expectations derived from crunching that same subset via traditional means.
Hope that helps! I found Spark very easy to get up and running with, you can do a lot of experimentation in your IDE, and when you want to try a cluster, it ships with some convenience scripts that make it very easy to start a cluster on AWS.
[1]: http://spark.apache.org/docs/1.2.0/programming-guide.html#in...
1 cascading.org/ 2 https://github.com/deusdat/guacaphant
http://engineering.ooyala.com/blog/open-sourcing-our-spark-j...
Would strongly advise you to consider hadoop. We also used storm and found it to be much stable.
Databricks makes a lot of noise though.
I am partly asking this because you clearly feel strongly about this topic (your account was created an hour ago, most likely to comment on this).
Spark on YARN has been very stable and zero hassle for us. Run one command and it is deployed around the cluster and up and running. For performance Spark absolutely destroys stock MapReduce/Tez especially if like us you have a cluster with lots of basically unused RAM.
And Spark SQL, Spark Shell are both fantastic additions to the Hadoop ecosystem.