MapReduce and Spark
vision.cloudera.com
vision.cloudera.com
All of these are just productivity gains; not to mention the performance gains you get when you go from MapReduce to Spark.
http://spark.incubator.apache.org/docs/latest/quick-start.ht...
The DAG used by spark represents how one job/partition of data depends on another job/partition and what methods (e.g. filter) need to be applied on the parent data to get the child data. This is useful when a node goes down and that portion of data has to be recomputed. Note that users can choose to persist some intermediate results to hdfs to avoid recomputation in case of failure.
"Performance: Launches ~1000x faster, runs ~10x faster"
"Launch scaling: Hadoop (~N), MR+ (~logN)"
"Wireup: Hadoop (~N2), MR+ (~logN)"