Spark: Open Source Superstar Rewrites Future of Big Data
wired.com
wired.com
I know this because it's my full-time job to actually get stuff done inside Hadoop.
Spark may be a great system, but this article doesn't do much to settle the issue. When you read fluff like "sweeping software platform", "famously founded the Hadoop project", "great open source success stories" and machine learning described as "crunching and re-crunching the same data -- in what's called a logistic regression", it's time to move on.
It would be difficult for an article to convey this. Having used both Hadoop and Spark, practicality and expressiveness are precisely what made me fall in love with Spark. You can do so much with so little code. Don't take anyones word for it, see for yourself. Download the code and run the interactive shell, takes 2 minutes. It was totally mind-blowing for me.
Wired is a mainstream technology magazine.
This is where marketing and branding becomes the primary factor influencing adoption, and not technical merit. Hadoop gathered so much momentum and hype as part of the Big Data buzz in the past few years that it's only now beginning to percolate through to telecommunication carriers and other larger/slower-moving enterprises.*
* I work primarily with wireless carriers; can't say much about broadband, although I'd hazard a guess and say that the majority are only now allocating experimental budgets to see how Hadoop can help them manage their Big Data.
Among the things I love most about the Spark and the eco-system:
* repl -- so great for running short little experiments or even full blown jobs. saves you time compiling small changes and really get to know your data quickly.
* caching -- processing in-memory opens up many possibilities beyond doing iterative, machine learning jobs. quantifind, for example, demoed a system that allows them to run ad-hoc queries using Shark on the fly across GBs of data (think OLAP, a bit) in seconds or less.
* scala -- makes for very succinct code using closures and built-in operations; check out some examples here: http://spark-project.org/examples/
And some of the upcoming projects are also very cool. Tachyon, for example, will enable users to share data with a very robust in-memory file system. A teammate and I, for example, could have used it recently because we were simultaneously running different analysis against the same data, so we had to cache duplicate instances on two clusters.
Spark make doing MR easy. I've used other frameworks on hadoop MR, but nothing compares with the ease with which you can express computations using it. And it does both batch and real-time/streaming. It is a very well thought-out project.
Additionally, I believe that HDFS keeps 3 copies of the data around on 3 different nodes for redundancy. So there is the overhead of that network traffic.
Sounds like, with that configuration applied, the in-memory performance difference between Hadoop and Spark should not be nearly as large.
The programming abstraction treats all data as collections (RDD in Spark terminology), and allow programmers to apply bulk transformation on these collections. Some examples of operations you can apply include the traditional map and reduce, the relational filter, join, outerJoin, and more advanced ones like sample. This abstraction makes it much easier to write distributed programs. As the Wired article mentioned, a distributed program written in Spark often looks identical to a single node program. This substantially reduces the amount of code one needs to write for distributed programs, and the best part is the code really expresses the algorithm (rather than cluttered with JobConf setup).
And the scheduler and the engine itself are aware of the general DAG of operators, so they can schedule and run those operators better. For example, if you have multiple maps, the execution gets pipelined; if you are joining two collections that are partitioned the same way, the execution avoids an expensive shuffle step.
There are many other benefits too. I'd encourage you to give a try. Thanks!
Disclaimer: I am on the Spark team at UC Berkeley.
It doesn't look like there's been any direct comparison of the two, though it looks like there's overlap. (I've wanted to start a streaming data processing project, and this looks like it would be good to consider for it.)
Anyway, the takeaway that I think you were downvoted for omitting is that if you can bypass the serialization, startup, and shutdown steps associated with the traditional Hadoop iterative process you'll get a huge speedup.
* Here is an example of a simple job written using Scala for hadoop (and uses Mahout libraries) - https://github.com/sandys/distributed-scala-mahout/tree/wiki...
* you can embed pig inside jython
* you can write UDFs using jruby or jython.
* I didn't try to figure out how, but I'm pretty sure you can build a standalone job jar using jruby and warbler
There might be a different way of hooking into Spark through Scala or python, but I'm pretty sure that fundamental support is not the big advantage here.
Pig has a repl - but as I quickly realized from playing around, that you end up mucking around with classpath problems (and questions of having your Jars in distributed cache) once you attempt to build UDFs involving a few different libraries.
Plus, I haven't used and so cannot comment on the ecosystem of Cascalog/Clojure - which is as functional as you can get.
Then again, while I absolutely love doing work with big data, I've been having a bit of an "existential crisis" since the NSA leaks :(
Book recommendation: "Who owns the future" by Jaron Lanier.
Can someone explain this to me?
Nobody ever got fired for building with Java.
Considering there are already lots of BigData tools in Java ecosystem (hadoop, hive, pig, mahout etc.) Scala looks like a very reasonable choice.
I wish I found Scala as easy to use as I do Go. I do not enjoy the syntax, nor do I consider running on top of the JVM a selling point.
I came from a Java, but worked in Python for many years, so that probably explains my bias.
It's written on the JVM for the simple reason that if they wrote it in Go or Erlang, no Enterprise would adopt it as there isn't a CTO at a non-tech Fortune 500 that has every heard of Erlang or GO, and wouldn't know the first thing about trying to hire developers for it. Remember, the jobs written for Map Reduce are done in the same language ( typically ) as the MapReduce code itself.
Of course, there are tasks for which an Erlang implementation would be faster, but as others have mentioned, most organizations would prefer to write Scala.
(Yes, I find that icky too. :-))
At Twitter we don't program using Hadoop directly; we mostly use either Scalding or Pig, languages that compile down to Hadoop code. https://dev.twitter.com/blog/scalding http://www.slideshare.net/kevinweil/hadoop-pig-and-twitter-n...
I believe this is how many other companies use Hadoop as well.
The benefit here is that it's possible to write new backends for Pig and Scalding that compile down to Spark or anything else. And then you have backwards compatibility with all your old big-data code.
I worked on an entire team of developers where I was the only one who understood the raw Java MapReduce API. Almost everyone else on my team got by with learning HiveQL, and a very minimal understand of MapReduce design flow.
Your belief is absolutely correct.
This implies that the confiscated servers were originally not located in the US. I wonder which country they were located and on what legal basis the US could confiscate servers there?