Dataflow/Beam and Spark: A Programming Model Comparison
cloud.google.com
cloud.google.com
Seems as if they're trying to ride the wave of the recent upsurge in interest in dataflow in general (with sub-fields such as Flow-based programming, and implementations like Akka streams etc). That's OK, but hijacking a term for the whole field, is not.
One question though. Is there a reason you can't use the Spark Window functions? https://databricks.gitbooks.io/databricks-spark-reference-ap...
Take it with a grain of salt.
They claim this isn't about the length of code, yet they select the most verbose way to use Spark and proudly display how long it is.
I mean sure, lack of event-time based processing is known limitation of Spark (and a pretty annoying one - though it is supposed to be worked on) but there are ways to write about it without code made to look bad on purpose.
EDIT: come to think of it, this whole article is "spark streaming can't do event time" written in thousands of words with contrived examples attached.
The primary argument is demonstrated through color coding different logical bits, which end up being clearly portable and elegantly distinct in dataflow.
This is demonstrated in two ways:
1. The "juicy value add" code that does the aggregation is labeled yellow, and doesn't change across all the samples with Dataflow. With Spark, it needs to be rewritten for every use case. Similarly, for all colors.
2. In Dataflow all the colors are separate. This makes expressing your logic easier. In Spark, the colors mix in dramatic ways with every demonstrated use case.
As Tyler said, all this is described in the blog post itself, but I don't blame you for missing it, since it's a really long post :)
The Datasets thing is more interesting. There is no doubt that the way they unify the Dataframe/RDD programming model is better, but it is so new (1.6 only) we certainly haven't migrated to it yet. The documentation isn't huge, either: http://spark.apache.org/docs/latest/sql-programming-guide.ht...
Spark streaming really feels operationally immature compared to a lot of other stream processing frameworks, even Dataflow. The criticism is both unsurprising and warranted.
In spite of these shortcomings, it has pretty good Kafka integration, mostly uses the same paradigm as batch and plays well with hadoop infrastructure. Makes it a decent choice for many use cases
Regardless, most of these stream processing frameworks are still very much in early days and lack a lot of the sophistication you find in custom in house systems (such as found at... Google ;-). The open source world will no doubt catch up and overtake those systems, but right now there is still enough of a gap that it is rather painful.
Spark streaming pains:
1. Backpressure & stragglers. Duh. 2. Setup & tear down is still rough, even compared to Storm. 3. The whole context singleton thing means you need a new VM for each job, which annoys the #@$@#$ out of me. 4. Error handling isn't just unclear, it's kind of disastrous. 5. You can feel its "batch" heritage in lots of places, not just the stragglers. For some that is a feature, for me, a bug, even though with Storm I use Trident. 6. When a job runs amuck, it's a pain to recover from it. Storm is no picnic either, but it is indeed better.
Spark Streaming is a fully open-source project; although the Dataflow SDK is also OSS, my understanding is that the released version can only handle bounded datasets. Support for streaming (which is the major innovation, IMO) is only available in the form of stubs that call out to Google's paid, proprietary Dataflow service.
It's totally fine to compare an open-source project with a proprietary alternative, but I think it's odd that this article opens by talking about how the Dataflow SDK is being opened, and then spends all its time talking about proprietary features.
But assuming it's true. it's very welcome news! I'll be keeping a close eye on future releases.
As I understand it, Spark has event-time support coming soon as well. I think basic stuff is landing in 1.7. Not sure precisely what they have planned, but I can only imagine that Spark will also become an excellent platform for executing streaming Beam pipelines in due time. In the meantime, the streaming runner for Spark can either target those features which Spark does support well (i.e., processing-time windowing, in this case), or try to emulate those it doesn't (such as how it was done in the article).
Who do they write these things for anyways? It's not like we're college admissions or professors, just give us the straight deal. Unless you're pitching to schools and naive undergrads of course.
How come Google Cloud Storage can be used instead of HDFS? I'm comparing google/amazon/azure right now. Both Amazon and Azure have 2 types of storage options - the regular object storage (S3a and Blobs) and block storage (S3 and Data Lake Store). S3 and DLS can act as the file system for Hadoop themselves (meaning you can let the data sit there and fire up clusters just for processing when needed), but they cannot interface with tools like the regular storage.
Meanwhile, Google's storage is like regular object storage, but you can run map/reduce (dataproc) and Spark on it.
Technically, Cloud Dataproc clusters have both HDFS (on PD) for write/read-intensive operations (and scratch space) along with the GCS connector. GCS is not the default file system, however.
The Hadoop FileSystem interface doesn't really force an specific underlying implementation, and you can even use "local" filesystems without any issue, in fact, IIRC, MapR-FS is just an optimized NFS drive, i.e. a shared network drive.