Google Dumps MapReduce in Favor of New Hyper-Scale Analytics System
datacenterknowledge.com
datacenterknowledge.com
This is not even true. They have recently published research that involved using MapReduce in their own systems. Example: http://research.google.com/pubs/pub41376.html
1. Google's infrastructure is evolving atop MapReduce (FlumeJava/MillWheel).
2. Google's PR decided to call it "not using MapReduce anymore" because in marketing, "beyond <current fad>" sounds really cool.
3. The rest of the world's PR/press/marketing fall for Google's clever PR.
Either way, it is great to see Google making its core technology accessible as part of their PaaS =)
Thanks for the clarifications :)
If "hyper-scale Analytics" counts as "clever PR," then eating paste counts as "clever food choice." It's not worth my time to decode the BS.
But what do I do when people start describing themselves as "developer evangelists" ?
That has to be some kind of sell signal, right ?
Citation: http://dl.acm.org/citation.cfm?id=1806638
The key word is pipeline. If you have some analysis that runs in several stages, you'll be taking the output of one stage, and connecting it to the next. If you want to compose multiple phases, chained together, raw MapReduce isn't going to help you very much with the chaining.
What's described in the paper is a way to do the chaining in a nice way. The system will take care of writing the raw MapReduces for you. But it'll also do a lot of work on the interconnections between your stages as well.
For example, Spark provides the primitives needed to build GraphX (http://amplab.github.io/graphx/, http://spark.apache.org/graphx/), which is essentially GraphLab on Spark.
Buzzword... overload!
Here are the relevant papers...
* FlumeJava (iterative, data-parallel pipelines like Spark): http://pages.cs.wisc.edu/~akella/CS838/F12/838-CloudPapers/F...
* MillWheel (fault-tolerant stream processing like SparkStreaming): http://research.google.com/pubs/pub41378.html
Pointers to the IO blog posts...
* "Reimagining developer productivity and data analytics in the cloud" http://googlecloudplatform.blogspot.com/2014/06/reimagining-...
* "Sneak peek: Google Cloud Dataflow, a Cloud-native data processing service" http://googlecloudplatform.blogspot.com/2014/06/sneak-peek-g...
The Dataflow-specific talks at Google IO 2014...
* Big data, the Cloud way: Accelerated and simplified https://www.youtube.com/watch?v=Y0Z58YQSXv0
* The dawn of "Fast Data" https://www.youtube.com/watch?v=TnLiEWglqHk
* Predicting the future with the Google Cloud Platform https://www.youtube.com/watch?v=YyvvxFeADh8
* Keynote (starts at Urs Hölzle's segment on Google Cloud) https://www.youtube.com/watch?v=wtLJPvx7-ys#t=6932
The rest is buzzwords propping up sweeping ridiculous conclusions.
I dont think there's a lot of companies where data would reach this huge. Anyone has any idea on how large a typical warehousing database is?
While I couldn't quickly find anything that speaks to any kind of average or normal size of a data warehouse, this article mentions Facebook's being around 300PB: https://code.facebook.com/posts/229861827208629/scaling-the-...
Seriously.