Map Reduce: A really simple introduction
ksat.me
ksat.me
I would create an n-ary tree (n is arbitrary, maybe 100) from coworkers. The 'leafs' would produce an array of ten numbers based on the blogs assigned to them. Then each node in the tree would take n arrays and sum them up into one array, until me, the root have the final array of ten numbers.
First: this is much simpler than what is described in the blog. Second: the 'grouper' in his example is totally a bottleneck. (And in fact the reducers are also bottlenecks in his case. While my system scales practically to infinity.)
More complex logic happens of course - but problems that "fit" for MR tend to follow a similar pattern to counting.
I just do counting in MapReduce, but it was to use Expectation-Maximization on mixtures of Gaussians to cluster a data set of 16 million documents. Although computing the Guassian's isn't really trivial, all you do at the end is sum up and normalize.
It's very cool and isn't much more than counting (replace the map step). In fact the paper I took it from (and the whole Mahout on Hadoup thing) is a bunch of machine learning algorithms that are just big summations at the reduce step.
As an aside, it would be cool to have several "stories" like this for simple map-reduce. I've found that sometimes it takes multiple attempts for the concept to get across.
Is anyone aware of a good primer on the subject?
Generally speaking, it sounds like a lot of the divide and conquer CS algorithms could potentially use this approach, assuming other constraints (IO and network volume) aren't an issue...
Reduce - Gather and aggregate/collate the results from the workers
That's it. That's the big thing to understand. The fact that everyone keeps looking for some new way to describe MapReduce I think adds to the mystique and the belief that it is somehow profoundly complex.
Aside from the many typos, I don't see how this is useful. MapReduce is an absurdly simple concept. Once you litter any explanation with analogies or scenarios, you distract the comprehension of the reader.
Am I to think about parsing text on a blog? Bad CEOs? Google and programs? What?
In fact as a more general rule I would say that when people try to explain something, it is lazy, almost always confusing -- and usually an indication of the writer's own ignorance -- when said explanation resorts to contrived scenarios or analogies.
It's like if the UPS shipper had to deliver an elephant and twelve penguins, one of which suffered from gastroenteritis. By driving the truck using the hybrid energy recovery system, just how much conflict-gas that was fueled by the death of thousands of soldiers would a Catholic priest yield?
I feel like seeing the essential (as you've done) and then finding examples that are specific implementations is how I learn best.