Big Data Lambda Architecture
databasetube.com
databasetube.com
Not every problem can be reduced into a completely cache-able batch job, trivially parallelizable across all of your data. 'Big Data' isn't about breaking up your batch processing into three layers, it's about being smart enough and knowledgeable enough in compsci, statistics, calculus, text processing, regexes, machine learning, business analysis, et cetera, to design an effective system which harvests useful insights from a large bank of atomic, messy, inconsistent data, with an appropriate level of availability and consistency.
The real work is not in using/configuring Hadoop—it's about figuring out what information would bring greater-than-marginal value to a business, and how to compute that efficiently from an existing corpus of data.
There's no silver bullet. Remember?
EDIT: I think the following is particularly disingenuous: "The lambda architecture solves the problem of computing arbitrary functions on arbitrary data in real time by decomposing the problem into three layers"
This is such a ridiculous promise, that it put me in a strongly skeptical mood for the rest of the article.
Also, it seems like "delta architecture" would be a better name since the "speed layer" is all about updating views with deltas anyway and the architecture clearly cannot do what it promises.
The view update problem is by far not easy even if you take a subset of a well-defined language like SQL. And the article really seems to assume arbitrary queries, i.e. executing even an arbitrary program on the input data. Just a simple example: "If x is there then remove y" and you have other rules inserting and removing x and y. Then you easily run into a query which given the data does not have a single unambiguous solution and in some cases (loops with an odd number of negations) it is not even clear at all what the solution should be and if it can be defined.
I'll repeat myself: most big data problems fit into the mold of query = function(data), map reduce is a practical substrate for building algorithms to compute these functions, and this paper presents a practical architecture to implement these types of systems.
When he says arbitrary functions, I think he's basically comparing it to the limitations of SQL solutions that are traditionally used for data warehousing. SQL is abused beyond belief and also slow beyond belief for certain types of queries.
MapReduce provides a pretty good place to start from. Sure, it's awkward for certain things. But if you're going to rearchitect a distributed system from scratch for each computation you do, you're not going to make progress very quickly.
See this paper for the opposite viewpoint: http://news.ycombinator.com/item?id=4520057
He provides solutions to common problems which can be expressed with MapReduce.
The lambda architecture is more powerful than what's being advocated in that article (plain mapreduce), although I'm not sure if it's really a thing people are using or just a subject of a book.
I have read the recent draft of the "Big Data" book by the author, which describes the architecture that the article discusses in better detail. Honestly, if you are a beginning practitioner in this field, you can't really go wrong by reading it.
Note that you can do a lot with streaming algorithms (it's not just counting). Also the reduced memory usage (orders of magnitude) makes the complexity of random writes not such a problem as you have less need to go outside a single machine.
[1] Slides on streaming algorithms: http://noelwelsh.com/streaming-algorithms/2012/11/22/streami...