MapReduce vs. SQL: It’s Not One or the Other
gigaom.com
gigaom.com
Also, if this gets more points than the earlier link to the paper itself I will be sad.
EDIT: sadness.
It's all in the headline. If the paper was titled "Researchers confirm it: MapReduce doesn't scale" (and it if wasn't a PDF) it would have a much higher score.
I've never heard of Greenplum and Teradata. I had however used a number of RDBMSes for five years at three different jobs. I read a book and took a class on relational databases. Persistent storage and concurrent access were central features of any database I've ever used or learned about. The databases the authors of the article use seem to be of the same kind. For example, they first load the data into a database and then use SQL to query it.
There are no equivalent concepts in the MapReduce framework. It doesn't make any sense, for example, to talk about loading data into MapReduce. A map reduction is typically a process that takes input from some place in some form (could even be a relational database), performs some computation on it, and writes the output in another (possibly identical) form. I don't see how you can say the two are competitors, any more than Perl is a competitor to MySQL.
I actually know quite well what MapReduce is typically used for, and I strongly doubt relational databases are used for precisely the same (or even somewhat similar) tasks.
They are parallel databases, which is what the article is comparing MapReduce against. The basic idea is that you partition the data over a set of machines, and similarly divide and parallelize query processing.
MR is basically a specialized parallel DB that executes a single query of the form:
SELECT map(...), reduce_agg(...)
FROM ...
GROUP BY map(...)
Again, MR's architecture differs in various ways from the design of a typical parallel DB, but they are certainly comparable.It doesn't make any sense, for example, to talk about loading data into MapReduce
MR input data is typically stored on a distributed filesystem (e.g. HDFS, GFS). Loading the data onto that filesystem is analogous to loading the data into a parallel DB; in both cases, the input data is partitioned over a set of nodes. For good performance, you want a copy of the data to reside close to where the subsequent computation is going to take place. But certainly MR and DBs differ in this respect: a DB typically assumes it "owns" the data, whereas MR is more amenable to processing data stored in an arbitrary external format. This difference is not fundamental, though.
I strongly doubt relational databases are used for precisely the same (or even somewhat similar) tasks.
You would be mistaken, then. If you want to do analysis on large volumes of data using a cluster of machines, both parallel DBs and MR are viable choices. The whole point of the paper is to take a set of queries, execute them using both Hadoop and parallel DBs, and compare the resulting performance and user experience.
Wow, that's a lot of cluelessness in a single sentence.