MapReduce and RDBMS: Practice and Theory
grigory.us
grigory.us
Actian Vortex has shown that you can have a fully ACID compliant database running on top of Hadoop all while providing exceptional performance. http://wwwcdn.actian.com/wp-content/uploads/2014/06/AAP-Hado...
Apache Spark is getting close to being able to do both, but still as a developer building a data stack, I would not inspect terabytes of data every single time if 80% of questions can be answered by looking at data once and saving summarized results in relational format.
I thought Hadoop vs RDBMS was a fight settled may be 4-5 years ago! Amusing to see it being raised at this time.
On the other hand, we both build and use summary tables with Spark. (in a relational format to boot, and using Spark SQL).
I think you would benefit from re-evaluating the assumptions you made 4-5 years ago.
I know they're orders of magnitude slower than e.g. Vertica today, but I wonder if there is a fundamental reason for that, or is it just the implementation?
Spark is starting to amass all the heavy hitters in big data so no doubt it will get much faster over time.
Columnar MPP needs to be carefully defined to answer the question. Some important data models that fit in these architectures operate on data types that are not meaningfully order-able at a mathematical level i.e. you can't sort them. A lot of columnar implementations, and virtually all in open source, assume sortability as a property of the represented data types.
tl;dr: Columnar MPP sometimes is not far outside what you can express with Parquet/Spark/MapReduce/etc, just much faster, but there are data models supported in some advanced Columnar MPP systems that are not usefully expressible with that stack. It depends on the platform and the use case.
I also don't understand what you mean by orderable. Spark does not require records to be oderable. Maybe you can elaborate?