Tenzing: A SQL Implementation On The MapReduce Framework
research.google.com
research.google.com
Tenzing is not ACID compliant - specifically, we are atomic, consistent and
durable, but do not support isolation.It seems to me that that whole industry (DW & ETL) is a dinosaur whose lunch is about to get eaten by some upstarts.
Isn't ETL just an acronym that means "I wrote this Perl script to populate the database"?
How on earth is that even an industry?
Where things get complex is in the Transform aspect of some jobs. Mapping disparate schemas is complex, often messy work. Especially when one (or both) sides of the ETL job have poor/no primary keys, foreign keys, or even are just "mostly standard" CSV files [shudder].
Also: some ETL jobs can get quite large. I know one guy who had to create an ETL system that continuously moved data from one 1200-table system into some other system. Crazy.
It may be difficult to understand how this is an industry coming from a web development/startup angle (big supposition there) but there are literally thousands of companies with lots of databases varying in age, size and complexity that need integrating, and plenty of companies competing for that work as either implementors or software providers. A perl script might do the job but most products focus on performance, reuse, ease of maintenance and compatability across many different database/file types.
I also get the impression that Exadata is a pretty impressive feat of engineering and, if you need to do what it's optimized for and are prepared to pay a few million per rack, it's a very good option.
Your second comment is true, however the DW industry has in the last year figured this out and started to embrace the "Big Data" movement. Informatica (the largest player in the DW space according to Gartner) added HDFS connectors to its latest release, for instance.
"Tenzing has read-only support for structured (nested and repeated) data formats such as complex protocol buffer struc- tures. <...> The engine itself can only deal with flat relational data, unlike Dremel [17]"
And from section #5.4 I assume that currently they use Dremel query engine, but are in the works of creating another one.
So that's 10 queries/employee/day. That screams "experimental". Still, this would be very nice.
and this was quite a while ago