Hadoop Is About Scalability, Not Performance
manamplified.org
manamplified.org
I've been working a lot with Hive lately--Facebook's creation that lets you basically type SQL and the Map/Reduce jobs are automatically generated. Hive vastly speeds up development, as you don't need to write custom Map/Reduce jobs to do the tasks you had in mind when working with straight Hadoop.
If you're looking to learn more about Hadoop and Hive, Cloudera puts together an awesome starter kit available here: http://www.cloudera.com/distribution
One thing completely puzzles me about how hadoop is positioned, on the one hand there is this accent on reliability and scalability, then on the other there is the single meta data node.
How does that work in practice, surely there must be some way to have multiple meta-data stores ?
It includes Hive pre-installed and Pig as well. It also has a list of tutorial exercises an sample databases for you to practice on.
I'm not really a Hadoop expert to be able to address the data store issue. I'm basically one of two researchers at our company that is testing out these new technologies for select tasks (parsing log data, machine learning algorithm implementations, quick statistical munging, etc). In our configuration we have 7 nodes and roughly 3 TB capacity (although we're scaling as our needs increase). For our needs it works quite well. There was an instance where one of our nodes went down while in the middle of a big processing job, and Hadoop kept chugging. Quite amazing.
Another thing to note is that Hadoop becomes incrementally more useful as the size of the files increase--which to my limited DB knowledge, is not the same for relational DB's.
Right; the typical configuration is to have one database that does transaction processing, and another separate database (a "data warehouse") that collects data from multiple operational DBs and runs analytic queries over it. Hadoop is basically just an alternative analytic query processor in this configuration.
Hadoop becomes incrementally more useful as the size of the files increase
I think the same would apply to parallel DBs: you are basically talking about just partitioning the data over the storage nodes, which is a common feature in both systems.
A non-zero price tag makes many applications impossible (uneconomic in practice). So that's a good reason to use a non-relational engine for them.
The reality of SQL implementations with the leading vendors is infrastructure overhead. Even scaling with many server inside one application is not really recommended:
http://wedonotuse.blogspot.com/2006/11/brief-intro-to-market...
Going to a truly shared infrastructure takes you out of the comfort zone so far you might as well dump SQL databases, since people only pick them because of familiarity.
1- See http://databeta.wordpress.com/2009/05/14/bigdata-node-densit...