I'd be really interested at moving forward with the tech but for the time being I do have to stay where I'm - in my primitive cubicle cell.
I'd be really interested at moving forward with the tech but for the time being I do have to stay where I'm - in my primitive cubicle cell.
e.g. At a national bank, determine whether distance to the nearest branch or ATM correlates with deposit frequency or average customer relationship value. Your inputs are a) 10 billion timestamped transactions, b) 50 million accounts, c) 200 million addresses and dates for which they began and entered service, and d) a list of 25,000 branch/ATM locations and the dates they entered service.
This is fairly straightforward to describe as a map/reduce job. You could do it on one machine with a few nested loops and some elbow grease, too, but the mucky mucks might want an answer this quarter.
I know the feeling, though: I keep wanting to try it, but haven't been able to find a good excuse in my own business yet.
And if they don't, the cost of supporting this "straightforward" Hadoop infrastructure, both in terms of hardware, engineering and support, is so massive that the little elbow grease for a simple I-know-what-it-does solution may well be worth it.
In other words, I share alexro's concerns. If you're buying into the M/R hype and process your blog logs in the cloud, that's one thing. But legitimate business use-cases are probably not as common as people may expect/hope.
That's not true. A small cluster and support from a company like Cloudera is much less than an new Oracle install.
If you're just looking for a quick example: imagine you want to parse 6,000 webserver logs (each being 100+GB) to determine the most popular pages, referrers and user agents. The map phase extracts all fields you want and the reduce phase counts all the extracted unique values.
"mapreduce" is about processing at scale. You can always do the same thing on a smaller set of data locally (and usually with just sed, awk, sort, and uniq).
Clearly, Riak does not expect each and every query to run for hours and involve TBs of data. MapReduce is just a straightforward way to process distributed data.
With a HDFS cluster you can cost effectively dump undifferentiated data from existing sources into a "data lake" without worrying about complex and highly selective transforms, and then use tools like Pig and Hive to do ad-hoc interrogations of the data.
Most data warehouse implementations fail to a large extent because of the ETL problem - Hadoop could help solve that in a big way.
Further reading (no connection to me) http://www.pentaho.com/hadoop/
Hadoop has a long, long list of companies using it[2]. Cloudera has a similar list[3]
[1] http://open.blogs.nytimes.com/2007/11/01/self-service-prorat...