Building a Data Intensive Web App with Hadoop, Hive, & EC2
cloudera.com
cloudera.com
Full source code on Github: http://github.com/datawrangling/trendingtopics/tree/master
Dataset on Amazon Public Data Sets: http://developer.amazonwebservices.com/connect/entry.jspa?ex...
Cloudera have taken ~$11M in VC funding so far[1]; is there really ~$110M in profit to be made off Hadoop consulting, training and support in the medium term? I wonder.
One possibility is that they're using consulting to build revenue and mindshare in the short-term, and using the capital they've raised to launch something more substantial in the longer-term (say, running their own cloud/hosted Hadoop service).
[1] http://ostatic.com/blog/hadoop-centric-cloudera-gets-6-milli...
HN insights will be valuable, thank you!
I was reading this website - http://www.metabrew.com/article/anti-rdbms-a-list-of-distrib...
I have not tried HBase and HyperTable myself yet, but the blog post says that they still have latency issues. What are your views?
Also, after processing the raw log data with Hadoop, we only need to store/lookup 3M records in the MySQL presentation layer, which is well within the capabilities of a tuned RDBMS. Many Rails sites are backed by MySQL, so I thought linking Hadoop/Hive to a common data workflow would make for a good example.
I've been hearing that recent improvements to HBase 0.20 could make it a contender: http://stackoverflow.com/questions/1022150/is-hbase-stable-a... and some high volume sites like Mahalo are already using it. That said, there are other alternative data stores (Cassandra, Voldemort, Tokyo Tyrant) that might be worth exploring if a database isn't cutting it for you.