How I came to love big data
37signals.com
37signals.com
Not to mention the massive cost-savings from using the right technology with a small footprint versus using a brute-force approach and a large cluster of machines.
Aren't those ... largely opposite? To say "very adhoc" i would anticipate that as largely meaning "not predictable". Also...can you perhaps quantify "very large"? How many terabytes? I've done a decent amount of work with the "new school" OLAP approaches (hadoop / mapreduce, etc.) and found them to work quite well especially in certain cases such as time series (think weblog analysis) where sequential scanning is a simplistic approach.
I guess I'm comparing it against better-suited-for-brute-force approaches where someone is analyzing a log file for really random things that tend to be a one-time thing. "Show me hits to this particular resource from IP addresses which match this pattern and the user-agent contains Safari and the response time is larger than 300ms and the response size is less than 100ms!". While you could fit this data easily into a DM, you'd need to plan ahead for that sort of querying. If it's a non-frequent occurrence, it'd make more sense to process it sequentially (even if its across 10,000 machines in a hadoop cluster).
When I left the company, our Greenplum cluster (so a bit of both worlds: RDBMS cluster that automatically parallelizes queries across multiple nodes and aggregates the results) was around 500T. This approach was scaled up from a single MySQL instance though, which were seeing around 5 million new rows per day for one particular business channel.
I'm not suggesting that "new school" approaches do not work nor not necessarily work quickly. What I am suggesting is this: MR is a very naive approach that is only "fast" due to executing the problem in parallel across many nodes. If you have datasets which are going to be queried often in similar use cases, one should take the past ~30 years of innovation in RDBMS' instead of masking the difficulties by throwing a lot of CPU (and therefore money) at a problem and solving it in the most inefficient manner possible. It pains me to see people coming up with overly complex "solutions" to basic OLAP needs on Hadoop-based or even NoSQL platforms instead of simply using the right tool for the job.
That said, the one-off cases where it doesn't make sense to build out a schema, ETL pipeline, and managing a database because it's a very niche or one-time need: That's where the real value of Hadoop/MR comes into play.
It really depends on what you want to do. If you want to compute a statistic over a large amount of data or build a predictive model, some sort of sampling may create virtually identical results with orders of magnitude less time/memory. In modeling, sampling away certain classes may actually be essential for algorithmic stability.
Sometimes sampling is wrong. If you have to create statistics over your entire dataset (let's say, number of transactions aggregated by user), sampling may not be worth the loss of precision, and it won't really save you much time either. In general, you never want to sample away data in a way that the confidence in the statistic you are calculating becomes low.
Now, whether or not the inferential premises of statistics hold up on website data (and population data) that typically is neither random nor representative, that's another story.
I definitely think Clojure has a future in big data - the repl, immutability and the strong links to the Java ecosystem make it an obvious choice for developing interactive (in so far as this is possible) big data applications.
Also, it's fun to code, which should never be discounted as a reason for a language to succeed.
Suppose you are interested in user behavior. You want sample(group_by(user_ident, all_user_events)), not group_by(user_ident, sample(all_user_events)). This involves running a group_by on the full data set.
In short, not all problems are big data problems.
Regardless, Impala sounds like it could be pretty sweet!
I had to smile when I read that. Working with data, sometimes optimization or redesign can yield significant performance gains. (Especially when reworking some of my colleages queries or code...)