How to process a million songs in 20 minutes
musicmachinery.com
musicmachinery.com
This looks like a fun project, but I can't help but feel like Hadoop experience reports are a little late to the party at this point. Is there anyone out there who doesn't immediately think MapReduce when they see numbers at scale like this? If anything, the tool is overused, not neglected.
The cost though... that was shocking! It comes out to about $42/month to host the data and a $2 per analysis run. That's very affordable if you're producing results someone might want to pay for.
Maybe the headline is a bit misleading but I find the post interesting and valuable.
however, in paul's case here he's really just using MR as a quick way to do a parallel computation on many machines. There's no reduce step, it's just taking a single average from each individual song and not correlating anything or using any inter-song statistics.
Do you get back info such as how long each instance took to do its job?
And since you pay per minute on EC2, you can either chose to have 10 instances calculate a minute or 1 instance for 10 minutes. Either way, you pay for 10 minutes of processing time.
p.s. simplified but probably somewhere in a reasonable ballpark
For fast iterative development this isn't an issue - you can re-use the job flow (the instances you spun up), so you can launch the job several times over the course of that hour.