I would use map/reduce to solve this problem. You can run a single node psuedo-cluster in a few minutes by using Cloudera's package repositories (CDH3, https://wiki.cloudera.com/display/DOC/Hadoop+(CDH3)+Quick+St...). http://en.wikipedia.org/wiki/MapReduce