Does analyzing this actually require joining in large data sets -- that is, larger than will fit on a single machine?
I'd always assumed that the records involved weren't very large, but I don't know much about the problem space, so I'm not sure if other data gets joined in in a way that benefits from cluster-based analysis.