Introducing Apache Mahout
ibm.com
ibm.com
In the more traditional data-mining areas of clustering, latent variable discovery and supervised classification, Mahout does scale and does deployment very well. I am a Mahout committer, but I use R all the time, often times for prototyping or for small analyses. I would hate to have to deploy an R solution, however. Sampling is a fine solution for the first 80% of gain and if you are in a startup situation, that may well be enough for you. On the other hand, efficient deployment is usually pretty important as well.
As always, you mileage may vary.
I would recommend that you pop over to the Mahout mailing list for better feedback. I doubt that the Hacker News community knows as much about Mahout as the people who develop and use Mahout.
Does any body know of any large scale data mining use of Apache Mahout?
Often times you want to run an experiment by training a classifier, testing it on a development data set, tweaking it and starting over. In the mean time, you just sit and wait. Reduce that cycle from a couple of hours even days to a shorter amount of time is probably the most appealing aspect of a project like Mahout.
We're not using Mahout, but I keep an eye on it since it might be an improvement to our current solution.