Machine Learning Showdown: Apache Mahout vs. Weka
blog.algorithmia.com
blog.algorithmia.com
Your point about being difficult to use is exactly the problem that Algorithmia solves.
None of these (Mahout and Weka) are mainstream anymore. For large-scale classification, people are using packages like VW[1] . And for small-scale experimentation, SciKit or R.
Because you want to coalesce all your model updates for every pass over the data, the long startup-time for hadoop jobs actually plays into this. You can have hundreds of millions of samples, in a sparse space of hundreds of millions of parameters, and do it faster on a single node using VW than in a hadoop cluster of equivalent nodes running Mahout.
For large scale, distributed stats I'd go with SparkR.
Looking at the graph number of trees vs accuracy, I would have expected that the line would asymptotically reach a maximum accuracy given more and more trees; however for weka it looks quite wavy and for mahout it even looks as if there's an optimum and more trees are worse.
Or is it just noise and I'm interpreting too much?
http://imgur.com/mYRdIC0 http://imgur.com/cLQn74O http://imgur.com/bg2UaJN http://imgur.com/66oUGdM http://imgur.com/oU92V09 http://imgur.com/qvILooJ http://imgur.com/URyRiOB http://imgur.com/AybUI20 http://imgur.com/kugzG2D
Example: http://imgur.com/mRRz1L3