Benchmarking Random Forest Classification
about.wise.io
about.wise.io
If you look at the implementation for ski-learn, each tree emits a normalised probability vector for each prediction, those vectors are simply multiplied together to get the aggregate prediction, so its not very difficult to do yourself.
Although regardless, you are applying a batch learning technique anyway. You want an incremental learner for big data.
Although I'm a big believer in streaming/online machine learning, it's not necessarily the best solution. There are many cases when batch is the better option, especially for big data. Anything historical, really.
We have been working hard to reduce computing times and memory footprint (though, there is still a lot of improvement on that side).
(Unfortunately, I cannot run your benchmarks myself, because the compiled version of WiseRF requires a newer version of glibc than the one on my cluster, and crashes.)
For me, the promise of in-the-cloud machine learning is that I can call 'train' method, and specify one single hyperparameter: training budget (i.e. $). Perhaps also the max time before I am returned a trained model.
That's it. Can you do that?
Would love to hear about your use cases & get you on the beta.
-Joey Richards, Chief Scientist @ wise.io