Machine Learning as a Service
about.wise.io
about.wise.io
I do like Machine Learning as a Service as a loss leader. e.g. customer walks up to the door, can't really get the problem cracked with an out of the box solution, but instead you sell/him her on an expensive long-term consulting project. i.e. the IBM Model.
Does anyone know how one of the pioneers in the segment, Numenta, is fairing? They've been around for a while and seem to have recently changed their name to Grok Solutions.
Deep Learning as a service seems like something that could work as their are less knobs for the user to fiddle with. That being said, it does not seem like Deep Learning is quite there yet.
I have to wonder how founders/founders-in-the-making react when faculty members from their alma mater, from their own department no less, enter their space. Must be a little bit like having Google enter your niche.
Out of interest: what is/was the companies name?
Interesting analogy... given that Google is already in this niche (with their prediction API)
I don't think a web service would be broadly applicable. Perhaps in certain domains it would make sense, but bandwidth costs and duration would be a huge factor in most solutions.
The other option is doing it all in-house, the storage, the servers, oh yeah, and that whole ML thing. Not a bad thing, but not necessarily your core-competency.
In case dealing with terabytes of data it would probably make sense for the MLaaS operator to run points-of-presence in major clouds.
Thanks for the interest and feedback!
Joey Richards, Chief Scientist, wise.io
Obviously this offers an on-site option that Google does not, which might open up other realms of options. I'm mostly curious in how / how well / what range of problems they're capable of.
- Benchmarking against Weka, R, and Python is not exactly pitting your product against stiff competition. Skytree (skytree.net) is another company in the same space with the same focus. Benchmarking against them would be interesting.
- I thought it was amusing that in a company of < 20 people the five founders thought it necessary to adopt such grandiose titles: CEO (fine), CTO, Director of Engineering, Chief Scientist, Director of Data Science (exactly how are the hairs split between these four?)
(1) We chose to do our original benchmarks against R, Weka and sklearn because these are the tools that the vast majority of people currently use. You'd be amazed how many companies use Weka! That said, we do benchmark favorably against the other competition. We will be publishing a series of blog posts with these benchmarks. Stay tuned!
(2) Our titles in fact do mark a clear delineation between our respective roles and responsibilities, and this is well understood within the company. Perhaps the titles are a bit grandiose, but we have a very big vision for this company.
The Myna page explicitly states it uses multi armed bandit algorithms, the algorithm fits the properties of the following claims. Wise's use of 'patent-pending', 'machine learning technology' and 'deploy machine intelligence' gives the impression of hiding shortcomings with jargon.
As for whether my comment was misplaced or not -- I guess that's up to the community to decide if they care to do such a thing. I do admit it was a bit snarky, but I also genuinely do find the titles a bit amusing.
Tying in to another post on the front page right now, I do think that generalists are advantageous in early stage companies, and titles tend to be more appropriate with specialisation.
I've found that ML algorithms can be like race cars - you really have to know their quirks to get performance out of them. The opposite would be analogous to a luxury car - almost everything is taken care of by the computer - you don't shift gears and don't open the lid.
So, is this WiseRF a race car or a luxury car?
It's the variations on the standard RF that have interested me. Rotation Forests, where features are partitioned randomly and rotated (by PCA or random projection) before decision boundaries are drawn. Extremely Randomized Forests, where the node splits are completely random and not based on best possible Gini/Entropy gain along some feature. There's even an interesting use of deterministic annealing out there for incorporating unlabeled data points in an attempt at semisupervised learning.
These each have their own parameters to tune, but have had slightly different performance on different problem domains. Even models require different data-- most decision trees have a quite natural way of imputing missing values, but something like a Rotation Forest can handle neither missing values nor categorical data (unless you map it to m binary features). And that complexity spooks me away from Machine Learning as a Service, where one could start failing to understand his or her models. (Plus then I'd probably be out of a job.)
The typical parameters are the following: - Number of trees - Percentage of data used to train each tree - Maximal depth of the tree - Minimum information gain (although this can usually be set to 0 and use only the depth) - Minimum sample size (same as the minimal information gain) - Number of thresholds to try for continuous data
Tuning is also difficult because of the non deterministic nature of the algorithm. If you compare to sets of parameters it requires more evaluation in order to be sure if one set is better because of the better choice of the parameters or because the algorithm chose the right data samples. This effect decreases with the number of trees, but the training time is increased in this case.
I think random forests are really good for very large datasets, but in my experience for smaller datasets boosting (for example JointBoost or GentleBoost) can give better results.
a) genetic recombination in the form of sex might be "bagging across genes" that prevents tight co-adaptation (= overfitting on the evolutionary timescale)
b) the brain uses noisy discrete firing instead of continuous communication because it allows for "bagging across network topologies".
The bulk of his talk is about how dropout in neural nets can be used to accomplish some of the same performance benefits that model averaging gives on other algorithms (at a much smaller relative computational cost).
Here's the talk: https://www.youtube.com/watch?v=DleXA5ADG78
With the MLaaS platform, all of the model optimization is taken care of under the hood (we also allow users to do their own parameter selection / tuning if desired). Our super fast implementation, WiseRF (10-100x faster than RF in sklearn or R) enables us to efficiently explore the hyperparameter space.
Thanks again for your questions and comments.
Joey Richards, Chief Scientist, wise.io
(2) Soon, we'll be publishing a series of blog posts to benchmark WiseRF against competing implementations. Look for that next week.
Thanks for your interest!
There seems to be an emphasis on efficiency, although I don't think that most freely available machine learning libraries are fundamentally poorly implemented. One problem with these libraries is that documentation can sometimes be scarce. Another problem, for the 0.01% of companies which actually have "big data", is that they might not scale, whatever that means.
Regardless of the library used, one of the bigger problems may be that machine learning, if it's worth it at all, is inherently fickle and tricky. To make an overly broad conjecture: if an externally provided machine learning solution works well, either your data didn't require that much domain knowledge to understand (it was "obvious") or some external/outsourced firm has a deeper understanding of your data than you do. More of the former type of analysis might not necessarily be a bad thing, though.
Even in a project where ML is more important, I think it's usable...
Thanks for your interest. You can try it yourself by downloading a trial edition of the software. The page http://about.wise.io/wiserf/ mentions the algorithm we're using.
As for the stats, you can reproduce the numbers yourself if you're interested. The ML codes are written in C++ with a very strong emphasis on performance and a low memory footprint.
Give it a try and let us know what you think.
Damian Eads Director of Engineering, wise.io