Neglected machine learning ideas
scottlocklin.wordpress.com
scottlocklin.wordpress.com
For example, the recent AAAI 2014 conference had a bunch of papers on online algorithms for various problems. [1] Likewise, COLT had five or so papers on online learning. [2] Same with KDD [3], SODA [4], and the many other conferences this year that accept papers about ML.
And learning in the presence of noise? Unsupervised learning? Feature engineering? I am literally doing multiple research projects in all of these areas right now! The only way I can imagine that you think they're neglected is that you just don't know where to look for them, because these topics are all over the place in my world. For example, one common term for "feature engineering" is "representation learning," and this was a big topic at this year's SDM conference, specifically w.r.t. data mining in networks.
Why can't you find a book you like for topic X? Maybe it's because researchers have little incentive to write books. You folks in industry could fix that. What with all your ridiculous market valuations of various mobile apps, surely you could scrape together enough funding to convince the experts in their field to write a book.
[1]: http://www.aaai.org/Conferences/AAAI/2014/aaai14accepts.php [2]: http://orfe.princeton.edu/conferences/colt2014/the-conferenc... [3]: http://www.kdd.org/kdd2014/program.html [4]: http://www.siam.org/meetings/da14/da14_accepted.pdf
The unlabeled data helps discover the structure of the data which in turns helps the supervised learning.
There is a lot of 2-phase approach to this problem, by first discovering feature with unsupervised learning and then using those features in supervised learning, but in many cases it's possible to do both at the same time.
If you have a generative model for your data, you can treat the labels as potentially missing values and learn the joint distribution.
Deleted comment
Most cookbook techniques such as ridge regressions, cross-validation, etc have a Bayesian interpretation as a prior on the parameter.
Bayesian techniques allow you to use all the data available.
That said, sometimes they are computationally expensive, and it's better to approximate them by using a test set.
It comes down to verbiage really - some general form of label propagation vs censored data.
Some algorithm make the estimation easy and the distribution quite implicit, others make the distribution explicit but training is harder.
Umm is there a typo there ? Otherwise it is so wrong that I cannot even begin to pick the flaws.
In fact the key revolutionary idea in classification theory, courtesy Vapnik, was that one needn't estimate the density at all.
I cannot over-emphasize the revolutionary part, prior to Vapnik, that over-arching consensus was that one should try to learn the the density and then threshold it to obtain the classifier. The motivation for this is the result that the thresholded conditional probability is the optimal classifier for a {0,1} loss.
What Vapnik showed was that it is better to learn the classifier directly (he called it structural risk minimization, SRM, for short), without any parametric assumptions whatsoever. The main reasoning is that density estimation is freaking hard, and depending on a freaking hard task as a pre-requisite is backwards. Learning to classify is actually easier than learning densities. It was one of those paradigm shifts that come to a field only once in a while.
Even if we let SRM be, even in traditional statistics there have been methods as old as time that do not assume any parametric model. Those into non-parametric statistics (I am not one) would be very upset with your claim :) The problem with parametric models is that you almost never know the parametric family. You can of course test, but how many would you test, there are infinite number of such families. Non-parametric methods dont make parametric assumptions, (they do make some and considerably weaker assumptions. Assumptions are necessary for learning). The drawback is that they are more expensive, similarly if you happen to know the correct form of the family parametric model would be hard to beat.
The idea of boosting comes to mind. A quintessential ML concept, widely applicable, but really nothing to do with density estimation.
All the PAC learning results are another example. They show very strong results over very wide problem classes, but don't have to do with density estimation. They apply, for example, to problems where no density exists (just a probability measure), much less a smooth-enough density to estimate.
Another example of ML research that comes to mind is time and space efficient learning. There is a whole body of work on probabilistic analysis of data structures used for ML (I'm thinking partly about Alex Gray's work, http://www.cc.gatech.edu/~agray/). The work fits solidly into ML, but it's not about density estimation, it's about making provably fast algorithms for large problems.
Yes, "parametric" is often meant to imply a finite parametrization, but that's a dumb terminology.
PAC and MDL arent as much as at odds that you make it up to be. If you are familiar with PAC-Bayes and PAC-MDL theorems you will know what I mean. My main issue was with the ridiculous (or with your definition of whats 'parametric' then a vapid) claim that all of ML is parametric.
Secondly, wanted to highlight that its a bad way to go about ML, by estimating data distribution first and then find the decision function. This is the precise reason why SVMs broke out of the then state of the art. The two stage 'plug-in' approach used to be the preferred way then. A decision function might be Bayes with respect to some density, but thats besides the point, it is ill advised to match the density and then derive the decision function, except for very special and narrow cases.
Even if you throw PAC out, parametric stats (am using the standard definition here) is a tool with very narrow scope. If it happens to be in the current scope then by all means use it.
BTW you make good point about transductive learning, it got compared to semisupervised, but they are not the same thing. Both use unlabeled data, but for transductive, the points were you seek labels have to be given ahead of time. This enables very strong guarantees. I think its potential has not come to fruition yet.
EDIT: @murbard2 > Re: See what I did there.
Color me violently unimpressed. That is a not a parametric distribution.
I continue to stress that plugin estimates are in general not a good idea. Even die-hard Bayesians would not use it, they would rather estimate the MAP or the mean aposteriori decision function directly, rather than fit data distribution and then derive the optimal decision function from it.
I find Bayesian inference to be a good idea, except that it is costly unless you use conjugate priors (these are motivated more by convenience than by data) and its "turtles all the way down" problem.
Its not that you have advocated plug in estimators, but given your claim about everything about ML is about fitting parametric densities, it is likely to mislead readers into believing that the two step way is a good idea.
You like SVM? Fine, then why do large margins matter at all? "Intuition" ?
Give me f: A -> B and I'll give you P : (A x B) -> R+ See what I did there?
Much of this stuff is being actively worked on though. If I could give one practical tip. Read KDD conference papers. Those are very applied and usually very accessible demonstrations of what techniques are out there, what problems they are typically applied to and importantly how well they worked.
Excellent post.
That's not a surprise; "online learning" is a pretty misleading name. Here are some better alternatives:
* incremental learning / training
* stream-based learning / training
So, let's use these terms instead -- then it will be easier to find in your favorite search engine.Here is one example using the term "stream-based learning":
http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.1.79...
"online learning" algorithm
That gives me 10 relevant hits out of 10 on the first page.See: http://en.wikipedia.org/wiki/Online_machine_learning
> Online machine learning is a model of induction that learns one instance at a time.
I would contend that there can be a "significant" seeming amount of literature on field X but field X may still wind-up not pursued in the larger scheme of things.
Often what happens is a single individual or small circle, gets interested in a given field and researches it among things for as long as the funding persists and then once the funding dries up they move on. Or one person has tenure, keeps researching but everyone else moves on because it doesn't look like a way to keep getting funding.
Even more, as the author mentions, a big question is what approaches are taught as the way to do it (and I guess it again comes down whether you're aiming for just machine-learning/a-better-heuristic-statistics-for-big-data or if you are aiming for moving towards intelligent algorithms, even if intelligence means just flexible adaptivity).
Yes, you can find lots of results if you search for "online learning", say. Otoh, for whatever given algorithm that has mindshare currently, is there a quest to find an online version? My sampling of the literature says no and I happen to agree with the author that online processing could be an important piece of artificial intelligence advances.
Edit: found them thanks to Google image search: http://www.darkroastedblend.com/2014/01/machines-alive-whims...
Agreed, very nice attribution. It wasn't added in response, I noticed it 8 hours ago (last night for me).