The main trick in machine learning
edinburghhacklab.com
edinburghhacklab.com
In order to make it tractable, you pick a finite model space, train it on finite data, and use a finite algorithm to find the best choice inside of that space. That means you can fail in three ways---you can over-constrain your model space so that the true model cannot be found, you can underpower your search so that you have less an ability to discern the best model in your chosen model space, and you can terminate your search early and fail to reach that point entirely.
Almost all error in ML can be seen nicely in this model. In particular here, those who do not remember to optimize validation accuracy are often making their model space so large (overfitting) at the cost of having too little data to power the search within it.
Devroye, Gyorfi, and Lugosi (http://www.amazon.com/Probabilistic-Recognition-Stochastic-M...) have a really great picture of this in their book.
over-constraining your model space means having too few parameters in your model But fixing your data size, the "power" of your search goes down when you increase the number of parameters.
So it is no so much an issue of avoiding errors, but of choosing the right number of parameters for your model.
It seems like you can "mis-power" your model also.
For example, the Ptolemaic system could approximate the movement of the planets to any degree if you added enough "wheels within wheels" but since these were "the wrong wheels", the necessary wheels grew without bounds to achieve reasonable approximation over time.
That would be an example of over-constraining your model (i.e. imposing the arbitrary constraint of a stationary Earth).
A system of Ptolemaic circles can approximate the paths taken by any system. So the system really isn't absolutely constrained to follow or not follow any given path.
You could claim you have constrained your model not be some other better model but that, again, seems like a poor way to phrase things since a more accurate model is also constrained not to be a poor model.
Even specifically, the Newtonian/Keplerian system has the constrain of the sun being stationary as much as the Ptolemaic system has the constraint of the earth being stationary.
Edit: As Eru points out, the Ptolemaic system basically uses the Fourier transform to represent paths. Thus the approximation is actually completely unconstrained in the space of paths, that is it can approximate anything. But by that token, the fact that it can approximate a given path explains nothing and the choices that are simple in this system are not necessarily the best choices for the given case, estimating planetary motion.
* The model could encompass the behavior of the input in a smooth fashion if it's basic parameters are relaxed.
* The model would tend to start finding models that are wildly different from the main model at the edges (space and time) if its parameter are relaxed, even if the model would eventually find the real model with enough input and training.
one has to handle these two conditions differently, right?
It's not really arbitrary--given the understanding at the time, there was no ability to measure the motion of the earth. In particular, stellar parallax which was understood as a contra-indication and too small to measure just yet. So a non-stationary Earth went against what they knew at the time rather strongly.
That said, relativity comes back and makes choosing a frame of reference arbitrary in the end, though some are easier to do physics in than others.
Machine learning is almost like learning chess in that there are certain obvious mistakes that noobs continue to make. And like chess there are multiple levels of thinking and understanding that are almost impossible to teach to someone that doesn't have lots of experience. Hopefully more blog posts like this will help people get past the novice level.
Regarding technical content:
N-fold Cross validation [1] can be a more effective approach to having a single held out or validation set. You split your data into N groups, say N = 10. Then you use groups 2-10 as a training set to make predictions on group 1, then groups 1,3-10 to make predictions on group 2, etc. Recombine the prediction output files and use the measured error to tune and tweak your predictor. It's more work and can still lead to overfitting, but it's generally better to overfit the entire training set than it is to overfit one held out sample.
[1] http://en.wikipedia.org/wiki/Cross-validation_%28statistics%...
Of course the basic topic of validation gets pretty deep fairly quickly too. Out out of bag scores anyone?
1) Have lots of data
2) Accept the possibility that your problem domain cannot be generalized.
I always find, whether in academic literature or in message boards, a desire to fit every round peg into a square hole. The reality of real world data is that sometimes, it's just a 50/50 coin toss. This might be because the features that really indicate some sort of pattern can't be defined or they can and the data can't be reliably retrieved, or the humans running things have a poor understanding of the problem domain to start with.
TL;DR: There's no magic
(I'm not disagreeing, just referring to a different kind of "magic")
Everything else matters, but when your ML doesn't work it's 100% a feature selection problem. Which usually means it's 99% a problem of getting lots of domain expertise jammed up against a lot of ML experience and mathematical understanding. It's also a bear.
1. Feature selection.
2. intelligent data massage. Real world data has usually noise that humans can easily identify as irrelevant or erroneous.
3. logit regression.
Starting with simple, well understood algorithms first should be the second lesson after knowing about validation sets. In those cases where they are not enough, they set the baseline for comparison against other algorithms.
Personally, I wish people emphasized more the importance of a general understanding of econometrics when doing machine learning. In most of the introductory courses I've seen, the link between both field is never made explicit, despite the obvious analogies (coincidentally, there was an article by Hal Varian on the front page two days ago that discussed how both fields could benefit from sharing insights [1]). Understanding the idea behind minimizing generalization error is one thing, but I find that thinking in terms of internal/external validity and experiment design often gives people a more intuitive understanding of validation procedures, both regarding why and how we should do it. The same goes for understanding effect size, confidence intervals, causality (and causality inference), and so on.
One great one from my Machine Learning professor was an assignment where we were required to normalize our data to [0,1]. After doing this and then going through the typical cross-validation cycle, he had us try and figure out where we contaminated our validation sets. As it turns out, we all normalized our data before splitting it up, which meant that training data influenced testing data.
It's a simple fix, but if you've done that and gone to run a large convolutional neural network for a week only to find that you made a stupid error like that, it can be pretty painful. (Especially since the bad generalization error might not be obvious until you use it the model in production)
I used to work in algorithmic trading (the kind which aims build consistent viable portfolios, not the HFT arms race).
This of course relies heavily on building your model, which can be anything from some simple linear regressions to more advanced techniques more commonly associated with the buzz word of machine learning, this applies to all predictive methods. You begin searching the training data to find optimal model parameters and then verifying performance on the validation set. The number ONE mistake I saw most was that when you get bad results on the CV set, going back to step 1.5 instead of just throwing the whole model out. To take your same core idea, tweak it slightly, add/remove a few parameters and restart the process. Unfortunately doing this enough times and your CV set starts to become the training set. Thus leaving your true validation set the day you turn it on live in production with real money.
It's never a good feeling to see your positively skewed returns in your training, testing and "CV" set morph into essentially a mean zero random distribution in production. This was quite an important lesson to learn for me.
In Andrew Ng's "Machine Learning" offering on Coursera he talks about having three sets of data:
1. Training data. He uses this for fitting most model parameters.
2. A second set for "more general" analyses -- judging the effects of additional data, regularisation parameters, neural-network topology etc. Performance on this data is used to decide which model to use and how to use it.
3. A third set to estimate how good the choice of model is.
The theory is that the parameters in #1 are fitted to the training data, and the model choice is "fitted" to the data in #2. Even though we think (hope?) that the inferences made in those two steps will generalise reasonably well, we should still expect measures of fit from those analyses to be optimistic. We need a set that has not been used for calibration to reliably estimate how good our model will be on data in the field.
I still think there is something quite fundamental, though, about validation sets and other related resampling-based methods for estimating generalisation performance (cross-validation, bootstrap, jackknife and so on).
The built-in picture you get about predictive performance from Bayesian methods comes with strong caveats -- "IF you believe in your model and your priors over its parameters, THEN this is what you should expect". Adding extra layers of hyperparameters and doing model selection or averaging over them might sometimes make things less sensitive to your assumptions, but it doesn't make this problem go away; anything the method tells you is dependent on its strong assumptions about the generative mechanism.
Most sensible people don't believe their models are true ("all models are false, some models are useful"), and don't really fully trust a method, fancy Bayesian methods included, until they've seen how well it does on held-out data. So then it comes back to the fundamentals -- non-parametric methods for estimating generalisation performance which make as few assumptions as possible about the data and the model they're evaluating.
Cross-validation isn't the only one of these, and perhaps not the best, but it's certainly one of the simplest. One thing people do forget about it is that it does make at least one basic assumption about your data -- independence -- which is often not true and can be pretty disastrous if you're dealing with (e.g.) time-series data.
Bayesian model averaging entails P(X) = P(X|M1)P(M1) + P(X|M2)P(M2). It assumes that either M1 or M2 is true. No conclusions can be derived from that. It might be useful from a purely predictive standpoint (maybe) , but it has no place inside the scientific pipeline.
There is a related quantity which is P(M1)/P(M2). That's how much the data favours M1 over M2, and it's a sensible formula, because it doesn't rely on the abominable P(M1) + P(M2) = 1
Model averaging can be quite useful when you're averaging over versions of the same model with different hyperparameters, e.g. the number of clusters in a mixture model.
You still need a good hyper-prior over the hyperparameters to avoid overfitting in these cases though, as an example IIRC dirichlet process mixture models can often overfit the number of clusters.
Agreed that model averaging could be harder to justify as a scientist comparing models which are qualitatively quite different.
Yeah, but in this case, there's a crucial difference: within the assumptions of a mixture model M, N=1, 2, ... clusters do make an exhaustive partition of the space, whereas if I compute a distribution for models M1 and M2, there is always M3, M4, ... lurking unexpressed and unaccounted for. In other words,
P(N=1|M) + P(N=2|M) + ... = 1
but
P(M1) + P(M2) << 1
Is the number of clusters even a hyperparameter? Wiki says that hyperparameters are parameters of the prior distribution. What do you think?
I quite like "Bayesian reasoning and machine learning" too: http://web4.cs.ucl.ac.uk/staff/D.Barber/textbook/090310.pdf
Generally we're in danger of overfitting when the cardinality of our data is comparable to or less than the cardinality of our parameters (including meta-parameters like which model to select).
What I just described is a perspective derived from Bayesian model selection. But Bayesian model selection encompasses other types of model selection; it need not be considered a separate path.
The expected risk is the sum of empirical risk (training set error) and the structural risk (model complexity).
In many instances, having low empirical risk comes at the cost of having high structural risk, which is overfitting.
http://infolab.stanford.edu/~ullman/mmds.html
> There are some who regard data mining as synonymous with machine learning. There is no question that some data mining appropriately uses algorithms from machine learning. Machine-learning practitioners use the data as a training set, to train an algorithm of one of the many types used by machine-learning prac- titioners, such as Bayes nets, support-vector machines, decision trees, hidden Markov models, and many others.
There are situations where using data in this way makes sense. The typical case where machine learning is a good approach is when we have little idea of what we are looking for in the data. For example, it is rather unclear what it is about movies that makes certain movie-goers like or dislike it. Thus, in answering the “Netflix challenge” to devise an algorithm that predicts the ratings of movies by users, based on a sample of their responses, machine- learning algorithms have proved quite successful. We shall discuss a simple form of this type of algorithm in Section 9.4.
On the other hand, machine learning has not proved successful in situations where we can describe the goals of the mining more directly. An interesting case in point is the attempt by WhizBang! Labs1 to use machine learning to locate people’s resumes on the Web. It was not able to do better than algorithms designed by hand to look for some of the obvious words and phrases that appear in the typical resume. Since everyone who has looked at or written a resume has a pretty good idea of what resumes contain, there was no mystery about what makes a Web page a resume. Thus, there was no advantage to machine-learning over the direct design of an algorithm to discover resumes.
It has been found that, for most problems, a simple model which well represents previous experience should be accepted instead of a more complex one with marginally better representation. I would claim that the reason this generally works is an empirical discovery, as opposed to a mathematical result, but probably has philosophical implications in its success.
A man is looking around at the ground under a street lamp. You ask him what he is looking for, and he says "I'm looking for my keys. I dropped them somewhere in that parking lot over there." "Then why are you looking inder this street lamp?" you ask. He answers: "Because this is the only place I can see!"
For example if I have a linear model, Y = a + b * X, I will choose a and b to minimize in-sample fit. Choosing a and b to maximize out of sample fit goes against all theory.
However, if I want to choose which parameters go into my model, maximizing out of sample fit would be a good approach.
So at the end of the day, there is not a huge philosophical difference between using in-sample and out-of-sample fit, only different approaches to the same problem. In both cases, the assumption is (usually) the the data is i.i.d., and in both cases, you are choosing some coefficients/parameters/hyperparameters with the intent of maximizing out of sample fit, but using different methods.
To me, even the statement "if I have a linear model" makes very little sense from the perspective of ML. Contrast with "if I think I'm dealing with a situation where a linear model might offer a good fit".
Regarding "maximizing out of sample fit would be a good approach", I think ML is always and just-about-only concerned with maximizing out-of-sample fit, for if it wasn't, the solution would be a lookup table.
I'm not trying to imply that you're wrong, rather that I think the 'gulf' is real. Or maybe I'm misunderstanding your point. For example, I feel that mjw's comment in this thread captures my view, which I think is more ML centered: https://news.ycombinator.com/item?id=6878336
Is that comment also in accord with your view, and it's me that's on the wrong side of that gulf?
I've been getting into ML lately for my startup, it's a personal finance system that will learn your habits and use that to predict things in the future. It's been overwhelming attempting to move into this domain of software engineering (so much so that I am currently just hard coding certain important patterns and using basic statistical modelling instead) but it is absolutely fascinating!
TL;DR: Randomly, a certain percent (5-10%) of data is 'hidden' and never used for building/refining your model, but is only used to evaluate how well your model fits (or explains) that unseen data. This is absolutely, fundamentally essential to prevent over-fitting your data!!
EDIT: Think that you are solving a huge jigsaw puzzle, but made of thousands of jello pieces. You randomly hide a 100 or so pieces and try to solve the puzzle. Having used all the pieces (except the hidden 100), you think the puzzle forms a Treasure Map. Now, you take the previously hidden pieces and try to fit those into the puzzle and if after using the hidden pieces your puzzle still looks like a Treasure Map, you may have found a (mostly) correct solution. But, if you are unable to fit those hidden places in a way that still keeps the Treasure Map intact, you must question if you did in fact find the correct solution or if there is another, slightly different, solution that may be (more) correct because it will account for the hidden pieces a little better?
Understanding the motivation behind validation is an absolutely fundamental concept, and lack of coherence on the topic shows an inherent lack of understanding of the goal of building the model in the first place; GENERALIZATION.
This is synonymous with one checking in code that has no issues locally, without testing in the stack or a production environment.
I work and hire in this space and it's actually a bit shocking how widespread this lack of understanding is. Asking a candidate how to evaluate a model, even at a basic level, is this field's version of FizzBuzz. Just like Fizzbuzz, a lot of candidates I've encountered who are "trained" in machine learning or statistics fail miserably, and my peers seem to have similar experiences.
These issues are expected, given how popular data science is these days. We all win when more people are getting their hands dirty with data, but it's extraordinarily easy to misuse the techniques and reach misleading conclusions. This can potentially lead to people pointing fingers at the field and it's decline. The only thing we can do is correct the wrongs and do our best to limit incompetence that only serves to tarnish the field.
The real trick (for most algorithms) is to select the correct features to train against. This really is more of a black art than an exact science, so I think labeling it a trick is justified.
Validation is not a magic bullet, we need to be critical of any part of the model that is given as truth, otherwise we might end up fitting a solution to the wrong problem.
More generally I think that textbooks should emphasize the need for the scientific method and stress that any model (or theory) is only as good as its ability to explain the entire problem domain.
Of course there are many different theories, but that's my favourite.
It's very interesting in the sense that the totality of brains over time is essentially a sort of supervised learning with huge amounts of input data.
The brain contains/is the model. It is trained by a range of inputs and by definition it generalizes outside those inputs.
If you're asking how does the brain minimize out-of-sample error? It does that by the virtue that it's model isn't too complex for the training set, just like what you do in machine learning. If the brain had a model that was too complex it would overfit and poorly generalize just like machine learning would do with a too complex of a model...
Some people would have trouble handling something that had \lim_{a \to x} (some complicated f(a,x,y)) where y is a constant even though they could handle it with standard notation.
For another possible example, take something you've written recently, replace all the variable names with things like Integer, Double, and the function names with For, While (within the syntax of the language) and then try reading it.
Besides this, there's the jesus-in-toast, man-in-the-moon, face-on-mars business. The brain overfits everything, but it never stops training. It's in constant reinforcement learning.
I have seen new PhDs read about it "in theory", but not internalise it for practice, and then they go off an do Bayesian structure learning without a validation set. This DOES happen.
This post is to hammer into the brains of any beginner thinking about machine learning that understanding the validation set's purpose is the most important thing to internalise first.
e.g. Machine learning is easier than it looks: https://news.ycombinator.com/item?id=6770785