Prediction Games
argmin.net
argmin.net
I guess it seems like parameters need to be "counted" differently or there's something misunderstood about what a parameter is, or whether and how it's being adjusted for somewhere. Some of the gradient descent literature I've read, makes it seem like there are sometimes adjustments for parameters as part of the optimization process, so talking about "overfitting doesn't mean anything" is misleading.
It just seems like something where there's a lot of imprecision in terms that is critically important, no definitive explanations for anything, and so forth.
The results are the results, but then again we have hallucinations and weird adversarial probe glitches suggestive of overfitting (see also e.g., http://proceedings.mlr.press/v119/rice20a). I might even suggest the definition of overfitting in a DL context has been poorly operationalized. Sure you can have a training and a test set, but if the test set isn't sufficiently differentiated from the training set, are you going to identify overfitting? I can take training and test sets with a traditional statistical model and if I define the test set a certain way, minimize overfitting results.
I guess I just feel like a lot of overfitting discussions tend to feel kind of handwavy or misleading and I wish they were different. The number of parameters has never really been the correct metric when talking about overfitting, it just happens to align nicely with the correct metric in conventional models.
On the contrary, if a printing press controller “overfits” to the printing press it’s installed on, that is actually pretty desirable!
So what are you actually trying to prevent when you want to prevent “overfitting”, and why?
If it was smarter, it could figure things out just from the language specification, but we're not there yet.
Overfitting is typically not a concern. You train on the last N days of user interactions, and because of the volume of data there isn’t time for the model to see an interaction twice.
So you don’t need a test set. Your performance metrics may go up or down in a day depending on data drift.
How are hallucinations suggestive of overfitting?
Overfitting is a tactical term, not a strategic one, and is heavily coupled to the specific implementation.
Scamming seniors over the phone is a strategy, pretending to be their grandson is a tactic.
So we should have done what, exactly, Ben?
It worked great for them. Current masterpiece from Netflix has 13 Oscar nominations! Every AI company should learn and apply this lesson!