...because you need a control to evaluate how well your product is doing? I know it's a young field, but boy, do some folk love removing the "science" from "data science"
I'm not claiming that's what happened here, nor am I interested in nitpicking "what counts as 'science'". I'm just saying this is a reasonable thing to do.
But this is, of course, 1000 times more expensive to do. And if you only train 100, or 10, or 1 model, then the deduction becomes increasingly unstable.
So from a practical point of view, it's probably not feasible, because you would put those resources into something else instead that has more ROI.
Your suggestion of running 1000 training runs with different subsets of data sounds excessive and unnecessary to me.
It really depends upon the data. A smaller set of data that mostly consists of "truth" might be better than a larger dataset that also has many "lies".
Perhaps what you mean is that the model might be more representative, rather than _better_.
Your control in an online environment is the current baseline. You don’t need to save the test set anymore, you can push it online and test it directly.
I do not consider myself to be insane.
Doing this does complicate decisions for releasing subsequent model updates, as the production model can't be directly compared against new iterations any more. Instead a pre-production model would need to be used, that has not seen the test set. However, if data drift is likely, then re-using the old test set wouldn't be useful anyway.
More complication arises when users expect that things which worked previously in one way - continue working in this way. Users don't really care that their traffic was in the test set. In an even more extreme case, many industrial problems have a high correlation between the traffic today and the traffic next week, An optimal solution for such a situation would be to complete a full memorization today's traffic and use that for next week. In many cases, an overfit model can effectively perform this memorization task with fewer parameters/infrastructure than an actual dictionary lookup.
I'm simplifying now, but you can think of epochs as "how many times we train over the entire dataset? 1 time? 10 times?"
Correspondingly, you can think of dataset size as "how many Wikipedia pages we include in the dataset? 1 million? 10 million?"
Now let's think about overfitting.
What happens when you increase epochs is the model is more likely to overfit your data.
What happens when you increase dataset size is the model is less likely to overfit your data.
> Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.
If it performs about as well in instances it has never seen before (test set) then it's not overfit to the test.
2. Open ai is capped profit. They are also not a publicly traded company. researchers are researchers regardless of who they work for. Training on test data is especially stupid for commercial applications because customers find that out quick and any reputation is gone.
Hardware companies, which live and die on benchmarks, do this all the time. Meanwhile, it does appear that OpenAI is underperforming consumer expectations, and losing users quite quickly at this point, despite doing incredibly well on benchmarks.
Also, this isn't about profit. It's about market cap and it's about prestige. Those are not correlated to profit.
I don't know what you're talking about. GPT-4 is the best model out there by significant margin. That's coming from personal usage not benchmarks. A 10% drop in traffic the first month students are out of school is not "losing users quickly" lol.
ChatGPT didn't gain public use waving benchmarks around. We didn't even know what they were until GPT-4's release. The vast majority of its users know nothing about any of that or care. So your first sentence is just kind of nonsensical.
Anyway whatever. If that's what you believe then that's what you believe. Just realize you have nothing to back it up.
Basically the only people running benchmarks that could have been gamed on GPT-4 were other researchers, not companies, customers or users looking to use a product.
Normal users are certainly not running benchmarks and companies running benchmarks are running ones on internal data, which just defeats the whole point of gaming these research benchmarks.
Citation Needed.
But not an expert or OP!
For a particular model you try to minimally do this by separating a test and validation set, but on a meta-meta level, it's easy to see it happening.
"corroborate", you find queries of the same level which would give satisfactory output upon good performance but fail in a faulty overfitted model.