The more data we use to train a model (or as you said, the more patches we use), the better it’s performance will be.
But that’s definitely not needed most of the time in real life for an average person, just like it’s not needed for an average developer anymore.
Time will tell I guess.
Here is hope they use something like category theory mixed with philosophy to put it on a secure foundation
In that case, just make new problems. If it is being 'patched' to pass specific known problems, then the new ones would fail.
If it is able to answer them, then maybe it is actually analyzing them and working out the solution.
Not sure how you can assume there was no underlying improvement, and these are cases of feeding it the answers.
Compare
> And it's only fixed for the stated case, but if you reverse the genders, GPT-4 gets it wrong.