Double Descent in Human Learning
chris-said.io
chris-said.io
If you force the model to fit your training data perfectly then it is no wonder that you can only start to optimize the size of your parameters after you have enough leeway to easily fit the entire training set.
However another way to achieve the same effect is to make the assumption that the parameters should be small(ish) explicit. Fitting the whole training set is nice but ultimately pointless if you're just fitting noise. If the actual data is a simple polynomial + random noise then fitting all the noise will give a worse estimate.
(There’s a math-ier follow up linked as wel).
I can imagine some tasks where people have the same problem. You're overfit to a very specific task, but outside it "when you have a hammer everything looks like a nail" and you end up doing something dumb.
Other signs of non-overfitting: Abstraction laddering, task decomposition, novel (i.e. unseen) joke explanation
I was playing Fallout: New Vegas on Wine, and for some reason, the music on my pipboy wasn't playing. I searched the internet for the Wine errors from the terminal to no avail, and as a last resort asked ChatGPT. It gave me step-by-step instructions on how to fix it, and it worked.
If that doesn't demonstrate that LLMs have some kind of internal model of the world and understanding of it, then I don't know what will.
I haven’t played this game so I don’t know what I’m searching for but this was my first result. Seems on the money, no?
I've been through this forum, many reddit posts and other sites - none of the solutions worked. What worked was that ChatGPT figured out that I need to add the following line to ~/.wine/system.reg:
[Software\\Wine\\GStreamer]
"DllOverrides"="mscoree,mshtml="
And install 32-bit version of gstreamer good plugins: sudo apt-get install gstreamer1.0-plugins-good:i386
If you happen to find these exact instructions anywhere on the internet, please share, as that will be enough to convince me that LLMs aren't anything more than glorified search engines.Otherwise, I can't help but be skeptical. If nothing else, it's plausible that LLMs have some kind of internal representation of the world.
How about this guy? You sure the first instruction is necessary?
I do think there is some decent ability to piece things together. But this example seems too niche.
The fact that it is useful for finding patterns that may be different than humans tend to find is not an indication of understanding of the underlying data.
It is no different than clustering in traditional stats. While those found patterns are sometimes incredibly useful, clustering knows nothing outside of the provided dataset.
As other have mentioned, Google's search results are actually really bad at finding novel results these days due to many factors like battling SEO tricks etc...
But while the results of LLMs is impressive, there is no mechanisms for it to have an 'internal model of the world' in their current form.
It may help to remember that current LLMs would require an infinity of RAM to be even computationally complete right now.
Without invoking your own self-awareness as an argument, how do you know that other people "understand" stuff, and aren't merely "finding patterns"? In other words, in what way do you define "understanding", such that you can be sure that LLMs have no such thing?
> there is no mechanisms for it to have an 'internal model of the world' in their current form.
How do you know that? We don't even know why humans have an internal model of the world. What if internal modelling of the world is just sufficiently-complex pattern-matching?
Predicting the "next token" requires an "internal model of the world". It might not be how we do it, but without something that acts like it I'd be very interested in how you think it comes up with its predictions.
Let's say it needs to continue a short story about a detective. The detective says at the end: "[...] I have seen every clue and thought of every scenario. I will tell you who the killer is:". Good luck continuing that with any sort of accuracy if you don't have some abstract map of how "people" act. You can see how I can think of a lot of examples that require something that acts as a model of the "world".
There's a definite structure and pattern to everything we do. This (to an LLM) hidden context gives rise to the words we write. To re-invent them, like it has to do, it must basically conjure up all this hidden state. I'm not saying it gets it right, I'm just saying that there is no other way than to model the world behind the text to even get into ballpark-right territory.
Anything that is computationally compute needs an infinite amount of RAM. This is not unique to LLMs or even to machine learning.
That doesn’t explain why it generalizes well, though.
By the way, the linked Colab notebook is missing a "self." in front of the "lamb" in the unused regularization branch of the fit function.
They show it happens across a variety of architectures and real world datasets