- GPT style language models end up internally implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples): https://arxiv.org/abs/2212.10559
There's more. People are just in denial.
- GPT style language models end up internally implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples): https://arxiv.org/abs/2212.10559
There's more. People are just in denial.
GPT models are constructed with pretrained gradients which are applicable in a large set of situations. It’s just an optimization technique, albeit a clever one.
Quoting from the paper:
In summary, we explain ICL as a process of meta-optimization: (1) a Transformer-based pre- trained language model serves as a meta-optimizer; (2) it produces meta-gradients according to the demonstration examples through forward computa- tion; (3) through attention, the meta-gradients are applied to the original language model to build an ICL model.
There's no difference.
The forward computation is computing gradients. These new gradients are applied to the model, to build the icl model via "attention".
What I said: "implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples):"
Perhaps you're not getting it. Forward computation is the neural network processing an input. The paper is saying gradients are built here. These new gradients are then applied through the transformers "attention" step. The icl model is the model with the attention applies.
It is literally just another perspective of what transformers do.
Precision in speech is critical when discussing complex subjects.
It took me a couple minutes to "figure out" just what you mean and even now I don't fully get it.
Are you in actuality complaining about the word "implement"? You're pedantically arguing that the word doesn't belong?
That's the least ludicrous intent for all the possible intents behind your reply.
Even your pedantism can't win here though. You are in fact wrong, implement is an appropriate word here, the only thing I understand here is how mistaken you are.
That aside, you’ve been arguing that these models understand things and citing these papers as evidence. They are not. They are evidence that the ability of these models to generate text based off of their existing training set can easily be finetuned in a number of ways to add training sets after the initial zero-shot learning.
That’s all these models do, they generate text based upon some training set. If we define understanding as the ability to extrapolate beyond what one has been told, they are expressly not doing that. Your papers explain this quite well.
Edit: to see more concrete examples of this, look into the unfortunately named “hallucination” ability of LLMs. Once you realize that they only know what they were told and are unable to logically extrapolate the point becomes clearer. I hope that helps.
Not only am I extremely well versed in the definition and philosophy of the word science (likely much more well-versed than you), but you are completely and utterly wrong about the compliment part. A scientist is a human, if I call a scientist pedantic during a normal discussion then the scientist will take it as an insult. Do you think a scientist has conditioned his mind into a sort of emotionless robot who can needlessly branch off onto a debate about the definition of the word "science" and "pedantry" when the topic is actuality "machine learning"? No. A scientist can both be pedantic and stupid, being a scientist does not preclude one from being human.
>That aside, you’ve been arguing that these models understand things and citing these papers as evidence. They are not. They are evidence that the ability of these models to generate text based off of their existing training set can easily be finetuned in a number of ways to add training sets after the initial zero-shot learning.
I posted two papers. You're conveniently ignoring the first and naively mistaken about the second.
Part of "understanding" is the ability to formulate new theorems from previously known facts, in order to do this one must "understand" how these facts compose to form new statements. This is what's happening in the fine tuning. It is a demonstration of understanding... that it knows how disparate knowledge composes to form new knowledge. The very definition of understanding.
>That’s all these models do, they generate text based upon some training set. If we define understanding as the ability to extrapolate beyond what one has been told, they are expressly not doing that. Your papers explain this quite well.
Of course. You cannot extrapolate anything beyond what you Observe as well. Can you literally form new knowledge out of thin air? No. You have three things: Existing knowledge, knowledge through observation, and knowledge through composition of existing knowledge.
Without introducing new knowledge, LLMs can be coerced to compose existing knowledge to form new knowledge. Additionally, in the ICL step they can be introduced to new knowledge and form additional compositions there. This has been demonstrated repeatedly.
>Edit: to see more concrete examples of this, look into the unfortunately named “hallucination” ability of LLMs. Once you realize that they only know what they were told and are unable to logically extrapolate the point becomes clearer. I hope that helps.
It's obvious chatGPT makes stuff up. Every one who has worked with LLMs in depth is fully aware of this. It's an obvious thing, you don't even have to "look it up" everyone knows about it.
This claim is made DESPITE the fact that LLMs hallucinate. It's obvious these models are imperfect and it's obvious they have huge deficiencies. But when it doesn't hallucinate, when the answer is Novel, creative, correct and unmistakably not existing in any training set, then we know the model understood the query you gave it.
I mostly agree with your position but have a quibble with this characterization. Knowledge can also be generated from randomization and enumeration. For instance, we could enumerate all Turing machines that might satisfy some property, or we could randomly permute some Turing machine as in genetic algorithms to find some new behaviours.
You might be tempted to categorize these under "composition", but I think they have different properties from composition, which is typically understood to be a finite deterministic map. Enumeration is potentially unbounded, and random mutation is non-determinstic.
Simply add a new input neuron with a random seed or add a tiny bit of noise to some weights or seed the tokens yourself in your query.
"<Query Seed: 4334> hello chatGPT, how are you?"
In this case chatGPT can deliberately randomize the response through understanding the intent of what you mean by "Query Seed".
In the brain, if such randomness existed, it would largely be modelled as a similar mechanism. A seed value (or multiple seeds at different places) is either inserted near the input step or happens at a branching logic step. Additionally, the query to the human brain can also be seeded.
True randomness requires that this seed number comes from quantum properties of particles that expresses itself in a sort of macro level random number. There is no other known source of true randomness in nature, though we can get perceptually identical results through just seeding with timestamps.
Either way it's trivial to add and not critical to what we mean by the word "understanding" because it's both easy to add and we aren't even sure if we have such randomness is in our brains. If it existed and if a human had this mechanism removed from his brain we would still say that this human is capable of "understanding" things.
Randomness is trivially addable. Additionally it's arguably not a source of source knowledge.
Its just a selection parameter. Which sources of knowledge do I use for composition and in what way do I compose it out of available compositions? The randomness parameter can influence these steps.
If you think there are other sources of knowledge, then what are they?
For example, if I have trillions of atoms randomly composed together with the objective to form a perfect cube, it will take eternity for a valid solution to arise.
If I have pre-existing cubes of atoms already pre-formed into cubic lego bricks and I have these components compose randomly, then it is far more likely that I get a cube.
LLMs, can do this with additional reinforcement training. chatGPT specifically is effective because it has this on top of the original GPT-3 model.
This is essentially random selection of known datums. It happens at the genetic level as well on a higher level.
The generation of raw novel datums in genes, however, is something that cannot happen in human brains. The mechanism for natural selection cannot exist within the brain itself, unless you're calling the trial by error process "natural selection." But again a human doesn't select a completely random strategy to "trial" he will select a strategy out of a known set. This is again, selecting a random datum.
Keep in mind, actual pure randomness generates useless noise 99.99999% of the time.