This article predates GPT-3 and GPT-2, it even predates the essay "The Bitter Lesson" <http://www.incompleteideas.net/IncIdeas/BitterLesson.html>.
It might be true long-term, but it's certainly not written with the current advances in mind.
This article predates GPT-3 and GPT-2, it even predates the essay "The Bitter Lesson" <http://www.incompleteideas.net/IncIdeas/BitterLesson.html>.
It might be true long-term, but it's certainly not written with the current advances in mind.
it will still have no awareness
We dont know how awareness works, so we're not in a position to say what has and hasn't got it.
Think about the milliseconds in which ChatGPT uses its 4000 token length to analyse approx 2000 words all at same time, in a process encompassing a massive number of GPUs processing a mind boggling number of parameters. How are we to say what it going on for those milliseconds? There could be some sort of abstract analogue to awareness happening there for a burst of a few milliseconds. These LLMs are all about emergent effects from huge scale. Similarly, me and you are made out of 30 trillion microscopic dumb biological robots, none of whom know who we are or care, but nonetheless, we have awareness.
This isn't even a debate anymore between two guys on the internet. There is literal science behind this and the difference between us is one person is behind on the science.
But forget the science. There is even chatGPT output that literally shows it understood what it was told.
It is clear chatGPT does not have perfect understanding of the world. It does create really stupid output. But this is ignoring the fact that there are tons of answers it gives that show unmistakably that it knows what you asked it.
Context dependence is literally the novel thing with transformers, they are context dependent statistical models to generate next words, make it huge and feed it a page of context and it can map those to a page of output to match that context.
- GPT style language models end up internally implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples): https://arxiv.org/abs/2212.10559
There's more. People are just in denial.
GPT models are constructed with pretrained gradients which are applicable in a large set of situations. It’s just an optimization technique, albeit a clever one.
Quoting from the paper:
In summary, we explain ICL as a process of meta-optimization: (1) a Transformer-based pre- trained language model serves as a meta-optimizer; (2) it produces meta-gradients according to the demonstration examples through forward computa- tion; (3) through attention, the meta-gradients are applied to the original language model to build an ICL model.
There's no difference.
The forward computation is computing gradients. These new gradients are applied to the model, to build the icl model via "attention".
What I said: "implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples):"
Perhaps you're not getting it. Forward computation is the neural network processing an input. The paper is saying gradients are built here. These new gradients are then applied through the transformers "attention" step. The icl model is the model with the attention applies.
It is literally just another perspective of what transformers do.
Precision in speech is critical when discussing complex subjects.
It took me a couple minutes to "figure out" just what you mean and even now I don't fully get it.
Are you in actuality complaining about the word "implement"? You're pedantically arguing that the word doesn't belong?
That's the least ludicrous intent for all the possible intents behind your reply.
Even your pedantism can't win here though. You are in fact wrong, implement is an appropriate word here, the only thing I understand here is how mistaken you are.
That aside, you’ve been arguing that these models understand things and citing these papers as evidence. They are not. They are evidence that the ability of these models to generate text based off of their existing training set can easily be finetuned in a number of ways to add training sets after the initial zero-shot learning.
That’s all these models do, they generate text based upon some training set. If we define understanding as the ability to extrapolate beyond what one has been told, they are expressly not doing that. Your papers explain this quite well.
Edit: to see more concrete examples of this, look into the unfortunately named “hallucination” ability of LLMs. Once you realize that they only know what they were told and are unable to logically extrapolate the point becomes clearer. I hope that helps.
Not only am I extremely well versed in the definition and philosophy of the word science (likely much more well-versed than you), but you are completely and utterly wrong about the compliment part. A scientist is a human, if I call a scientist pedantic during a normal discussion then the scientist will take it as an insult. Do you think a scientist has conditioned his mind into a sort of emotionless robot who can needlessly branch off onto a debate about the definition of the word "science" and "pedantry" when the topic is actuality "machine learning"? No. A scientist can both be pedantic and stupid, being a scientist does not preclude one from being human.
>That aside, you’ve been arguing that these models understand things and citing these papers as evidence. They are not. They are evidence that the ability of these models to generate text based off of their existing training set can easily be finetuned in a number of ways to add training sets after the initial zero-shot learning.
I posted two papers. You're conveniently ignoring the first and naively mistaken about the second.
Part of "understanding" is the ability to formulate new theorems from previously known facts, in order to do this one must "understand" how these facts compose to form new statements. This is what's happening in the fine tuning. It is a demonstration of understanding... that it knows how disparate knowledge composes to form new knowledge. The very definition of understanding.
>That’s all these models do, they generate text based upon some training set. If we define understanding as the ability to extrapolate beyond what one has been told, they are expressly not doing that. Your papers explain this quite well.
Of course. You cannot extrapolate anything beyond what you Observe as well. Can you literally form new knowledge out of thin air? No. You have three things: Existing knowledge, knowledge through observation, and knowledge through composition of existing knowledge.
Without introducing new knowledge, LLMs can be coerced to compose existing knowledge to form new knowledge. Additionally, in the ICL step they can be introduced to new knowledge and form additional compositions there. This has been demonstrated repeatedly.
>Edit: to see more concrete examples of this, look into the unfortunately named “hallucination” ability of LLMs. Once you realize that they only know what they were told and are unable to logically extrapolate the point becomes clearer. I hope that helps.
It's obvious chatGPT makes stuff up. Every one who has worked with LLMs in depth is fully aware of this. It's an obvious thing, you don't even have to "look it up" everyone knows about it.
This claim is made DESPITE the fact that LLMs hallucinate. It's obvious these models are imperfect and it's obvious they have huge deficiencies. But when it doesn't hallucinate, when the answer is Novel, creative, correct and unmistakably not existing in any training set, then we know the model understood the query you gave it.
I mostly agree with your position but have a quibble with this characterization. Knowledge can also be generated from randomization and enumeration. For instance, we could enumerate all Turing machines that might satisfy some property, or we could randomly permute some Turing machine as in genetic algorithms to find some new behaviours.
You might be tempted to categorize these under "composition", but I think they have different properties from composition, which is typically understood to be a finite deterministic map. Enumeration is potentially unbounded, and random mutation is non-determinstic.
Simply add a new input neuron with a random seed or add a tiny bit of noise to some weights or seed the tokens yourself in your query.
"<Query Seed: 4334> hello chatGPT, how are you?"
In this case chatGPT can deliberately randomize the response through understanding the intent of what you mean by "Query Seed".
In the brain, if such randomness existed, it would largely be modelled as a similar mechanism. A seed value (or multiple seeds at different places) is either inserted near the input step or happens at a branching logic step. Additionally, the query to the human brain can also be seeded.
True randomness requires that this seed number comes from quantum properties of particles that expresses itself in a sort of macro level random number. There is no other known source of true randomness in nature, though we can get perceptually identical results through just seeding with timestamps.
Either way it's trivial to add and not critical to what we mean by the word "understanding" because it's both easy to add and we aren't even sure if we have such randomness is in our brains. If it existed and if a human had this mechanism removed from his brain we would still say that this human is capable of "understanding" things.
Randomness is trivially addable. Additionally it's arguably not a source of source knowledge.
Its just a selection parameter. Which sources of knowledge do I use for composition and in what way do I compose it out of available compositions? The randomness parameter can influence these steps.
If you think there are other sources of knowledge, then what are they?
For example, if I have trillions of atoms randomly composed together with the objective to form a perfect cube, it will take eternity for a valid solution to arise.
If I have pre-existing cubes of atoms already pre-formed into cubic lego bricks and I have these components compose randomly, then it is far more likely that I get a cube.
LLMs, can do this with additional reinforcement training. chatGPT specifically is effective because it has this on top of the original GPT-3 model.
This is essentially random selection of known datums. It happens at the genetic level as well on a higher level.
The generation of raw novel datums in genes, however, is something that cannot happen in human brains. The mechanism for natural selection cannot exist within the brain itself, unless you're calling the trial by error process "natural selection." But again a human doesn't select a completely random strategy to "trial" he will select a strategy out of a known set. This is again, selecting a random datum.
Keep in mind, actual pure randomness generates useless noise 99.99999% of the time.
Just for reference, I've been on a walk with my cat just yesterday ;-). Bengal cats do like going for a walk. My cat follows me without a leash. And he even begs to go for a walk once in a while.
The best way to understand them isn’t to look at what they do well but where they fail. A great example is how ChatGPT can initially make a few chess moves that seem reasonable, but very quickly it stops making valid moves. It’s not operating from some model of the game but rather imitating sequences of moves it’s seen before. The best analogy isn’t cognition but someone trying to make a better version of “lorem ipsum” for whatever promo you’re giving it.
https://thegradient.pub/othello/
https://arxiv.org/abs/2210.13382
with enough training data it could probably model chess - not well enough to win but well enough to make legal moves
You think that I am misunderstanding whats going on and anthropomorphising ChatGPT. I know how ChatGPT works, my position is that we might be overestimating our selves, and underestimating the power of emergent phenomenon.
People won't believe you if you show actual evidence. You have to throw them a scientific paper written by an "expert" lol. And even then they will find a hard time changing their viewpoint.
Its so strange why people are trying to downplay it all when even the science is showing they're wrong.
They have to throw accusations around of anthropomorphisation. Seriously? It's very easy to identify the bias of anthropomorphisation. Anyone can easily tiptoe around that bias with a simple argument. Clearly what's going on with chatGPT is much more complex then that.
I recommend people stop using that word in this topic. It's akin to accusing someone they have brain damage. Clearly they don't.
I propose that if you gathered enough chess transcripts like this:
e4 e5 Nc3 Nc6 f4 exf4 d4 d5 Bxf4 Bb4 exd5 Qxd5 Kf2 Qh4+ ....
And fed them into a blank GPT then it would learn chess like a language and be able to make legal moves most of the time. It would do this by inferring and modelling the board and the rules. It wouldn't be a great player but it would be able to make moves.
This is bascially what the Othello paper I linked above is all about. They used GPT-2 I think. Chess is harder but I reckon could be done with a bigger model and more training data.
Anyway, my point was less about the game than how its failure show what’s going on better than it’s successes. The GPT approach is optimized for chat bots and its successes have more to do with exploiting how we approach communication than anything that can turn into AGI.
"It’s not operating from some model of the game"
I'm saying:
"it would model the rules of chess, if you fed it enough game transcripts"
Instead it’s modeling something else which is somewhat related to the game. Aka someone playing tick tack toe who moves on top of their opponent’s move isn’t playing tick tack toe.
Oh? don't anthropomorphize the thing we are supposed to "chat" with? It's basically what they're designed for except when they rudely tell you "I'm just an LLM!"
That isn't true.
https://ai.googleblog.com/2022/11/characterizing-emergent-ph...
While one can correctly argue that this is a consequense of increasing size, no one realized that increasing the size of LLMs would give them reasoning ability in 2018.
Indeed, before Minerva[1] came out expert forecasters predicted[2] a 12% improvement in the ability on the MATH benchmark[3]. Minerva improved it by 50% and PaLM improved it even more.
[1] https://arxiv.org/abs/2206.14858
[2] https://bounded-regret.ghost.io/ai-forecasting-one-year-in/
It's not my work.
And no, they probed this theory. The OCWCourses benchmark was new questions they generated from MIT Open Courseware that didn't exist prior to this benchmark.
> You can only reason when you have awareness within which to reason, which these systems do not.
I don't understand this objection at all.
Reasoning seems roughly the same as logical inference, and logical inference is a computational process.
Well that isn't true for any sensible definition of "idea of what words mean". If you examine embeddings produced by LLMs you find similar words in similar locations in the multi-dimensional space.
> They just statistically mash up patterns of words from massive input with no clue whatsoever what those patterns of words actually mean in the real world.
This is simply not true. As multiple papers (eg [1]) have shown these models can do tasks that require chained reasoning and give results that can only be explained by this ability.
There's even a hint of a possible explanation in that paper:
> For certain tasks, there may be natural intuitions for why emergence requires a model larger than a particular threshold scale. For instance, if a multi-step reasoning task requires l steps of sequential computation, this might require a model with a depth of at least O(l) layers.
Your understanding might have been correct for LSTMs trained before 2018 or so, and it was a reasonable model in the BERT era of early transformers. But it needs to be updated now with more recent results.
[1] https://ai.googleblog.com/2022/11/characterizing-emergent-ph...
It is akin to someone looking over the shoulder at your math homework and deriving understanding just from that.
This isn't a mistake. This is what the current science says.
What makes me curious is that chatGPT even has output that displays complete understanding and is so novel it is impossible for the output to be anything else other then actual understanding. Yet people are still running around downplaying it all.
The proof is right there. You can interact with chatGPT. The science is also there. Yet people just want to downplay it.
Perhaps it's fear of the future and what it means for humanity?
What exactly is "reasoning"? Seem, like the ability to consider options and choose one that fits the situation best. One might even call it "prediction". I generally argue that these big brains of ours (and science too) are intended to predict the future of various hypotheticals in order to choose the path forward with the best outcome.
Another frightening thing about chatGPT is that it has no "motivation". A lot of people think that AI will need some kind of motivation in order to be useful. How much of what we do as humans is actually done on autopilot? And then there's the NPC meme. I'm not sure these concepts (reason, thought, motivation) are so well defined that we will be certain when AI attains them or not.
The AI community is only beginning to realize the emergent affect of LLMs. Yes the underlying model is not new, and yes the underlying model just looks like a word predictor. But it is becoming obvious to experts that there are high level macro structures within the network at play here.
I don't know why people are constantly downplaying the AI. Yeah it does create stupid output. But if you just focus on that then you'd be completely ignoring the actual science around it.
What do you mean with "true understanding" here? Asking a computer to solve say, multiplication of large numbers, and it will provide the correct expected answer, with speed and and accuracy that outperform any human out there.
Should we say that such a calculator application "understand" arithmetic?
>I don't know why people are constantly downplaying the AI.
Maybe that’s in a balance with "everybody seeing the raise to conscious understanding of AI" each time a new credibly-human-like-output is generated?
Just like pareidolia: "yes there is obviously something that evokes a visage in that rock" doesn’t have to be interpreted as a sign that Nature loves to purposefully sculpt human faces.
From an evolutionary perspective, we see faces everywhere just like we see sign of understanding in surrounding agents because it makes our species more efficient in social cooperation, which to my mind is by a very large margin our best asset.
https://science.howstuffworks.com/life/inside-the-mind/human...
https://www.amazon.com/SuperCooperators-Altruism-Evolution-O...
Optical illusions are a form of bias within our visual system that natural selection has optimized for a very biased use case. We see this bias within ourselves and we can extend the knowledge further.
If this bias exists in our visual system it must also exist in other systems. Our logic, judgement, emotions and morals all have in built biological biases calibrated for survival in a prehistoric world that is very different for the modern world we live in.
Does awareness of these biases eliminate our biases? Can knowledge of our weaknesses allow us to side step the bias and judge something with utter clarity?
You assume the answer is yes, and you assume I'm biased so you illustrate to me the concept of an optical illusion and hope that armed with this knowledge I can see my own bias and as a result escape it.
The problem is, as you can guess, is that I'm extremely knowledgeable about my biases. I came to my conclusions about chatgpt fully aware and fully mindful of these concepts.
So then you can conclude that my bias is highly sophisticated as it is also built on an extremely recursive and sophisticated awareness of bias itself.
If such a sophisticated form of bias can exist, then who is to say that you're not the one who is biased? You literally try to frame your own viewpoint as a contrast to the concept of biases within optical illusions. You could be contrasting your own bias with the concept of bias itself.
So I am telling you this. You missed something in the latest AI hype cycle. Your bias is so sophisticated it is aware of bias itself and it is blinding you to the fact that this iteration of AI hype was different from the last.
When you query chatGPT it can answer with sentences that can only be novelly formulated from true understanding of a topic. Literally, and I tell you this fact as well fully aware of my own biases and the human capacity to "see faces everywhere."
So if you claim I'm biased and I claim you're biased? Who is actually biased?
Perhaps not the person with actual scientific papers and studies backing his arguments. I have this.
Perhaps if I present to you these scientific papers you will become more self aware about the sophistication of your own biases.
Maybe you will become more aware about how you revell in your superior knowledge about how humans have "evolved" to see faces everywhere in the same way that they have a tendency to anthropomorphize what you think is essentially a statistical word predictor.
You may become more aware of how you apply that knowledge to construct a scaffold of delusion that blinds you to the actuality of what chatGPT is capable of.
But you may not want to become aware of it. You may say you don't need to see these scientific papers , or when given those papers you will pour yourself over every detail attempting to find a flaw to prove your argument right.
A truly unbiased person will flip his opinion in an instant when given contrary logic and facts. He will abandon years of belief instantly if new logic presents that his beliefs were wrong. Are you that person? Doubt it. Humans don't work this way.
Such is the nature of your bias. Or it could be the horrible display of the insane sophistication of my own biases. At this level, both of us have no choice but to follow our biases to the bloody end.
Either way my claim is as simple as this. The way chatGPT understands some parts of the world is completely isomorphic to how we colloquially use the word "understanding".
At the very least perhaps this reply can help you understand that people who have contrary opinions to you aren't just mindless sheep who fall for classic and cheap human biases. Get off that high horse.
>A truly unbiased person will flip his opinion in an instant when given contrary logic and facts. He will abandon years of belief instantly if new logic presents that his beliefs were wrong. Are you that person? Doubt it. Humans don't work this way.
We fully agree on that, it seems. :)
>The way chatGPT understands some parts of the world is completely isomorphic to how we colloquially use the word "understanding".
Here too, I am open to agree if we precise which colloquial part of using this word we are referring to. It’s not like we always attach the same exact meaning to each word we use in whatever situation. Some words like deictics or even nouns like "thing" are only given meaningfulness in the situation at hand.
>what chatGPT is capable of.
I thing that is the point where our perspective diverge. I’m not disputing what a software collection is capable of. Only what it implies in term of consciousness experiment. Telling that some software can exceed any human at delivering some outputs doesn’t seem to be a topic that anyone dispute, does it? It doesn’t mean the software collection goes through the same mean of sentience to achieve this performance.
You can send me the links that worth a read according to you, thanks. :)
Now I just hope I’m not replying to a chatbot. :D
- GPT style language models end up internally implementing a mini "neural network training algorithm" (gradient descent fine-tuning for given examples): https://arxiv.org/abs/2212.10559