That's true in some sense, but definitely not if you apply it exclusively to language. The "response" you select is a step in a plan to achieve a goal. Part of that plan may be to utter a sentence, or it may be to stop, duck and roll.
If the plan is to utter a sentence, the way you generate the sentence is NOT to emit the most likely token to fit all of the speech input you ever heard before.
Is it? Or are you just a bird trying to explain how flying works by dropping poop on my head?
> If the plan is to utter a sentence, the way you generate the sentence is NOT to emit the most likely token to fit all of the speech input you ever heard before.
This is completely incoherent to me.
I will do this even if I have never before heard anyone utter this sentence, or anything similar to it - except for the words themselves, which I need to have learned from hearing them spoken once or twice in my early life.
This same kind of planning is visible in any animal you care to study long enough - even in insects, possibly even in jelly-fish. It doesn't exist to even a shallow degree in a conversation with GPT-3.
> > If the plan is to utter a sentence, the way you generate the sentence is NOT to emit the most likely token to fit all of the speech input you ever heard before.
> This is completely incoherent to me.
The way LMs generate sentences is to keep picking the most likely next token (letter) that matches the prompt + the text they generated so far, up to some depth. That is, the basic step that happens when you give a prompt to LaMDA, say, "Do you want a glass of water?" is that it will predict the most likely token is "Y". Then, you prompt it again with "Do you want a glass of water? Y", and it will predict "e". "Do you want a glass of water? Ye" -> "s". "Do you want a glass of water? Yes" -> ".". So, the program around LaMDA will show you the output "Yes.". (there are more steps after this, and they way it decides that it generated enough tokens is not trivial, and tokens are not exactly single letters etc; but this is the basic way it actually works). The likelihood function is based on all of the text it has gone through in the one-time training step.
This is definitely NOT the way humans output language.