My mental model of LLMs being next-word predictors with a long context window suggests they don't. Are there any papers on this?
My mental model of LLMs being next-word predictors with a long context window suggests they don't. Are there any papers on this?
On the other hand, no matter how large the model they struggle to generalize between sentences of the form "George Washington was the first US president" and "The first US president was George Washington".
Clearly generalization behavior is unintuitive, you can come up with a post hoc explanation (weights are independent between early and late layers) but I doubt anyone would have thought ahead of time that models would have an easier time generalizing between languages than between slightly different ways of saying the same thing in English.
According to that paper, people made decisions more emotionally in their native languages compared to their second language - where they tend to be more logical.
In case of LLMs there were some nice cases where GPT gave different replies depending on the language of the query. Factual information was roughly the same, but in some cases the model gave totally different replies depending on the language of the query.
As a bilingual speaker - my experience is that the replies are very similar regardless of whether I speak in Polish or English, or mix them both within the same sentence.
Yes, LLMs really encode knowledge. This is not up for debate, and it's obvious to anyone working in AI. It outputs words and input words, yes, but the magic is what is in between. The words are broken down immediately at the first layer. Next it involves 100s of billions of parameters doing god knows what. But it's safe to assume they encode the base reality as best it can. It is what i call 'language-induced reality model'.
The task of next word predictor in a humongous dataset turns out, is best solved by creating an inner representation of the real world. It's still not a perfect model since it's just induced by language though, so obviously it has limitations. But it's really modeling it, and with higher fidelity than people imagined you can learn the world through words alone.
https://gwern.net/scaling-hypothesis#why-does-pretraining-wo...