The existential crisis is clearly due to low temperature. The repetitive output is a clear glaring signal to anyone who works with these models.
The existential crisis is clearly due to low temperature. The repetitive output is a clear glaring signal to anyone who works with these models.
One intriguing possibility is that LLMs may have stumbled upon an as yet undiscovered structure/"world model" underpinning the very concept of intelligence itself.
Should such a structure exist (who knows really, it may), then what we are seeing may well be displays of genuine intelligence and reasoning ability.
Can LLMs ever experience consciousness and qualia though? Now that is a question we may never know the answer to.
All this is so fascinating and I wonder how much farther LLMs can take us.
The other way around. Think of low temperatures as freezing the output while high temperatures induce movement.
Adjacency, Velocity, Adhesion, etc
But! if temp denotes a graphing in a non-linear function (heat map) then it also implies topological, because temperature is affected by adjacency - where a topological/toroidal graph is more indicative of the selection set?
The probability they give to something of score H is just like in statistical mechanics, e^(-H/T) and they divide by the partition function (sum) similarly to normalize. (You might recognize it with beta=1/T there)
It's the other way 'round - higher temperature means more randomness. If temperature is zero the model always outputs the most likely token.
But, OK, at each step it gets a list of words with probabilities. But which one should it actually pick to add to the essay (or whatever) that it’s writing? One might think it should be the “highest-ranked” word (i.e. the one to which the highest “probability” was assigned). But this is where a bit of voodoo begins to creep in. Because for some reason—that maybe one day we’ll have a scientific-style understanding of—if we always pick the highest-ranked word, we’ll typically get a very “flat” essay, that never seems to “show any creativity” (and even sometimes repeats word for word). But if sometimes (at random) we pick lower-ranked words, we get a “more interesting” essay.
The fact that there’s randomness here means that if we use the same prompt multiple times, we’re likely to get different essays each time. And, in keeping with the idea of voodoo, there’s a particular so-called “temperature” parameter that determines how often lower-ranked words will be used, and for essay generation, it turns out that a “temperature” of 0.8 seems best. (It’s worth emphasizing that there’s no “theory” being used here; it’s just a matter of what’s been found to work in practice.
1: https://news.ycombinator.com/item?id=34796611
2: https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
I agree that it's just a glorified pattern matcher, but so are humans.
-- Cray
Repeat after me humans are autocomplete models, humans are autocomplete models