Same with statistics and markov chains, people for years tried to generate chat bots with those, but they never worked well.
Same with statistics and markov chains, people for years tried to generate chat bots with those, but they never worked well.
I think Juergen Schmidthuber has developed a lot of ideas around compression being the basis for consciousness and understanding.
There was the paper that showed that when showing a language model Othello moves it ends up building an internal representation of the board.
And now I was reading this abstract:
```Theory of mind (ToM), or the ability to impute unobservable mental states to others, is central to human social interactions, communication, empathy, self-consciousness, and morality. We administer classic false-belief tasks, widely used to test ToM in humans, to several language models, without any examples or pre-training. Our results show that models published before 2022 show virtually no ability to solve ToM tasks. Yet, the January 2022 version of GPT-3 (davinci-002) solved 70% of ToM tasks, a performance comparable with that of seven-year-old children. Moreover, its November 2022 version (davinci-003), solved 93% of ToM tasks, a performance comparable with that of nine-year-old children. These findings suggest that ToM-like ability (thus far considered to be uniquely human) may have spontaneously emerged as a byproduct of language models' improving language skills.```
The rules of chess are much simpler than any structure of the real world we might hope for it to understand—and yet it failed to learn the rules of chess.
Based on this, I am pretty sure it has not learned any meaningful structure.
That misses the point. Intelligent beings (humans) can learn the rules of any board game given enough time. We don't need special training. What your parent comment says is it can't even learn the rules, let alone be good at it.
The rules of English are ridiculously more complicated than chess, and it has that figured out just fine.
Or, you could store the function itself.
To see the difference, compare for yourself what would happen for very large or very small inputs in both types of “understanding”.
f(x) = x * (sin(sin(x)))^2
Ask it to give you the values of f for integers of x between -10 and 10. I tried it 5 times and it was never close.
I chose this for several reasons. One, it’s very unlikely that it’s memorized the answers somewhere on the internet. Two, it’s pretty chaotic if you look at a graph of it, so interpolation won’t work. (It is bounded by x and 0 for all values but for large absolute values of x it varies wildly.) And three, memorizing values won’t get you anywhere since it becomes much more chaotic as x increases.
I also disagree that approximations are simpler. As you can see, the actual function is only sine and multiplication. To approximate this function would be far harder.
And yes, in your example, of course it's simpler to just store the function itself - provided that you know in advance what it is, which, to remind, GPT does not. But when we're dealing with f(question)=answer of chatbots, the function that GPT ends up approximating is decidedly not simple.
A realistic example might be the Mod/RM byte in x86 instruction encoding. There are underlying regularities, but a lookup table could quite possibly be smaller than the code required to generate the correct Mod/RM byte given operands. So you can understand the Mod/RM encoding without thereby being able to compress anything.
The building blocks are the same - neural nets. The language outputs are the same enough to fool a human.
So what’s that elusive secret sauce that makes you ‘aware’ and other things not?
And unlike your other examples how are you going to convince the machine it’s not aware when the only physical difference is a wet neural net versus a dry one?
The fact that I am aware given there’s no physical evidence to indicate why I should be implies that everything is aware at some level.
Why do you think a LLM doesn’t know what a tree or car is?