In short, it's the good old religious debate about souls, just wrapped in techno-philosophical trappings.
What that something more is is even less defined with even fewer theories as to what it is than there are around the woo and mysticism of human intelligence. And as LarsDu88 points out in a separate thread, there are alternative explanations for what we're seeing here besides "We've created some sort of weird internal 3D engine that the diffusion models use for generating stuff," which also meshes closely with the fact that generations routinely have multiple perspectives and other errors that wouldn't exist if they modeled the world some people are suggesting.
If there's something more going on here, we're going to need some real explanations instead of things that can be explained multiple other ways before I'm going to take it seriously, at least.
But the other side doesn't see it that way, specifically not the "something more" part. It's "just math" all the way down, in our brains as well. The "emergent phenomena" are not undefined in this sense - they're (obviously in LLMs) also math, it's just that we don't understand it yet due to the sheer complexity of the resulting system. But that's not at all unusual - humans build useful things that we don't fully understand all the time (just look at physics of various processes).
> which also meshes closely with the fact that generations routinely have multiple perspectives and other errors that wouldn't exist if they modeled the world some people are suggesting.
This implies that the model of the world those things have either has to be perfect, or else it doesn't exist, which is a premise with no clear logic behind it. The obvious explanation is that, between the limited amount of information that can be extracted from 2D photos that the NN is trained on, and the limit on the complexity of world modeling that NN of a particular size can fit, its model of the world is just not particularly accurate.
> we're going to need some real explanations instead of things that can be explained multiple other ways before I'm going to take it seriously, at least.
If we used this threshold for physics, we'd have to throw out a lot of it, too, since you can always come up with a more complicated alternative explanation; e.g. aether can be viable if you ascribe just enough special properties to it. Pragmatically, at some point, you have to pick the most likely (usually this means the simplest) explanation to move forward.
You might not, but you don't have to look far in these very comments to be met with woo and mysticism on the emergent phenomena side, either.
> The obvious explanation is that, between the limited amount of information that can be extracted from 2D photos that the NN is trained on, and the limit on the complexity of world modeling that NN of a particular size can fit, its model of the world is just not particularly accurate.
I think the more obvious solution is that it's not modelling the world, because... why would it be? It seems obvious to me that this is a significantly more complex way to handle their given task than every other explanation that does not require them create a model of the world.
> Pragmatically, at some point, you have to pick the most likely (usually this means the simplest) explanation to move forward.
I agree completely with this, which is why the idea that this is modelling the world is absolutely bizarre to me, particularly when our understanding of their training is that it's all weighting for pixel nearest neighbors via ranking how well it does at denoising 2D images.
It's not like de-rendering is something new, either. ShaderMap has existed since the mid 2000s and could get similar results from 2D images without any AI/ML, and we have other models that can generate it without anyone suggesting they model the world, e.g. https://arxiv.org/abs/2201.02279
People like LeCun don't think that even models where the stated goal is world simulation are able to do it, e.g. sora, either - https://twitter.com/ylecun/status/1758740106955952191 - so it seems even more unlikely when that isn't the goal that this has happened due to some emergent phenomenon.
Will that baby automatically become a complex tool user with a better comprehension of the world around it than the monkeys in its group?
How responsible is the 'training data' in it's environment responsible for human intelligence?
It's also not just the qualia of sensation, but also that of the will. We all 'feel' we have a will, that can do things. How can a computer possibly feel that? The 'will' in the LLM is forced by the selection function, which is a deterministic, human-coded algorithm, not an intrinsic property.
In my view, this sensation of qualia is so out-there and so inexplicable physically, that I would not be able to invalidate some 'out there' theories. If someone told me they posited a new quantum field with scalar values of 'will' that the brain sensed or modified via some quantum phenomena, I'd believe them, especially if there was an experiment. But even more out there explanations are possible. We have no idea, so all are impossible to validate / invalidate as far as I'm concerned.
1st evolutionary algorithm and then the constant input we receive from the World being the training data and we having reward mechanisms rewiring our neural networks based on what our senses interpret as good?
But also, the very notion of qualia suffers from the same problem as other vague concepts like "consciousness" - we cannot actually clearly define what they are. All definitions seem to ultimately boil to "what I feel", which is vacuous. It is entirely possible that qualia aren't physically real in any sense, and are nothing more than a state of the system that the system itself sets and queries according to some internal logic based on input and output. If so, then an LLM having an internal self-model that includes a state "I feel heat" is qualia as well.
Is it "possible"? Absolutely. However, I have no means by which to measure, as you say, where at least with humans and animals I posit that their shared behaviors do indicate the same feeling, so I have some proof.
With an LLM, the output is a probability distribution.
Moreover, if it is as you say it is, then computers have qualia as well, which is scary because we would be committing a pretty ethically dubious 'crime' (at least in some circumstances).
Again, I just don't see it. Anything is possible, as I said, but not everything is as likely, by my estimation. And yes, that is entirely how I feel, which is as real as anything else.
The images we get from these neural networks are trained on looking pleasing, for some definition of pleasing. That's why they look good on the whole, but get into that uncanny valley the moment you go inspecting. Similar to dreams.
Whereas obviously real human perception is (usually) grounded in reality.
“Understanding” would mean that they be able to train themselves, which they are as yet unable to do.
We are tuning weights and biases to statistically regurgitate training data. That is all they are.
Training themselves is necessary. All learning is self learning. Teachers can present material in different ways but learning is personal. No one can force you to learn either.
"To understand" is one of those poorly defined concepts, like "consciousness", it is thrown a lot in the face when talking about AI. But what does it mean actually? It means to have a working model of the thing you are understanding, a causal model that adapts to any new configuration of the inputs reliably. Or in other words it means to generalize well around that topic.
The opposite would be to "learn to the test" or "overfit the problem" and only be able to solve very limited cases that follow the training pattern closely. That would make for brittle learning, at surface level, based on shortcuts.
The weasel word here is "reliably". What does this actually mean? It obviously cannot be reliable in a sense of always giving the correct result, because this would make understanding something a strict binary, and we definitely don't treat it like that for humans - we say things like "they understand it better than me" all the time, which when you boil it down has to mean "their model of it is more predictive than mine".
But then if that is a quantifiable measure, then we're really talking about "reliable enough". And then the questions are: 1) where do you draw that line, exactly, and 2) even more importantly, why do you draw the line there and not somewhere else.
For me, the only sensible answer to this is to refuse to draw the line at all, and just embrace the fact that understanding is a spectrum. But then it doesn't even make sense to ask questions like "does the model really understands?" - they are meaningless.
(The same goes for concepts like "consciousness" or "intelligence", by the way.)
The reason why I think this isn't universally accepted is because it makes us not special, and humans really, really like to think of themselves as special (just look at our religions).
Our capacity to make mistakes does not necessarily equate to a lack of understanding.
If you’re doing a difficult math problem and get it wrong, that doesn’t necessarily imply that you don’t understand the problem.
It speaks to a limitation of our problem solving machinery and the implements we use to carry out tasks.
e.g. if I’m not paying close enough attention and write down the wrong digit in the middle of solving a problem, that could also just be because I got distracted, or made a mistake. If I did the same problem again from scratch, I would probably get it right if I understand the subject matter.
Limitations of our working memory, how distracted we are that day, mis-keying something on a calculator or writing down the wrong digit, etc. can all lead to a wrong answer.
This is distinct from encountering a problem where one’s understanding was incomplete leading to consistently wrong answers.
There are clearly people who are better and worse comparatively at solving certain problems. But given the complexity of our brains/biology, there are myriad reasons for these differences.
Clearly there are people who have the capacity to understand certain problems more deeply (e.g. Einstein), but all of this was primarily to say that output doesn’t need to be 100% “reliable” to imply a complete understanding.
> The weasel word here is "reliably". What does this actually mean? It obviously cannot be reliable in a sense of always giving the correct result, because this would make understanding something a strict binary
What did you mean by "it cannot be reliable in a sense of always giving the correct result", and why would that make understanding something a strict binary?
I do agree that some people understand some topics more deeply than others. I believe this to be true if for no other reason than watching my own understanding of certain topics grow over time. But what threw me off is that someone with "lesser" understanding isn't necessarily less "reliable". The degree of understanding may constrain the possibility space of the person, but I think that's something other than "reliability".
For example, someone who writes software using high level scripting languages can have a good enough understanding of the local context to reason about and produce code reliably. But that person may not understand in the same way that someone who built the language understands. And this is fine, because we're all working with abstractions on top of abstractions on top of abstractions. This does restrict the possibility space, e.g. the systems programmer/language designer can elaborate on lower levels of the abstraction, and some people can understand down to the bare metal/circuit level, and some people can understand down to the movement of atoms and signaling, but this doesn't make the JavaScript programmer less "reliable". It just means that their understanding will only take them so far, which primarily matters if they want to do something outside of the JavaScript domain.
To me, "reliability" is about consistency and accuracy within a problem space. And the ability to formulate novel conclusions about phenomena that emerge from that problem space that are consistent with the model the person has formed.
If we took all of this to the absurdist conclusion, we'd have to accept that none of us really understand anything at all. The smaller we go, and the more granular our world models, we still know nothing of primordial existence or what anything is.
Other way around. If a kid understood the concepts of multiplation, they could train themselves on the next logical steps like exponents.
The consequences of this in the terms of AI would mean they build on a series of concepts and would quickly dwarf us in intelligence.
We don't know, because these concepts were discovered before writing. We do know that far larger jumps have been made by individual mathematicians who never made it past 33 years of age.
I don't know many people that would suggest that knowing about ones family structure, such as their mom/dad/uncle, and how their history relates to them, is required to be _completely reconstructed_ every time they interact with their environment somehow from first principles.
Online reinforcement learning can have large merits without resorting to stating it's required for learning. Just as self awareness can occur independently of consciousness can occur independently of intelligence is independent of empathy and so on. They are all different, and having one component doesn't mean anything about the rest.