Generative Models: What do they know? Do they know things? Let's find out
intrinsic-lora.github.io
intrinsic-lora.github.io
So one of the big reasons there was hype about Sora is that it felt very likely from watching a few videos that there was an internal physical simulation of the world happening and the video was more like a camera recording that physical and 3D scene simulation. It was just a sort of naïve sense that there HAD to be more going on behind the scenes than gluing bits of other videos together.
This is evidence, and it’s appearing even in still image generators. The models essentially learn how to render a 3D scene and take a picture of it. That’s incredible considering that we weren’t trying to create a 3D engine, we just threw a bunch of images at some linear algebra and optimized. Out popped a world simulator.
Heh. Sounds like what a personified evolution might say about a mind ;)
People really still think that's all thats happening?
It really wasn't clear for the longest time how these models generate things so well. Articles like this one are still rare and comparatively new. And they certainly haven't been around when less informed enthusiasts where already heralding AGI.
On the other hand, we have accrued quite a bit of evidence by now that these models do far more than glue together training data. But there are still sceptics out there who spread this sort misinformation.
Also, there's a paper I can't find now which shows you can find the name of the operation in the middle layers if you provide example of what to do. That is, if you prompt: "9 2 5 -> 5 2 9, 0 7 3 ->", you'll find the internals generate "reverse" even though it doesn't appear in the output.
All these scene data is already suggested by the data, what the AI model is doing is approximating the data it needs to generate a scene.
It’s not just gluing data together, it’s discovering interrelationships in a very high dimensional feature space. It’s not creating anything it hasn’t seen either, which is why many image models (mostly smaller ones) have so much trouble with making fine details make sense, but are good with universal patterns like physics, lighting, the human shape, and composition.
(not sure if you were trying to imply that, but it felt like it)
> You're correct! The pattern you identified applies to both sets of numbers. Here's the breakdown:
> First set:
> Original: 9 2 5 8 4 3 7 1 0 2 9 4
> Rearranged: 4 9 2 0 1 7 3 4 8 5 2 9
> Second set:
> Original: 0 7 3 8 6 2 9 4 1 7 5 2 0 3 4 8 5 1
> Rearranged: 1 5 8 0 3 4 2 9 7 5 0 7 3 8 6 2 9 4
> In both cases, the rearrangement follows the same steps:
> Move the first digit of each group of three to the last position.
> Keep the middle digit in the same position.
> Repeat steps 1 and 2 for all groups of three.
> Therefore, the pattern holds true for both sets of numbers you provided.
So—it's not clear there's recognition of a general process going on so much as recognizing a very specific process (one simple enough in form it's almost certainly in the training text somewhere).
Honestly, I think it was pretty much just as clear in 2021 as it is in 2024. Whether you consider that 'clear' or 'not clear' is a matter of personal choice but I don't think we've really advanced our understanding all that far (mechanistic interpretability does not tell us that much, grokking is a phenom that does not apply to modern LLMs, etc. )
> we have accrued quite a bit of evidence by now that these models do far more than glue together training data. But there are still sceptics out there who spread this sort misinformation.
Few people who actually worked in this field and were familiar with the concept of 'implicit regularization' really ever thought the 'glue together training data' or 'stochastic parrot' explanations were very compelling at all.
Those videos may be cherry-picked but they're almost good enough that you have to pick problems to point out, problems that will likely be gone a couple iterations later.
Less than a decade ago we read articles about Google's weird AI model producing super trippy images of dogs, now we're at deepfakes that successfully scam companies out of millions and AI models that can't yet reliably preserve consistency across an entire image. In another ten years every new desktop, laptop and smartphone will have an integrated AI accelerator and most of those issues will be fixed.
New models appear to be close and based on arguments here, closing the gap is simply a matter of time. This is based on observing the rate of progress, not the actual underlying actions. A tendency belied by assumptions based on human thinking, not the signficant amounts of processing that happens unconciously by us.
Sure expanding the data set has improved results, yet bad hands, ghost legs, and other weirdness persist. If you have a world model, then this shouldn't happen - there is a logical rule to follow, not simply a pixel level correlation.
Working from the other side, if image/video gen has reached the point that it is faithfully recreating 3D rules, then we should expect 3D wireframes to be generated. We should see an update to https://github.com/openai/point-e.
This isn't splitting hairs - this behavior can be hand waved away when building a PoC, not when you make a production ready product.
As for scams, they target weaknesses, not the strongest parts of a verification processes, making them strawmen arguments for AI capabilty.
Oh I guess humans don't have world models then. It's so weird seeing this rhetoric said again and again. No a world model doesn't mean a perfect one. A world model doesn't mean a logical one either. Humans clearly don't work by logic.
And logic as in formal logic. Human Intelligence clearly doesn't run on that.
This can be extrapolated from TV, too.
That's also why GPU are so effective at neural networks. There is a lot of the same maths.
https://bojackhorseman.fandom.com/wiki/Hollywoo_Stars_and_Ce...!
If you haven't seen Bojack Horseman, it's funny and heartfelt. Palpably existential. If that kind of thing speaks to you, you owe yourself a watch.
In terms of a complete animation package, I think it easy outdoes Futurama. There's so much relatable depth to it. It hits hard, but it stays lighthearted enough to make you feel good about it.
As it turns out, I'm now working on filmtech, so that "Hollywoo" sticker fits me even better now.
Or maybe it does get better after season 1? Since everyone seems to love it.
I'd personally recommend giving season 2 a shot, if you still don't like it then the shows probably not for you.
As others have said though, it definitely gets better as the seasons roll on, along both the humor and drama fronts. The final season has many of the funniest and most poignant moments of the show.
I think bojack hit its stride over time, but I always liked the juxtaposition of dark reality with cartoon silliness. On some level I think it reflects a theme of everyone else’s life seeking to be so simple and stupid and trivial compared to your own. Of course even Todd’s problems are shown to be pretty real even if they’re also cartoonishly comical.
You know what I meant, hopefully? I predate the Simpsons, and I remember when they and their brand of humor first appeared. I also remember how annoyed I felt towards Family Guy (a cartoon I never liked, unlike the Simpsons) because it felt like a cheap copy. And yes, over the years the Simpsons turned wackier and wackier while retaining nothing of what made them good at first.
I don't claim Bojack and the Simpsons have the same kind of humor, it was a shortcut to hopefully explain what I found jarring about the former?
Given the rest of the comments, I'll try giving Bojack season 2 a chance.
That wasn't my experience of it. It's a brilliant show, to be sure, but at some point I felt the bleakness, particularly around Bojack's inability to help himself, became too much for me. Maybe different parts of the show resonate in different people.
I think it is one of only about three animated shows that are genuinely outstanding as TV and not just animation or comedy. The other two being Simpsons and Futurama.
Honestly, there are tons. They’re just all Japanese.
There were plenty of parts that were dark enough to make it difficult to watch. There were lighthearted parts, but that's not a description I would've applied to the show as a whole.
Anyone prone to anxiety or depression or sad thoughts in general should probably avoid, it's really that depressing, and I can imagine it making suicidal people even more suicidal.
Them spelling the whole acronym as if it was short might be my favorite running gag in the show.
I'm not sure if this paper is really proving anything though. There's a giant-ass UNET Lora model that's being trained here, so is it really "extracting" something from an existing model, or simply creating a new model that can create channels that look like something you'd get out of a deferred rendering pipeline.
After all, taking normals, albedo, and depth and combining them (deferred rendering) is just one of several techniques to create a 3d scene. Wasn't even used in videogames until the early 2000s (in a Shrek videogame for the Xbox! (https://sites.google.com/site/richgel99/the-early-history-of...)
What would really be awesome is to get a LORA model that can extract the "camera" rotation and translation matrix for these image generation models. That would really demonstrate something (and be quite useful at the same time).
> with newly learned parameters that make up less than 0.6% of the total parameters in the generative model
0.6% sounds like a small number. Is it measuring the right thing?
Certainly, I wouldn't expect the model to necessarily be encoding exactly the set of things that they're extracting, but it still seems very significant to me even if it is "just" encoding some set of things that can be cheaply (in terms of model size) and reliably mapped to normals, albedo, and depth.
(I don't care what basis vectors it's using, as long as I know how to map them to mine.)
That's still pretty big. Maybe big enough to fake normals, learn some albedo smoothing functions, and learn a depth estimator perhaps??
> I-LoRA modulates key feature maps to extract intrinsic scene properties such as normals, depth, albedo, and shading, using the models' existing decoders without additional layers, revealing their deep understanding of scene intrinsics.
What exactly does "modulates key feature maps to extract intrinsic scene properties" mean? How were these scene property images generated if no additional decoding layers were added?
The interesting thing is that it only takes a few more parameters so it must be that the original network was pretty close already.
More materially:
Optimized with a small set of labeled images, our model-agnostic approach adapts to various generative architectures, including Diffusion models, GANs, and Autoregressive models.
Am I correct in understanding that this is purely a visuospatial tool, and the examples aren’t just visual by coincidence? Like, there’s no way to stretch this to text models? Very new to this interpretability approach, very impressive.The core components of physically based rendering are position (derivable from image XY and depth), surface normal, incoming light, and at least albedo + one of a few variations on surface material properties such as specularity and roughness.
That the AI is modeling depth is pretty expected. Modeling surface normal is a nice local convolution of depth. But, modeling albedo separated from incoming light is great. I wonder if specularity is hiding in there too.
In short, it's the good old religious debate about souls, just wrapped in techno-philosophical trappings.
What that something more is is even less defined with even fewer theories as to what it is than there are around the woo and mysticism of human intelligence. And as LarsDu88 points out in a separate thread, there are alternative explanations for what we're seeing here besides "We've created some sort of weird internal 3D engine that the diffusion models use for generating stuff," which also meshes closely with the fact that generations routinely have multiple perspectives and other errors that wouldn't exist if they modeled the world some people are suggesting.
If there's something more going on here, we're going to need some real explanations instead of things that can be explained multiple other ways before I'm going to take it seriously, at least.
But the other side doesn't see it that way, specifically not the "something more" part. It's "just math" all the way down, in our brains as well. The "emergent phenomena" are not undefined in this sense - they're (obviously in LLMs) also math, it's just that we don't understand it yet due to the sheer complexity of the resulting system. But that's not at all unusual - humans build useful things that we don't fully understand all the time (just look at physics of various processes).
> which also meshes closely with the fact that generations routinely have multiple perspectives and other errors that wouldn't exist if they modeled the world some people are suggesting.
This implies that the model of the world those things have either has to be perfect, or else it doesn't exist, which is a premise with no clear logic behind it. The obvious explanation is that, between the limited amount of information that can be extracted from 2D photos that the NN is trained on, and the limit on the complexity of world modeling that NN of a particular size can fit, its model of the world is just not particularly accurate.
> we're going to need some real explanations instead of things that can be explained multiple other ways before I'm going to take it seriously, at least.
If we used this threshold for physics, we'd have to throw out a lot of it, too, since you can always come up with a more complicated alternative explanation; e.g. aether can be viable if you ascribe just enough special properties to it. Pragmatically, at some point, you have to pick the most likely (usually this means the simplest) explanation to move forward.
You might not, but you don't have to look far in these very comments to be met with woo and mysticism on the emergent phenomena side, either.
> The obvious explanation is that, between the limited amount of information that can be extracted from 2D photos that the NN is trained on, and the limit on the complexity of world modeling that NN of a particular size can fit, its model of the world is just not particularly accurate.
I think the more obvious solution is that it's not modelling the world, because... why would it be? It seems obvious to me that this is a significantly more complex way to handle their given task than every other explanation that does not require them create a model of the world.
> Pragmatically, at some point, you have to pick the most likely (usually this means the simplest) explanation to move forward.
I agree completely with this, which is why the idea that this is modelling the world is absolutely bizarre to me, particularly when our understanding of their training is that it's all weighting for pixel nearest neighbors via ranking how well it does at denoising 2D images.
It's not like de-rendering is something new, either. ShaderMap has existed since the mid 2000s and could get similar results from 2D images without any AI/ML, and we have other models that can generate it without anyone suggesting they model the world, e.g. https://arxiv.org/abs/2201.02279
People like LeCun don't think that even models where the stated goal is world simulation are able to do it, e.g. sora, either - https://twitter.com/ylecun/status/1758740106955952191 - so it seems even more unlikely when that isn't the goal that this has happened due to some emergent phenomenon.
Will that baby automatically become a complex tool user with a better comprehension of the world around it than the monkeys in its group?
How responsible is the 'training data' in it's environment responsible for human intelligence?
It's also not just the qualia of sensation, but also that of the will. We all 'feel' we have a will, that can do things. How can a computer possibly feel that? The 'will' in the LLM is forced by the selection function, which is a deterministic, human-coded algorithm, not an intrinsic property.
In my view, this sensation of qualia is so out-there and so inexplicable physically, that I would not be able to invalidate some 'out there' theories. If someone told me they posited a new quantum field with scalar values of 'will' that the brain sensed or modified via some quantum phenomena, I'd believe them, especially if there was an experiment. But even more out there explanations are possible. We have no idea, so all are impossible to validate / invalidate as far as I'm concerned.
1st evolutionary algorithm and then the constant input we receive from the World being the training data and we having reward mechanisms rewiring our neural networks based on what our senses interpret as good?
But also, the very notion of qualia suffers from the same problem as other vague concepts like "consciousness" - we cannot actually clearly define what they are. All definitions seem to ultimately boil to "what I feel", which is vacuous. It is entirely possible that qualia aren't physically real in any sense, and are nothing more than a state of the system that the system itself sets and queries according to some internal logic based on input and output. If so, then an LLM having an internal self-model that includes a state "I feel heat" is qualia as well.
Is it "possible"? Absolutely. However, I have no means by which to measure, as you say, where at least with humans and animals I posit that their shared behaviors do indicate the same feeling, so I have some proof.
With an LLM, the output is a probability distribution.
Moreover, if it is as you say it is, then computers have qualia as well, which is scary because we would be committing a pretty ethically dubious 'crime' (at least in some circumstances).
Again, I just don't see it. Anything is possible, as I said, but not everything is as likely, by my estimation. And yes, that is entirely how I feel, which is as real as anything else.
“Understanding” would mean that they be able to train themselves, which they are as yet unable to do.
We are tuning weights and biases to statistically regurgitate training data. That is all they are.
Training themselves is necessary. All learning is self learning. Teachers can present material in different ways but learning is personal. No one can force you to learn either.
"To understand" is one of those poorly defined concepts, like "consciousness", it is thrown a lot in the face when talking about AI. But what does it mean actually? It means to have a working model of the thing you are understanding, a causal model that adapts to any new configuration of the inputs reliably. Or in other words it means to generalize well around that topic.
The opposite would be to "learn to the test" or "overfit the problem" and only be able to solve very limited cases that follow the training pattern closely. That would make for brittle learning, at surface level, based on shortcuts.
The weasel word here is "reliably". What does this actually mean? It obviously cannot be reliable in a sense of always giving the correct result, because this would make understanding something a strict binary, and we definitely don't treat it like that for humans - we say things like "they understand it better than me" all the time, which when you boil it down has to mean "their model of it is more predictive than mine".
But then if that is a quantifiable measure, then we're really talking about "reliable enough". And then the questions are: 1) where do you draw that line, exactly, and 2) even more importantly, why do you draw the line there and not somewhere else.
For me, the only sensible answer to this is to refuse to draw the line at all, and just embrace the fact that understanding is a spectrum. But then it doesn't even make sense to ask questions like "does the model really understands?" - they are meaningless.
(The same goes for concepts like "consciousness" or "intelligence", by the way.)
The reason why I think this isn't universally accepted is because it makes us not special, and humans really, really like to think of themselves as special (just look at our religions).
Our capacity to make mistakes does not necessarily equate to a lack of understanding.
If you’re doing a difficult math problem and get it wrong, that doesn’t necessarily imply that you don’t understand the problem.
It speaks to a limitation of our problem solving machinery and the implements we use to carry out tasks.
e.g. if I’m not paying close enough attention and write down the wrong digit in the middle of solving a problem, that could also just be because I got distracted, or made a mistake. If I did the same problem again from scratch, I would probably get it right if I understand the subject matter.
Limitations of our working memory, how distracted we are that day, mis-keying something on a calculator or writing down the wrong digit, etc. can all lead to a wrong answer.
This is distinct from encountering a problem where one’s understanding was incomplete leading to consistently wrong answers.
There are clearly people who are better and worse comparatively at solving certain problems. But given the complexity of our brains/biology, there are myriad reasons for these differences.
Clearly there are people who have the capacity to understand certain problems more deeply (e.g. Einstein), but all of this was primarily to say that output doesn’t need to be 100% “reliable” to imply a complete understanding.
> The weasel word here is "reliably". What does this actually mean? It obviously cannot be reliable in a sense of always giving the correct result, because this would make understanding something a strict binary
What did you mean by "it cannot be reliable in a sense of always giving the correct result", and why would that make understanding something a strict binary?
I do agree that some people understand some topics more deeply than others. I believe this to be true if for no other reason than watching my own understanding of certain topics grow over time. But what threw me off is that someone with "lesser" understanding isn't necessarily less "reliable". The degree of understanding may constrain the possibility space of the person, but I think that's something other than "reliability".
For example, someone who writes software using high level scripting languages can have a good enough understanding of the local context to reason about and produce code reliably. But that person may not understand in the same way that someone who built the language understands. And this is fine, because we're all working with abstractions on top of abstractions on top of abstractions. This does restrict the possibility space, e.g. the systems programmer/language designer can elaborate on lower levels of the abstraction, and some people can understand down to the bare metal/circuit level, and some people can understand down to the movement of atoms and signaling, but this doesn't make the JavaScript programmer less "reliable". It just means that their understanding will only take them so far, which primarily matters if they want to do something outside of the JavaScript domain.
To me, "reliability" is about consistency and accuracy within a problem space. And the ability to formulate novel conclusions about phenomena that emerge from that problem space that are consistent with the model the person has formed.
If we took all of this to the absurdist conclusion, we'd have to accept that none of us really understand anything at all. The smaller we go, and the more granular our world models, we still know nothing of primordial existence or what anything is.
Other way around. If a kid understood the concepts of multiplation, they could train themselves on the next logical steps like exponents.
The consequences of this in the terms of AI would mean they build on a series of concepts and would quickly dwarf us in intelligence.
We don't know, because these concepts were discovered before writing. We do know that far larger jumps have been made by individual mathematicians who never made it past 33 years of age.
I don't know many people that would suggest that knowing about ones family structure, such as their mom/dad/uncle, and how their history relates to them, is required to be _completely reconstructed_ every time they interact with their environment somehow from first principles.
Online reinforcement learning can have large merits without resorting to stating it's required for learning. Just as self awareness can occur independently of consciousness can occur independently of intelligence is independent of empathy and so on. They are all different, and having one component doesn't mean anything about the rest.
The images we get from these neural networks are trained on looking pleasing, for some definition of pleasing. That's why they look good on the whole, but get into that uncanny valley the moment you go inspecting. Similar to dreams.
Whereas obviously real human perception is (usually) grounded in reality.
It looks like they actually only tried depth/normal maps for real images, not albedo/lighting maps, but it certainly seems possible. Honestly depth maps often look impressive at first glance but typically if you actually try to use them for anything (e.g. reprojection, DoF blur) their hidden flaws instantly become super apparent, and they aren't as useful as you might think. Gaussian splatting reconstructions already do reprojection better anyway. OTOH albedo maps look weird, but relighting is extremely useful and important. I'm not excited about yet another way to generate approximate depth maps, but I am excited about relighting.
I understand this paper links to open source model where that can be verified, but maybe this is one secret sauce of these more advanced models?
Or, are you asking how the researchers extracted it?
> I-LoRA modulates key feature maps to extract intrinsic scene properties such as normals, depth, albedo, and shading, using the models' existing decoders without additional layers, revealing their deep understanding of scene intrinsics.
I think this is awesome but maybe it’s not really that surprising given how this “generate and then finetune” approach has already worked so well?
Then you could have it start to create a vocabulary and an understanding around what all these parameters and weights do, since to us they're just a bunch of seemingly random floating point numbers.
At which point we may be able to start creating NNs from first principles, without even requiring training.