Often these systems will have really bizzare artificats, people with 3 arms, etc. However at the same time when you glance at the output without looking carefully you will sometimes miss these artifacts even though they should be absolutely glaring.
Often these systems will have really bizzare artificats, people with 3 arms, etc. However at the same time when you glance at the output without looking carefully you will sometimes miss these artifacts even though they should be absolutely glaring.
I suspect also the reason the images look OK at a glance is because the images as a whole also represent patterns in the model so they actually come from "real life" / artist created images and thus have some sense of cohesion. But making the AI have all the right patterns so it never makes a mistake at all scales of the image while also being able to combine the pattern with real understanding of what they are conceptually is the real trick but until then it will be a "salad bowl collage" thing at random intervals.
The closest thing to the brain it looks like to me is simply the hierarchical nature of it which seems similar to v1/v2/the vision system in humans but I've only been told that, I'm no neuroscientist.
FWIW I don't think there is anything particularly wrong in the model architectures or training data that in some fundamental way makes it impossible to always get 2 arms. After all, lots of other tricky things are almost always correct. I suspect it's a question of training time and model size mostly (not trivial of course as it's still expensive to re-train to check modified architectures etc). It's also a matter of diffusion sampling iterations and choice of sampler at inference time, for the case of SD.
I also don't think there's anything wrong with the model architectures in themselves or the data, nor that it is impossible, only that it is hard and as you say I think it needs a lot of data and clever engineering to fix mistakes. It may even be possible to fix most mistakes, over time, which would be pretty impressive imo, but the absolute limits of what a model can produce/"contain" with our hardware is kind of an open question though interesting.
There's simply nowhere for the computation to go.
But it rarely would put out say 8 arms. And the repeat artifacts are miles ahead of earlier stuff like clip draw or disco diffusion. So it does seem to have some idea of what's going on, just isn't perfect yet. It gets much worse without the 512x512 resolution, if you push both dimensions it loses scene coherence a lot more.
However, where it struggles I find is with finer details, and also _placement_ of things like arms, eyes, and relationships between them. This I think is because it only has a general idea of the shape of persons but no data for the exact specifics like where the arms, legs, eyes and so on should be placed in a very realistic anatomical way, and this is where I think the challenge is - the gap between a general pattern of a person and an extremely specific but also general one where it can modify it and transform it like a real human artist can. I'm not sure that's in the data exactly
This is a fundamental misunderstanding of what it is doing.
You can see in work like https://twitter.com/lintool/status/1579830653126086656 that the model does have an understanding of what parts of the visual model represent as concepts.
That is thoroughly confused to the point of uselessness.
The reason you get structural issues is because it's hard for the architecture to express large scale structure, but they get better and better at it simply by scaling up the network.
But now you get SD, dalle and others which add more information not just by scaling, but also by mapping sentences/words to pre-existing images that already have cohesion. That way when you write in sentences to the text prompt, the model has more semantic information about what an eye is, but (IMO) only _indirectly_ because it will map a sentence to images that match that sentence. The question is always what information is actually contained in the training set and what is missing from it and when it creates an image where is the information from etc.
In some ways, that means I think that meaning to us as humans, is different from scaling which is almost like pixel resolution except resolution of patterns and differentiation of patterns. Meaning in this sense is things like creating a doorway with no actual door, but still the doorway itself looks super realistic is rendered. You can fix it by scaling and increasing the differentiation of patterns I guess, but you can never fix all instances completely with scaling. That's why in some ways I think meaning is sort of orthogonal to scale, however on a philosophical level, they should converge but that's for another topic.
I may have missed something in my thoughts here because this is sort of difficult to talk about without writing a book eventually.
And you're right that this is pretty unfounded intuition. Humans often seek meaning in things without meaning, so it might be unfounded. At some point all i can really do is shrug and say it feels "spooky" to me.
I’m no NN guy, but to me all it seems as basically underconstrained and unrelated to “understanding”. It’s like these e.g. woodwork, magic trick, dancing, guitar, etc teachers who fail to message a way to do something and can only tell “look”, then just do it, ask you to repeat, and get annoyed when you fail again.
Turns out nobody quite knows how to draw a bicycle. They get the gist but the details don't make sense.
I strongly suspect that if we do ever fully map the "architecture" of the brain, the result will be a massive graph that's not readily understandable by humans directly. This is already the case in biology. We'll end up with a computational artifact that'll help us understand cause and effect in the brain, but it'll be nothing like a tidy diagram of tensor operations like in state of the art ML papers.
There’s some image I see on occasion that’s 100% garbage. If you focus on it you cannot make out a single thing. But if you glance at it or see it scaled down, it looks like a table full of stuff.
I don't know if AGI is down the road diffusion models have taken us. I'm not even really sure what most people mean by AI when they talk about it. But stable diffusion et al are clearly super human. I'm not sure that AGI is down the trail cut by diffusion models, but if it's ever accomplished, these models will almost assuredly represwbt some of the learnings required to get there.
Very well then I contradict myself,
(I am large, I contain multitudes.)