But now you get SD, dalle and others which add more information not just by scaling, but also by mapping sentences/words to pre-existing images that already have cohesion. That way when you write in sentences to the text prompt, the model has more semantic information about what an eye is, but (IMO) only _indirectly_ because it will map a sentence to images that match that sentence. The question is always what information is actually contained in the training set and what is missing from it and when it creates an image where is the information from etc.
In some ways, that means I think that meaning to us as humans, is different from scaling which is almost like pixel resolution except resolution of patterns and differentiation of patterns. Meaning in this sense is things like creating a doorway with no actual door, but still the doorway itself looks super realistic is rendered. You can fix it by scaling and increasing the differentiation of patterns I guess, but you can never fix all instances completely with scaling. That's why in some ways I think meaning is sort of orthogonal to scale, however on a philosophical level, they should converge but that's for another topic.
I may have missed something in my thoughts here because this is sort of difficult to talk about without writing a book eventually.