I suspect also the reason the images look OK at a glance is because the images as a whole also represent patterns in the model so they actually come from "real life" / artist created images and thus have some sense of cohesion. But making the AI have all the right patterns so it never makes a mistake at all scales of the image while also being able to combine the pattern with real understanding of what they are conceptually is the real trick but until then it will be a "salad bowl collage" thing at random intervals.
The closest thing to the brain it looks like to me is simply the hierarchical nature of it which seems similar to v1/v2/the vision system in humans but I've only been told that, I'm no neuroscientist.