12 months ago I'd have agreed. But these are from raw StableDiffusion, no face enhancement (unless noted). Just created these right now:
https://imgur.com/a/dBtVtg3 (Note this has face enhancement on and I censored this so it is SFW)
(Still not great at hands yet!)
The rest of your comment is just your misinterpretation of what the models are doing. Modern diffusion models do have large scale knowledge across the whole scene.