To learn logical concepts just from images seems entirely impractical, like we can't rely on having enough images such that models can understand words coherently as language. You could draw a picture of a sign that says "children crossing" not because you can understand and remember exactly what an image of such a sign would look like, but because you have an understand of English and the character set that would let you reproduce it. If you tried to learn to create the same sign in Arabic you'd either need to see a huge number of signs to learn from or (more likely) build a language model for Arabic.
The kind of abstract understandings that we know we can train in language models just aren't learned by image transformers at this scale (or likely any practical scale). A language model could easily understand: "A red cube is stacked on top of a blue plate, a green pyramid is balanced on the red cube" and infer things like the position of the pyramid relative to the blue plate, image models quickly fall over with such examples.
An interesting nascent (and hacky) example of the benefits of combining models is people are using language models like GPT-3 to create better prompts for image models.
Also, I seem to recall that at least some models deliberately harmed generation of human faces (e.g. by selection of training data) to draw away attention from the deepfake/fakenews usecases and the related ethical,political and PR issues; I would assume that if any of them wanted to actually try and make specifically faces look good, that would be purely a matter of some engineering work without any breakthroughs needed - I mean, we have evidence from face-specific models that the same technical architecture can do decent faces.
https://replicate.com/laion-ai/erlich
It still has issues of course, but a lot better at spelling than DALLE2.
The major breakthrough that happened 6 months ago is that someone put their api behind a website for people to play with
I think a lot of this has been solved in DALL-E already. It's pretty good at right-looking faces, and fingers. Text not so much... but it does appear that whatever OpenAI are doing behind the scenes, that's getting better too.