You can clearly see [1] how the glyphs are created by approximately merging several letters in one place to create a new symbol, while maintaining the overall structure of words and paragraphs within a well-composed layout in the page. The periodic table [2] seems to be doing the same, where the layout is the structure of colored boxes and two sizes of letters within them.
Image composition seems to be doing something similar, learning to associate concepts with their visual representations at the right level, and merging already-seen examples in the right proportions to create a novel image. Cats and vampires are represented by their distinctive features, arms and legs are correctly positioned as parts of the body according to the action they perform, and instructions of style (either by artists like "Pieter Brueghel", art styles like "digital" or "mosaic", or even camera settings like ISO exposure [3]) are translated into lower level pixel representations of color, shapes and shadowing.
My hypothesis is that if you included examples where the letters are taught one by one, like in kindergarten primers, it may be able to learn those concepts as well and generate better painted text (although I'm not sure it could make the jump to "understanding" the relation between their role as input instructions and as output image text).
[1] https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/855...
[2] https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/157...
[3] https://www.bramadams.dev/projects/dalle-tricks#let-there-be...
These fall into The Valley for me. I don’t know why, but they do.