> I'd love to read a less technical writeup explaining the challenges faced and why it took so long to figure out spelling. Scrolling through the paper is a bit overwhelming and it goes beyond my current understanding of the topic.
I'm not an ML researcher but I can answer this. Note that this is not information from the paper, just my own findings from following the "scene".
We actually figured out spelling not long after diffusion models came out, the imagen paper that came several months before SD1 explained how they did it, which is the same technique SD3 and Dalle3 use, instead of using CLIP's text encoder, they use T5.
The reason why image models can't spell is the same reason why language models also have difficulty spelling. Tokenization. Simply speaking instead of seeing each letter individually, we split a sentence into sub-words, most commonly called tokens and that's what the model sees. The model never gets to see each letter individually. But it turns out that if you make the model big enough and feed it enough data, it actually learns how each token is spelled out.
Clip is both small (200-500M params) and trained on limited text data (only image captions). T5 is trained on a large corpus of data and is also huge (~5.5B params). This makes T5 the obvious choice if you care about spelling.
So why use CLIP in the first place? Simple: it's much easier to train on clip embeddings than T5 embeddings, not only are they smaller, but because clip is trained on text/image pairs and due to backpropagation, the text embeddings also contain a lot of visual semantic information. This simplifies a lot of the work the text->diffusion attention modules need to do. Another reason is that T5 is absolutely massive, it's 5x larger than the image part of the model and 10-20x larger than CLIP.
If you take a close look at diagram (a) in the paper you'll actually see that SD3 uses both CLIP and T5, not only that but they trained it in a way that makes the encoder used optional, so you can use the CLIP models only if you don't care about spelling and image composition (CLIP is also bad at understanding prompts), which is useful because most GPUs can't handle T5 on it's own. Tho I suspect someone will distil T5 so it becomes 5-10 times smaller than it currently is at a minimal loss on how good it is for prompting.