This is very cool, but it's gimmicky. All of the rich text could simply be a modifier before or after the word (such as an adjective or phrase). Given most LLM work is plain text, this benefit isn't as neatly transferable as prompt engineering.
For example, the method was originally tested in Stable Diffusion 1.4. But we can easily apply it to Stable Diffusion-XL (or any finetuned model like ANIMAGINE-XL) even though the new model has a different text encoder and U-Net weights.