Expressive text-to-image generation with rich text
rich-text-to-image.github.io
rich-text-to-image.github.io
We can already weigh parts of prompts, we can already specify colors or styles for parts of the images. And even if we could not, none of this needs rich text.
In the beginning I even think their comparisons are dishonest. They compare "plaintext" prompts with "rich text" prompts, but the rich text prompts contain more information. What? Like, seriously, who is surprised the following two prompts give different images?
(1) "A girl with long hair sitting in a cafe, by a table with coffee on it, best quality, ultra detailed, dynamic pose."
(2) "A girl with long [Richtext:orange] hair sitting in a cafe, by a table with coffee on it, best quality, ultra detailed, dynamic pose. [Footnote:The ceramic coffee cup with intricate design, a dance of earthy browns and delicate gold accents. The dark, velvety latte is in it.]"
the worst part is "Font style indicates the styles of local regions". In the comparison with other methods section they actually have to specify in parentheses what each font means style-wise, because nobody knows and (let's be frank) nobody wants to learn.
So why not just use these plaintext parentheses in the prompt?
I really stopped myself from immediately posting my (rather negative) opinion, but after over an hour, it hasn't changed. As far as i can see, this isn't useful, rich text prompts are a gimmick.
Some examples here: https://youtu.be/ihDbAUh0LXk?si=i3LFfkDXIDKKvne3&t=91
It's even worse with styles; midjourney can't do a guitar in one style and the rest of the image in another style. You really only get one style per image.
How about a plain-text interface like this?
> A girl with [long hair](orange) sitting in a cafe, by a table with [coffee](^1) on it, best quality, ultra detailed, dynamic pose. [^1](Ceramic coffee cup with intricate design, a dance of earthy browns and delicate gold accents. The dark, velvety latte is in it.)
It could be a text description of a fighter or noble wearing coats or armour. And then substitute in different style description of coats and armour depending on the family, class, race or other attributes suitable for the world you're trying to generate.
With the same seed, and an extremely similar prompt, why would you get an entirely different image?
If I take seed 9999999 (just example) and my prompts are
(1) "very large gothic church at dusk, spooky, horror, red roses" and
(2) "very large gothic church at dusk, spooky, horror, white roses"
then with all models I tested over the last year or so, you get _very_ similar images, with different colored roses, and (at most) very minor changes eleswhere. this only seems to work if you keep in mind the prompt being parsed left to right, so changes further to the beginning of the prompt have larger effects. Again, of course, you need the same seed.
But, with this said, why would that be any different with plain/full/rich text. Apologies if I am somehow blinkered and asking something really obvious.
For example, in Figure 11 of the paper (https://arxiv.org/pdf/2304.06720.pdf), you can see that full-text "rustic cabin -> rustic orange cabin" does not turn the cabin orange.
For coloring, the core benefit of our method is that it allows precise color control. For example, it can generate colors with rare names (e.g., Plum Purple or Dodger Blue) or even particular RGB triplets that we cannot describe well with texts.
You can examples in Figure 4 here: https://arxiv.org/pdf/2304.06720.pdf
RE: plaintext - The "plain-text" result is just a baseline. We call the "plaintext parentheses in the prompt" full-text (i.e., expanding the rich text info into a long sentence). We show many "full-text" in the paper https://arxiv.org/pdf/2304.06720.pdf. You can see in Fig 11 that full-text results cannot change the color, style, and do not respect the description. More examples in Figure 13, 14, and 15.
The main issue of using full-text is that it cannot preserve the original plain text image, thereby requiring many rounds of prompt tuning/engineering. We also compared with two other image editing methods, Prompt-to-prompt and InstructPix2Pix. But they could not handle localized editing well. You can see some example comparisons for Color (Figure 4), Style (Figure 5), Footnote (Figure 8), and Font Size (Figure 9). https://arxiv.org/pdf/2304.06720.pdf
RE: Style - Yes, you can specify what styles you want by just describing it.
It could be a simple word like "Ukiyo-e" or "Van Gogh", or some detailed descriptions. You can check out some examples in the video: https://youtu.be/ihDbAUh0LXk?si=wVF9LIF1NVqLtNDC&t=59
The particular font family used is just a "label".
Glad that you posted the comments! I hope this clarifies. Happy to answer any questions.
You can find the segmentation information on the bottom-right of the rich-text generation result.
Not sure how this could be communicated better.
All of the techniques that they are showing have already existed for awhile in places like Automatic1111/ComfyUI or its extensions (i.e. regional prompting, attention weights). Having it connect so seamlessly with rich text is awesome and is a cool UI trick that might make normies notice it.
Also, related, but NLP is extremely undertooled on the prompt engineering side. Most of the techniques here would work just fine on any LLM. If you don't believe me, read this: https://gist.github.com/Hellisotherpeople/45c619ee22aac6865c...
The stability of the overall image during local changes makes me think that maybe this could be a key to video generation (because the biggest problem with existing diffusion-based approach for video is their instability from frame to frame).
Prompt weighting alone can fix undesired aspects of an output, especially with SDXL and its dual text encoders.
Our results show that Prompt-to-Prompt would generate artifacts and other heuristic methods are not very effective.
You can see some examples in Figure 21, 22, 23. https://arxiv.org/pdf/2304.06720.pdf
For example, I'm wondering if a prompt written in Comic Sans should be turned into a comic-style illustration, or does it come out as a simplistic and childish drawing? Is a gothic font meant to imply a style of architecture, old Germanic peoples, or goth music and style?
See also https://design.tutsplus.com/articles/the-psychology-of-fonts...
For example, the method was originally tested in Stable Diffusion 1.4. But we can easily apply it to Stable Diffusion-XL (or any finetuned model like ANIMAGINE-XL) even though the new model has a different text encoder and U-Net weights.
(As a side note: using decorative typefaces was an unconvincing example.)
The second half of the video provides an overview of the method. https://www.youtube.com/watch?v=ihDbAUh0LXk