61 karma · joined July 1, 2014
The second half of the video provides an overview of the method. https://www.youtube.com/watch?v=ihDbAUh0LXk
For example, in Figure 11 of the paper (https://arxiv.org/pdf/2304.06720.pdf), you can see that full-text "rustic cabin -> rustic orange cabin" does not turn the cabin orange.
For coloring, the core benefit of our method is that it allows precise color control. For example, it can generate colors with rare names (e.g., Plum Purple or Dodger Blue) or even particular RGB triplets that we cannot describe well with texts.
You can examples in Figure 4 here: https://arxiv.org/pdf/2304.06720.pdf
You can find the segmentation information on the bottom-right of the rich-text generation result.
For example, the method was originally tested in Stable Diffusion 1.4. But we can easily apply it to Stable Diffusion-XL (or any finetuned model like ANIMAGINE-XL) even though the new model has a different text encoder and U-Net weights.
Some examples here: https://youtu.be/ihDbAUh0LXk?si=i3LFfkDXIDKKvne3&t=91
Our results show that Prompt-to-Prompt would generate artifacts and other heuristic methods are not very effective.
You can see some examples in Figure 21, 22, 23. https://arxiv.org/pdf/2304.06720.pdf
RE: plaintext - The "plain-text" result is just a baseline. We call the "plaintext parentheses in the prompt" full-text (i.e., expanding the rich text info into a long sentence). We show many "full-text" in the paper https://arxiv.org/pdf/2304.06720.pdf. You can see in Fig 11 that full-text results cannot change the color, style, and do not respect the description. More examples in Figure 13, 14, and 15.
The main issue of using full-text is that it cannot preserve the original plain text image, thereby requiring many rounds of prompt tuning/engineering. We also compared with two other image editing methods, Prompt-to-prompt and InstructPix2Pix. But they could not handle localized editing well. You can see some example comparisons for Color (Figure 4), Style (Figure 5), Footnote (Figure 8), and Font Size (Figure 9). https://arxiv.org/pdf/2304.06720.pdf
RE: Style - Yes, you can specify what styles you want by just describing it.
It could be a simple word like "Ukiyo-e" or "Van Gogh", or some detailed descriptions. You can check out some examples in the video: https://youtu.be/ihDbAUh0LXk?si=wVF9LIF1NVqLtNDC&t=59
The particular font family used is just a "label".
Glad that you posted the comments! I hope this clarifies. Happy to answer any questions.