HNHacker News
TopNewBestAskShowJobs

jbhuang0604

61 karma · joined July 1, 2014

https://jbhuang0604.github.io/
submissionscomments
jbhuang0604··on Expressive text-to-image generation with rich text
Thank you!
jbhuang0604··on Expressive text-to-image generation with rich text
The UI part is basically a way to organize the user's intention. In the backend, we develop method for extracting "token maps" (i.e., which spatial regions correspond to specific words) and use region-based diffusion to achieve these localized editing results.

The second half of the video provides an overview of the method. https://www.youtube.com/watch?v=ihDbAUh0LXk

jbhuang0604··on Expressive text-to-image generation with rich text
Yes! I am particularly excited about this feature.
jbhuang0604··on Expressive text-to-image generation with rich text
Yup, it could be similar, but it mostly only works for very simple prompts (e.g., one subject in the image).

For example, in Figure 11 of the paper (https://arxiv.org/pdf/2304.06720.pdf), you can see that full-text "rustic cabin -> rustic orange cabin" does not turn the cabin orange.

For coloring, the core benefit of our method is that it allows precise color control. For example, it can generate colors with rare names (e.g., Plum Purple or Dodger Blue) or even particular RGB triplets that we cannot describe well with texts.

You can examples in Figure 4 here: https://arxiv.org/pdf/2304.06720.pdf

jbhuang0604··on Expressive text-to-image generation with rich text
Yes, if you go to the huggingface demo: https://huggingface.co/spaces/songweig/rich-text-to-image

You can find the segmentation information on the bottom-right of the rich-text generation result.

jbhuang0604··on Expressive text-to-image generation with rich text
Got it! Thanks for the feedback! This is definitely something we can improve.
jbhuang0604··on Expressive text-to-image generation with rich text
Thanks for the comment! Our method is model-agnostic. It can be easily adapted to any LLM (aka the text-encode) and any text-to-image models.

For example, the method was originally tested in Stable Diffusion 1.4. But we can easily apply it to Stable Diffusion-XL (or any finetuned model like ANIMAGINE-XL) even though the new model has a different text encoder and U-Net weights.

jbhuang0604··on Expressive text-to-image generation with rich text
Yes, I think the "footnote" showcases this well. You can use it to interactively explore your visual imagination.

Some examples here: https://youtu.be/ihDbAUh0LXk?si=i3LFfkDXIDKKvne3&t=91

jbhuang0604··on Expressive text-to-image generation with rich text
Thanks! BTW, we recently implemented the Automatic1111 extension so you can install and use it directly with A1111.

https://github.com/songweige/sd-webui-rich-text

jbhuang0604··on Expressive text-to-image generation with rich text
We compared with three other prompt term weighting (aka font size) methods: - Prompt-to-Prompt - Parenthesis - Repeating

Our results show that Prompt-to-Prompt would generate artifacts and other heuristic methods are not very effective.

You can see some examples in Figure 21, 22, 23. https://arxiv.org/pdf/2304.06720.pdf

jbhuang0604··on Expressive text-to-image generation with rich text
Yes, the particular font family was just a "label". We don't really interpret what that font style would look like.
jbhuang0604··on Expressive text-to-image generation with rich text
Exactly! You get to have full control over the detailed contents you wish to generate.
jbhuang0604··on Expressive text-to-image generation with rich text
Yes, you can expand the rich text information into a long sentence. We call this full-text in the paper. The issue of using "full-text" is that it's hard to edit the image interactively. Every time you change the text, you get an entirely different image.
jbhuang0604··on Expressive text-to-image generation with rich text
Thanks a lot for the comment! (one of the authors here)

RE: plaintext - The "plain-text" result is just a baseline. We call the "plaintext parentheses in the prompt" full-text (i.e., expanding the rich text info into a long sentence). We show many "full-text" in the paper https://arxiv.org/pdf/2304.06720.pdf. You can see in Fig 11 that full-text results cannot change the color, style, and do not respect the description. More examples in Figure 13, 14, and 15.

The main issue of using full-text is that it cannot preserve the original plain text image, thereby requiring many rounds of prompt tuning/engineering. We also compared with two other image editing methods, Prompt-to-prompt and InstructPix2Pix. But they could not handle localized editing well. You can see some example comparisons for Color (Figure 4), Style (Figure 5), Footnote (Figure 8), and Font Size (Figure 9). https://arxiv.org/pdf/2304.06720.pdf

RE: Style - Yes, you can specify what styles you want by just describing it.

It could be a simple word like "Ukiyo-e" or "Van Gogh", or some detailed descriptions. You can check out some examples in the video: https://youtu.be/ihDbAUh0LXk?si=wVF9LIF1NVqLtNDC&t=59

The particular font family used is just a "label".

Glad that you posted the comments! I hope this clarifies. Happy to answer any questions.

jbhuang0604··on 3D photographic inpainting from a single source image
Yes, you are right. Existing single-image depth estimation models are not far from perfect. We hope to see future development in that direction to further improve the visual quality of 3D photos.
jbhuang0604··on Multi-view Wire Art
Thanks for sharing. Yes, what we found is that multiple input images often are not consistent with each other (i.e., no physical volume that can satisfy the constraints from all the views). We thus start with a set of consistent voxels and connect them so that the projections are as similar to the inputs as possible.
jbhuang0604··on Multi-view Wire Art
This is interesting. Imagine that we have a series of multi-view wire structure (i.e., adding a time dimension), then we probably can project three different animations.