The problem, from the paper:
> Several methods alleviated this issue by incorporating explicit text position and content as guidance on where and what text to render. However, these methods still suffer from several drawbacks, such as limited flexibility and automation, constrained capability of layout prediction, and restricted style diversity.
Looking at the diagram provided, they use GPT-4 to suggest the position following the text prompt.
I see it as very useful for making sure to have the text in the right position without doing manual work trying to find the right position. I'm not an expert, but doesn't this method add another cost and overhead for calling Text-to-Image models?