This is very interesting, I would love to play around with this.
However, it looks like the quality of the result really depends on the prompt, or in other words, on the quality of the training dataset captions.
Would there be a way to "fine-tune" the network on a smaller, better-captioned dataset to have better control over the output?
Can I make my own dataset and use it with few-shots learning to expand the capabilities of existing networks? Making it better at generating say, spaceships or fantasy animals?