If the Twitter user claimed that the text prompts themselves were generated by asking GPT-3 for "An interesting sequence of text prompts to feed an image generation AI" or something like that, I would have believed them.
Presumptively, I imagine it's harder to create a model that generates images matching a certain human-language input prompt than to create an image generation model with no language component and have it pick its own scenarios internally. I don't think the former is done to palm off "most of the work" to humans, but rather because people want an easy way to see create own ideas so there's more demand for it.
> You have to re-train it from scratch every time you want it to remember something truly new, there's no feedback loop to do that.
As far as I'm aware, this isn't true. Deep learning is perfectly compatible with fine-tuning an existing model using new data. OpenAI/MS have been doing this with Codex to improve it based on Copilot telemetry and code from new languages/libraries.