We need to backpropagate from output string back through the network to generate the input string.
LLMs are much more complicated structures than imagenets, so I’m sure there are a lot of challenges involved.
Also I imagine the input (output?) might be complete ghibberish.