I think long term LLMs should directly generate Abstract Syntax Trees. But this is hard now because all the training data is text code.
That's interesting. Is there research into adding memory or has it been proven that it provides no pragmatic value over any context it outputs?