But in brief, the short-term evolution of LLMs is going to involve something like letting it `eval()` some code to take an action as part of a response to a prompt.
A recent paper, Toolformer: https://pub.towardsai.net/exploring-toolformer-meta-ai-new-t... which is training on a small set of hand-chosen tools, rather than `eval(<arbitrary code>)`, but hopefully it's clear that it's a very small step from the former to the latter.
You can stop it from being recursive by passing it through a model that is not trained to write JavaScript but is trained to output JSON.
You might say that it doesn't preserve state between different sessions, and that's true. But if it can read and post online, then it can preserve state there.
Feedback loops are an important part.
But let's say you take two current chatbots, make them converse with each other without human participants. Add full internet access. Add a directive to read HN, Twitter and latest news often.
Interesting emergent behaviour could emerge very soon.
A sobering intuition pump: https://www.lesswrong.com/posts/kpPnReyBC54KESiSn/optimality...