Currently LLMs are forward passes only, for each token. You get a final set of data that is statistical in nature. I.e basically a fancy look up table of sorts.
To get more capability, you have to post process that data. For example, you can get the current LLMs into situations where you ask it a question, and it gives an answer, then you ask a clarifying question, and then it gives you something slightly different, then when you ask it to reconcile the two answers, it will say something like "my apologies...", and gives you the final answer.
You could do some clever prompt engineering to automate this whole process (i.e have scripts that call the LLM with certain prompts based on answers), but thats a manual step that someone has to code. Ideally, this process should exist in LLM, and be learned.
Now, you could just train a bunch of LLMs on different concepts and then train a selector on which LLM to use, and do everything with forward pass, but then you have a bunch of redundant models that are taking up space, without any way for those models to talk to each other.
Thats why figuring out RNNs (where a layer could feed data back into a previous layer) training in the context of LLMs would be pretty much the only way to do this.