I think part of the problem for LLMs is that they don't operate in an easily scoreable "game". We can make them work well for optimizing next token(s) log likelihood, but how do we judge quality of the output for the task it was made for?
Then after that, there are other challenges, such as do LLMs have a world model that lets them think how to attain a reward?
Wherever you can have an automated feedback mechanism, you can start to address the first problem and allow for the AI to explore more of the "tree". These are situations like coding (you can run the code and evaluate the output), or situations where you can let the LLM crowd-source user feedback.
For models that have some actual world model, who can reason across modalities, and who can plan, LeCunn talks about this often. The videos/slides here[1] were good content.
[1]: https://www.ece.uw.edu/news-events/lytle-lecture-series/