LLM's are trained on human knowledge and taste. They are actually pretty good at deciding if a conjecture would be found "interesting" by the mathematical community or not.
Note that I am saying LLM, and not chatbot or agent. But even a chatbot can often still reasonably rank a list of mathematical statements by vague properties like "interestingness".
How to RL this is a bit of an open question, but there are interesting conjectures of how to do it.