if we somehow are able to take a "test-driven-approach" with them.
Upfront, we tell it what criteria must pass. Connect it to a "test/task" runner of some sort per prompt. Then, it keeps trying behind the scenes to come up with an answer that passes the criteria (to prevent humans from having to waste iterations on 'no, you got it wrong, try again')
However, I will admit... most of the time when the LLM can't solve it, retrying and retrying over and over slightly different ways really doesn't typically end up in success. If the task is "too complex" for the LLM, I have yet to find a way to break it down context wise and overcome the limitation.