When the answers can be checked by another program and best-of-n results get better with increasing n, yes.
I think TDD code development is this, at least in principle.
I think TDD code development is this, at least in principle.
Some things are easier to verify than to generate.
LLMs can often make a distribution of generated outputs that is much closer to a good answer than other methods.