This must be one of the best applications for LLMs, as you can always automatically verify the results, or reject them otherwise, right?
However I could see this being useful for verified software development, which usually involves a huge amount of tedious and uninteresting lemmas, and the size of the proofs becomes exhausting. Having an LLM check all 2^8 configurations of something seems worthwhile.
But generating useless code, or proofs, just to discard them is hardly a consequence and externality free effort.
10 trillion llms, powered by a Dyson sphere, that still output slop is still worse than one unappreciated post doc.