LLMs are trained to predict tokens on highly mediocre code though. How will it exceed its training data?
The question is, do we have good enough feedback loops for that, and if not, are we going to find them? I would bet they will be found for a lot of use cases.
/end extreme over optimism.
I think you can have LLMs do that too, and then generate synthetic training data for "high-effort code".
Part of the problem is that better code is almost always less code. Where a skilled programmer will introduce a surgical 1-3 LOC diff, an incompetent programmer will introduce 100 LOC. So you'll almost always have a case where the bad code outnumbers the good.