You're going to get a wide range of responses on this but given the improvement in these models in just one year, and the number of bugs they're detecting which humans could not, I suspect that even median vibe-coded software is going to surpass median human coded software soon - if it hasn't already.
The important thing to remember here is there perfect isn't on the table. The benchmark is existing human-introduced bugs vs LLM-introduced bugs. Many developers have encountered odd bugs which a human would not have introduced, while forgetting about all the bugs caught which humans introduced. Or their opinion is formed by models from six months ago.