The LLM vendors are all competing on how well their models can write code, and the way they're doing that is to refine their training data - they constantly find new ways to remove poor quality code from the training data and increase the volume of high quality code.
One way they do this is by using code that passes automated tests. That's a unique characteristic of code - you can't do that for regular prose, or legal analysis or whatever.
"Even if you describe your problem (prompt) to a high standard, there is no way it can deliver a solution of the same standard."
My own experience doesn't match that. I can describe my problems to a good LLM and get back code that I would have been proud to have written myself.
Which is fine, as long as people are aware of it.