I had a LLM based prototype called "cleanroom" which would convert a program into a spec and then back into a program.
The results were disgusting: the spec would encode all sorts of irrelevant implementation details, and then the new version would reimplement them faithfully, and be 3x more bloated than the original. The exact opposite of what I was going for!
I didn't put much effort into it, maybe it was solvable with prompting (or more likely, more human effort on the spec phase), but it looks like the LLM has the same problem as the human, it can't know what the intention was, and it can't know what's relevant, what's essential and incidental.
But basically, what I needed wasn't a spec but user stories. (And probably multiple prototype outputs to choose from...)
I should definitely give it another crack though...
---
P.S., Spoiler for next ten years: software as biology (esp. crossbreeding, mutation, selection pressure...)