I don't know why I have to say this again, but I am telling you that the workflow you underlined in the article is good. But the more you argue about it, the more it sounds like you want people to take a more generalized conclusion than what you actually demonstrated.
> "Attempting" implies a high risk of failure.
Of course it does? Your safeguards also imply a high risk of failure. You have restrictions that just rollback everything the LLM "attempts" to do.
That is not to say that the overall workflow is failure prone, but obviously you have setup an apparatus that allows the LLM to just shotgun attempts, whether it understands it or not. And sometimes it's not going to be able to find any solution. So it's not really appropriate for people to leave with the impression that anything measurable can be successfully optimized with LLMs.