In one case I used a made-up metric (since I didn't know the exact name or if it existed) and it somehow optimized that too.
In one case I used a made-up metric (since I didn't know the exact name or if it existed) and it somehow optimized that too.
It's also not terrible on token usage for smaller projects.
> "Attempting" implies a high risk of failure.
Of course it does? Your safeguards also imply a high risk of failure. You have restrictions that just rollback everything the LLM "attempts" to do.
That is not to say that the overall workflow is failure prone, but obviously you have setup an apparatus that allows the LLM to just shotgun attempts, whether it understands it or not. And sometimes it's not going to be able to find any solution. So it's not really appropriate for people to leave with the impression that anything measurable can be successfully optimized with LLMs.