Then they can start attempting to optimize it. They can also spin round and round making the numbers worse because they don't actually know what to do.
Then they can start attempting to optimize it. They can also spin round and round making the numbers worse because they don't actually know what to do.
In one case I used a made-up metric (since I didn't know the exact name or if it existed) and it somehow optimized that too.
It's also not terrible on token usage for smaller projects.
> "Attempting" implies a high risk of failure.
Of course it does? Your safeguards also imply a high risk of failure. You have restrictions that just rollback everything the LLM "attempts" to do.
That is not to say that the overall workflow is failure prone, but obviously you have setup an apparatus that allows the LLM to just shotgun attempts, whether it understands it or not. And sometimes it's not going to be able to find any solution. So it's not really appropriate for people to leave with the impression that anything measurable can be successfully optimized with LLMs.
You need a measurement that can falsify hypotheses and reject branches that won't work.
Also, if all you have left in your project are performance issues that are hard to identify without flailing around (even with Fable/Astra) despite sampler/profiler reports, then you're doing really well and I wouldn't assume you're going to fare much better than the sota models in terms of stabs in the dark.
"Claude, if this idea doesn't measure as an improvement (use X benchmark and a T-test), discard it and try the next idea."