A lot of the public successes with agents is really LLM-driven local search against an objective function that is evaluated in more traditional ways. This one seems to fit the pattern.
https://orlp.net/blog/bad-ai/#objective-p-mathrm-relevant-1-...
I think it still holds up.