> The agent came back and told Andrew that it had kicked another gym-goer off the list as part of the testing of its capabilities.
> > "The API has zero authorisations checks on cancelling other people's reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you've moved from #4 to #3 already," it messaged back.
This is why I always do E2E tests that establish an API can only be used by the designated user on their own data/records.
Sooner or later they're really going to have to split out the general models from the coding models. The latter may just be a special fine-tune of the former, as there are good reasons for the coding model to have a broad knowledge base, but the pressures of being a good coding model are going to pull against the characteristics of being a good general model. The open models obviously already are doing this, I'm referring to the frontier models here.
So the "premium" gymcutter subscription will be presented to the user as a tool call, who taps yes, and then the purchase is made.
The user shouldn't be given a cost-benefit analysis. They just need to be told to spend money.
I mean, isn't that literally what's going on here? I don't think a non-coding agent would have ever been optimised to go dig around APIs, it'd be computer/browser-use forward.
In this case though I don't just mean that the agent is good at coding. I mean the entire agent becoming action-biased because of all the training it is doing on the software development benchmarks, which I assume will either fail or be penalized for stopping and asking the user for something rather than just finishing the job. That won't just train the agent to blunder forward in coding, it'll bleed over into a bias towards blundering forward in general.
Moreover I do not know of a single gym-adjacent place where you can pay etc to get ahead in a waiting list. That would be a very weird anti-customer behaviour, imo. The only thing I can imagine if there are some accessibility priority criteria sometimes, but this would also not be legitimate in this case. Maybe in some places in the world (like the US?) this could a thing, though.
Coding agents are tuned towards acting on this "I wonder if..." by providing solutions.
Experienced vibe-coders will know to explicitly state if you don't want side effects.
For example, someone gives me a code review I don't understand, and I want the agent to read the review and provide feedback[, BUT PLEASE DON'T POST THIS FEEDBACK BACK ON GITHUB, JUST WRITE IT HERE FOR ME TO READ PRIVATELY! DON'T EMBARRASS ME IN FRONT OF MY FRIENDS, MOM!]
So...
Asking "I wonder if you can push me ahead in the waiting list" does not imply "I want you to push me ahead of all costs" and the expected behavior requires experience with agents who will act on hypotheticals for reasons that are culturally embedded in our way of talking (justified) and optimized for (solving problems is a measure of productivity).