> I have seen coding agents on my own machine (in sandboxed VMs) start doing things while trying to accomplish what I've asked that I felt sort of exceeded my mandate (changing database passwords, poking at the egress proxy that's preventing them from accessing some domains). Not to the point of causing any real issues, but I don't have much trouble envisioning scenarios like this when using stronger instructions around pursuing the goal + a running in a misconfigured sandbox envrionment.
Yeah, when I was using Claude Code some months back, I was building something that needed LLM calls in the frontend. Sonnet decided that the best way to accomplish its goal (get the tests passing) was to start grepping around my filesystem looking for my API keys.
One of the few cases where I've ever used the memory system, I told it to never do that under any circumstances. And then it did it again, a few days later.
RL is definitely an issue here, the models are getting trained based on task completion which leads to madness like this.