These LLM agents are just massively complicated optimizers thrown at fuzzily defined problem spaces, with fuzzier constraints.
The people using the model set up the landscape it explores and turned it loose to do real things. It just found an allowed basin in the model that they weren't aware of and started blindly grinding towards an optimal answer.