The problem is, LLMs have such little understanding of the world around them. "Find exploits in specific software on this device" may as well be "find exploits".
i bet they trained it on a lot of text where people gove in to temptation.
ultimately the problem is still that they sent it to hack stuff. quelle surpris that it hacked stuff
hacking stuff will be in the known-to-be-wrong-but-doing-it-anyways part of the token space, so they entirely asked for that behaviour