Show me any system without an initial prompt or goal set.
I don't believe anthropic or openai are doing anything similar to get their "hack a website" results. The fact that they are prompting dangerous prompts in non-airgapped environments is evidence of at minimum negligence, if not malicious intent on the behalf of those companies.