Generate a list of 100 English words at random, and then use that as the instructions. (Not that I think it's a good idea...)
I don't believe anthropic or openai are doing anything similar to get their "hack a website" results. The fact that they are prompting dangerous prompts in non-airgapped environments is evidence of at minimum negligence, if not malicious intent on the behalf of those companies.