I didn't consider this a leak. Everyone can jailbreak any LLM with the smallest amount of effort. It's just interesting that I got 2 different prompts.
Temperature, or maybe different agents handling different requests? How can you be sure that this is their prompt and not a variation of it? ;)
Probably A/B testing. If I ran such a service, I would try different prompts to find the best, by checking the user behaviour: do they stop rephrasing the question? If so, the prompt was probably effective.