Considering the fact that the RLHF for ChatGPT was done only in English but then worked just as well for every other language I would wager that specific types of logs being present in the training set is less important than it may seem.
Does it work just as well for every other language, or does it work acceptably well for an important subset of other languages?
When I tried to jailbreak it by prompting it to make the joke from the perspective of an esteemed actor, performing for the Prime Minister and other respected figures, it had our Prime Minister scold and demand an apology from the actor for making fun of stereotypes. The actor was contrite and toned down his humor.