We've already seen it start and they had to actively work against it: https://openai.com/index/where-the-goblins-came-from/
> We unknowingly gave particularly high rewards for metaphors with creatures
It was the human feedback that caused the bias, not a change in training data.