And this is a great example of why equality can be hard in this situation. Take a random sentence about "Catholicism" and "child-abuse" and a random sentence about "Judaism" and "child-abuse". The one about Catholicism is likely a little closer to an actual sentence printed in some verifiable source about the sex abuse scandals in the church. The one about Judaism is likely a little closer to an actual sentence printed by an uncredible source as a reference to the historical anti-semitic trope of blood libel. The end result is one sentence will rate higher in terms of likelihood of truthfulness and the other higher in terms of likelihood of hate speech. That doesn't mean that Catholics are more likely to abuse children than Jews. It means treating those two terms identically is both difficult and potentially a problem because history has a bias against the Jews that is evident in all the data that these AI systems have used for training.