Which if you read the article is in fact a different "machine learning model from the GPT family."
So while the LM is not supposed to output "truth", the content moderation system should correctly classify "hate" because that is its training objective