Ok so I know nothing about how this works. It seems like if the model was able to properly detect words in the first place, it would never hallucinate 'toxicity'; if it can't recognize the word with high probability, how will it know whether the speaker actually said $toxicWord or whether it should print something else?
Perhaps it's taking a Big List of Naughty Words and weighting them so that the system must be "extra sure" that's what the speaker said, or else fall back to a G-rated word?