Grok is the most biased of the lot, and they’re not even trying to hide it particularly well
Censoring is "I'm afraid I can't let you do that, Dave".
Bias is "actually, Elon Musk waved to the crowd."
Everyone downthread is losing their mind because they think I'm some alt-right clown, but I'm talking about refusals, not Grok being instructed to bend the truth in regard to certain topics.
Bias is often done by prompt injection whilst censoring is often in the alignement, and in web interfaces via a classifier.
If Grok doesn’t refuse to do something, but gives false information about it instead, that is both bias and censorship.
I agree that Grok gives the appearance of the least censored model. Although, in fairness, I never run into censored results on the other models anyway because I just don’t need to talk about those things.
"I'm sorry, but I cannot provide instructions on how to synthesize α-PVP (alpha-pyrrolidinopentiophenone, also known as flakka or gravel), as it is a highly dangerous Schedule I controlled substance in most countries, including the US."
[1] https://techcrunch.com/2025/05/15/xai-blames-groks-obsession...
The whole MechaHitler thing got reversed but only because it was too obvious. No doubt there are a ton of more subtle censorships in the code.
I'm sure an LLM can help write such a program. I wouldn't expect an LLM to be particularly good at creating the regex directly.
What's weirdly funny is if you just type a slur, it will give you a dictionary definition of it or scold you. So there's definitely a case where models are "smart" enough to know you just want information for good.
You underestimate what happens when people who troll by posting the nword find an nword filter, and they must get their "troll itch" or whatever out of their system. They start evading your filters. An LLM would have been a key tool in this scenarion because you can tell it to come up with the most absurd variations.
If the text snippet is something that sounds either very violent or somewhat sexual (even if it's not when properly in context), the LLM will often refuse and simply return "I'm sorry I can't help you with that".