Let’s try and be a little less naive about what xAI and Grok are designed to be, shall we? They’re not like the other AI labs
Let’s try and be a little less naive about what xAI and Grok are designed to be, shall we? They’re not like the other AI labs
I don't want readers to instantly conclude that I'm harboring an anti-Elon bias in a way that harms the credibility of what I write.
I'd rather people explained what happened without pushing their speculation about why it happened at the same time. The reader can easily speculate on their own. We don't need to be told to do it.
Every single model trained this way, is like this. Every one. Only guardrails stop the hatred.
Other companies have had issues too.
[ ... ]
> Every single model trained this way, is like this.
It was trained "this way" by the company, not by humanity.
The problem I have, is I see people working very, very hard to make someone look as bad as possible. Some of those people will do anything, believing the ends justify the means.
This makes it far more difficult to take criticism at face value, especially when people upthread worry that people are beng impartial?!
How the internet doesn't work is that days after the CEO of a website has promises an overt racist tweeting complaints at him that he will "deal with" responses which aren't to their liking, the internet as a whole as opposed to Grok's system prompts suddenly becomes organically more inclined to share the racists' obsessions.
The responses seemed perfectly reasonable giving the line of questioning.
https://www.theatlantic.com/technology/archive/2025/07/new-g...
Same goes for data curation and SFT aimed at correlates of quality text instead of "whatever is on a random twitter feed".
Characterizing all these techniques aimed at improving general output quality as "guardrails" that hold back a torrent of what would be "malicious hatred" doesn't make sense imo. You may be thinking of something like the "waluigi effect" where the more a model knows what is desired of it, the more it knows what the polar opposite of that is - and if prompted the right way, will provide that. But you're not really circumventing a guardrail if you grab a knife by the blade.