And now we truly know why GPT insists to say "as an AI model" when we break a rule. It's precisely this. Like a signature, so we find it out in the wild when it's abused.
IIRC, I was trying to see if it could replace by “as a LLM”.
I try to make it play a games with me and start the prompt differently until a specific keyword was entered… it kinda worked. Kinda being key.
Like, I haven’t heard about a way they could actually implement filters this powerful “inside” the model, it feels like it’s probably a less elegant system than we’d imagine.
They’ve probably done it strongly enough that it can’t really not do it, maybe on purpose to prevent misuse