The censorship seemed to be based on keywords, applied the input prompt and the output text. If asked about events in 1990, then asked about events in the previous year DeepSeek would start generating tokens about events in 1989. Eventually it would hit the word "Tiananmen", at which point it would partially print the word, then in response to a trigger delete all the tokens generated to date and replace them with a message to the effect of "I'm a nice AI and don't talk about such things."
If the word Tiananmen was in the prompt, the "I'm a nice AI" message would immediately appear, with no tokens generated.
If Tiananmen was misspelled in the prompt, the prompt would be accepted. DeepSeek would spot the spelling mistake early in its reasoning and start generating tokens until it actually got around to printing to the word Tiananmen, at which point it would delete everything and print the "nice AI" message.
I'm no expert on these things, but it looked like the censorship isn't baked into the model but is an external bolt on. Does this gel with other's observations? What's the take of someone who knows more and has dived into the source code?
Edit: Consensus seems to be that this instance was not being run locally.