a) censored at output by a separate process
b) explicitly trained to not output "sensitive" content
c) implicitly trained to not output "sensitive" content by the fact that it uses censored content, and/or content that references censoring in training, or selectively chooses training content
I would assume most models are a combination. As others have pointed out, it seems you get different results with local models implying that (a) is a factor for hosted models.
The thing is, censoring by hosts is always going to be a thing. OpenAI already do this, because someone lodges a legal complaint, and they decide the easiest thing to do is just censor output, and honestly I don't have a problem with it, especially when the model is open (source/weight) and users can run it themselves.
More interesting I think is whether trained censoring is implicit or explicit. I'd bet there's a lot more uncensored training material in some languages than in others. It might be quite hard to not implicitly train a model to censor itself. Maybe that's not even a problem, humans already censor themselves in that we decide not to say things that we think could be upsetting or cause problems in some circumstances.