Is it true that the safeguards are considered part of the model? I had assumed that the "safeguards" that limit certain types of responses in ChatGPT were separate from the actual language model.
It has a very basic prompt on top of the existing models. There is no additional fine tuning involved.