> Once a model is open-weight, safeguards that do exist can be removed
Safeguards trained into the model (ie exist in the weights) can’t be removed.
Safeguards trained into the model (ie exist in the weights) can’t be removed.
There's a subreddit for people wanting to sex-talk to various models. It just so happens that the same prompt they use to 'jailbreak' SOTA models for sex talks also works if you want to have model write malware, or tell you how to design a highly illegal device.