Is it? OpenAI released a gpt oss safeguard. You give it a policy it gives you a Rating
Messages comes in rate it and reject with hitting the model. Then you don’t need to fill the prompt with “please don’t do this”
Messages comes in rate it and reject with hitting the model. Then you don’t need to fill the prompt with “please don’t do this”