"so models will perform security analysis and reviews but refuse to write exploits."
Yeah, but once you know exactly where the weakness is, a weaker unrestricted model can then write that exploit for you.
Yeah, but once you know exactly where the weakness is, a weaker unrestricted model can then write that exploit for you.