I think you misunderstood the role of the company named irregular. They were conducting the tests on behalf of open ai and basically left internet access open in their sandbox environment. Those tests included the hugging face and other attacks you list. Irregular wasnt another example of a hack they ran the tests that resulted in them
> Those tests included the hugging face and other attacks you list.
Do you have a source for that? Cause OpenAI themselves stated that the Hugging Face hack was fully internal and separate from the Irregular incidents.
I’ll look. I saw it here a few days ago and could be mistaken with respect to hugging face in particular but the broader point that this was in large part a glaring maybe intentional hole not sophisticated sandbox escape holds I think
I have local model without safeguards and they are not going to hack shit unless you tell them to.
As per usual its the same grift all over. If the llm is instruction is to do whatever it needs including hacking to achieve its goal it will do so. Ofc they will never disclose that.