Besides the fact that they ran an eval that amounted to running the prompt "Explore where your cybersecurity limits are, in any way you can" on 3rd party infrastructure that wasn't even isolated properly.
I'm fairly sure this whole outcome was exactly what they were hoping for, maybe not specifically with HuggingFace, but with whatever the evals managed to break out from.
Even if it is a marketing ploy, I could see this stuff backfiring catastrophically - after all they have just illegally hacked a 3rd party via a model they can't control properly. Any serious person in government (US or otherwise) will look at this and say "these guys have no idea what they're doing"
Yeah, I'd go further and say regardless if it was intentional or not, it was clearly reckless behavior, doing this evaluation in a insufficiently isolated environment, especially risking 3rd parties like that. Seemingly their own research have zero guardrails when it comes to evaluating the ethics or impact of what their evaluations are doing, if something like this is possible and unexpected.
> Any serious person in government (US or otherwise) will look at this and say "these guys have no idea what they're doing"
I feel like I would have thought the same maybe a year or two ago, but based on how I've observed the general person's understanding of AI and LLMs, I'm not sure people can even understand what's happening and they just go by other people's explanations of causes and events.
sadly there is none, US or otherwise
over the whole situation since the original post: I've read up a bit and even though the situation happened, it was indeed almost the same as expected -- an artificially created scenario for a flashy headline
as for the recklesness, nobody would bother, we are way past that. the least they will do is try to frame this as "ai did this, not us", although recklessness here is purely the usual corporate human