> Cherry on top: That was part of an attempt to jail-break itself via self-prompt injection.
We continue to see so-called "prompt injection" "attacks" in the wild that override a user's intended program with an attacker's [0], and/or the LLM producer's intended "safety" instructions with the user's. The fact that this sort of program hijacking is possible at all is strong evidence of negligence. Why?
OpenAI and Anthropic both claim that they're working on very dangerous Internet-connected tools. So very dangerous that the production of and access to said tools needs to be tightly regulated, they claim. If one actually believes that the computerized tool one is working on is very dangerous, one generally doesn't design that tool so that it blindly executes instructions handed to it by complete strangers on the Internet. That's akin to connecting the sole activation switch for a biosphere-evaporating firebomb to the Internet.
The major LLM producers are so obviously negligent and -as a bonus- have openly admitted to committing cybercrimes [1] that would get people like you and me fined out the ass and jailed for ages if we did them. The tragedy is that they're making so much money for the rich and powerful that -much like the architects of the 2008 housing crash- they'll never see any meaningful punishments for their actions.
[0] One recent example is <https://agentic.tracebit.com/context-bombs/>, but there are so, so many more to choose from.
[1] ...the "cyber" prefix is so stupid...