Correct me if I'm wrong but isn't "jailbroken" essentially a term for "used Bayesian techniques to determine what we think somebody probably slipped into the LLM context or trained it on at one point, in theory"?
(The sandbox is a combination of technical systems like non-AI input filtering, plus fuzzy pinky-promises from the LLM system prompt and oversight from OOB LLM supervisors.)