For one, the prompt involves the model simulating its own output, which clearly has a flavor of Universal Turing Machine to it.
Then the token smuggling technique leans on the ability of the model to statically simulate the execution of code. Therefore a perfect automated filter that relies on analyzing code in prompts would be impossible. (However the filter only needs to be better than the LLM in practice)
I wouldn't be surprised at all if you could make some sort of formalized argument proving that it would be impossible to prevent all jailbreaks.
But the human programmed guard rails act this way since the more powerful human LLM can figure it out. So for now we will still need humans!
I don’t think anyone has put together the halting problem for LLMs directly yet though. You could imagine a halt token but any simulated LLM should be less powerful. Interesting thought experiment. Can chatgpt create an algorithm to solve the digits of pi and execute it? Might try this.
Google has a paper about DNN architectures and the Chomsky hierarchy for generalizing to distribution shifts. This is interesting in that specific architectures should limit what a transformer LLM can do.
I imagine this is an active research area.
If someone told you "i can guarantee Fred Smith here will never, ever say anything inappropriate. He's not capable of it." (Fred being a regular old human.) You'd say "Well, no, you can't guarantee that. You may have given Fred all the best training in the world. You may have selected Fred from 10,000 other candidates as the least likely to ever say anything inappropriate. Fred may have strict instructions not to. But he still could."
It may be the same with LLMs.