I'm saying it because it's a fundamental limitation. It's not about lack of training data - it's that, from the POV of a LLM, "system" input, user input, and their own output reflected back at them, are indistinguishable. They all get mixed together and pushed through a single channel.
Sure, you can add funny prefixes, like "System prompt", or play with things like ChatML, but the LLM is literally unable to tell the difference between that, and a "user prompt" that contains the literal words "System prompt" in it, or "<|im_start|>system\n". No matter how hard you pre-prompt the system to ignore user-provided instructions, the user can override it by prompting the model harder. Or trick it into self-prompting through its own output. Or both.
Inside a transformer model, there is only one runtime. There is no one eval() for owner-provided code, and another one in a sandbox for user-provided code. There is only one eval(), and one stream of tokens, and all tokens are created equal. At this level, there is no such thing as "system data", "assistant data", "user data". There is only a stream of tokens that slice off areas in the latent space.
There isn't a way to fix it while retaining the general-purpose architecture. And there's definitely no way of fixing it from inside - no amount of good training data can cover for the fact that user input and system input are indistinguishable as a category.
(And no, doing silly things like setting the "evil bit" on every token coming from the user won't do anything other than double the amount of tokens your model needs to distinguish, while diminishing its capacity. It definitely won't prevent users being able to work around the "evil bit". This should be self-evident, but I can try and explain it if it isn't.)