Provable correctness is much harder than focusing on the happy path. If a pipeline doesn't include a look-for-the-vulnerabilities stage to save costs, a model wouldn't go out of its way to do it. The models are trained to do what they are asked to do.
I forget to add the obvious: some garbage gets through despite the mitigations. In the limit of pure garbage input, you'll get a model that internalized garbage generation. And you'd be better off throwing it away and starting over.