I’m sure the lawyers will eventually figure out a way to train an LLM on them.
Technically, yes, it's impossible to guarantee that it won't just regurgitate source material (which is mostly around the tails of the data distribution), but the whole point of training is to build generalized intelligence.
PS: you speak of "pre-training" and "post-training", so I'm curious what you think is the main part of the training (?)