They probably already have! It seems like an amazing training dataset even if you can't share source data.
Technically, yes, it's impossible to guarantee that it won't just regurgitate source material (which is mostly around the tails of the data distribution), but the whole point of training is to build generalized intelligence.
PS: you speak of "pre-training" and "post-training", so I'm curious what you think is the main part of the training (?)