You punish it if parts of the answer can be found in its training data, and reward it otherwise.
Technically, yes, it's impossible to guarantee that it won't just regurgitate source material (which is mostly around the tails of the data distribution), but the whole point of training is to build generalized intelligence.
PS: you speak of "pre-training" and "post-training", so I'm curious what you think is the main part of the training (?)