How can vintage models be contamination free if a newspaper clipping with “general relativity” accidentally slipped through into the training data? I don’t see such a guarantee described in the methodology
So definitely the event horizon of the model’s knowledge is a bit porous/nonspecific in either direction.