If one of the researchers forgot to check "don't train on this" a single time, the pretrain is going to know what they were working on. The question of whether they remembered to check "don't train on this" every single time is super concrete, and embarrassing to _both_ sides if the answer is "no, he didn't read the eula."
These models will absolutely remember a brilliant insight that appeared a single time in the pretrain corpus, because if they couldn't-- they would get a slightly worse loss. I don't know what this wishy-washing "well maybe we trained on it but we didn't read it" is supposed to mean.
(Example of Fable knowing the content of a deeply unimportant LessWrong post I wrote: https://claude.ai/share/a907b46c-bf7b-4fca-9c71-8582cf8507cc a working Navier Stokes solution would be way more salient)