Those are well known problems, that people talk about on different contexts. They would have to review their entire training set.
> Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly
What exactly is the source of your belief that the Putnam would not be in the test data? Didn’t they train on everything they could get their hands on?
So it's like 99.9999999% wrong to assume something public isn't on the train set, such as Putnam problems in this case. This is about it.
this whole notion of putnam as test being trained on is a fully invented grievance
read the entire thread in this context
I agree that having putnam problems on OpenAI training set is not a smoking gun, however it's (almost) certain they are on training set, and having them would affect performance of the model on them too. Hence research like this is important, since it shows that observed behavior of the models is memoization to large extent, and not necessarily generalization we would like it to be.
OAI uses datasets like frontiermath or arc-agi that are actually held out to evaluate generalization.
Every decent AI lab does this, else the benchmark result couldn't be trusted. OpenAI publishes results of ~20 benchmarks[2] and it is safe to assume they have made reasonable attempt to remove it from training set
https://kskedlaya.org/putnam-archive/
I would expect all llms to be trained on it.
The short version is that llm trainign data is the lowest quality data you are likely to see unless you engage in massive potential copyright infringement.