"Solve (1503+5171)*(9494-4823)" reliably gets the correct answer from ChatGPT
"Write a poem about the solution to (1503+5171)*(9494-4823)" hallucinates an incorrect answer though
That suggests to me that they've papered over the models inability to do basic math, but it's a hack that doesn't generalize beyond the simplest cases.
i’m able to reproduce your failure on 4o
1. The part of the network that does complex math and the part that write poetry are overlapping in strange ways.
2. Most of the models nowadays are assumed to be some mixture of experts. So it's possible that saying write the answer as a poem activates a different part of the model.
The poem thing probably causes them to not decide to use those tools.
The short version is that llm trainign data is the lowest quality data you are likely to see unless you engage in massive potential copyright infringement.
> Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly
What exactly is the source of your belief that the Putnam would not be in the test data? Didn’t they train on everything they could get their hands on?
So it's like 99.9999999% wrong to assume something public isn't on the train set, such as Putnam problems in this case. This is about it.
this whole notion of putnam as test being trained on is a fully invented grievance
read the entire thread in this context
I agree that having putnam problems on OpenAI training set is not a smoking gun, however it's (almost) certain they are on training set, and having them would affect performance of the model on them too. Hence research like this is important, since it shows that observed behavior of the models is memoization to large extent, and not necessarily generalization we would like it to be.
OAI uses datasets like frontiermath or arc-agi that are actually held out to evaluate generalization.
Every decent AI lab does this, else the benchmark result couldn't be trusted. OpenAI publishes results of ~20 benchmarks[2] and it is safe to assume they have made reasonable attempt to remove it from training set
https://kskedlaya.org/putnam-archive/
I would expect all llms to be trained on it.
Those are well known problems, that people talk about on different contexts. They would have to review their entire training set.
Basically if something appeared online or was transmitted over the wire should no longer be eligible to evaluate on. D. Sculley had a great talk at NeurIPS 2024 (same conference this paper was in) titled Empirical Rigor at Scale – or, How Not to Fool Yourself
Basically no one knows how to properly evaluate LLMs.
(There was also a much earlier piece of software that would generate semi-intelligible Kant or Hegel one sentence at a time, though that was through a series of a priori generation rules and a large at the time dictionary of stock phrases. I wonder what ever happened to that.)
The problem is just that people keep insisting that those things are intelligent.