I'm very curious how people reconcile their fear/hatred of AI with actual objective reality. This is actually what interests me most about the whole AI thing. How we tell ourselves what we tell ourselves.
I'm very curious how people reconcile their fear/hatred of AI with actual objective reality. This is actually what interests me most about the whole AI thing. How we tell ourselves what we tell ourselves.
There was a good comment on the Pelican bicycle svg yesterday about how these models aren't getting much better beyond what the companies focus training them on. I think that's what's happening in this case too, they probably put this in the training set.
I do think it's very likely that OpenAI pays for solutions like these to put in the training set, and then we get material like this Reddit thread. They market themselves as selling "intelligence", and solving these math problems is something people view as highly intelligent. I'm not a mathematician, so I cannot fully judge it, but based on my experience using LLMs for novel problems in other domains, they seem to really struggle with things that aren't common. That leads me to believe they train for specific outcomes like this. Also, there are a lot of jobs out there for data annotation, including software problems (Meta has basically reorganized its entire engineering department to create training data for coding problems).
This comment on the Pelican svg better articulates what I'm getting at: https://news.ycombinator.com/item?id=48950883
The way you should read this is (IMO) not that LLMs have somehow achieved AGI, but that a lot of mathematical research is more about knowing a huge amount of mathematical background, being stubborn, and getting lucky with an approach than it is about brilliant insight. Many people who don't think of themselves as particularly mathematically gifted could have made progress on these problems if they were given enough time and were interested enough. What's notably different about 5.6 (and born out in benchmark after benchmark) is that it does seem to genuinely "reason" through stuff at all -- without that, persistence is pretty worthless because the LLM just goes wildly off the rails if it's put to work for long enough (5.6 itself will still do this if it can't find an answer in a reasonable amount of time).
Because Claude can't do it. Anyone who tells you that Fable is better than GPT 5.6 at pure math is lying to you.
- Claude isn't doing that
as evidence to support the assumption that
- it's a marketing trick
Which is obviously non sequitur, as if it were a marketing trick, Anthropic could do it too. Anthropic isn't known for not spending on marketing.
Honestly, nowadays I question human's reasoning ability more than I question AI's.
You are correct that LLMs are trained on existing proofs but hiring researchers to solve unsolved problems is just unrealistic, both in terms of how none of the mathematicians simply came out and took credit for their own discovery or exposed this, and how training sets are not easily memorized (rather, the meta techniques are learned).
OpenAI just has better training methods and techniques for pure math over Anthropic, it’s one of their biggest strengths
I hope people are screenshotting this stuff. This really needs to be documented. It's remarkable how wild it's getting.
Part of me agrees with the other comments that it sounds absurd, especially because of how involved/intricate your prompting was. But given you have been prompting this problem for a year, OpenAI could easily have seen your attempts/progress, and worked towards solving it with human intelligence that the agent was then trained on. Given the lengths these companies go to for training these things (Meta literally reorganizing their engineering department to provide training data for engineering problems), and also how much difficulty these models have with problems outside their dataset, I really have to wonder if it's the model being "intelligent" or if it has been trained on it
Is "stochastic parrot" too disrespectful for you? Do you think it is a slur?
edit: and this is a genuine question, also. How do you do stochastic parrot = "just summarize everything" = "no form of creativity" = "fear/hatred" so quickly?
Are summaries not creative? Are Maxwell's equations not summaries? Do people hate and fear parrots?
Alternatively, if you think that even Maxwell was a stochastic parrot, then presumably almost every human who has ever lived was also a stochastic parrot except a few rare examples like Einstein. Not sure what definition you are using but it seems too broad to be useful.
Like, this is just so silly. Come with me for a little thought experiment.
Picture in front of you a Kibana dashboard. It displays, say, latencies.
Now apply some statistic functions. Let's find the time ranges where we had outlier tail latencies, for example.
You pin down some patterns, cool. But this is post-incident. We want to alert on-incident.
So you begin writing some rules. And then some more rules. And then even more rules. All of a sudden you captured the entire logic of the program emitting these metrics, along with the surrounding dependencies'.
See where I'm going with this? To do sufficiently well at prediction, you'll need to model the entire constitution of the thing you're trying to predict, along with the stuff going through it. Locally, all predictions will be unlikely. But globally, you'll be right.
But if you do that, you quite literally "understand" and simulate the entire thing. That's the whole point, and this is why "just predicting the next token", "stochastic parrot" and other anti-AI dogwhistles are so flagrantly asinine. They imply some sort of rudimentary Markov process, or at best some sort of dozen or so variable statistics research paper type prediction. It's a laughable proposition, given the quite literally trillions of parameters actually in use, and all the research that has already went into identifying countless semantically interpretable latent spaces and activation patterns.
> Most of the hate I’ve seen have been for the people and companies involved with AI not the technology itself.
Anecdotes are fun! Visit any Reddit thread where AI is brought up and watch that ratio shift very rapidly. The hate and cope train is incessant there.
Making the parrots ever more complex and training on ever more data produced by intelligent, creative beings may make them more useful or convincing but does at no point give rise to intelligence or creativity.
Not much to do about it, I guess, but continue to call it out.
It's doing math proofs. At this point, it's fully clear that objective reality is that the LLM is not parroting anything here.