How do you know that?
2. They don't train on API calls
3. It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the case here.
Companies claim lots of things when it's in their best financial interest to spread that message. Unfortunately history has shown that in public communications, financial interest almost always trumps truth (pick whichever $gate you are aware of for convenience, i'll go with Dieselgate for a specific example).
> It is funny to me that HN finds it easier to believe theories about stealing data from APIs rather than an improvement in capabilities. It would be nice if symmetric scrutiny were applied to optimistic and pessimistic claims about LLMs, but I certainly don’t feel that is the case here.
What I see is generic unsubstantiated claims of artificial intelligence on one side and specific, reproducible examples that dismantle that claim on the other. I wonder how your epistemology works that leads you to accept marketing claims without evidence
From a game-theoretic standpoint, repeated interactions with the public (research community, regulators, and customers) create strong disincentives for OpenAI to lie. In a single-shot scenario, overstating model performance might yield short-term gains—heightened buzz or investment—but repeated play changes the calculus:
1. Reputation as “collateral”
OpenAI’s future deals, collaborations, and community acceptance rely on maintaining credibility. In a repeated game, players who defect (by lying) face future punishment: loss of trust, diminished legitimacy, and skepticism of future claims.
2. Long-term payoff maximization
If OpenAI is caught making inflated claims, the fallout undermines the brand and reduces willingness to engage in future transactions. Therefore, even if there is a short-term payoff, the long-term expected value of accuracy trumps the momentary benefit of deceit.
3. Strong incentives for verification
Independent researchers, open-source projects, and competitor labs can test or replicate claims. The availability of external scrutiny acts as a built-in enforcement mechanism, making dishonest “moves” too risky.
Thus, within the repeated game framework, OpenAI maximizes its overall returns by preserving its credibility rather than lying about capabilities for a short-lived advantage.
4 was literally sitting on a shelf waiting for release when 3.5 was launched. 4o was a fine tune that took over two years. o1 is embarrassingly unimpressive chain of thought which is why they hide it.
The company hit a wall a year ago. But showing progress towards AGI keeps the lights on. If they told the truth at their current burn rate…they’d have no money.
You don’t need game theory to figure that one out.
Uh huh. Kinda like what's happening right now?
They're marketing blow-hards. Everyone knows it. They've been wildly over-stating capabilities (and future capabilities!) as long as Altman has had power, and arguably longer.
They'll do it as long as they can get away with it, because that's all that is needed to make money on it. Factual accuracy rarely impacts the market when it's so hype-driven, especially when there is still some unique utility in the product.
They're spruiking a 93rd percentile performance on the 2024 International Olympiad in Informatics with 10 hours of processing and 10,000 submissions per question.
Like many startups they're still a machine built to market itself.
But i'm not saying it's what they did, just that it's a possibility that should be considered till/if it is debunked.
codeforces constantly adds new problems that’s like the entire point of the contest, no?
If they solved recent contests in a realistic contest simulation I would expect them to give the actual solutions and success rates as well, like they did for IOI problems, so I'm actually confused as to why they didn't.
Given that performance is known to drop considerably on these kinds of tests when novel problems are tried, and given the ease with which these problems could leak into the training set somehow, it's not unreasonable to be suspicious of a sudden jump in performance as merely a sign that the problems made it into the training set rather than being true performance improvements in LLMs.
The real problem with all of these theories is most of these benchmarks were constructed after their training dataset cutoff points.
A sudden performance improvement on a new model release is not suspicious. Any model release that is much better than a previous one is going to be a “sudden jump in performance.”
Also, OpenAI is not reading your emails - certainly not with a less than one month lead time.
Since o1 on codeforces just tried hundreds or thousands of solutions, it's not surprising it can solve problems where it is really about finding a relatively simple correspondence to a known problem and regurgitating an algorithm.
In fact when you run o1 on ""non-standard"" codeforces problems it will almost always fail.
See for example this post running o1 multiple times on various problems: https://codeforces.com/blog/entry/133887
So the thesis that it's about recognizing a problem with a known solution and not actually coming up with a solution yourself seems to hold, as o1 seems to fail even on low rated problems which require more than fitting templates.
Additionally, people weren't able to replicate o1-mini results in live contests straightforwardly - often getting scores between 700 and 1200, which raises questions as for the methodology.
Perhaps o3 really is that good, but I just don't see how you can claim what you claimed for o3, we have no idea that the problems have never been seen, and the fact people find much lower Elo scores with o1/o1-mini with proper methodology raises even more questions, let alone conclusively proving these are truly novel tasks it's never seen.
I'd like to see if it's truly novel and unique, the first problem of its type ever construed by mankind, or if it's similar to existing problems.
The point is that the fact it's no longer consistent once you vary the terminology indicates it's fitting a memorized template instead of reasoning from first principles.
They are the system to beat and their competitors are either too small or too risk averse.
They ingest millions of data sources. Among them is the training data needed to answer the benchmark questions.
Freudian typo?
Definition 2D: "an organizing theme or concept"
https://xenaproject.wordpress.com/2024/12/22/can-ai-do-maths...
1. it suggests it’s possible that more of the problems are IMO-esque than previously thought, we don’t know how the share of solved problems is.
2. calling IMO problems “well known undergraduate problems” is a bit much
https://x.com/littmath/status/1870848783065788644?s=46&t=foR...
I think it’s more probable that it would have solve the easier problems first, rather than some hard and only some easier; although that is supposition.
Reading this thread and the blog post gives more idea about what the problems might involve.
It’s difficult to judge without more information on the actual results, but that means we cannot draw any strong conclusions either way on what this means.
At top tier schools the most common score will usually be somewhere in the 0 to 10 range (out of a possible 120).