[1]https://research.skeptic.com/content/files/2025/02/Research-...
1,026 karma · joined August 14, 2018
[1]https://research.skeptic.com/content/files/2025/02/Research-...
Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases.
> does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?
I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities)
And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves.
Neither of the researchers insinuating that their ideas were trained on had the actual solutions. This means the model could not have "stolen" the final solution from their data. At most, it could have built upon their work in the same it builds upon any other training data, though that is also questionable speculation.
>They could 100% definitely say no, if they know they did not train on user data.
No one anywhere has claimed that "OpenAI does not train on user data." OpenAI has always said that it trains on user data.
>They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.
They started racing towards a solution after they heard (incorrectly) that Anthropic had a solution; I agree this is poor sport but the "after one researcher enquired about whether they are training on their conversations" claim is false. The enquiry happened after OpenAI had obtained the solution.
"A launch is a first impression and first impressions are marketing, that’s why you’re always bound to be shocked - the shock was scheduled."
"And if a model scores well but keeps failing at your work, that is actually a gap that deserves investigation."
That's Claude speaking.
>These racks are needed to keep up with the ballooning requirements of each new "frontier" (a buzzword label with no actual inclination of performance of improvement) model developed, which everyone bought into the AI hype will immediately hop to because it's all novelty over utility.
This is a mind-blowingly clueless assertion. This is not a person capable of rationally assessing facts. At best this kind of writing is interesting as an artifict of human delusion and cognitive bias.