1,390 karma · joined March 13, 2025
OpenAI should be allowed to produce whatever it wants but it just can't claim that it has actually solved without the due process like peer review. If for example OpenAI solves a new conjecture, OpenAI should be free to publish it in their blog or arxiv in whatever way they desire. It can be slop, it can be non-slop. No one should police it.
Mathematicians are free to use it or discard it. They shouldn't externalise their concerns and restrict labs.
Mathematics is seen today as the noblest and most aristocratic of professions. Turns out, AI disrupts it because access to capital/compute now decides the results. Mathematicians don't like this corruption - understandable.
Its like guild of accountants opposing the calculator and require a responsible release. haha
Sure and there are regulations on who can own guns and licenses. And that's literally how Mythos is handed out today - to trusted partners. What is your opinion about that?
Also, companies are not allowed to release stuff like Nuclear technology, chemical weapons etc. These literally can't be regulated and a single adversary can use it destructively.
"but but they are just tools ok" doesn't work when the capability is big enough.
Look, it might be easy sitting behind a computer asking OpenAI to release whatever new model because you personally don't think you will be affected. But the society needs to be careful on what it hands potential adversaries.
Who said anything about hacking? You can use it, but it is very restricted. But a much more capable model must have much higher restrictions on alignment. Obviously you wouldn't need to spend too much on alignment on a 3B gemma model. This must be obvious, I'm not sure why you bring this up?
> They didn't need to do that - they can just release when ready - nobody is expecting them to release a new model on a certain date.
Again a conspiracy theory? People expected a model during Dev Day. And they are just setting expectations. Whats with you guys and conspiracy theories?
Why would you expect something else?
No, I don't think a company can get away from releasing a powerful tool that can cause disruption if used poorly. This is just not how society works.
The reality is that labs will get blamed for things people do using it. And they have - look at any news about US Army using Claude etc
This understanding comes from the fact that people think either models are fundamentally unsafe or they are completely safe. The reality is that there are different levels of alignment possible. So it is really a tradeoff between effort/time taken to align the model and speed of release.
You are replying to a post where they are doing exactly that. But you analysed it as if they are secretly in a cash crunch and trying to keep the lights on.
You and I both agree that delaying release is the rational response.
What's the necessity for the conspiracy theory, when the simple explanation is that they are delaying the release for safety reasons?
Hmm,
> Remember the relentless “AI is beating everyone in Math Olympiads” and “it’s all over for humans, no point in learning STEM? Well…
> The 2025 Math Olympiad problems are out and naturally people tested the leading LLMs on those… and… spoiler .. they all fail, badly.
>That’s it LLMs are all about storage in training data which enables recombinatorial retrieval. Don’t expect them to solve or respond to new things like a human.
> They fail with novel and unknown tasks because their ability to generalize is vastly overstated by the people invested in the technology.
> The value of LLMs is in their data. No data, no answer. You have a novel piece of work that’s never been done before? OpenAI will insist it’s worth nothing because it’s a spec in the Ocean, yet that’s not true at all.
https://www.linkedin.com/posts/georgzoeller_proof-or-bluff-e...
Today you can solve IMO problems with ~$5.
Its up to you to verify that he has dismissed the technology and that he has been wrong about its performance in mathematics.
Edit:
Even without lean I can solve IMO problems using Sol/Fable today.
You made a categorical statement about LLMs
> That’s it LLMs are all about storage in training data which enables recombinatorial retrieval. Don’t expect them to solve or respond to new things like a human
> They fail with novel and unknown tasks because their ability to generalize is vastly overstated by the people invested in the technology.
It was not contextual but categorical. And it is false today.
So what’s the trick in this honestly? OpenAI shouldn’t care about alignment and just release it? I’m genuinely asking.
“It’s all hype”
“Government bailouts”
“Labs are fundamentally unprofitable”
And the same people who say all this will again claim OpenAI should be held responsible and arrested for when agents hack other systems.
It’s a fundamentally unthoughtout analysis. If you believe that OpenAI should be more careful and prevent future hugging face incidents, surely their decision to delay models is rational? Does it really require tinfoil analysis?
What would you say when their models cause new incidents like hugging face one?
Some tasks are reasoning shaped by nature and you can't just throw a big model at it.
For truthers AI is simultaneously a machine that hallucinates 100% of the time but also so powerful that it can cause nuclear levels of destruction. Because something something “wrong hands”.
How were the sandboxes poor?
I... I can't believe people still think this way. A whole class of security measures are no longer useful and this guy thinks it is "pedestrian".
This kind of isolation - artifactory having access to only a specific IPs - is what is used in most places, including my university and at my own job.
LLMs basically rendered this kind of isolation almost useless by bringing in the ability to chain together never before found vulnerabilities to escape the containment. This was not possible before.
How is this pedestrian??! The author might say "oh well they should have airgapped actually" in a smug way. But you can only know this retrospectively. Even then, isn't it noteworthy and a meaningfully different class of cyber attack?
https://openai.com/index/gpt-5-1/
It says literally the thing you wanted from system 2. Its almost exactly that.
This is what you said btw:
"it's that the system itself decides how to reason based on the nature of the problem it faces"