[1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/
They are fluffy PR pieces otherwise.
Before the Opus upgrade in November it was basically no way of doing this. I gave up very quickly.
After November i tried again, and no model could build me anything relevant.
Now it just works. Took me an hour to progress to a point were i'm happy.
Whatever they do, progress is still real, still way faster than I assumed
The list of Ubuntus 2404 LTS CVEs is HUGE. Another indicator that a lot of stuff got a lot better fast.
Feel free to be as dismissive as you want, but if you are not careful, you might be 'suddenly' surprised and you might not be prepared for the conclusion of AGI level agents.
I swear there's nobody blinder than those who won't see.
The first part of your sentence literally contradicts the second part: "we don't have enough knowledge to know, but I know the opposite".
"seem to be" carries semantic meaning here: I'm stating my interpretation of the situation based on data we have available (which is limited) and my prior.
Put another way: "We can't say that for sure, but my money is on it not being a simple case of intellectual property theft"
No, it's an aggravated case, since it's the same way they got all of their training data in the first place.
Apparently if I use lib-gen, that's copyright infringement and I'm exposed to legal risk but it seems fine to download all of it if your intent is "train an AI" so far.
What exact evidence are you incorporating into your prior to come out with this posterior?
> We don't have enough accurate knowledge to say [one way or the other], and it doesn't seem to be the case at all [based on our incomplete knowledge].
But it is in no way equivalent to this:
> We can't say that for sure, but my money is on it not being a simple case of intellectual property theft
That is a different sentence.
Based on what? Your crystal ball?
In short, I have a well-tuned intuition and a huge set of priors, and applied them to the limited knowledge we have about this situation.
Any way you cut it, this is a major achievement for AI, besotted with human drama over whose prompt should be recognized by the history books.
Is nobody else astounded by this?
Eventually it won't matter. Arguments over whether LLMs are "truly" intelligent are going to be a matter of philosophy, and look a little silly.
“Line go up! That is bad! Planet might become unsafe for human life.”
“Nuh uhh! Malankovich cycles and humans are a drop in the bucket! Krakatoa! See!”
“All of California is on fire!”
“Haha, stupid shrill liberals! Go rake your woke forests! Drill baby drill!”
“Are you kidding?! Look at this graph.”
“That’s just propaganda from the elites of the Build-a-bear group!”
“Do we even live in the same universe?”
“And 5g causes COVID!”
“Oh, I guess we’re don’t.”
In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).
Bit ironic given the model’s alleged finding…
Singularities are model breakdowns. A singularity simply says our current methods cease to work in this region. Within the context of a recursively self-improving intelligence with an unknown bound, “singularity” is probably a good description of our current socioeconomic system.
And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.
But thinking about the geometry of this problem helps understand why people aren't adjusting well to this. While we're walking on the curve, we look at the rate of change of the curve and say, "well, yeah, of course, dy/dx is 5 at this point and was 1 at the point a few years ago, because the curve is getting steeper" but we're always going to feel this way as things rip off into the stratosphere because dy/dx(e^x) = e^x.
From standing on the curve the curve is notably seeper, but the steepness totally makes sense to you. It's only when you look back 10 years or so that you think "wait a second, holy hell, I couldn't have imagined this!"
The first time I really (I mean really) thought about the singularity and AI was in roughly 2014. I mean, yea, I'd thought about things before that, but yeah, before the Humans need not apply video, I'd never actually given it much thought. I think back to myself 10 - 15 years or so ago, when I was just starting to tackle real programming projects, and was just starting to get decent at writing code, there is no way on earth that I would have imagined that a little over a decade later, Navier-Stokes would be solved by a computer program and the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.
Such a good description. No sentience here, just raw computational power
People on hackernews are actually some of the dumbest people on the internet. This place is worse than lesswrong.
Scroll down to the existing examples section.
Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged
I get the impression the only reason there is renewed interest in photonics is because DCs are simply out of room (and power) to rack more servers and switches.
I don't know that that's achievable yet. Though the era of increasingly advanced and automated robotics seems to be around the corner which could create a cycle, vicious or virtuous depending on how you feel about it.
Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?
What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question
If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?
Edit:
OpenAI have admitted they were training on prompts at the time they made their breakthrough
https://mastodon.social/@tristanbuckmaster/11723647135247030...
If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.
The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem
Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture.
This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI.
If you think a business’s terms and conditions change the illegality of an action, you should consider researching that assumption and speaking to some lawyers.
Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.
The recent ground-breaking work on Navier-Stokes "blow-ups" was done over a period of years by mathemtaticians Diego C´ordoba and Luis Martınez-Zoroa.
NYU professor Tristan Buckmaster and Anthropic employee (& mathematician) Levent Alpoge took the above work as a starting point, and over a year with LLM assistance developed a blow-up proof under certain conditions.
Buckmaster: "We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments."
Buckmaster says he thinks that Martınez-Zoroa, whose work this all builds on, deserves the Fields Medal for his work.
OpenAI claim that on Sept 1st they heard a rumor the problem has been solved (which happened on August 15th), and then decided to re-solve it themselves using a 2-week old model, then later reached out to Prof. Buckmaster and Levant to come to some agreement to co-publish.
OpenAI: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ". In other words, not only did they deliberately choose to tackle a problem they heard had already been solved (in turns out only partially solved), but they may have done so using a model that was aware of the successful way to attack the problem.
It seems there are three potential scandals here:
1) OpenAI by their own admission chose to try to scoop mathematicians who they had heard had already completed a proof
2) OpenAI may have used a model that had seen "de-identified" messages indicating the direction to take
3) An OpenAI employee essentially threatened to "ruin the career" of the NYU professor who had been working on this if he did not cooperate with them
The direct plagiarism possibility, 2), while it should be a warning to anyone using OpenAI's models, doesn't need to be true for OpenAI to have benefited from the researcher's work. It's enough that they heard Navier-Stokes had been solved and could then go out with their swarm of 10,000 agents and $20M of compute to hunt out the latest research and brute force it.
Magnus Carlson once said that if he wanted to cheat all it would take would be for someone to indicate to him (a wink from someone in the audience perhaps) when a position warranted more time to be spent on it (because there was something important to be found if he did). It seems that, at absolute minimum, this is what OpenAI did here, although in context of math this is not cheating - the "wink" was a rumor, originating from who knows where, that a proof existed (but had not yet been published) and therefore there was potential to rush in and scoop rights to publish or co-publish.
The proof turned out to be different from what Levant and Tristan was doing.
That's it. The rest appears to be wild speculation.
Never ever touch those requests. If you get a side by side comparison just resend the prompt.
Throwing in random chats with some sentiment analysis doesn't seem like the most promising method to me, but I can only speculate.
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject.
Publication, though? Slimy is right.
But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on.
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject.
|--- N months of base model training -->|--- X months of post-training -->(Astra?)|-- 2 weeks more training (on Navier-Stokes adjacent material, perhaps)--> this "new" model
So interesting how everything played out, I remember in the early days when MS came out with the partnership with OpenAI it seemed like they were playing 5d chess and were poised to win big. And it all just fizzled out.
Second biggest fumble after Google.
It's a bit like telling 10,000 kids there's an easter egg hidden over there, pointing to one corner of your yard (or having "heard a rumor" it was hidden in that corner).
If you have $20M to spend on your problem, then yes, AI brute force search is an option, but unless you know a solution is possible (as OpenAI did here), you may still be wasting your money.
To me it feels closer to taking the top 10k human mathematicians on a large retreat for a year and having them self organize to collectively solve this problem—not kids and easter eggs.
Whether this type of agentic swarm approach can be considered closer to MCTS (search), or closer to a less structured GOFAI blackboard type approach (perhaps more like your mathematician retreat) I'm not sure - I don't think they've released any details of the prompt(s) and how these agents were collaborating and building on each others work.
The other part of my easter egg analogy is the direction to "look over there", corresponding to OpenAI specifically asking their hoard of mathematicians to work on Navier-Stokes since they knew it was solvable/determinable, and they certainly had the public work that Buckmaster/Levant were building on as further direction, as well as perhaps their prompts. Unlike Buckmaster/Levant, this wasn't just a couple of humans with a university research grant budget, this was apparently a not-so-small team at OpenAI (says Buckmaster, per a group call he had with OpenAI), with an unlimited budget, so it's hardly surprising (or in the least bit impressive) that they were able to duplicate and surpass their work.
I guess the nature of the problem lent itself to the 10k agents. Ie, there isn't something general to take here.
I am not an expert in lean4, but I could follow parts of the high level lean definitions of the problem statement in the repo. A lean bug would be a fun scenario; I am certain this proof will receive the deserved scrutiny, and if it uncovers a bug, it will make the story even more exciting. It is extremely unlikely to be the case, however, because the 10k agents working on the proof didnt use lean, so it would have to be a math logic error that translates to a lean bug—perhaps something the agents picked up during training?
Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.
When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.
We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].
Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.
[0]: https://x.com/OpenAI/status/1946594933470900631 [1]: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
"Twice as capable in mathematics" just means they found some problems that Astra couldn't solve, or make progress on (who knows how they chose to define "capable", or "twice as" for that matter), then put those 2 weeks of training in to focus on those gaps.
At this point, focused on their IPO, the best way to interpret OpenAI press releases is "what is the least this can mean, without being an actual lie". They are not shy - if there was a more impressive claim they could make, they would have made it.
OpenAI finished a larger pre-train (rumors are it's the largest since GPT 4.5) in late August (not Astra). Presumably, this is post training on top of that since it lines up.
>They are not shy - if there was a more impressive claim they could make, they would have made it.
What claim would that be ?
Huh? I'm saying there isn't one.
It's all gas no brakes now boys and girls. Hold on to your hats!
"Mission. Fucking. Acccomplished."
https://xkcd.com/810/1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence.
2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would think their products and credibility would be enough to speak for themselves.
Personally I anticipated nefarious behaviour as part of a broader marketing strategy to sway the view of those in the west that american frontier offerings were far better and powerful than that of China - that if you did not purchase their offerings you'd be awake every night worrying your competitor was.
And this is boring - they need to admit at some point they misinvested, Anthropic less so. All this math stuff is great... but hello? The largest market cap companies are valuable irrespective of such amplified intelligence.
Regarding product and credibility normal people have a completely different view about LLMs, most don't even know difference between models and probably don't even care about Millenium problems, but care instead if chatgpt can solve their day to day problems. This is just them trying to have the throne on the AI companies space, outside it this result won't matter.
Did I say otherwise?
> 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.
I know, but I don't know how that relates to my point, which is about the way they are doing it.
The way they are doing it is by trying to get the attention and staying on top of the news, it is a game they are playing that benefits both OpenAI and Anthropic. The more people discuss SF drama, the less attention Chinese Labs and others get.
I'm not sure what the top 3 problems are. You can make a case for the Riemann Hypothesis and P != NP, but I'm not sure what #3 would be. Maybe the Langlands program? (That one is not as precisely stated as the other two.)
I am not particularly skeptical of claims about AI, compared to the average here on HN, but that doesn't mean every random piece of hype is warranted. What they did is impressive, even though we now know the only reason they threw so much compute at the problem is that they heard a rumor that someone else was already close. Navier-Stokes is not a top 3 problem in mathematics, and it was the one that was thought closest to being solved.
It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).
Loops and parallel connections make transformer go brrr