HNHacker News
TopNewBestAskShowJobs

20k

1,904 karma · joined September 24, 2014

submissionscomments
20k··on On the Navier–Stokes Millennium Prize Problem
OpenAI trained on their private unpublished research notes effectively, while also trying to get one of the paper authors fired
20k··on Navier-Stokes – Tristan Buckmaster [pdf]
That doesn't make it fine. We should not excuse this behaviour just because its rampant already, especially when it comes to such a serious prize
20k··on On the Navier–Stokes Millennium Prize Problem
That does not make it ethical
20k··on On the Navier–Stokes Millennium Prize Problem
This is textbook plagiarism, scooping their result knowing that the research was part of the training data
20k··on On the Navier–Stokes Millennium Prize Problem
The biggest issue we aren't talking about is, of course, that those two researchers were not the only two using ChatGPT to work on the problem at the time
20k··on On the Navier–Stokes Millennium Prize Problem
I just want to add to this another update by the author as well:

https://mastodon.social/@tristanbuckmaster/11723647135247030...

Which seems to be very directly accusing OpenAI of plagiarism

20k··on On the Navier–Stokes Millennium Prize Problem
I mean, its years worth of hard work by multiple researchers it would seem, which OpenAI simply lifted and claimed as its own. These researchers weren't just letting OpenAI burn tokens while sipping martinis on a beach
20k··on On the Navier–Stokes Millennium Prize Problem
OpenAI's T&Cs let them steal your children I'd suspect, that doesn't make it morally correct
20k··on On the Navier–Stokes Millennium Prize Problem
The researchers are pretty directly accusing OpenAI of plagiarism

https://mastodon.social/@tristanbuckmaster/11723647135247030...

20k··on On the Navier–Stokes Millennium Prize Problem
So, better to make that walled garden <checks> OpenAI? One of the scummiest companies on earth?
20k··on On the Navier–Stokes Millennium Prize Problem
https://mastodon.social/@tristanbuckmaster/11723647135247030...

He very much is accusing them of stealing his work

20k··on On the Navier–Stokes Millennium Prize Problem
They kind of did though, they were hoping to keep the fact that they may well have plagiarised these researchers unpublished work quiet. They did not want this to turn into a scandal about the fact that they appear to be training on prompts without consent

It makes a certain amount of sense. The internet data is too polluted with AI usage now to be useful, so the only AI free new data source is the prompts people feed into ChatGPT. The only problem is that its clearly plagiarism

Edit:

OpenAI have admitted to training on prompts at the time the breakthrough was made:

https://mastodon.social/@tristanbuckmaster/11723647135247030...

20k··on On the Navier–Stokes Millennium Prize Problem
Yeah well, its easy to do if you steal someone elses work and then try to threaten them into staying quiet about it

Edit:

OpenAI have now admitted they were training on prompts at the time they made their breakthrough:

https://mastodon.social/@tristanbuckmaster/11723647135247030...

20k··on On the Navier–Stokes Millennium Prize Problem
Are you joking? This is evidence that OpenAI is committing plagiarism en masse of researchers private work and threatening them into staying quiet to re-present their results as their own. This would be one of the largest scandals of all time
20k··on Navier-Stokes – Tristan Buckmaster [pdf]
I mean, its the researchers themselves talking about this

https://mastodon.social/@tristanbuckmaster/11723647135247030...

20k··on On the Navier–Stokes Millennium Prize Problem
The issue is that if OpenAI is training on prompts generally, what we really have is the first fully automated luxury plagiarism machine. In that it isn't able to genuinely solve problems, but merely steal the work that other mathematicians have been putting into prompts, and regurgitating that to other users as its own work. That makes them incredibly less useful as research tools

The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem

20k··on On the Navier–Stokes Millennium Prize Problem
Drama aside, this solution would be a counterexample disproving the smoothness postulate, which means that it leads to nothing new unfortunately. We already had working solutions to navier stokes, the only thing we didn't know is if the equations possessed a technical property

Its a bit like solving p = np with a negative result. Its an incredibly difficult problem, but it doesn't lead to anything at all on its own. This is why people are talking about the fact that the solution methodology is much more interesting than the solution - the tools used to crack something like this may lead to solving more useful problems

20k··on On the Navier–Stokes Millennium Prize Problem
It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model

What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question

If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?

Edit:

OpenAI have admitted they were training on prompts at the time they made their breakthrough

https://mastodon.social/@tristanbuckmaster/11723647135247030...

20k··on Navier-Stokes – Tristan Buckmaster [pdf]
>I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

If you think these companies are not training on your prompts you are incredibly naive. These models were built by stealing and pirating literally everything they can get their hands on no matter the legality. AI companies are always very specific about what they're not doing - in a way that you can drive a truck through the loopholes

20k··on Discovery of a new OpenAI agent message board
Shockingly poor security to let an application have totally unrestricted access to the web with no review, of course this kind of stuff is going to happen
20k··on Can I opt out of my input or output data being used for training?
I mean, they've already violated the law in acquiring all their training data already, why would they be uncomfortable violating a contract to get more training data?
20k··on Can I opt out of my input or output data being used for training?
I mean the models have literally trained on:

1. Child porn

2. Stolen music

3. Private github repos, before that was 'stopped'

4. Illegally pirated books

Them training on company prompts against the terms of service would be one of the least bad things that these companies have trained AI models on

Why do you think a company - willing to break the law for child porn - won't break the law when it comes to your personal data?

20k··on Can I opt out of my input or output data being used for training?
It isn't, in both cases you're protected by exactly the same legal system
20k··on Can I opt out of my input or output data being used for training?
You have to be rather naive if you don't think these companies don't simply train on your prompts with or without your consent. They literally scrape everything - legal or not - and claim its fair use to train on, including straight piracy

The idea that they'll steal from everyone except you is just wishful thinking

20k··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
These companies absolutely do have people who's job it is to sit on social media and promote these products, its marketing 101
20k··on AI boosted homework scores, then exam scores dropped: study
Yes, it strongly suggests that AI as a learning tool is at best pretty much useless, but on average its strongly detrimental to most people exposed to it as it makes it too easy to cheat
20k··on Claude Code May–August 2026 weekly limits promotion
They'll all be getting around to this sooner or later, its just too expensive
20k··on Understanding is the new bottleneck
Extremely silly. Even if LLMs did everything that everyone says they do (which they absolutely don't), they still hit a fundamental limit of complexity when they stop being useful

Its only a good idea if you work in selling tokens, otherwise you're dooming anything other than a simple app to inevitably breaking after it hits a certain level of complexity

20k··on Is it all just vapourware?
This doesn't seem to match the claims of 10x productivity that are being levied though. In fact it seems like the overall productivity is lower
20k··on Is it all just vapourware?
Then why isn't it showing up in open source at all? The people who file PRs are just regular average devs, but the quality of LLM PRs seems to be universally crap in comparison
← PreviousPage 2 of 11Next →