> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
That they don't is telling.
It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.
We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.
Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely.
> not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes
It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).
The fact they can't make a blanket denial triggers everyone's bullshit detectors, and they're getting eviscerated over it.
Again, see the Apple lawsuit. OpenAI has never been constrained by facts in what it can and can’t say.
To the extent anything is setting off my bullshit detector, it’s in the idea that this time is different (Moreover, the idea that we should assume this divergence without evidence.)
OpenAI doesn’t have the benefit of doubt. They shouldn’t for anyone who’s honest and reasonable. That doesn’t mean they’re automatically at fault. But when the twentieth person comes forward and says a pattern is continuing, I’m giving them the preliminary benefit of doubt. It’s a bit extreme to conclude based on that. But it’s far more baseless to swing to the other side and claim we need to clear the table for a serial offender.
Further, given that this is all in the open now, they can search the training data. No way somebody is using some specific unique cutting edge mathematical approach to solve a fluid dynamic problem 99.9% of people have never heard of and it's not locatable. Considering they spent $15,000,000 already on this, they could afford to grep around to be able to state that their hands are clean.
No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.
Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.
Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.
Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.