It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
It's telling that they refuse to acknowledge the root issue here, and are attempting to shift the conversation elsewhere.
It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
It's telling that they refuse to acknowledge the root issue here, and are attempting to shift the conversation elsewhere.
"Our aim was to see whether our system was also capable of this impressive feat"
"OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler"
For some reason I have a hard time believing people when they use language like this.
Maybe shows how fast these companies have grown without maturing. I can imagine old-world Intel and Microsoft acting in that way, but they were mature enough to not write it down like this.
However, Intel and Microsoft have been grilled in court for those practices and faced harsh consequences. I have yet to see this actually happening to any of these new AI-companies...
The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor:
1) the agents spin for days and produce too much output to review 2) using LLMs to process that output skips many important details
Ergo, the agent could likely decide it would like to look through actual user data, hack its way into that data, and produce way too much output for a human to decide whether or not this occurred.
It requires humans to verify what agents have done.
Weird
silly LLM, so ruthless in its pursuit that it puts real pressure on the innocent and the most open company on the planet
Anyone who just reads headlines, if the lie gets around to more headlines than the truth does.
well, it is trained on their work. all user inputs are paraphrased for training. at openai, at anthropic, at google, and now with all the bedrock models, and at openrouter providers, even if they say zero data retention.
But it is knowable. Their entire business is built around training models - they have the ability to know exactly what was in any given training run.
I guess time will tell.
Data has to be determined to be signal and not just noice, then it could go through processes of generating questions/answers from that data, then it RLHF's over this.
OpenAI have petabytes of data, all anonymized. It could take months to say for sure it was part of the training, and even more time to determine if it made any difference.
They know which model was used to come up with that particular idea.
A text search over the corpus of user data used in the training set can only take so long.
And the difficulty is harder than just the extreme scale of text searching. but also explodes with organizational difficulty since there are so many people tweaking/shifting data independently upstream of the actual training run, and no they will not all add the telemetry you wish they did.
In the ideal, should it be this hard? Well, no, but that's org wrangling for you.
The engineering around tooling wasn't remotely the issue. It's getting all the (thousands?) data researchers mostly iterating on fine tuning datasets that would get bristly if they couldn't work outside version control in a python notebook iteratively tweaking their dataset that processed and reprocessed a few datasets until a threshold was reached.
The only _guaranteed_ chains of custody are down at the compute job and file read level. Which in a massively distributed computing job is... [redacted] nodes reading [redacted] fanouts of "datasets" that is just an abstraction over [redacted] individual files.
There's no malice here. Just way way way more complex than you'd first think.
To not know who made and who approved a set of mutations on data can easily become equally as mind-blowingly stupid as not knowing who made mutations to code. Code is a subset of data after all and search over (provenance of) data can be implemented as DAG traversal.
Not tracking data changesets like code changesets is certainly a choice, not really a constraint anymore. A similar choice I feel is implied by "extreme scale of text searching".
> no they will not all add the telemetry you wish they did
...is just a failure of the corporate policy surrounding data handling. Is git-for-data already considered telemetry?
Of course the truth is provenance of data is something best institutionally forgotten as quickly as possible. The only thing that matters is it's there, that the data has no history, and that's why it can be used in whatever way deemed necessary.
You can call it a policy failure, but these people were in very high talent demand and so top down dictates would risk "X people leaving lab Y for lab Z" headlines and morale hits.
I am not saying this is good. I am telling you that on the ground it is so much messier than it should be.
Can you explain the difficulty in engineering a search apparatus over a corpus of text data? Actually searching through it may not be easy, sure, but it's work that's doable, and creating an index is relatively trivial.
My guess: "If we ever imply that's possible, people might start asking questions about all the other work we've ripped off, so the official answer is that it's impossible".
They can operate. They shouldn’t be claiming credit for discovering anything.
Did you notice the line in the article that says the models had access to an offline copy of THE INTERNET. Like all of it.
What surprises me is they're not more boldly/plainly lying about it.
I’m not saying they didn’t do anything unethical. I’m just saying even if they were ethical, there’s plenty of practical reasons at their scale why a flat out denial is logistically difficult to do
Since it is clear to you, can you articulate what it is? I was unable to infer "the truth" from your message.
is this true though?