1,904 karma · joined September 24, 2014
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Which seems to be very directly accusing OpenAI of plagiarism
https://mastodon.social/@tristanbuckmaster/11723647135247030...
He very much is accusing them of stealing his work
It makes a certain amount of sense. The internet data is too polluted with AI usage now to be useful, so the only AI free new data source is the prompts people feed into ChatGPT. The only problem is that its clearly plagiarism
Edit:
OpenAI have admitted to training on prompts at the time the breakthrough was made:
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Edit:
OpenAI have now admitted they were training on prompts at the time they made their breakthrough:
https://mastodon.social/@tristanbuckmaster/11723647135247030...
https://mastodon.social/@tristanbuckmaster/11723647135247030...
The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem
Its a bit like solving p = np with a negative result. Its an incredibly difficult problem, but it doesn't lead to anything at all on its own. This is why people are talking about the fact that the solution methodology is much more interesting than the solution - the tools used to crack something like this may lead to solving more useful problems
What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question
If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?
Edit:
OpenAI have admitted they were training on prompts at the time they made their breakthrough
https://mastodon.social/@tristanbuckmaster/11723647135247030...
If you think these companies are not training on your prompts you are incredibly naive. These models were built by stealing and pirating literally everything they can get their hands on no matter the legality. AI companies are always very specific about what they're not doing - in a way that you can drive a truck through the loopholes
1. Child porn
2. Stolen music
3. Private github repos, before that was 'stopped'
4. Illegally pirated books
Them training on company prompts against the terms of service would be one of the least bad things that these companies have trained AI models on
Why do you think a company - willing to break the law for child porn - won't break the law when it comes to your personal data?
The idea that they'll steal from everyone except you is just wishful thinking
Its only a good idea if you work in selling tokens, otherwise you're dooming anything other than a simple app to inevitably breaking after it hits a certain level of complexity