Do be aware that it does include copyrighted content so distribution is piracy.
Do be aware that it does include copyrighted content so distribution is piracy.
Also, most image-text dataset pairs contain far worse than that. You might want to check out LAION-5B and what stanford researchers have found in there. Technically, anyone who even touched that could in theory be in some serious, serious trouble. I find it quite remarkable that nothing has happened yet.
- OpenAI and others will just settle with MPAA, RIAA and the likes for a revenue stream (a single digit billion a year, likely) + some kind of control over what people can and cannot do with the AI + the access to the technology to produce their own content.
- artists will see peanuts from the deal, and the big names are going to be able to stop doing any kind of business with artists which are just expenses in their eyes. They will have been replaced by machines that where trained using their art with no compensation whatsoever.
IP is already predatory capitalism, AI will definitely be weaponized against the workers by the owners of the means of “production”.
It's impossible to create such a list while evading all such material.
These hashes is exactly how researchers later discovered this content, so it’s clearly not hard.
Full paper: https://stacks.stanford.edu/file/druid:kh752sm9123/ml_traini...
It is a pretty hard problem.
And this isn’t just “one engineer”. Companies like StabilityAI, Google, etc have used LAION datasets. If you built a dataset you should expend some resources on automated filtering. Don’t include explicit imagery as an intentional choice if you can’t do basic filtering.
That's an amplification of copyright, original expression is protected, but not the ideas themselves, those are free. And don't forget when we actually get to use these models we feed them questions, data, we give corrections - so they are not simply replicating the training set, they learn and do new things with new inputs.
In fact if you think deeply about it, it is silly to accuse AI of copyright violation. Copying the actual book or article is much much faster and cheaper, and exact. Why would I pay a LLM provider to generate it for me from the title and starting phrase? If I already have part of the article, do I still need to generate it with AI? it's silly. LLM regurgitation are basically attacks with special key, entrapments. They don't happen in normal use.
While I don't think it's because you're wrong, per se, it's just that none of this drama really matters.
Somehow people are just not able to get this through their heads. Stable diffusion is like 12GB or something and you have people convinced it's a tool that is cutting and pasting copyrighted works from an enormous image archive.
Good to know I can avoid copyright on a book just by zipping it up!
LLM's are not compressing petabytes of information down to a few gigabytes.