https://twitter.com/Rahll/status/1739003201221718466
It's frankly impressive how well this image is embedded in the weights of their model, down to tufts of hair. And it's far from the only one.
1. Models are trained with 100% uncopyrighted or properly licensed input data
2. Every output of the ML model is evaluated to make sure it's not too close to training data
3. Copyright law is changed to have a specific cutout for AI
#1 is the approach taken by Adobe, although it generally is harder or more expensive to do.
#2 destroys most AI business models
#3 has been done in some countries, but seems likely that if done in the US it would still have some limits.
For example, I could train a model on a single image, song, or piece of written text/code. Then I run inference, and get out an exact copy of that image, song, or text. If there are no limits around AI and copyright, then we've got a loophole around all of copyright law. I don't think that the US would be up for devaluing intellectual property like that.
4. A ruling comes down that enshrines what all the big companies have been doing (with the blessings of their armies of expensive, talented, and conservative legal teams) as legitimate fair use
You are forgetting the massive asterisk that you need to provide multiple paragraphs of the original article in order to get verbatim output from chatgpt. In what world are people doing that to avoid paying the NYT?
https://arstechnica.com/tech-policy/2023/12/ny-times-sues-op...
And this is a screenshot of their session whith copilot
https://cdn.arstechnica.net/wp-content/uploads/2023/12/Scree...
Its funny that the behavior was patched if OpenAI believes it isn't copyright infringement.
It supports an argument that GPT shouldnt produce outputs that are extremely similar, not that the content can not be used as an input.
as well as all the output it ever generated
Does that mean that models that can not produce copies of X length ARE fair use?
not necessarily
"sufficient but not necessary" I believe is the term
it is a fringe case that rarely occurs, and only with a lot of user prompting.
legal discovery could almost certainly compel the LLM host to provide access to the output of the weights themselves without the "gatekeeper" present
If there is a case to be made, I think it has to be around the original use of the works, the transcription process. Not the weights, or the output
If the archive can't produce the original work, it's not infringing.
If you printed the binary of Harry Potter and sold it as a painting, that would be fair use. It doesn't matter if the data is encoded in it, if it is not used for extraction.
Think of Andy warhol's Campbell Soup. Nobody is going to confuse the art for a can of soup and try to eat it. That's not being sold as a label for other soups. However, the original Campbell Soup data is absolutely encoded there.