I hope a court establishes some rules of engagement here, even if it’s not this case.
I hope a court establishes some rules of engagement here, even if it’s not this case.
This would be a huge blow to open-source and research developers and I'd even argue it could help openAI to get a bit of a moat ala regulatory capture.
Google won that suit under fair-use as a massive searchable database was found to be transformative as well as the non-commercial nature.
So; if your web scraping companies goal is to allow people to bypass a paywall I suspect you'll have trouble in the future. If your web scraping company instead say allows people to do market analysis on how many people need a piano tuner in NYC and it doesn't do that by copying a NYT article doing original research I think you'll be fine.
FWIW, I can’t replicate on either GPT 3.5 or 4, but it may be that OpenAI has added new measures to prevent this.
That said, the meme image of Willy Wonka comes out of stable diffusion 1.5 almost perfectly with surprising frequency. Then again, this is probably because it appeared hundreds or thousands of times in the training set in all sorts of contexts because it's such a popular meme. There is a tension between its status as an integral part of language and its nature as a copyrighted screen grab.
However, I had good luck reproducing poems on GPT 3.5, both copyrighted and not copyrighted, because the choice of words is a lot more "specific" so to speak, and therefore higher temperature isn't enough to prevent complete reproduction of the originals. See https://chat.openai.com/share/f6dbfb78-7c55-4d89-a92e-f4da23... (Italian; the second example is entirely hallucinated even though a poem with that title exists, while the first and third are recalled perfectly).
I’m more surprised that it can repeat 100 articles; if that behaviour is consistent in larger sample sizes and beyond just NYT dataset (which might be repeated on the web more than other sources, causing overfitting), that would be impressive.
You could imagine at some point a large enough GPT5 or 6 or 7 will be able to memorize verbatim every corner of the web.
It's more like, is the new work a distinct expression, e.g. satire or commentary, based on the original.
You can reproduce the original verbatim and still be transformative by adding an element of critique.
Example: https://www.dmca.com/articles/akilah-obviously-vs-sargon-of-...
I think all this is doing is making us realize that we have built a massive economic system on a fundamentally flawed idea of ownership over ideas, and the only two solutions will be to tear up the rule book, which will be extremely painful, or double down, which will be fatal.
in japan, where they said anything goes for ai
so its best to not to lose a competitive edge with things that people openly publish on the internet, if you put it out there for everyone to see then expect other people to use it
its about a precedent. If you don't keep up with international competition, you lose.