How would that work from a legal perspective, though?
Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?
How would that work from a legal perspective, though?
Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?
The only question is copyright, but I find it hard to argue that LLM training is not sufficiently transformative in 99% of cases.
It won't, with the LinkedIn vs HiQ precedent.
Does Google have a special agreement with Reddit (and all other sites?) or is it legally "fair use" to reproduce web pages that are available freely online?
I think there are probably issues to address with scraping it blindly:
- Can Reddit imprint its data somehow? A watermark?
- Can Reddit prove that certain type of information appeared on Reddit first and thus that serves as proof its data was used without authorization?
If OpenAI can't work around this, I'm not sure they would be willing to cross any lines in terms of copyright, they've already done it with ChatGPT and I am guessing rules are only going to get stricter on this topic.
To me this feels like its opening up the door for the elimination of copyright as any algorithmic layer interjected between scrapped data and end users could claim to be "inspired".