Reddit founder wants to charge Big Tech for scraped data used to train AIs
marketwatch.com
marketwatch.com
> User Content Transmitted Through the Site: With respect to the content or other materials you upload through the Site or share with other users or recipients (collectively, “User Content”), you represent and warrant that you own all right, title and interest in and to such User Content, including, without limitation, all copyrights and rights of publicity contained therein. By uploading any User Content you hereby grant and will grant Y Combinator and its affiliated companies a nonexclusive, worldwide, royalty free, fully paid up, transferable, sublicensable, perpetual, irrevocable license to copy, display, upload, perform, distribute, store, modify and otherwise use your User Content for any Y Combinator-related purpose in any form, medium or technology now known or later developed.
Another relevant bit:
> Except as expressly authorized by Y Combinator, you agree not to modify, copy, frame, scrape, rent, lease, loan, sell, distribute or create derivative works based on the Site or the Site Content, in whole or in part, except that the foregoing does not apply to your own User Content (as defined below) that you legally upload to the Site. In connection with your use of the Site you will not engage in or use any data mining, robots, scraping or similar data gathering or extraction methods.
I guess the one question I'd have here is whether or not these LLM's are being created by scraping content in a manner that ignores robots.txt files. While not strictly illegal, this is relatively hostile behavior, and there may be a case that it amounts to unauthorized access.