Why? What is the goal of a scraper, and how does disabling the source of the data benefit them?
The next scraper doesn’t get the data. People don’t realize we’re not compute limited for ai, we’re data limited. What we’re watching is the “data war”.
Seems like there’s a fuck ton. All of Wikipedia, GitHub for code, etc.
I can understand targeting certain sites like Reddit, etc. but not random websites
If you look closely even Google does this. This is probably why many popular sites started getting down ranked in the last 2 years. Now they're below the fold and Google can present their content as their own through the AI box.
Yea, but, the FTC doesn't want it to be.
It feels a lot like they're stuck for improvements but management doesn't want to hear it.