Presumably the crawlers don’t already have an LLM in the loop but it could easily be added when a site is seen to be some threshold number of pages and/or content size.
Presumably the crawlers don’t already have an LLM in the loop but it could easily be added when a site is seen to be some threshold number of pages and/or content size.
It becomes an economic arms race -- and generating garbage will likely always be much cheaper than detecting garbage.
My point isn’t that I want that to happen, which is probably what downvotes assume, my point is this is not going to be the final stage of the war.
I don't follow that at all. The post of yours that I responded to suggested that the scrapers could "just add an LLM" to get around the protection offered by TFA; my post explained why that would probably be too costly to be effective. I didn't downvote your post, but mine has been upvoted a few times, suggesting that this is how most people have interpreted our two posts.
> it can learn which pages are real and “punish” the site by requesting them more
Scrapers have zero reason to waste their own resources doing this.