An AI Scraping Tool Is Overwhelming Websites with Traffic
vice.com
vice.com
Please make this tool “opt-in” by default - https://news.ycombinator.com/item?id=35681085 - April 2023 (41 comments)
Not respecting robots.txt is idiotic and website maintainers are right to be pissed, but that's not a good faith proposal. People don't have to opt-in to Google crawling their site with a special header, Google just crawls everything unless you add noindex[1].
[1] https://developers.google.com/search/docs/crawling-indexing/...
Why exactly? So megacorporations can keep owning everything of value?
+I don't see how downloading all the images on a website is disruptive.
If you can't handle each of your images being downloaded once something is really wrong with your website
Also, the article claims the scraping generated so much traffic that it got reported to the owner as an attack. I’m not claiming anything other than what’s being reported. They should have at least throttled their application to reduce the load.
If the owner didn’t notice the scraping we might not even be having this conversation.
Because when the megacorporations were just corporations the web was young.
Now the web is old, and what was interesting and novel at the small scale is tedious and taxing at the large scale.
It's much the same as how it's now a death march to write a new standards-compliant browser, or how you can't really found a new country unless you want to start a bloody war or carve a chunk out of Antarctica.
Sometimes, the rules have changed by the time new things come along.
the issue isn't "downloaded once" but "downloaded at once". agressive crawlers often turn into dos attacks
This tool is scraping sites, it has webmasters reporting actual disruption, it doesn't have robots.txt support. When people complained (eg in https://github.com/rom1504/img2dataset/issues/48), the author's stance was basically "PRs welcome". It looks like a third party recently contributed a PR to make it respect robots.txt (https://github.com/rom1504/img2dataset/pull/302), albeit without `Crawl-Delay` support, which is not merged yet.
I have seen the same thing with other recent AI tools (eg https://github.com/m1guelpf/browser-agent/issues/2) and I think it's important to defend the robots.txt convention and nip this in the bud. If a bot doesn't make a reasonable effort to respect robots.txt and it causes disruption, it's a denial-of-service attack and should be treated as such. No excuses.
I guess it’s a commentary on the AI gold rush and another product for Cloudflare to sell...
I haven't done it specifically for images. I imagine a tool like this is crawling pages in order to find image links, although it could be using e.g. search engine results for this. Corrupting the data is also an option.
It's a blunt tool but if people don't honor robots.txt I don't need to concern myself with their pearl-clutching notion of morality. (I don't feel honor bound to list every possible evasion in robots.txt anyway.)
His contention is that denying content to AI tools deprives people of their right to better AI tools...
If anything picks up a URL and uses it later, that is definitely a web crawler.