AI haters build tarpits to trap and trick AI scrapers that ignore robots.txt
arstechnica.com
arstechnica.com
Nepenthes is a tarpit to catch AI web crawlers - https://news.ycombinator.com/item?id=42725147 - Jan 2025 (263 comments)
The article completely misses the point that AI scrapers are not a "future threat of AI domination". They already do damage by DDOSing site's networking infrastructure and inflicting very real costs to a site hoster.
Even when the data is completely free, like in case of Wikipedia or OpenstreetMaps, scraping it is unethical and should be illegal. Most of the open data resources have procedures, which allow downloading of the data in the archived form, without need for scraping. They are built with sharing in mind.
So the arguments the article tries to use (what if it is for public good?) has no sense. 1) it is not 2) there are many ways to fetch the open data properly and respectfully.
It is very likely that voting, or voting with your wallet, or probably any kind of activism, would have more impact than withdrawing from the (online) public life.
Both for me. They should be spanked for their behavior and lack of respect. I don’t want compensation though because I write open-source applications, but I want them to respect the license (which they don’t obviously).
Also I don’t understand why you feel anybody is withdrawing from the internet. It’s only a tarpit and I’m sure most of those who react don’t have ChatGPT subscriptions.
I already don't pay for any ai services or touch any models. but Increasingly services that used to be helpful for me are throwing them in - YouTube premium has some sort of ai summarize things, etc. how do I signal that I don't want companies scraping content there?
One of the first persons that tried to scam users out of money by asking for help with his tuition was doxxed by his own ISP after the aforementioned ISP got so much hatemail it crashed their servers when his message was posted to every message board by a script which caused prolific BBS users to download it possibly several hundred times, paying for the privilege of each message.
Copyright is - like it or not - the way we regulate commercial intellectual "property". I can see different IP doctrines, and I don't necessarily defend the current one. It's not derived from real property rights, but rather in ensuring economic incentives for people to make stuff that otherwise wouldn't have been made, such as pharmaceuticals and hollywood movies, to simulate property rights. It's an imperfect solution which is there to ensure economic incentives and balance, and most importantly, it's the one we got.
But then, multi-billion dollar corporations feed your copyright protected (you thought) works straight into their supply chain, wouldn't you be pissed? It's no a small part either, but their models would be extremely nerfed without copyrighted data. Forget AI, forget tech, just look at it from a purely economic ecosystem perspective. Crying "fair use" during a highway robbery probably don't sit right with many, I hope.
This comes across as snark, but I will assume you are well meaning. I have put code, guides, and videos on the internet for other people to consume for free in the hope that those people find that stuff useful.
I did not put stuff on the internet for it to be hoovered up and frankly stolen by massive AI companies to enrich themselves. If they are going to use my things in their commercial product than yes, they should be compensating me for that.
> Is this because they are multi-billion dollar companies, or because they behave poorly
It's also both of these.
> It is very likely that voting, or voting with your wallet, or probably any kind of activism, would have more impact than withdrawing from the (online) public life.
The only true control I have is withdrawing. I don't give these companies money, I don't live in a country that can meaningfully legislate against them, and I would consider withdrawing a form of activism.
I refuse to support these AI companies in any way (as long as they continue to be bad actors) and I have taken down all Youtube videos I've created, my personal website, and I have moved all my code to a self hosted, private Git service in order to deny them my work.
It's a glorified librarian at best
ML is an application of AI.
As an engine is an application for a car, you don't call an engine a car. ML is programmed information, and theres nothing artificial about that.
What is AI about it? Show me something AI from ML.
All algorithms are defined, all tokens are defined. All guardrails are defined.
All of it has been programmed by humans that which isn't artificial.
If it was AI then it would generate itself for itself.
If you don't think that ML is AI, I am fine with that.
"Cloud Servers" anyone? Glorified dedicated servers.
I wouldn't classify it as intelligent though. It's still reading a scripture of words.
It is immaterial that these AI companies ignore contractual obligations (TOS), and are in fact performing attacks on said sites (DDOS is an attack).
In the last 3 months, there have been 4 or 5 small project that I regularly frequent where their sites that have been knocked offline as a result of this type of bad behavior, where they definitely are not following the robots.txt.
The article is just bad shilled journalism.
I mean I think in the minds of AI evangelists (particularly of the quasi-religious "LLMs will bring forth a benevolent god-like superintelligence" variety), those are essentially the same thing.
(Yeah, it's ridiculous characterisation, but given the source it shouldn't be _surprising_ characterisation.)
Most crawlers use some form of timeout mechanism, usually informed by some priority scheduling. This deals reasonably well with crawler traps.
Since Nephentes-like traps are getting so common now (and in particular, not always behind robots.txt), I added a clause to Marginalia's crawler that prevents it from extracting links from pages that are less than 2 Kb and take more than 9 seconds to load. It's 4 lines of code and means the crawler doesn't get stuck at all.
I totally get the frustration though. My sites get an insane amount of bot traffic as well. I think roughly 1% of the search traffic to the html endpoint is human, and that's while providing a free API they could use instead. ... I just don't think this is going to fix anything.
> That's likely an appealing bonus feature for any site owners who, like Aaron, are fed up with paying for AI scraping and just want to watch AI burn.
Website owners aren't "haters" if bots ignore robots.txt, consuming resources that translate to expenses and a bad experience for legitimate website visitors.
Why would the website owner have to commission much larger server(s), pay more for this traffic and get nothing in return? At least search engines send human visitors your way.
It's not "AI haters". It's exploitation hating.
A little note in robots.txt offering commercial terms would also be available.
Calling people who don't want certain things hoovered up by an LLM "AI haters" is a level of manipulation I'd think was only reserved for someone with a vested interest in the tech. Just encourages devious behavior instead of a more diplomatic approach of respecting people's wishes (read: robots.txt).
Might I suggest https://tvtropes.org/pmwiki/pmwiki.php/Literature/BelindaBli...
As a bonus, beyond introducing the magic robot to the wonderful world of really poorly written erotica, it's also very anachronistic (widespread use of fax machines, pagers, smartphones, LinkedIn, and East Germany appear to coexist at a single moment in time), so will cause further confusion.
Ultimately instead of going down this path, I decided to just start charging for access to the service (it was long overdue)[2].
Users who are logged out can still see old cached content (which is a single DB read op), but to aggregate new content requires an account. I feel like this is a good (enough) middleground solution for now.
[1]: https://kulli.sh
[2]: https://lgug2z.com/articles/in-the-age-of-ai-crawlers-i-have...
if depth > 5 and if sem_hash(content) in hist: return
I think it will probably make it harder for screen readers, unfortunately.
wouldn't this lower their page rankings? that's the kind of shenanigans of the old days with meta key word stuff and what not
That's kind of the point. AI has just been more terrible news after terrible news.
It is NOT a good thing. Unless you know, you like being covered in oil and lit on fire...
25 years ago, if you had blocked the googlebot scraper because you resented google search, it would only have worked to marginalize the information you were offering up on the internet. Avoiding LLM training datasets will lead to similar outcomes.
What benefit is gained by allowing AI companies to train on your content? LLMs work on a token by token basis.