This is how I stopped crawlers from my sites. I used robots.txt, but backstopped that with detecting undesirable IP addresses and user agent strings. Those would get an error page instead of the actual page.
It does mean you need to keep an eye on your logs to spot new crawlers coming around, though.
“No person shall circumvent a technological measure that effectively controls access to a work protected under this title.”
I'm guessing that because robots.txt does not actually attempt to prevent access, it also doesn't "effectively control access" and so isn't covered by this.
The clear intent of the law is to prevent actual cracking of access controls.
In my view, an "access control" is a mechanism that prevents access to unauthorized people. An advisory page like that is not preventing access, it's merely asking people not to access.
I think I remember that there have been court cases that support this interpretation of the DMCA, but I'm not sure. Now, if that notice page had a login form or only allowed certain IPs to proceed, that would be an access control.
- "Ask HN: Prevent ML like GPT from using public posts like this one?" https://news.ycombinator.com/item?id=33980566
- "Ask HN: Does ChatGPT respect Robots.txt?" https://news.ycombinator.com/item?id=35027823
Restricting access helps a little.
But in the end the kingdom of humanity is all good and bad. Your life will only be worse if all you think or worry about is the bad.
PS: for reasons unrelated to the topic I do believe that podcast might curtain soon.
I've stopped adding new stuff to the public areas of my websites. Not as a result of GPT specifically, but as a result of a dramatic increase in the amount of scraping being done to support things that I don't approve of. AI training in general being one of them.
(I expect that the AI will still probably have similar kinds of problems than it does now regardless of whether or not it is allowed, though.)
This isn't clear one way or the other, though recent rulings indicate that US courts are firmly poised to declare web scraping of public content fully legal. IE there is nothing you can do to stop it other than making it not public.
Legal team is still checking, but technically you can copyright and define usage license for your content. But if you write about common subjects or non copyright material you may be better off the internet