Using Nginx to block Meta, Twitter and ChatGPT access to your sites
gist.github.com
gist.github.com
Worth noting that most of these bots are 'good bots' (i.e. they will obey robots.txt). So you can avoid the nginx resource usage entirely by adding suitable robots.txt entries.
I think using nginx tests like this could have negative effects on showing OpenGraph metadata (including images).
If choosing this approach however I would probably respond with a 403 code to mark forbidden as bots are more likely to continue making attempts if they think the server would come back online.
In the main site config redirect anyone not using HTTP/2.0. GoogleBot still doesnt use HTTP/2.0 so this will block Google. Bing is OK though. One could instead use variables to make this multi-condition and make exceptions for their CIDR blocks. Point "auth." DNS record to the same IP and ensure you have a cert for it or a wildcard cert.
# in main TLS site config:
# replace apex with your domain and tld with its tld.
if ($server_protocol != HTTP/2.0) { return 302 https://auth.apex.tld$request_uri; }
Then in your "auth" domain use the same config as the main site minus the redirect but then add basic authentication. Anyone not using HTTP/2.0 can still access the site if they know the right username/password. If you get a lot of bots then have an init script copy the password file into /dev/shm and reference it from there in NGinx to avoid the disk reads. # then in the auth.apex.tld config.
# optionally give a hint replacing i_heart_bots with name_blah_pass_blah
auth_delay 2s;
location / {
auth_basic "i_heart_bots"; auth_basic_user_file /etc/nginx/.pw;
}
This will block some API command line tools, most bots good or bad, some scanning tools. Some bots will give up prior to 2 seconds so you will get a status 499 instead of 401 in the access logs. Only do this on silly hobby sites. Do not use in production. Only people wearing a T-Shirt like this one [1] may do this in production.One may be surprised to find that most bots use old libraries that are not HTTP/2.0 enabled. When they catch up we can replace this logic using HTTP/3.0 and UDP. Beyond that we can force people to win a game of tic-tac-toe or Doom over Javascript.
[1] - https://www.amazon.com/Dont-Always-Test-Production-Shirt/dp/...
Once it’s online, it’s online.
The deadenders who felt it was worth it will keep trying for at least a while; the new exploiters will tend to give up sooner. robots.txt is a courtesy. Not everybody puts stuff on the internet with a working theory that your experience is more important than theirs.
- I don't want my content on those sites in any form and I don't want my content to feed their algorithms. So I do not care for opengraph or previews.
- Using robot.txt assumes they will 'obey' it. But they may choose not to. Its not mandatory in anyway.
- Yes, they can fake UA. This does not mean I should not take any measures to block them just becasue they can fake.
If they wanted to scrape your site, nobody can prevent them.
ChatGPT can't be an impolite Internet citizen (spoofing UA's) and claim to be using AI for the good for humanity, so they're not going to be dishonest with their user-agent.
That reads an awful-lot like "Google can't be evil and claim that their motto is 'Don't be evil', so they're not going to be evil" but here we are. The profit motive eventually undoes any principled claim by a company.
Like anything else in IT security it's never "set and forget" permanently; the effectiveness of things like that decay over time and must be periodically re-evaluated.
But if something can be used to your advantage now, even if for a while, then why not use it.
and check your logs to see who is not complying
[0] https://webmasters.stackexchange.com/questions/137914/spike-...