Dark Visitors – A list of known AI agents on the internet
darkvisitors.com
darkvisitors.com
Hopefully it's pretty self-explanatory but I made this website as a simple resource for people who want to stay up to date with the ever-changing cast of AI user agents.
Feel free to sign up with the Google Form to get notified when this list is updated. And if you know of any agents I'm missing, please submit them. Thanks!
This is almost certainly the coral web grounding found as an option on coral.cohere.ai
1: https://taosecurity.blogspot.com/2008/02/review-of-dark-visi...
I'm also curious, will adding "Common Crawl" to the user-agent disallow list in your robots.txt actually do anything?
Added context: I run a tracker of vetted AI agents with verticalized use cases https://staf.ai
YMMV
Isnt that generally the point of putting something on the internet?
There is a typo in one of the classification texts:
not currently classified as artificailly intelligentI'm not sure how often it's updated. It seems fairly complete from other sources I've been looking at (mainly robots.txt files for well-known publishers).
I assume they'd have visibility into a huge amount of this
An updating list of IPs for these agents wouldn't go amiss. I might have access to some site logs I could mine for that.
We have very few tools to protect ourselves here, and need to make the most of the ones we have.
Personally, I think that such efforts are pointless -- all it takes is one crawler to make it through your defenses and you may as well have not done a thing. The alternative, though, is to remove your sites from the open web entirely. Either way, Commmon Crawl needs to be excluded if you want to avoid your stuff being used to train AI.
https://apnews.com/article/dungeons-dragons-ai-artificial-in...
a variety of stakeholders involved - not just "individuals doing pointless things" .. new tech is making winners and losers right now.
Honestly, it seems like we've been here before. I hope y'all have got your right-click-disable scripts in place too.
It's going to be an interesting half-decade.
Would be nice if they had a fully filled robots.txt for download (I could only find the example file).
This example blocks everything, which is probably not what you want, but it's meant to be a starting point.
Love the design and name, too.
For chatgpt user:
https://platform.openai.com/docs/plugins/bot
And for gtpbot:
https://platform.openai.com/docs/gptbot https://openai.com/gptbot.json
Modern browsers already have LLMs built in so anti-scraping is a folly.