Do really think LLM vendors that download 80TB+ of data over torrents are going to be labeling their crawler agents correctly and running them out of known datacenters?
(I noticed Claude, OpenAI and a couple of others whose names were less familiar to me.)