245 karma · joined November 16, 2023
That's because it's from the blekko search engine.
https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main...
and the specific file that's every host we've seen in the latest 3 crawls is:
https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main...
Common Crawl is 300 billion webpages and 10 petabytes. I suppose your number is 1 of our 122 crawls.
Also, if your site has CC-BY-NC-SA markings, we have preserved them.
Also, if your site has CC-BY-NC-SA markings, we have preserved them.
The folks who crawl more appear to mostly be folks who are doing grounding or RAG, and also AI companies who think that they can build a better foundational model by going big. We recommend that all of these folks respect robots.txt and rate limits.
There's a bit of discussion of Common Crawl in Jeff Jarvis's testimony before Congress: https://www.youtube.com/watch?v=tX26ijBQs2k
I also suspect that there are a bunch of sub-contractors involved, working for companies that don't supervise them very carefully.
Common Crawl's CCBot has published IP ranges. We aren't a search engine (although there are search engines using our data) and we like to describe our crawler as a crawler, not a "scraper".
- it's a historical archive, the concept of "current" is hard to turn into a metric
- not only is our archive historical, it is included in the Internet Archive's wayback machine.
Our public web dataset goes back to 2008, and is widely used by academia and startups.