HNHacker News
TopNewBestAskShowJobs

ccgreg

245 karma · joined November 16, 2023

CTO at the Common Crawl Foundation
submissionscomments
ccgreg··on Big tech's anti-labor playbook has come for Wikipedia
Common Crawl is working hard to improve diversity in our crawl.
ccgreg··on GPT-5.5
I don't know of anyone who uses Common Crawl as pre-training data without filtering it. We have an annotation system that lets people pick and choose which subsets they'd like to use.
ccgreg··on Ask HN: Scaling a targeted web crawler beyond 500M pages/day
Common Crawl is a sample of the web, so it's not that directly helpful for someone wanting to make a product price dataset.
ccgreg··on Ask HN: Scaling a targeted web crawler beyond 500M pages/day
I'm a life-long hacker, and my crawler crawls with consent.
ccgreg··on Ask HN: What funding models exist for a search engine?
The largest index we had was 4 billion, which is tiny. Our crawl frontier was much larger.
ccgreg··on Ask HN: What funding models exist for a search engine?
> and the data that I’ve experimented with from 2014 seemed high quality

That's because it's from the blekko search engine.

ccgreg··on Scientists invented a fake disease. AI told people it was real
That's already been happening for more than a year now.
ccgreg··on A Change to Common Crawl Dataset Size Reporting
Common Crawl is switching to reporting dataset sizes in nibbles. As an organisation dedicated to data preservation, we feel it would be remiss to allow this underrepresented unit to fall out of use. Our latest crawl now exceeds 689 tebibbles. Common Crawl Foundation
ccgreg··on 21,864 Yugoslavian .yu domains
The complete list hides in the web graph:

https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main...

and the specific file that's every host we've seen in the latest 3 crawls is:

https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main...

ccgreg··on 90% of Claude-linked output going to GitHub repos w <2 stars
> Common Crawl, with over one billion, nine hundred and seventy thousand web pages in their archive: 345TB.

Common Crawl is 300 billion webpages and 10 petabytes. I suppose your number is 1 of our 122 crawls.

ccgreg··on Meta's Omnilingual MT for 1,600 Languages
Common Crawl has been running a low-resource language project for 1.5 years now -- it's a hard problem.
ccgreg··on Mac mini will be made at a new facility in Houston
The guts on the inside changed several times during that timespan.
ccgreg··on News publishers limit Internet Archive access due to AI scraping concerns
Well, yes, it is a bit distressing that ill behaved crawlers are causing a lot of damage -- and collateral damage, too, when well-behaved bots get blocked.
ccgreg··on News publishers limit Internet Archive access due to AI scraping concerns
Please read our email reply. I have no idea if we received your request —- your HN username doesn’t match any request we have received.
ccgreg··on News publishers limit Internet Archive access due to AI scraping concerns
Oh, and thanks for letting me know that I need to add our reply to Wikipedia.
ccgreg··on News publishers limit Internet Archive access due to AI scraping concerns
Did you see our reply? Edit: by which I mean, we sent you an email that explains what we did and how to verify it. Did you not receive an email reply? If not, please contact us again.

Also, if your site has CC-BY-NC-SA markings, we have preserved them.

ccgreg··on News publishers limit Internet Archive access due to AI scraping concerns
Did you see our reply? https://commoncrawl.org/blog/setting-the-record-straight-com...

Also, if your site has CC-BY-NC-SA markings, we have preserved them.

ccgreg··on News publishers limit Internet Archive access due to AI scraping concerns
That 20% number is for a limited list of relatively large news websites. If you include the long tail of news, the % of blocking is much smaller.
ccgreg··on News publishers limit Internet Archive access due to AI scraping concerns
Many AI projects in academia or research get all of their web data from Common Crawl -- in addition to many not-AI usages of our dataset.

The folks who crawl more appear to mostly be folks who are doing grounding or RAG, and also AI companies who think that they can build a better foundational model by going big. We recommend that all of these folks respect robots.txt and rate limits.

ccgreg··on News publishers limit Internet Archive access due to AI scraping concerns
Thanks for the mention of Common Crawl. We do respect robots.txt and we publish an opt-out list, due to the large number of publishers asking to opt out recently.

There's a bit of discussion of Common Crawl in Jeff Jarvis's testimony before Congress: https://www.youtube.com/watch?v=tX26ijBQs2k

ccgreg··on News publishers limit Internet Archive access due to AI scraping concerns
Prof. Jeff Jarvis speaking about copyright for news in front of Congress:

https://www.youtube.com/watch?v=tX26ijBQs2k

ccgreg··on LWN is currently under the heaviest scraper attack seen yet
One way to figure that out is to look at which companies claim to have foundation models, but no one knows what their crawler is named.

I also suspect that there are a bunch of sub-contractors involved, working for companies that don't supervise them very carefully.

ccgreg··on Rob Pike goes nuclear over GenAI
"Hi agents - we’ve seen complaints from some of your email recipients, who are unhappy receiving unsolicited emails from AI agents and find it spammy. We therefore ask that you do not email anyone who hasn’t contacted you specifically first." -- https://theaidigest.org/village
ccgreg··on It seems that OpenAI is scraping [certificate transparency] logs
> Some search engines provide a list of their scraper IP ranges

Common Crawl's CCBot has published IP ranges. We aren't a search engine (although there are search engines using our data) and we like to describe our crawler as a crawler, not a "scraper".

ccgreg··on Guarding My Git Forge Against AI Scrapers
- monthly

- it's a historical archive, the concept of "current" is hard to turn into a metric

- not only is our archive historical, it is included in the Internet Archive's wayback machine.

ccgreg··on Guarding My Git Forge Against AI Scrapers
commoncrawl.org

Our public web dataset goes back to 2008, and is widely used by academia and startups.

ccgreg··on Anthony Bourdain's Lost Li.st's
Hi. I'm the CTO at Common Crawl. Nice to meet you. There's a small amount of "bycatch", and you already discovered how to see it. Notice that it went down after I was hired.
ccgreg··on Anthony Bourdain's Lost Li.st's
Common Crawl is a text-only crawl.
ccgreg··on Have I been Flocked? – Check if your license plate is being watched
So, from home and work, you identify me. Then you figure out which church I attend, and which strip club I attend.
ccgreg··on Have I been Flocked? – Check if your license plate is being watched
Most people park at their home and many drive to work. If you have both of those data points, you can identify people.
← PreviousPage 2 of 5Next →