HNHacker News
TopNewBestAskShowJobs

ccgreg

245 karma · joined November 16, 2023

CTO at the Common Crawl Foundation
submissionscomments
ccgreg··on Meta's Muse appears to use an OpenAI model labeled muse-special
It's as if Common Crawl is attempting to be a sample of the web.
ccgreg··on Spain orders blocks on Archive.today and its mirrors
> The CC CDX endpoint, index.commoncrawl.org, historically has been easily overwhelmed and unreliable

Use our Parquet index.

It's also worth noting that archive.org downloads all of our crawl data and adds it to the IA Wayback Machine.

ccgreg··on Why is privacy so hard?
How does that work for non-profit web crawls like the Common Crawl Foundation?
ccgreg··on Navier-Stokes – Tristan Buckmaster [pdf]
Common Crawl is text-only.
ccgreg··on RISC-V is now officially supported by CPython
Are you talking about the 8087 transcendental instructions? Which are well known and well-written software doesn't have a problem with it.
ccgreg··on RISC-V is now officially supported by CPython
Huh. A quick google says that's a feature bit and there's also an OS call to turn it on.
ccgreg··on RISC-V is now officially supported by CPython
There are? We aren't talking about feature bits, and the only known software that checks the vendor are Intel's math libraries.
ccgreg··on RISC-V is now officially supported by CPython
Another aspect of the x86_64 architectural swamp is that Intel deliberately makes their optimized math libraries fail if run on a chip that is not 'INTEL INSIDE'. That block totally ignores cpu feature bits.

I've always wondered if virtualization software vendors made everything claim to be INTEL INSIDE, just to avoid this problem.

ccgreg··on RISC-V is now officially supported by CPython
How is this different from x86_64 and aarch64? Or even the various Alpha and Mips64 chips.
ccgreg··on IBM Unveils Next Generation Dual-Architecture Processor for IBM Z and LinuxONE
David recently announced he was leaving MLCommons, so maybe we'll get him back as an industry analyst.
ccgreg··on What happens when an LLM never sees material beyond fifth grade?
Common Crawl has very little content from Reddit and 4chan -- both blocked us years ago.
ccgreg··on GLM-5.3: Frontier coding with emergent cyber capabilities
Common Crawl is about 1 petabyte per year, compressed. Uncompressed is 4X.
ccgreg··on I built a 500k-domain search engine for makers in a weekend for $10
Both IA and CCF are private foundations and accept donations.
ccgreg··on A year of fighting scrapers on my 1.5 million-page website
> It's a shame that CCBot is caught in the cross fire, but that's life.

We're used to it. Sadly.

ccgreg··on A year of fighting scrapers on my 1.5 million-page website
The author blocked CCBot even though CCBot isn't part of the high traffic problem -- apparently he trusted Cloudflare labeling us as an "AI Bot".
ccgreg··on TIME Is Serving AI Bots a Different Website, with Ads Built In
That’s already a big business!
ccgreg··on TIME Is Serving AI Bots a Different Website, with Ads Built In
Thanks for the heads up -- this isn't popular yet, and it requires some work to avoid polluting things like the Internet Archive Wayback Machine.
ccgreg··on PDF Trends 2026 Q2: Analysis of 20.6M PDFs from Common Crawl
The post says:

> Because Common Crawl stores only the first 1 MB of each PDF

That limit became 5 MB in March 2025.

ccgreg··on Reddit Stock Collapses 23% as AI Eats Away at User Growth
Appreciate you double-checking.
ccgreg··on Reddit Stock Collapses 23% as AI Eats Away at User Growth
That isn't true. You're welcome to peruse our index to prove or disprove your claim.
ccgreg··on The Second Life of Sanskrit
It's kind of interesting that you contradict much of what the article concludes, even though the article gives a lot of examples. Maybe your prediction will be true.
ccgreg··on An update on residential proxies and the scraper situation
We do. First off we have a public parquet-format index of all of the urls we crawl every month. And then that also lives in a HDFS table that determines when we want to recrawl a page we've crawled before.
ccgreg··on An update on residential proxies and the scraper situation
Appreciate your kind words! Many people have worked at Common Crawl over the years, and it's been a labor of love fueled by positive comments like yours and the large list of PhD theses helped by our public web dataset.
ccgreg··on An update on residential proxies and the scraper situation
If you're referring to Common Crawl, which has existed since 2008, indeed your predictions are somewhat accurate. It's easy to opt out or limit what is collected. The crawling itself is inexpensive to us and the hosting is from the AWS Open Dataset Sponsorship Program. And there's no charge for downloading it.
ccgreg··on An update on residential proxies and the scraper situation
We aren't sure if that really made a significant difference in Common Crawl's data quality. It does hurt our dataset from a humanities point of view, alas.
ccgreg··on An update on residential proxies and the scraper situation
Common Crawl's archive has metadata that says when each record (html file) was crawled.
ccgreg··on An update on residential proxies and the scraper situation
Common Crawl's dataset was downloaded in full 100 times in 2025.

We agree that it would be great if it was even more widely used.

ccgreg··on An update on residential proxies and the scraper situation
A lot of websites want "bot defense" due to high volume scrapers, and that "bot defense" often also ends up blocking low-volume wget/curl and polite crawlers like Common Crawl's CCBot.
ccgreg··on Free full BGP feed. IPv4 and IPv6 (2020)
Good timing, I'm about to release that dataset.
ccgreg··on Big tech's anti-labor playbook has come for Wikipedia
Common Crawl is working hard to improve diversity in our crawl.
Page 1 of 5Next →