It's as if Common Crawl is attempting to be a sample of the web.
245 karma · joined November 16, 2023
Use our Parquet index.
It's also worth noting that archive.org downloads all of our crawl data and adds it to the IA Wayback Machine.
I've always wondered if virtualization software vendors made everything claim to be INTEL INSIDE, just to avoid this problem.
We're used to it. Sadly.
> Because Common Crawl stores only the first 1 MB of each PDF
That limit became 5 MB in March 2025.
We agree that it would be great if it was even more widely used.