> What is Common Crawl?
> Common Crawl is a 501(c)(3) non-profit organization dedicated to providing a copy of the internet to internet researchers, companies and individuals at no cost for the purpose of research and analysis.
> What can you do with a copy of the web?
> The possibilities are endless, but people have used the data to improve language translation software, predict trends, track the disease propagation and much more.
> Can’t Google or Microsoft just do that?
>Our goal is to democratize the data so everyone, not just big companies, can do high quality research and analysis.
Also DuckDuckGo founder Gabriel Weinberg expressed the sentiment that the index should be separate from the search engine many years ago:
> Our approach was to treat the “copy the Internet” part as the commodity. You could get it from multiple places. When I started, Google, Yahoo, Yandex and Microsoft were all building indexes. We focused on doing things the other guys couldn’t do. [2]
From what I remember reading once DuckDuckGo doesn't use Common Crawl though.
[2] https://www.japantimes.co.jp/news/2013/07/28/business/duckdu...