Common Crawl
commoncrawl.org
commoncrawl.org
If I have an idea for a search product, competing with Google isn't really the first roadblock my brain puts up. It's more like "sure brain, sounds swell; now, how do you propose to populate this engine of yours?"
The reason being initially you just need a lot of pages to work with. Anyone can write a simple
while(links) { get link }
crawler and just let it run for months on end without too many issues. Heck just some xargs and wget will get you by for a long time.
By the time you have your search indexing and spitting out results you are going to need your own dedicated crawler anyway to ensure you are crawling the pages you have identified as being most interesting.
I imagine this data set is not useful for those doing a search engine, but those wanting to calculate statistics on snapshots of the web, such as pages using jquery and the like.