Download the internet
dotnetdotcom.org
dotnetdotcom.org
# Information on how to block our crawler. (Hint, it doesn't involve legal action)
# Our purpose and goal. (Yes we have one and no it doesn't involve spam)
# Our technology. (Thanks open source!)
[...]
(1) Used a preexisting aggregate web content format. Their ad hoc format is simple enough, but can't handle content with NULLs, and loses valuable information (such as time of capture -- you can't trust server 'Date' headers -- and resolved IP address at time of collection).
They could use the Internet Archive classic 'ARC' format (not to be confused with the older compression format of the same name):
http://www.archive.org/web/researcher/ArcFileFormat.php
Or the newer, more involved and chatty but still relatively straightforward 'WARC' format:
http://archive-access.sourceforge.net/warc/
(2) Explained how the 3.2 million pages in their initial dump were chosen. (That's only a tiny sliver of the web; where did they start and what did they decide to collect and put in this dataset?)
(FYI, I work at the Internet Archive.)
The Nutch crawler is also reasonable for broad survey crawls. HTTrack is also reasonable for 'mirroring' large groups of sites to a filesystem directory tree.
(1) Add a scope rule that throws out discovered URIs with popular non-textual extensions (.gif, .jpe?g, .mp3, etc.) before they are even queued.
(2) Add a 'mid-fetch' rule to FetchHTTP module that early-cancels any fetches with unwanted MIME types. (These rules run after HTTP headers are available.)
(3) add a processor rule to whatever is writing your content to disk (usually ARCWriterProcessor) that skips writing results (such as the early-cancelled non-textual results above) of unwanted MIME types.
Followup questions should go to the Heritrix project discussion list, http://groups.yahoo.com/group/archive-crawler/ .
Even as a fairly large operation, you're might to be happy with a representative/well-linked set of 10 million, 100 million, 1 billion, etc. URLs -- which is only a subset of the whole web, hence a 'survey'.
A constrasting kind of crawl would be to focus on some smaller set of sites/domains you want to crawl as deeply and completely as possible. You might invest weeks or many months to get these deeply in a gradual, polite manner.
Use the version 1.x series, not the new version 2. 1.x is less convoluted and easier to use. There might be some http://en.wikipedia.org/wiki/Second-system_effect I think.
I notice on the Internet Archive file that you link it mentions that those files are no longer accessible -- are there similar places that you can grab spidered content?
If you just need fresh web content, it's not hard to collect for yourself quite a bit of broad material in a short period on a small budget, with an open source crawler.
The data from Dotbot might be good, or potential data feeds from Wikia Search/Grub.
Unless you are Google or Yahoo with thousands of servers, you could save some time by only processing pages that have actually been modified.
Creative ideas please!
It takes A LOT of machines to power a modern search engine which serves any real amount of traffic. One key component of an open source search engine would be a sort-of peer-to-peer distributed infrastructure. When I suggested this in an earlier thread, people were quick to point out the liability concerns here... but maybe it could work somehow... but then how do you get people to sign up for it?
That said, I think this is incredibly interesting stuff. I would really love to see open source, peer-served web utilities. For example, I'd want access to many of the components of a web search engine, not just the search results themselves. Things like a language model for spell checking or word segmentation. Or a set analysis tool for detecting synonyms.
It also turned up this error message, forever preserved in amber, so to speak: http://web.archive.org/web/20030802202553/www.permafrost.net...
It is interesting data, but when people build crawlers to index the entire www, especially if the data is intended as an SEO intelligence tool, certain issues arise. Some background on this particular issue: http://incredibill.blogspot.com/2008/10/seomozs-new-linkscap...
* Dotnetdotcom.org
* Grub.org/Wikia
* Page-Store.com
* Amazon/Alexa’s crawl and internet archive resources
* Exalead’s commercially available data
* Gigablast’s commercially available data
* Yahoo!’s BOSS API and other data sources
* Microsoft’s Live API and other data sources
* Google’s API and other data sources
* Ask.com’s API and other data sources
* Additional crawls from open source, commercial and academic projects
In my experience, the single most useful feature (main selling point) of the Linkscape tool is that it reports http status codes (for a price) so SEOs can detect 301 redirects, etc. AFAIK, dotnetdotcom.org has the only free, publicly available crawl data which also includes http status codes. Not sure about Exalead and Gigablast but I am pretty sure the other SEs don't release this information. To clarify: I don't have any proof, and things may have changed, but I've read some intelligent speculation (smarter than me) which claims that dotbot/dotnetdotcom.org provides the majority of the data (especially the unique info, like status codes) for the Linkscape tool.And that index is also open for download, though I haven't looked much into it.
# of Tubes Found Clogged 7188420And: are thoughtful enough to include typos, because (as we all know) some people appreciate the opportunity to find errors). [ "...discussion of girlfriend/boyfriend/husband/wife issues are stickily prohibited." ]
* http://tracker.thepiratebay.org:80/announce
Furthermore my comment will lead to more people, who might not have a DHT-supporting client, adding those open trackers.
Or, at least, whatever doesn't block robots. :/