How to write a crawler
emanueleminotto.it
emanueleminotto.it
- There is no mention of implementing a crawl-delay. You should always wait for several seconds (better yet, a minute) between requests to the same host.
- Do you follow redirects when requesting the robots.txt? You should! Some sites send you a redirect to a different URL even for robots.txt. In most cases it is just a slightly different hostname, like www.domain.com instead of domain.com. But it can redirect you to somewhere completely different in some cases.
- You probably don't want to crawl anything that ends with .jpg, .gif, and definitely not something like .avi, .wmv or .mkv. There are a LOT more file-extensions that you'll want to ignore.
I agree with cmiles74 that using a database is probably a bad idea. For a sizeable crawl (say a billion pages) this database will get pretty damn big. I doubt that you will be able to get decent performance out of anything with "SQL" in its name for such a use-case, unless you throw a ton of hardware at it. Building your own specialized solution for this would probably be a lot faster and less resource-intensive.
This a thousand times.
I learned this lesson AFTER getting several angry emails from Admins and getting outright banned from one site for not having any delay between the requests in the first crawler I built.
It'd probably be a good idea just to whitelist filetypes that you actually want to crawl, rather than trying to blacklist all the ones you don't.
We extract other things like the Opengraph image and social media links for author and store them independent of the content.
Speaking from experience its far better to lock your crawler down from the beginning then to do it after you kill someones site. I did expose a few slow pages for some sites, but not before crashing the whole thing.
The web is probably bigger than you think, Google says "when our systems that process links on the web to find new content hit a milestone: 1 trillion (as in 1,000,000,000,000) unique URLs on the web at once!" (July, 2008)
You might consider just crawling certain parts of the web, or using a search engine (api, like Yahoo! BOSS) to gather relevant links and crawl from there, using a depth limit. Just an idea.
Any numbers?
Now, depending on how you prune your crawl to get rid of "uninteresting" content (such as infinite calendars) and how you deduplicate the pages you find, you'll come up with vastly varying estimates of how big the visible web is.
Edit: on a side note, don't crawl the web using a naive depth-first search. You'll get stuck in some uninteresting infinitely deep branch of the web.
At some point querying the db to check what URLs you have can become quite heavy
Most of them use the same data-structures (e.g. b+trees) and especially the newer ones turn into copy-on-write systems like e.g. CouchDB
I guess if you want to have backups of that table, you probably would like to keep it at a manageable size, but why not just save the html blobs in a separate table?
The reason you don't want to hit the disk with each link (as both MySQL and PostGres usually do, barring caching) is that there can be hundreds to thousands of links on a page. A disk hit takes ~10ms; if you need to run hundreds of those, it's well over a second per page just to figure out which links on it are unvisited. Accessing main memory is about 100,000 times faster; even with sharding and RPC overhead for a distributed memory cache, you end up way ahead.
The reason to write the crawl text to an append-only log file is because disk seek times are bound by the rotation speed of the disk, which hasn't changed much recently, while disk bandwidth is bound by the rotation time of the disk divided by capacity, which has gone way up. So appends are much more efficient on disk than seeks are.
For example, when a tracking parameter is added to any URL within the site:
http://example.com/?cid=104484&pid=12002348&ref=1294902
http://example.com/?cid=104484&pid=12002348&ref=1294904
http://example.com/?cid=104484&pid=12002348&ref=1294905
http://example.com/?cid=104484&pid=12002348&ref=1294906
You can quickly get to billions of permutations for a single site. The canonical tag solves the problem when it's there, but I still haven't seen a simple solution to the problem when it's not.
If you are really sure that these pages are the same, try checking the body content (if two or more pages have the same MD5 of the content, those pages are the same) or look for a form that generate those URLs.
If that's too slow/space-intensive, try a bloom filter.
Some of the "little things" matter much more than your content analyzer or HTTP parsing - DNS performance and multi-homing being just a few that can have drastic effects.
Just as an example of how complex it gets, here's a brief overview of some of the features all crawlers should take into account: http://en.wikipedia.org/wiki/Web_crawler
It took a lot of man hours to build the content extractor we use for our search engine.
https://www.mashape.com/stremor/stremor-content-extractor
Also because not every site has a SiteMap, or well linked site structure you may have to turn to Social like FB and Twitter if you want to get everything.
> id is an incremental value, I choose 11 as a length
for this primary key but this value is defined
by the number of pages you’ll need to index
This is a bit confusing. You'd generally be better off with INT UNSIGNED as it doubles the range for auto-increment columns.Also the visited field would be better to be represented by a timestamp, choosing the right datatype does matter in large tables.
- by setting CURLOPT_ENCODING to '' you don't have to worry about (un)gzipping as curl ll do this for you (or it should)
- it might be a good idea to use url hash as url id (f.e. crc32)
- you should check content-length and content-type to avoid downloading huge files
btw your coding style is very disturbing. there shouldn't be spaces before or after ->
Multi-threaded. Built-in page delay (200ms default). Does HTTPS, headers, POST, cookies, follows robots.txt .
I prefer it to nutch for small to medium sized jobs.