How to crawl a quarter billion webpages in 40 hours (2012)
michaelnielsen.org
michaelnielsen.org
Given that there are plenty of existing, open-source crawling engines out there, I don't see how this decision is really accomplishing anything. Concretely, Apache Nutch[1] can crawl at "web scale" and is apparently the crawler used by Common Crawl.
There’s a more general issue here, which is this: who gets to crawl the web?
This, to me, is the most interesting issue raised by this article. In principle, there's no particular reason that, say, Google, has to dominate search. If somebody clever comes up with a better ranking algorithm, or some other cool innovation, they should be able to knock Google off their perch the same way Google displaced Altavista. BUT... that's only true if anybody can crawl the web in the first place... OR something like Common Crawl reaches parity with the Google's of the world, in both volume and frequency of crawled data.
The first scenario is definitely questionable. Sure, you can plug the Googlebot user agent string into your crawler, but plenty of sites are smart enough to look at other factors and will reject your requests anyway. (I know, I used to work for a company that specialized in blocking bots, crawlers, etc.)
It really is a bit of a catch-22. Site owners legitimately want to keep bad crawlers/bots from A. consuming excessive resources, and B. stealing content, from their sites. But too much of this will lock us into a search oligopoly that isn't good for anybody (except maybe Google shareholders).
I didn't say anything about Google's value prop. My point is, whatever their value prop is, it's built on top of their ability to crawl the web at mass scale, and very quickly. So anybody who wants to compete with Google, by being better at "ranking pages, removing spam, personalizing it to the viewer," or whatever, will need to be able to do crawl in a similar manner. IOW, crawling is part of the "price of admission".
Yes, but I think the comment you responded to were saying that it is an insignificant part these days.
Maybe so. In which case, I'd say I disagree. Great crawling ability is definitely a necessary, but not sufficient, condition for building a competitive search engine. And while the technical aspects of building a large scale search engine have been at least partly trivialized by OSS crawling software, elastic computing resources in the cloud, etc., what is at issue is the possibility of site owners blocking anybody who isn't (Google|Bing|Baidu|etc).
In this context, that's my concern: being blocked from crawling, if you're not already on the "allowed" list. Hence my reference to the question quoted above, from TFA.
In one aspect, it was literally their beginning, from which they were able to build a business and expand.
In another, much more on point aspect, it underlies the majority of their services, either directly of a few steps removed.
Like the foundation of a house, it may not always be the most visible aspect, and it may be taken for granted, but its contribution to the integrity of the whole can't be underestimated.
No, I don't. I didn't say anything even remotely like that. What I'm saying is that crawling is a prerequisite for all of the "... best treatment of that data to serve ads, personalise experience ..." stuff. That's why this topic matters: because if people aren't allowed to crawl (or can't do it effectively for technical reasons, but that's a subtly different issue), then they can't do anything else.
Just to be clear, I am most emphatically not saying that just building a crawler is enough to compete with Google. What I'm saying is that you can't build a competitive search engine without that "0.05%" bit because it's required to enable the other 99.95%.
Another possible approach is for sites to be self-indexing, with sharable indices, and validation (and penalties) for deceptive practices (either false term inclusion or exclusion).
The simple existence of a websearch and index protocol would wipe virtually all of Google's present value.
Both models are fairly isomorphic, modulo the crawl / search locus. The key is on agreed data structures, practices, and access.
I came here to state just this. What would code for the crawler in the article do that premade DDoS bots don't already do? (already does that task better too) Even in 2012 when the article was published, there were plenty of open source crawlers AND tools designed specifically for "burdening websites with traffic" that were available.
User-agent: google
Allow: /
User-agent: *
Disallow: /
well, yes it's good behavior to actually respect it, but well I've seen such robots.txt already which makes it really painful to create a competing search engine.Edit: In my scanning, I have to confess that wikipedia's robot file is the best. Fairly heavily commented on why the rules are there. https://en.wikipedia.org/robots.txt
The only problem is: nobody would care, people will still use Google because it's all they know. Some people even ignore the fact that underneath they use an operating system and a browser.
Don't you think that's what they used to say about Altavista?
250.000.000 pages come in at $580
There are 1.8b websites according to http://www.internetlivestats.com/total-number-of-websites/
Lets say on average each site has 10 pages (you have a dozen of huge blogs vs tens of thousands of onepagers), that would put the number at 18 billion pages.
Following that logic would mean the total web is 72 times larger than what was scraped in this test.
So for a mere $41.760 you too can bootstrap your own Google! ;-)
That 10x average seems to be a bit off considering our data, which is of course spotty since it's crawled by a third party.
But to give some numbers, in one of our experiments we filtered web sites from the archive for known entities and got 307,426,990 unique URLs that contained at least two of those entities (625,830,566 non unique) and in there were only 5,331,272 unique hosts. That archive contains roughly 3 billion crawled files (containing not only HTML, but also other MIME types) and covers mostly the German web over a few years.
There are a lot of hosts that have millions of pages. To name a few: Amazon, Wordpress, Ebay, all kinds of forums, banks even. For instance, www.postbank.de has over a million pages and they were not re-crawled nearly that often.
I assume a lot will be incorrect links and automatically generated pages as is often the case.
> It must be noted that around 75% of websites today are not active, but parked domains or similar.
So actually more like 0.5B websites. Feels quite tiny. Seems most activity online really is behind walled gardens like FB.
So the majority bandwidth will be videos. But as for unique users - social media and Google does make up the majority.
Try crawling ONE wordpress blog with less than 10 posts and you will be surprised just how many pages there are due to pagination on different filters and sorting options, feeds, media pages, etc.
I think the cost of fetching and parsing the data is much less than the cost of building an index and an engine to execute queries against that massive index.
Download latest list of urls from https://www.verisign.com/en_US/channel-resources/domain-regi...
tail -n+149778267 urls.txt | parallel -I@ -j4 -k sh -c "echo @;curl -m10 --compressed -L -so - @ | awk -O -v IGNORECASE=1 -v RS='</title' 'RT{gsub(/.*<title[^>]*>/,\"\"); {print;exit;} }'; echo @;" >> titles.txt;[1] https://news.ycombinator.com/item?id=4367933 [2] https://news.ycombinator.com/item?id=10865568
For example, what if you have a web site that generates a thousand random links on a page, which all load pages that generate another thousand random links, to infinity?
https://stackoverflow.com/questions/5834808/designing-a-web-...
You can do a lot of other things too, but the above already takes most of the sting out of such traps unless you use a huge number of domains.
Something I don't understand, how can a webpage "block" a request? is there a way from a basic http GET request to tell if it has been issued by a browser or something else?
Lot of webpages are generated dynamically. Consider something like "https://news.ycombinator.com/item?id=17462089"? does the crawler follows URL with parameters? what if parameters identifying the page are passed in the http request?
Sure. In the simplest case you just look at the User-Agent header. Your browser will send one thing, and something like GoogleBot will send something else. Now, if you're a website owner who wants to block bots, you can't just depend on that, because somebody writing a bot can trivially put any string in that header that they want. But there's a lot of other ways you can tell. A simple way to discriminate somebody who's just using curl or wget, for example, is to serve a page with some javascript in it and check if the javascript is executed or not. Usually you'd do something like this from a proxy that sits in front of your actual content, and throw out subsequent requests from a UA that fails the check. Of course identifying the UA consistently is yet another challenge.. if the thing handles cookies properly, you can use a cookie. You could try going by IP, that that's dicey in various ways. Etc., etc., yada, yada.
All in all, there's a constant arms race going on between the companies that want to block bots / crawlers, and the people who want to crawl/scrape content. The techniques on both sides are constantly evolving.
86 pages per machine is not very performant at all, just very simple parallelism will do.