Crawling the web at $2 per million pages
80legs.com
80legs.com
However, if you want to run crawling on your own infrastructure, I recommend Scrapy (http://scrapy.org/), a python crawling framework introduced on HN last year. Scrapy solves some of the more time-consuming problems involved in writing a crawler from scratch (multiple simultaneous requests, pipelined processing, raw caching, duplicate URL filtering) and comes with nifty development and administration tools. More importantly, it has an active and helpful set of core developers and good documentation. I am comfortable with both Python and Java but I chose Scrapy over 80legs because I can crawl for free on the machines I already have and I can afford to spend more time crawling from a single IP compared to 80legs which will let me crawl much faster but isn't free. Also, with Scrapy my bot can be 'naughty' - 80legs jobs obey robots.txt and limit the crawling rate per domain.
When my crawling needs outgrow my infrastructure, I am going to look at 80legs again.
You'll also be able to mashup third-party apps and libraries into your own code. This is something we hope to have in 2-3 months.
Yes, Java isn't as sexy anymore. But running your code on 50,000 computers sounds sexier to me than a REST API.
http://parselets.com/ + mysql + sphinx
I think 80legs can do some level of parsing now, but Parsely seems like a great way to describe the analysis for targeted crawls on predictable data. If it supported the combination you've described it would really open up some cool possibilities.
Hopefully guys that write parselets will be interested in becoming developers for the upcoming 80legs App Store. They'll be able to sell their parselets and earn 100% recurring revenue from 80legs users using their parselets.
http://en.wikipedia.org/wiki/SETI#SETI.40home
"Any individual can become involved with SETI research by downloading the Berkeley Open Infrastructure for Network Computing (BOINC) software program, attaching to the SETI@home project, and allowing the program to run as a background process that uses idle computer power."
Of course, the pricing structure would be difficult, as would setting up the infrastructure for it. My (uniformed, by-the-seat-of-my-pants) best guess would be that subverting the current folding method of jobs would be the best way to do it. Accept jobs, send out to client computers, clients get paid based off of how many they complete. The problem arises from: sandboxing the jobs, to make sure they can't steal client info, and making it relatively easy to program jobs, while still allowing users to do various things to increase their performance (like using GPUs and the like).
It may be a good start up project. But I'm not that determined to do this unfortunately.
On the other hand, if an existing company branched out into this, something like Amazon and MTurk, I could see it working. A company which already has a user base, which has software installed on clients could leverage this as a "Oh yeah, here, we think this is cool, try it".
You've got a URL, headers, and body content. Just extract what you want and crunch.
Plura supports desktop and web-based games.
If the game is hosted on a website, like a Flash game, the
developer only needs to include 1 line of iframe code. We
will soon be releasing a Javascript API for dynamically
controlling how this iframe is loaded, giving the developer
control over starting, stopping, and controlling CPU usage
in Plura. The iframe loads a Java web applet, which runs
completely in memory. This applet is forcibly restricted
from accessing the user's computer by the sandbox model
provided by Sun.
See also -
http://pluraprocessing.com/developer/index.htmlSo far as I understand, they have multiple models. One affiliate model is "Plura for Java Applets", where-as another is "Signed Java Applets"
I imagine there may be fewer options for unsigned applets, leaving the developer with less potential revenue every month, where as desktop application developers and signed java applications are left with the providers who don't need signed applications.
That said, I agree the Java dialog is ugly and scary ;( If I were an Affiliate, I'd want to avoid it. It breaks the user-experience of your site into some gaudy and jarring, not to mention unbranded and unrelated to the information the user is after.
If you're going to make a claim of 50,000 computers I'd want to see what you mean by that somewhere.
The reason why they didn't want it to be too expensive (ie. 5 times as expensive) is (1) competitors can easily equilibrate and steal market share if their idea works b/c of the economic inefficiencies in their pricing model and (2) this game plan is more a game of dependence rather than up front profits, so it makes sense to take very little profit up front to get user traction.
Side note: AWS guys have asked if we use them. We said "No, we'd be losing money then and wouldn't be able to scale to our size with you."
"Your parseLinks() and processDocument() methods must complete within a total 10 seconds per document processed"
... as a limit on processing leaves room for competition. One advantage of the BYO-Cloud solution is that you can pay for more intensive processing of the crawl.
Actually, the cost is not the biggest issue with the cloud. If you're talking about large-scale crawling, AWS will not adequately scale. You can't get enough nodes or enough bandwidth.
And it's the bandwidth aspect that makes web crawling not feasible on AWS. Yes, you have a few thousand nodes, but they're all going through a handful of external IPs, which will cause serious performance issues.. the worst case is that you'll get blocked entirely from the sites you're trying to crawl.
In other words, the bandwidth is not parallelizable on the cloud.
1. Yes, each instance can have its own IP, but by default, each account is limited to 5 IP addresses.
2. You can increase your limit, but my guess is that it's difficult to do so. You have to put forward a special request and have it approved.
You're right that blocking may not be a big issue, but crawling several different domains quickly will be hard.
Just so you know, we haven't encountered anyone doing large-scale crawling that considers AWS or the cloud in general to be a realistic option. The biggest reason is still the cost.. the outbound transfer rates just don't make sense at scale.
You are limited to some number of instances (20, 50?) and yes, you have to fill out a form to get more. The previous example with animoto shows how far you can go. I would wager that finding the funding for a large # of instances is more problematic than getting the approval.
I don't see why crawling several different domains quickly will be hard? There shouldn't be any difference between a bunch of instances on EC2 and a bunch of machines in a data center, from a technical point of view.
As far as the cost argument goes, of course I agree with you. If you can project a high level of CPU/bandwidth usage for an extended period of time then of course you should buy dedicated servers.
The only argument I was trying to make was that it is completely possible to do crawling on EC2 or any other cloud provider from a technical point of view, the only limitation is cost. I see the advantage of utility computing is that it offers a cheap way to handle bursty traffic, which you may certainly run into if your server utilization projections are off? I don't think you should use it as your primary set of servers if you can project some large volume of traffic.
I'm not arguing that it's impossible to do crawling on the cloud. I'm saying it's near-impossible to do it on a large-scale on the cloud. 3500 instances is pretty good, but will still be an order of magnitude slower than what 80legs is capable of.
Now, if you show me someone that has 10,000+ instances on the cloud, I may agree with you!
Are you supposed to implement your own cycle detection?
If not, how deep is the cycle detection that 80legs offers?
--
There are plenty of HoneyTrap[1] OSS projects which will quickly rack up lots of $$$ if the 80legs spider is backlisted.
For those who don't know, these projects create deeply linked pages, and sometimes create infinite cycles. They are trying to hinder spammers, but may hinder 80legs too.
--
[1] http://www.davidnaylor.co.uk/stopping-bad-robots-with-honeyt...
1. For each user crawl, we only allow the same url once, so any simple loops that involve the same urls are eliminated by this process no matter how large the loop
2. For more sophisticated "spider traps" that work with different urls and domains, they can have a limited effect on your crawls. Because of our per-domain rate throttling, the worst these traps can do is add a few cents per day to someone's crawl.
I'd pimp for Open Google instead.