Build a crawler to crawl million pages with only one machine in just 2 hours
medium.com
medium.com
Not close to a real comparison without the same URLs, but still fun to compare.
Did you happen to track memory usage at all? It would take a while to settle down, for sure.
I'm always interested in the amount of overhead Docker brings to the table. No biggie either way, thanks for sharing these details.
I was hoping to compare resource utilization and performance between 40 Docker instances each with 20 connections vs. 1 process.
It's not even clear whether or not the author actually hit any external websites: In order to have a quick test , I just build a nginx hello page in my cloud server. Then scale it up to a list of 1000000.
It's just a tutorial about how to use docker and celery to build a distributed system. I will have more test about the performance of multiprocess/threads/concurrencies or some other library that supports these technics. In this case, if you don't know how to or even don't want to build a test web server, some big sites like example.com could be your choice. BUT please be gentle to these public sites.
Thanks for your comment.
Of course, it matters what you're doing with the page content, and how you're managing your metadata, and all that.
You are right, asyncio and aiohttp is a great combination. I've used them both before. Though aiohttp cost less memory than requests, aiohttp does't support https proxy. This is the only one that isn't so perfect.
Besides , asyncio is something the same as concurrency of celery. I will have a test which one will have a better performance.
Thanks for your comment again.
Scrapy is good framework for crawler. I used it before, but it doesn't have some features I want. Writing a brand new one is easier for me.
Of course , I am open to compare the special performances of these two crawlers. Welcome to have a discuss more technical details here.
Thanks for your comment again.
On the other hand, if you are building a massive crawler you want to split the crawling and parsing into two separate functions and do the network I/O in golang/c and do the parsing with a Javascript headless browser like phantom.
I don't really see any reason to use python unless you haven't learned golang (the world's biggest crawler's own language).
However that ram usage though...ugh
You could get the same performance within 600mb by using 2 processes each running 20 threads.
But I guess hardware is cheap.
There is no avoiding the GIL within a single Python process (even with asyncio IIRC, though I've been using JS lately). Multiprocessing is usually the most efficient way to execute I/O intensive, independent parallel operations. Of course you can also run threads within each process.
I do wonder where the 300mb memory is coming from. Surely it can't all be python interpreter? It doesn't look like he's importing 300mb of modules, unless MongoClient really is that big. In that case he could create a separate worker process for persisting data, and only that worker process needs to load the MongoClient module.
One explanation for the memory overhead might be conntrack tables within the network namespace of the container. However I would expect that conntrack table to be on the host, where SNAT is performed. As an aside, the default Docker networking configuration is really not well suited to concurrent network requests, whether inbound or outbound. If you can avoid NAT (and therefore a conntrack table), that is preferable.
This stack could also benefit from tuning some kernel parameters, both within the containers and on the host. Great blog post with details: https://blog.packagecloud.io/eng/2017/02/06/monitoring-tunin...
Thanks again.
Whats your point? 20 threads will still run per GIL, and assuming a dual core cpu, 2 processes x 20 threads each will still run 40 workers.
Thanks for your comment.
While it's on topic.. anyone have any other recommendations for web crawlers? I'm particularly interested in finding unique identifiers (phone numbers, emails) and their contexts on gov-owned websites for a project.
I know the benchmarks cannot really be compared, since one involves a queue, and mongo, while the other does not. But, it seems like a prime use case for async.
from multiprocessing import Pool
from requests import get
urls = 1000 * ['http://localhost/hello']
def scrape(url):
return get(url).text
p = Pool(40)
results = p.map(scrape, urls)
~2.2 seconds on a dual core 2.2ghzI've used multiprocessing/threads/geven/asyncio before. And I will have a full test with these libraries.
Thanks again!
There are lots of factors involved which can completely skew benchmarks, for example, if you were scraping an average 10kb response instead of 'hello world' you would automatically be limited to 100req/s on a 10mbit pipe.
Thanks for your comment again. welcome to discuss more technical details about it.
- What is the relevance of Docker here? I'm pretty sure that celery+rabbitmq are enough to do a distributed scraper...
> and learn how to use docker and celery
Seems the OP was learning Docker at the time? I think it just comes down to the tools you're comfortable with.
It was just shoehorned.
The crux of a project such as this is maintaining a connection pool and managing it efficiently.
Also respecting robots.txt which the author barely mentions.
This is a "tool looking for a problem" kind of post.
Generally I'm getting fond of containers as a mechanism to encapsulate deployments e.g. in Python which have a lot requirements and which I've found finicky to make portable.
Full disclosure: I do something even worse, have containers which pull updates when I like with deploy keys, and run Celery etc in a virtualenv in the container... :P
The latter feels truly shameful but it does make it easy to keep the project contained even when running outside a container...