Because I already have a powerful distributed architecture. I was curious about the architecture of pyspider.
For example, how the queue is handled? Is it centralized? Is there a server managing it?
For example, how the queue is handled? Is it centralized? Is there a server managing it?
And yes for centralized queue which is in scheduler. It's designed to satisfy about 10-100 million urls for each project.
scheduler, fetchers, processors are connected with rabbitmq(alternatively). Only one scheduler is allowed. But you can run multiple fetchers or processors as needed.
phantomjs phantomjs_fetcher.js
and using it as proxy? The setup instructions are a bit unclear on this.But it works like a proxy, that any request with `fetch_type == 'js'` would be fetched through phantomjs and the response back to tornado_fetcher.