HNHacker News
TopNewBestAskShowJobs

binux

135 karma · joined November 17, 2014

submissionscomments
binux··on Show HN: A Python Spider System with Web UI
https://gist.github.com/binux/67b276c51e988f8e2c31
binux··on Show HN: A Python Spider System with Web UI
I'm working on a benchmarking suite https://gist.github.com/binux/67b276c51e988f8e2c31 and meet some problem...

pyspider comes from a vertical search engine project. we have two issues:

- 100+ websites, they may change the template or down sometime. We need a dashboard to monitor the changes and the fails.

- update in 5 minutes, when the website updated, we need follow that in 5 minutes. We are using a update time from index(list) page to tell the changed pages. And pages should been updated after about 30 days in case of we missed something. A powerful scheduler is needed.

obviously, I hadn't got the right way to do so with scrapy. I'm not very familiar with scrapy. So I can't say something pyspider can do but scrapy not.

binux··on Show HN: A Python Spider System with Web UI
Images not downloaded default. Both the fetcher and the phantomjs proxy is totally async.
binux··on Show HN: A Python Spider System with Web UI
I want to make it a http proxy in the beginning. But I found it hard to do so. Then I post every to it, but haven't change the name.

But it works like a proxy, that any request with `fetch_type == 'js'` would be fetched through phantomjs and the response back to tornado_fetcher.

binux··on Show HN: A Python Spider System with Web UI
The fetcher fit you already...
binux··on Show HN: A Python Spider System with Web UI
the architecture of pyspider: http://blog.binux.me/assets/image/pyspider-arch.png

And yes for centralized queue which is in scheduler. It's designed to satisfy about 10-100 million urls for each project.

scheduler, fetchers, processors are connected with rabbitmq(alternatively). Only one scheduler is allowed. But you can run multiple fetchers or processors as needed.

binux··on Show HN: A Python Spider System with Web UI
sorry :(
binux··on Show HN: A Python Spider System with Web UI
To make it more flexible and easy to reuse? I have implemented most features I need now.
binux··on Show HN: A Python Spider System with Web UI
http://demo.pyspider.org/debug/js_test_sciencedirect is a sample for this.

There is a phantomjs fetcher that can render the page as WebKit did. Furthermore, you can have some JavaScript running before/after page loaded to simulate a mouse click.

binux··on Show HN: A Python Spider System with Web UI
I have "organize the code using a single top-level package".
binux··on Show HN: A Python Spider System with Web UI
Yes, the scheduler, fetcher, processor is stand alone here, they are running in different process. But they are sharing some common libs. I haven't made a decision how to put them into a single package, and running together.

Any advice or project that I can refer to?

binux··on Show HN: A Python Spider System with Web UI
agree
binux··on Show HN: A Python Spider System with Web UI
Currently, yes and no.

pyspider is running original python code, something like portia is a code generator (Apologize if I'm wrong, I have not use it). So it can been made as another WebUI module.

But for flexible, I have no idea how to make it right currently. So, We have a css selector helper, but no plan for a complete tool.