Fast scraping in Python with Asyncio
compiletoi.net
compiletoi.net
`yield from` and asyncio makes everything much more simple.
Depending what you want to scrape, this might work better for you.
https://www.mashape.com/stremor/stremor-content-extractor-fa...
It will scrape articles even if they are in multiple divs. It even works on magazine layouts like TheVerge.com that doesn't work in Readability (py, api, or js).
If you need support down to Python 2.6 and similar functionality it's worth checking out:
Python used to be the clean and simple language for me, but all those ”hacks” to get async right just makes the final code ugly. The cleanness of Go wins me over instantly.
Even though the two languages are not directly comparable, the foundations of Go seems to me much better and solid. It reminds me of Rob Pike's essay “Less is exponentially more” http://commandcenter.blogspot.ca/2012/06/less-is-exponential...
I don't see anything "hacky" to get async right. It's a simple library. It's just not first class (then again, no reason it should be, isn't it not first class in Rust too?).
Those methods (gevent and eventlet) are the default and typical way of building concurrent python services.
That article conflates threads, shared state and explicit IO awareness and dispatching. Those can be different in each case.
Python doesn't have isolated heaps like Erlang but it is still possible to have lightweight threads with a queue and not share by convention data between lightweight threads.
Having or not having @coroutines or yields or deferreds or inlineDeferreds doesn't make a difference in that case.
In general, a set of jumping callback and errback functions are a lot worse concurrency mechanism than threads.
A callback chain is a poor man's threads, it is just spread non-locally across the whole code without a way to properly synchronize.
Threads done right, should make it easy to organize data locally. This is quite the opposite of what that article claims. A set of spread around callback functions, do the opposite.
Whether yields or inlineDeferreds are used are almost orthogonal to the above. All they do is help you localize data (which callbacks don't) but now you have to explicitly care where and when you do IO. You might not even care in real life but they force you to. That is stupid.
Oh and library fragmentation. There are now 3+concurrency frameworks for Python. They all look different. And because of this bubbling of IO primitives up to the surface of the API, make it almost impossible to share libraries between them.
One of the biggest advantages is the fact this is not multi-threading, so context switches can only happen at precise points (yield from). This makes it easier to reason about the code.
Copypasta of dead comment:
nashequilibrium 13 hours ago | link [dead]
rp = robotparser.RobotFileParser()
rp.set_url("https://news.ycombinator.com/news/robots.txt")
rp.read()
# Reads the robots.txt
rp.can_fetch("*", 'https://news.ycombinator.com/news')
>>>> True