Show HN: I wrote a tiny Python-based HN crawler with scrapy
github.com
github.com
It runs on cron, scrapes HN once every 30 minutes, and sorts it according to the Wilson score confidence interval for a Bernoulli parameter (the one from the 'How Not To Sort' article here: http://www.evanmiller.org/how-not-to-sort-by-average-rating.... )
See the results here: http://hn.elijames.org/
I've found that the HN articles I like tend to have high vote counts as compared to comment counts. Mostly this makes sense: an article from Scientific American has more votes than comments, whereas an article on Why PHP Sucks (which I find boring) would be highly controversial, which would have higher comments.
So I've treated comments as a negative signal. Mostly this works - I find myself skimming more effectively from my page than from HN itself.
Of course, by commenting here you are actually pushing this article downwards on your own page ;)
HN is slow enough as it is ;)
The comments to the right of the page all complain about the stability of the service and it all was just too frustrating to use while I wanted to get something up fast.
Also, keep in mind that the tool only scrapes one HN page per crawl (the home page). It's up to the consumer to be polite, but I hope that this tool is used responsibly.
HN is built using a homebrew stack, it's a miracle it performs as well as it does.
What you could do is cache the results and point your tool at the cache, that would already be much better. After all, if your conclusion is that the HN api is broken then maybe provide a better API rather than a tool that hits the source?
Surely it'd be sloppy oversight if the HN site couldn't be hosted on a homebrew stack?
People seem to imagine you need a constellation of mongo and load balancers and everything else just as a baseline for a helloworld web-app.
Paul rocks.
http://news.ycombinator.com/item?id=2120756
And comparing HN to a helloworld web-app is simplifying things a bit, the fact that it is sparsely designed does not mean there isn't a significant amount of work done under the hood.
One thing bothers me with your code tough, the following could be replace with built in Scrapy-tools.
Edit: The code got wrongly formatted in here, check it out on PasteBin http://pastebin.com/2QzWgWxN
# Without using BeautifulSoup.
for item in hxs.select('//td[@class="title"]/a'):
news_item = NewsItem()
news_item['title'] = item.select('text()').extract()[0]
news_item['url'] = item.select('@href').extract()[0]Is there an advantage to BeautifulSoup or is it just the tool that you're most comfortable with?
I'll be sure to mention this in the docs.
Edit: changes have been posted. Thanks for the suggestion!
It looks like he's using methods from an existing Python CLI interface project to scrape the site. Definitely a creative way to go, although I personally think starting from scratch wasn't too bad either.
btw you have a hardcoded path in there /Users/mvanveen/root/dev/news/out/
Edit: for personal use there is not much to take into account other than what you mention. It is not illegal to "see" a webpage, it's more of a question of what you do with the content. In fact, one could argue that scraping is better than refreshing in your browser (as you'll hit them with fewer GET requests if done properly - as you won't download js/css dependencies).
It's pretty tiny, feel free to fork/contribute!
1. Crawl all the links
2. Check for directory listings on all path combinations (those discovered in href's and src's)
3. Check robots.txt for any other discoverable pages
4. Brute force expected or possible URLs
5. Try and parse more links from any javascript (probably hard)
That's all I can think of.