Any good api to scrape HN other than this?
github.com
github.com
User-Agent: *
Disallow: /x?
Disallow: /vote?
Disallow: /reply?
Disallow: /submitted?
Disallow: /submitlink?
Disallow: /threads?
Crawl-delay: 30
So nominally you should feel free to set up a scraper that crawls one non-disallowed resource every 30 seconds.rp.set_url("https://news.ycombinator.com/news/robots.txt")
rp.read()
# Reads the robots.txt
rp.can_fetch("*", 'https://news.ycombinator.com/news')
>>>> True
It would be good to have a way to download ALL your stuff. Ask PG?
What is a safe limit to crawl this data, if I have to absolutely need that data? 30 mins between users? 1 hour between users?
There is rarely a need to scrape HN directly, but if you do make sure your bot is polite (especially with respect to rate limits).
Disclaimer: It's my own blog
edit: Uses HNSearch, so it doesn't violate the robots.txt and can be crawled faster
I'm not sure what you're trying to do though. I used beautifulsoup because I couldn't get lxml working on BB10, but if it was switched to using lxml it would be much faster.
Disclosure: Founder of Diffbot here.
You can use the twitter API and read from there
Certainly, if I have had access to it I know I could do some pretty useful sociology on HN's audience (= the pool of startup hire material).
If you have a fabulous idea for how to use the data contained on this site, I'm sure everyone will be impressed and interested to see it.
The current situation (PG and friends optimise a basic but very accessible website, and a handful of third parties build APIs on top) is much more manageable.