Ask HN: Does HN have an API and if not what's the etiquette for scraping?
I guess this question could open a more general discussion on the etiquette of screen scraping. So feel free to answer in general terms too.
I guess this question could open a more general discussion on the etiquette of screen scraping. So feel free to answer in general terms too.
I've worked with the API and it does a fantastic job. For what you're looking for , I think this API call will do the job:
User-Agent: * Disallow: /x? Disallow: /vote? Disallow: /reply? Disallow: /submitted? Disallow: /submitlink? Disallow: /threads? Crawl-delay: 30I think a year ago this wouldn't have been an issue if you spaced your requests out by several seconds, but during the daytime (United States daytime, that is) this site is hammered. Yesterday the site was brought to a crawl (no pun intended) by just from everyone commenting on that DDG/Google article.
If you spaced your requests out by 3 or 4 seconds and only crawled during the US night time then the servers shouldn't be affected too much. That doesn't account for the data transfer you use up though. Someone has to pay for that and there are already several sites (that I can think of off the top of my head) that crawl HN on a regular basis.
If it were me, I'd cache the results of an api call to HN with a last-retrieved timestamp and check against that for requests > 24 hours to retrieve again.
I like the look of http://api.ihackernews.com/ and so will poll this rather than scrape - thanks for the tip hardik988 and for all the other comments too.