Hacker News Dataset Update October 2016
aaron-hoffman.blogspot.com
aaron-hoffman.blogspot.com
I had a similar job I needed to do a few months ago and used AWS lambda to massively parallelize the work.
I was able to bring down what I estimated would take my laptop 30 days down to about an hour by sharding to a ton of small instances.
Might be worth a look if you plan on updating this with any regularity.
I have created an updated copy and made it available for download.
(This is the last 10MM entries, I can add the rest if people are interested.)
From the commit notes in that repo, the only changes from the initial release in 2014 are "minor README updates."
A more damning reason is that the official HN API in its current state is worse than the API it replaced! The Algolia API (https://hn.algolia.com/api) is still active, and can retrieve data with 1000 entries per page (vs. 1 at a time for the official API), and can also retrieve the comments plus text of a submission thread in a single HTTP request (the official API requires the user to perform a HTTP request to retrieve the text for each comment in a thread)
I was unaware of the algolia api, that will help for future tasks I'm sure. Thanks!
I've been using it for a project (collecting video lectures for https://www.findlectures.com) and it seems to work pretty well and seems to keep up to date.
See my comment in another thread on why this did not work.