Downloading All of Hacker News Posts and Comments
shitalshah.com
shitalshah.com
Not sure two giant JSON files is the best format for this, but I used jp: http://www.paulhammond.org/jp/ to browse through it.
* Discover geek friendly WordPress themes and plugins by analyzing CSS in stories posted on HN.
* My pet EVIL project: Extract self identifying statements from comments and create profile for HN users :).
* Find out abandonment rate of veteran users.
* Find out undiscovered great stories that didn't got in to frontpage because algorithm deficiency in HN (for example, get links posted by people with 10K+ karma but without upvotes.
I'd bet that you're not the only one who thought of that. And for less money I'd bet that someone has beat you to it.
If I find some free time I might download this and run some extracts as sytelus suggested.
Previously, Max Woolf worked on this http://minimaxir.com/2014/02/hacking-hacker-news/ via https://news.ycombinator.com/item?id=7291531
Fun fact: At 10,000 calls/hour and 1,000 objects per request, you can download all stories AND comments in less than an hour.
As an aside: I tried to use ML on Hacker News stories and have had exactly zero success. (i.e. the predictive models are not statistically significantly better than the NIR)
https://github.com/jaredsohn/hacker-news-download-all-commen...
I have a few other things in the works related to the hnsearch API (had them almost ready for release a few months ago but then I got distracted); this post is persuading me to finish them up soon. :)
I'm quite excited about this data release- there are many interesting ML models that could be trained here. One I hacked on previously on my own data was a comment ranker which uses a ranking loss to rank comments consistent with their observed order in data (which roughly reflects their number of upvotes, I believe). In principle I think it could be converted to a browser extension that gives a score for how well received your comment will be conditioned on the parent comment, as you write it in the text box. One of the main issues I ran into when I hacked a bit on it was space complexity, since you need to keep all the word embeddings (usually on order of 50-200D / word) around in memory of the extension, and there are many words.
P.S. Thanks for the data, you saved a good amount of our time.
Also, it's great to be able to filter through and find all of the highly-upvoted stories that I've missed out, and programmatically push them to Pinboard. Thanks for this.
> Content Category: "Suspicious"
Which is an odd catch-all category. It may be a keyword match from the domain. That sucks. Anyone got the github link?
To give you an example, this story itself was posted on last Friday evening PST (https://news.ycombinator.com/item?id=7825146). It got just one upvote and 0 comments. Exact same story with exact same page content on exact same domain was posted on Monday (today) afternoon PST and it got 80+ upvotes, 30+ comments and got on frontpage for more than 6 hours!
A lot of stories are like this. HN ranking algorithm isn't perfect.
No. Most stories in "new" get zero comments and are not spam.
Deleted comment
Edit: scratch that, can't get any peers/DHT response - can you check if you're announcing that torrent at all?