Analyzing Articles on Hacker News Using NLP
nbviewer.jupyter.org
nbviewer.jupyter.org
Why? There's still a lot of information in the posts that didn't receive significant interest, and unexplained filtering here seems more suspicious than anything.
I also wonder how many unique voters are present, especially regulars and those who vote before a link hits the front page. I bet there aren't that many. And I bet they are mostly INTJ. That is, content seems curated/controlled by a handful of users (ie, bubble). What can be done to buffer against bias? How can submissions automatically be surfaced/tested (shown on the front page to patterned/known users) even before receiving any votes?
I've always thought some percentage of the front page should be dedicated to (randomly chosen, though slightly weighted) new submissions. Some percentage of the page some percentage of the time to some percentage of users able to vote. Or maybe the new tab should be shown inline on the right. That is, I'm guessing most only see what's shown on the first page.
Apropos of nothing, I am working on building a model for predicting HN post performance.
TL;DR it is not easy.
I think I treated the title as a string rather than words and each column/feature was one letter. That is, if there were 256 possible characters and each title could only be 80 characters long, there'd be 256 * 80 * 2 columns (title + title reversed).
The trick with that approach is if you set the threshold high, it’s very easy to get a model with a high accuracy since classes will be imbalanced. (The vast, vast majority of HN submissions do not do well)
HN discussion: https://news.ycombinator.com/item?id=14400603
> This was the time when SpaceX successfully launched and landed its satellites at sea.
They didn't launch any of their own satellites, and they haven't landed any satellites at all :). I suppose the right word here would be "rockets".