Improving the Hacker News Ranking Algorithm
felx.me
felx.me
What worries me is the definition of quality you use. We look for submissions that we find valuable for us, not necessarily high quality. Interests are varied, and we all get value from different things. Quality might be very correlated with value, but it's not quite the same. And here comes the big issue: maybe more than 50% of the value we derive from HN comes from the comments. I feel many times we are upvoting submissions by sheer relevance, so we can have valuable discussions about a relevant topic. I don't think the given metrics are capturing this. I like the analysis and the proposal, but it's easy to see how there's some very important perspective missing.
I believe this creates a bias against longer articles. If a submission links to a longer piece, a user could take a long time to come back to HN to upvote the submission.
That's not to say this would be any worse than the current algorithm. By definition, a time-ranked frontpage and moving discussion will always favor shorter articles, or, if the content is long, produce plenty of comments that are only superficially related to the article.
As a submission ages, its rank wanes, and discussion gets more sparse: a piece that takes time to read, understand and digest will probably perform worse than a short one, since the discussion and votes will take place further into the future.
What I would try to do is calculate the expected value of the number of upvotes from the people who have seen it and use that as the metric.
So for any post there should be some number of views and some number of votes. The amount of time from when a user clicks the link to when a user clicks the upvote forms a distribution, because each user is different.
Then you assume that that distribution is a probability distribution for the users who haven't clicked on the post yet, and calculate the expected value using that.
This tries to mitigate the fact that there's a bit more delay in upvotes when reading longer posts.
This doesn't solve the problem of if there are too few upvotes to get a reliable distribution, or basically that it gives no advantage to longer posts when they are first posted. You can solve this with something called a prior distribution, and maybe scraping posted links to determine the length.
Another small issue is users upvoting an article after taking a break from reading, or maybe they are familiar with the link and don't even need to read it before upvoting. Remember we're trying to estimate how long would it take for users to read the article, and in those cases, the user isn't really doing that as much. So you could try to do some filtering or transforming before putting data points of that type into the distribution.
Part of HN’s charm is that niche content can surface to the top and many here are experts in their field, contributing great insights unavailable in normal media (example: pilots discussing possible causes after an airplane crash). That means the rest of us get to learn something new with perspectives from the best and many newbies become interested in more. That’s what top University classrooms feel like. I’m not sure what metric would quantify that quality (just # of comments as a metric encourages rage posting, maybe # comments with high upvotes?). The current system, although game-able and full of false negatives also has entropy built in - randomness that is more likely to give a wider range of content to the front page. That helps with genetic algorithms, so it might be of benefit, but missed if not quantified. If gaming the system counts - the most persistent also win some more.
Another option to consider is duration of stay on articles. From testing it seems Google factors in duration of stay on content as search criteria as well, which skews rankings towards longer pages. Those also tend to be ad-supported and likely within Google’s own advertising network. Within HN, more comments increase duration of stay, so that would likely increase provocative content, so may not be a good idea, but testing would surface that and more.
I assume the good thing about the snowball effect is that "making it to the frontpage" gives you guaranteed and pronounced exposure over a day or two - even though it's not maximizing the number of good links shown.
We're working on a more accurate simulation that will have many features of real Hacker News, like pagination, new-pages and so on.
Duplicate links are actually interesting since they are so easy to detect (identical link). Why not simply aggregate metrics for those things? The important thing is the link making it to the front page without creating a lot of duplicates. Somebody submitting a duplicate would effectively become an up vote for the already submitted link. Views and duplicate submissions are possibly more significant signals than up votes.
In search precision and recall are the two metrics people use for judging search quality. It's important to realize that hacker news is effectively a ranking algorithm and therefore is a search engine; even though it delegates actual search to Algolia. It sacrifices recall for precision: everything on the front page should be relatively high quality. But that's at the price of potentially high quality things never reaching it (recall).
Of course, with only 30 slots on the front page, there's only so much that can be on it. Especially if you consider that many users only drop by once or twice a day or so. So those slots stay occupied for quite long as well. Days in some cases. The choice as to what is right is highly subjective (i.e. the moderators decide) and biased towards the intentions of the site and that's intentional. But that doesn't mean it can't be improved.
Or maybe send the entire list of links of the day to the client and sort it using JavaScript provided by the user
Would be an interesting experiment, maybe they've tried it before?
You can imagine someone writing a clever query to extract hashed passwords or something from different tables. Also with the amount of traffic HN gets the front page is probably cached - running a query for every single request sounds expensive.
It would be very cool if possible tho.
I will probably go as far as saying that no web based forum has come even close to producing a online group discussion UX as good as what Usenet readers used to have.
The entire premise of this section is incorrect. There is no such thing as "enough votes to make it the front-page". There's no single threshold, nor is the value of a single vote constant. You're going to have a lot of other effects like specific domains having score penalties, votes being ignored due to them appearing to be fraudulent / non-organic, or the votes arriving too diffuse over a long time period. (Duplicates of the same URL will count as a vote for something like 12 hours after submission, so it's quite easy for a link to get a long tail of votes even after dropping out from the first page of /new.)
I would bet that most of what they claimed made it to the front-page never did. There used to be some HN stat-tracking sites around that had the full ranking history of submissions. Joining against that data would be a lot more credible.
> To achieve this goal, the new-page should expose every submission to a certain amount of views, to estimate its eligibility for the front-page.
That is obviously unworkable. /new is a slush pile, 90% of the submissions do not deserve even a single view based on the title. The proposal ensures that people visiting /new will only see the obvious garbage that nobody else has clicked on either.
The threshold of 2 upvotes is a necessary condition to be considered for the front-page (all pages, not just rank 1-30), not a sufficient one. I have this information from dang himself.
Garbage on the new-page will only be viewed, but never upvoted, so it drops quickly because of the view- and age-penalties in the formula and makes room for other submissions. But of course, this needs real-world experiments.
We define a view as a click by a registered user on a submission, so the user sees the submitted content. For Ask HN, the content would be the comments page.
I'm adding it to the article right now.
Seemingly worse for practical purposes, I don't think this data actually exists - adding clickthrough tracking to HN would be a huge change to the privacy profile of the site.
After all we got some very good feedback from HN and are incorporating it into our research. So thank you for it, this is valuable.
Growth of the whole community can be seen in the kaggle notebook: https://www.kaggle.com/felixdietze/hacker-news-score-analysi...
If you're referring to the front-page on HN, it could have been submissions of the second chance pool: https://news.ycombinator.com/item?id=26998308
Showing random, new posts on the home page provides equal viewership to the experiment set while removing the age bias. Many candidate posts will likely be low quality, but we can expect people to filter them out and participate less in them. High quality posts will organically attract people's attention and the algorithm can over time learn factors that differ between lower quality and higher quality posts.
Maybe he can report about these experiments himself.
We thought about modeling quality and agent preference with higher dimensional vectors, like in a recommender system. And to see if a user will upvote a specific content, you calculate the dot product. A simplified 1-dimensional user expectation and 1-dimensional quality can be seen as working on the scalars of the dot products itself.
There was a Master thesis which did some simulations with high-dimensional vectors and how many dimensions make sense. The result was that it basically doesn't matter.
In the end, we have to make a real-world experiment with the HN community to see if it works as expected.
However, to get the following requirement from the article:
> The algorithm should not produce false negatives, the community should find all high-quality content.
it might be better to estimate the upper confidence bound, like in Upper Confidence Bound bandit, rather than the lower confidence bound.
The biggest problem I have with the front page these days is it gets clogged with junk science and low-effort news posts that produce entirely predictable flame wars that just aren't very interesting. There's about two dozen domains that account for 95% of such stories. I go through and flag/hide them now, and that works okay, but doing it automatically would be better.
Another option would be an experiment on HN itself for a week or so.
It has happened to me a couple of times as well. That is clearly unexplainable.
Newsify is great but it’s been crashing a LOT with iOS 15 beta (not complaining, I know it’s beta). It’s a great reader for iOs devices.