How Hacker News ranking algorithm works
amix.dk
amix.dk
(= gravity* 1.8 timebase* 120 front-threshold* 1
nourl-factor* .4 lightweight-factor* .17 gag-factor* .1)
(def frontpage-rank (s (o scorefn realscore) (o gravity gravity*))
(* (/ (let base (- (scorefn s) 1)
(if (> base 0) (expt base .8) base))
(expt (/ (+ (item-age s) timebase*) 60) gravity))
(if (no (in s!type 'story 'poll)) .8
(blank s!url) nourl-factor*
(mem 'bury s!keys) .001
(* (contro-factor s)
(if (mem 'gag s!keys)
gag-factor*
(lightweight s)
lightweight-factor*
1)))))If you're just submitting interesting URLs and worry about them being shunted off the main page by gravity then there are no practical differences.
Sorry I can't be more transparent about how contro-factor is calculated, incidentally. Its purpose is to recognize flamewars.
function calculate_score($votes, $item_hour_age, $gravity=1.8) {
return ($votes - 1) / pow(($item_hour_age+2), $gravity);
} def calculate_score(votes, item_hour_age, gravity=1.8):
return (votes - 1) / pow((item_hour_age+2), gravity)I think I may have just reinvented the way reddit does it :P
(1) Wall-clock hours penalize an article even if no one is reading (overnight, for example). A time denominated in ticks of actual activity (such as views of the 'new' page, or even upvotes-to-all-submissions) might address this.
(2) An article that misses its audience first time through -- perhaps due to (1) or a bad headline -- may never recover, even with a later flurry of votes far beyond what new submissions are getting.
Without checking the exact numbers, consider a contrived example: Article A is submitted at midnight and 3 votes trickle in until 8am. Then at 8am article B is submitted. Over the next hour, B gets 6 votes and A gets 9 votes. (Perhaps many of those are duplicate-submissions that get turned into upvotes.) A has double the total votes, and 50% more votes even in the shared hour, but still may never rank above B, because of the drag of its first 8 hours.
(I think you'd need to timestamp each vote for an improved decay function.)
Recently, I have been implementing a ranking algorithm and I started in the wrong direction by taking all kinds of things into consideration. This is dangerous because you spend a lot of time on something that is mostly irrelevant.
The bottom line is that HN's algorithm works pretty well for the majority of cases. There are some edges, but solving them would not be trivial and would require a lot more work.
Of the two suggestions, replacing hours with artificial ticks adds very little complexity: it's the same formula, with one value replaced, and the accompanying factor adjusted. (The ticks might be submission-count, upvote-count, visit-count, or anything else trivially tallied -- it may not make much difference, except that over time greater activity could require adjusting the tick-deflator-factor.)
I'm not sure why I get down voted for this.
I'd be interested to know what the hourly fluctuation for HN is actually like, on account of having readers all over the world.
I'm in Australia, so your example "submitted at midnight" California time[1] means submitted at 6pm my time. Also 8am London time, 11am Moscow time. :).
[1] I'm going to go ahead and assume you're in California. ;)
It's true there's never total quiescence, but the pace of actions changes by a noticeable factor. (Without going to the data, I'd guess 5X from trough to peak over a day's cycle, and a somewhat smaller weekend-to-weekday difference. Holidays and nice bay area weather also play a factor.)
This would make your (excellent) idea easy to implement, because you could just use an autoincrement key as the timestamp, ignoring any sort of decay calculations.
I suspect that you have a nonstandard definition of "vast majority." I'd be surprised if even 1/4 of HN users are in California. There's a whole wide world out there! :)
Edit: Just to add to this - New York and Massachusetts are the next highest and combined make up about the same as California.
You probably mean something like a "large plurality"
Well, somewhat. You'd need a timestamped sufficient statistic, but not necessarily each vote.
Consider ACO. Given +1:
t0=last_change_timestamp
s0=score_after_last_change
s1=1+s0/exp(t1-t0)This is a pretty obvious algorithm, but the evil is in the details. First, since oknotizie is based in italy AGE is calculated in a special way so that nightly hours are calculated in a different way (every hour should be take into account proportionally to the traffic that there is in this hour).
Second, there is to do a lot of filtering. Oknotizie is completely built out of anti-spamming: statistical analysis on users voting patterns, cycles detection, an algorithm penalizing similarities in general in the home page, and so forth.
To run a simple HN style site is simple as long as the community is not trying hard to game it. Otherwise it starts to get a much more complex (and sad) affair.
http://blog.linkibol.com/2010/05/07/how-to-build-a-popularit...
The biggest missing ingredients are flagged posts dropping off quicker and posts that contain no URL dropping off quicker but there are quite a few other subtle tweaks.
The (very good) reason why the ARC sources do not give out the real ranking algorithm is to make it a bit harder to game the system.
So is Hacker News is a fork of news.arc, rather than straight news.arc? I figured it was, but never heard that officially (since people refer to news.arc as the HN "source code").
Edit: Also, the "lightweight" thing is interesting. There's something in place that sees if the post is a "rallying cry" or is mostly made of images. Additionally, if you link directly to an image file, or to some list of domains that have been deemed lightweight, that'll get marked as lightweight as well. Lightweight posts have a .3 factor, meaning that they're even more deflated than URL-free posts.
I did a fast test of how placements would look like with the vanilla HN algorithm using the current frontpage: http://paste.plurk.com/show/316811/
The code to generate the new sorting: http://paste.plurk.com/show/316812/
As you can see the rankings are not the same, but very similar...
19 points, six steps, that's averaging 3 points per step, so it's like a single flag counts for 3 'downvotes' on an article or something close to that.
That's all guesswork of course.
You can see it in action on maemo.org: http://maemo.org/news/
PHP sources: http://trac.midgard-project.org/browser/branches/ragnaroek/m...
points / comments
It works shockingly well. Its here if anyone would like to check it out: http://www.upthread.com/
(if
cond1 expr1
cond2 expr2
.. ..
condN exprN
else-expr)
With an even number of arguments, there is no else-expr (the same as CL/Clojure cond).The indentation is probably to distinguish between condition and conditional expression - the second column is what could be evaluated and returned from the whole expression.
[1] I too use Clojure
(let [really-long-name l
score (get-score x)
position (get-pos) x)
test test]
(...
....)))))
Sort of the same effect in the arc snippet in OP. > (if #t 'then 'else)
then
> (if #f 'then 'else)
else
'cond, in contrast, takes as many branches as you like, but they're parenthesized like so: > (cond (#f 'a)
(#t (display "I can have a body here!\n")
'b)
(#t 'c))
I can have a body here!
b
In arc, there's just 'if, which is like 'cond with a lot of implicit parenthesization and else-branches: arc> (if t 'then 'else)
then
arc> (if nil 'then) ; if no else-branch is given, nil is implied
nil
arc> (if nil 'a
t (do (prn "'do is like Scheme's 'begin or CL's 'progn.")
'b)
'else)
'do is like Scheme's 'begin or CL's 'progn.
bwhat are the peak use hours on HN? since everyone is a hacker, i'd assume it was evening hours on east and west coast US as heaviest load? not 9-5 ET?