You wrote: The titles are "user", "# of posts", "karma" & "karma/# of posts" ?
Yes, that's right. It's clear that "# of posts" is never over 50.
I suppose some readers here know Ruby, Perl, or Python, but haven't come through _The Unix Programming Environment_ route or done shell programming like we had to do before those languages existed, so I thought there may be some interest in an explanation of the pipeline. Bear in mind that the script interspersed here will have characters missing due to the site's software. The hex-encoded version in this thread can be run to get the verbatim script.
wget -qO - http://news.ycombinator.com/leaders |
First, we get the source of the leaders web page with wget.
tr '"' '\n\n' |
Because that has very long lines, and it's easier to pluck at most one user id per line, transliterate double-quote and greater-than sign characters into newlines.
sed -n 's/^user?id=//p' |
This then allows sed to only print lines that begin with /user?id=/ after deleting that same text, resulting in a list of user ids, one per line.
xargs -ri wget -qO - 'http://news.ycombinator.com/submitted?id={}' |
xargs takes each of these user ids in turn and runs a wget to fetch the "submitted" web page for that user id. All the web pages' HTML goes down the same pipe, concatenated together.
tr '' '\n' |
Again, we transliterate greater-than signs to split long lines at a convenient place.
sed -n 's/^\([0-9][0-9]\) points by .=\(.\).$/\2 \1/p' |
And sed gets used to filter out lines detailing the points scored by a user for a post, re-formating the information on the way, e.g. "42 points by ...id=ralph" becomes "ralph 42". There's one line containing user id and score for each of the scores listed on the one web page.
awk '{c[$1]++; s[$1] += $2}
END {for (n in c) printf "%-20s %5d %5d %7.2f\n", n, c[n], s[n], s[n] / c[n]}' |
All these (user id, score) tuples get fed to awk. It uses two associative arrays, AKA dicts or hashes, to build up information for each line it reads. Both arrays are indexed by user id; "$1" in awk refers to the first word of the input line. Array "c" maintains a count of the number of times a user id is seen. Array "s" accumulates the scores for each user id. Then at the end of all the input we loop through each index of "c", i.e. the user ids that have been seen, printing the user id, the frequency of that id, the sum of its scores, and the mean.
sort +3nr
These columns are given to sort which is told to sort on the fourth column, i.e. mean, numerically, in reverse. My sort(1) doesn't give the warning yours does, but you're right; I'm using old-school syntax that is deprecated but my fingers know how to type it. :-)
Cheers, Ralph.