danielha : 3.52
danw : 3.24
brett : 6.10
python_kiss : 4.24
mattculbreth : 5.34
sharpshoot : 7.06
jwecker : 3.88
staunch : 4.57
amichail : 2.30
Harj : 5.74
Alex3917 : 11.77
joshwa : 4.62
far33d : 4.84
nostrademons : 7.50
jamiequint : 7.26
Sam_Odio : 6.06
Elfan : 5.00
domp : 3.44
zaidf : 5.93
dfranke : 6.40
Readmore : 5.00
paul : 15.67
blader : 7.56
phil : 7.44
mattjaynes : 5.72
herdrick : 6.24
palish : 29.60
veritas : 4.08
bootload : 2.18
BioGeek : 10.06
require 'rubygems'
require 'active_support'
require 'net/http'
Net::HTTP.start('news.ycombinator.com', 80) do |http|
body = http.get('/leaders').body
users = body.scan(/user\?id=([^"]+)/)
users.each do |user|
body = http.get("/submitted?id=#{user}").body
points = body.scan(/(\d+) points? by/).map(&:first).map(&:to_f)
puts "%-15s : %6.2f" % [user, points.sum / points.size]
end
end
python -c 'print "77676574202d714f202d20687474 703a2f2f6e6577732e79636f6d6269 6e61746f722e636f6d2f6c656164657273 207c0a74722027223e2720275c6e5c6e27207 c0a736564202d6e2027732f 5e757365723f69643d2f2f7027207c0a7 8617267732 02d726920776765 74202d714f202d20276874747 03a2f2f6e6577732e79636f6d62696e6 1746f722e636f6d2f 7375626d6 9747465643f69643d7b7d27 207c0a747220273e2720275c6e27207c0a736 564202d6e2027732f5e5c285b302d395d5b302d 395d2a5c2920706f696e74732 a206279202e2a3d 5c282e2a5c 292e242f5c32205c312f7 027207c0a61776b20277b635b243 15d2b2b3b20735b24315d 202b3d20243 27d0a20202020454e4420 7b666f7220286e20 696e206329 207072696e74 662022252d32 307320253564202 535642025372e32665c6e22 2c206e2c20635b6e 5d2c20735 b6e5d2c20735b6e5d2 02f20635b6e5d7d27207c0 a736f7274202 b336e720a".replace(" ", "").decode("hex"),'
Cheers, Ralph.
wget -qO - http://news.ycombinator.com/leaders |
tr '"' '\n\n' |
sed -n 's/^user?id=//p' |
xargs -ri wget -qO - 'http://news.ycombinator.com/submitted?id={}' |
tr '' '\n' |
sed -n 's/^\([0-9][0-9]\) points by .=\(.\).$/\2 \1/p' |
awk '{c[$1]++; s[$1] += $2} END {for (n in c) printf "%-20s %5d %5d %7.2f\n", n, c[n], s[n], s[n] / c[n]}' |
sort +3nr
Cheers, Ralph.
http://news.ycombinator.com/comments?id=13271
Cheers, Ralph.
Neither are turned into links.
>=> <=< &=&
I did python -c '...' | cmp - orig to test it.
Cheers, Ralph.
...
Oh sorry. My fault cutting & pasting into idle :( but I did get a warning on sort sort: Warning: "+number" syntax is deprecated, please use "-k number"
...
Cool, worked. The titles are "user", "# of posts", "karma" & "karma/# of posts" ?
Yes, that's right. It's clear that "# of posts" is never over 50.
I suppose some readers here know Ruby, Perl, or Python, but haven't come through _The Unix Programming Environment_ route or done shell programming like we had to do before those languages existed, so I thought there may be some interest in an explanation of the pipeline. Bear in mind that the script interspersed here will have characters missing due to the site's software. The hex-encoded version in this thread can be run to get the verbatim script.
wget -qO - http://news.ycombinator.com/leaders |
First, we get the source of the leaders web page with wget.
tr '"' '\n\n' |
Because that has very long lines, and it's easier to pluck at most one user id per line, transliterate double-quote and greater-than sign characters into newlines.
sed -n 's/^user?id=//p' |
This then allows sed to only print lines that begin with /user?id=/ after deleting that same text, resulting in a list of user ids, one per line.
xargs -ri wget -qO - 'http://news.ycombinator.com/submitted?id={}' |
xargs takes each of these user ids in turn and runs a wget to fetch the "submitted" web page for that user id. All the web pages' HTML goes down the same pipe, concatenated together.
tr '' '\n' |
Again, we transliterate greater-than signs to split long lines at a convenient place.
sed -n 's/^\([0-9][0-9]\) points by .=\(.\).$/\2 \1/p' |
And sed gets used to filter out lines detailing the points scored by a user for a post, re-formating the information on the way, e.g. "42 points by ...id=ralph" becomes "ralph 42". There's one line containing user id and score for each of the scores listed on the one web page.
awk '{c[$1]++; s[$1] += $2} END {for (n in c) printf "%-20s %5d %5d %7.2f\n", n, c[n], s[n], s[n] / c[n]}' |
All these (user id, score) tuples get fed to awk. It uses two associative arrays, AKA dicts or hashes, to build up information for each line it reads. Both arrays are indexed by user id; "$1" in awk refers to the first word of the input line. Array "c" maintains a count of the number of times a user id is seen. Array "s" accumulates the scores for each user id. Then at the end of all the input we loop through each index of "c", i.e. the user ids that have been seen, printing the user id, the frequency of that id, the sum of its scores, and the mean.
sort +3nr
These columns are given to sort which is told to sort on the fourth column, i.e. mean, numerically, in reverse. My sort(1) doesn't give the warning yours does, but you're right; I'm using old-school syntax that is deprecated but my fingers know how to type it. :-)
Cheers, Ralph.
You forgot PG though, I guess he took himself off there -- he's averaging 8.34 per submission.