I tried the python programs under pypy and python3 (after running through 2to3) -- and got similar speeds as the author. I was a little surprised that my little awk script was slower than than pypy (completing in 6 to 7 seconds):
$ cat sum.awk
{ a[$2] += $3 }
END {
for (i in a) {
if (a[i] > max) {
maxk = i
max=a[i]
}
}
print "max_key:", maxk, "sum:", max
}
$ time gawk -f sum.awk -O <ngrams.tsv
max_key: 2006 sum: 22569013
real 0m7.041s
user 0m3.797s
sys 0m3.156s
(This under Linux subsystem for windows)According to the gawk profiler, and strace -c (count) -- the awk program mainly spends its time reading the file (without the loop at the end, looking for the max value, the runtime is essentially the same).
In fact, on the surface, pypy and python3 are quite similar on the syscall/strace front - with roughly 23k "read" calls -- awk did 375k. And adding cat in front sped it up by about two seconds:
$ time (cat ngrams.tsv |awk -f sum.awk )
max_key: 2006 sum: 22569013
real 0m3.969s
user 0m3.719s
sys 0m0.516s
$ time awk -f sum.awk ngrams.tsv
max_key: 2006 sum: 22569013
real 0m6.465s
user 0m3.609s
sys 0m2.859s