i am thought impressed with how fast pypy did
i am thought impressed with how fast pypy did
$ cat sum.awk
{ a[$2] += $3 }
END {
for (i in a) {
if (a[i] > max) {
maxk = i
max=a[i]
}
}
print "max_key:", maxk, "sum:", max
}
$ time gawk -f sum.awk -O <ngrams.tsv
max_key: 2006 sum: 22569013
real 0m7.041s
user 0m3.797s
sys 0m3.156s
(This under Linux subsystem for windows)According to the gawk profiler, and strace -c (count) -- the awk program mainly spends its time reading the file (without the loop at the end, looking for the max value, the runtime is essentially the same).
In fact, on the surface, pypy and python3 are quite similar on the syscall/strace front - with roughly 23k "read" calls -- awk did 375k. And adding cat in front sped it up by about two seconds:
$ time (cat ngrams.tsv |awk -f sum.awk )
max_key: 2006 sum: 22569013
real 0m3.969s
user 0m3.719s
sys 0m0.516s
$ time awk -f sum.awk ngrams.tsv
max_key: 2006 sum: 22569013
real 0m6.465s
user 0m3.609s
sys 0m2.859s (cat ngrams.tsv \
|parallel --pipe awk -f map.awk \
|awk -f reduce.awk )
max_key: 2006 sum: 22569013
This now runs in 17 to 18 seconds... :-/ $ cat map.awk
{ a[$2] += $3 }
END {
for (i in a) {
print "ignore", i, a[i]
}
}
$ cat reduce.awk
{ a[$2] += $3 }
END {
for (i in a) {
if (a[i] > max) {
maxk = i
max=a[i]
}
}
print "max_key:", maxk, "sum:", max
}
[ed: However, there are faster awks than gawk: $ time mawk -f sum.awk ngrams.tsv
max_key: 2006 sum: 22569013
real 0m2.826s
user 0m2.391s
sys 0m0.422s
mawk is (a little) faster than pypy on my machine.]
Which is a 15000-line Perl script, so to be expected.
Try --pipe-part instead:
parallel -a ngrams.tsv --pipe-part --block -1 awk -f map.awk |
awk -f reduce.awkThe (old) version of parallel packaged with Ubuntu 16.04 (linux subsystem for windows) - doesn't have --pipe-part -- but running from upstream, the speed is more reasonable:
$ time (./parallel-20170522/src/parallel -a ngrams.tsv \
--pipe-part --block -1 -j4 mawk -f map.awk \
| mawk -f reduce.awk )
max_key: 2006 sum: 22569013
real 0m2.265s
user 0m4.672s
sys 0m1.672s
(Tried a few variants with/without -jN -- and this seems typical for the fast end of the spectrum). $ time (cat ngrams.tsv \
| mawk -f map.awk \
| mawk -f reduce.awk )
max_key: 2006 sum: 22569013
real 0m3.472s
user 0m2.891s
sys 0m2.406s
[ed: btw, did a double-take when I saw your Gnu Privacy Guard id: 0x88888888 :-) ]https://stackoverflow.com/questions/27801945/surprising-resu...
Another likely optimization would be to use the csv module (which would parse the rows in C).