GNU datamash
gnu.org
gnu.org
Here's an example of datamash and R with timing.
time datamash sstdev 1 < data.txt
288891.28552648
0.76s user 0.01s system 99% cpu 0.775 total
time R --vanilla --slave -e \
"x <- read.table('data.txt', header=F); sd(x\$V1);"
288891.3
2.68s user 0.06s system 99% cpu 2.761 total
(The data.txt file is 1 million lines, each line a random number 1 to 1 million. The timing is on a MacBook Pro Retina 13" 2014) awk 'END{for(i=0;i<1000000;i++){ print int(rand() * 1000000) } }' </dev/null > data.txt
time datamash sstdev 1 < data.txt
288619.72189328
0.72s user 0.01s system 99% cpu 0.736 total
time R --vanilla --slave -e 'sd(scan("data.txt"))'
Read 1000000 items
[1] 288619.7
1.09s user 0.04s system 99% cpu 1.134 total
R read.table read performance is fairly slow by default because it has to infer the types of columns and check for inline comments, quotes ect.This seems like a better replacement for awk and bash one-liners to me than tasks I would use R for.
For instance counting unique elements.
#naive approach
time (sort data.txt | uniq | wc -l)
632209
13.09s user 0.04s system 101% cpu 12.984 total
#using hashing
time (awk '!a[$0]++' data.txt | wc -l)
632209
1.34s user 0.03s system 100% cpu 1.360 total
#R
time R --vanilla --slave -e 'length(unique(scan("data.txt")))'
Read 1000000 items
[1] 632209
1.20s user 0.04s system 99% cpu 1.244 total
#datamash
time datamash countunique 1 <data.txt
632209
0.83s user 0.01s system 99% cpu 0.840 total
Quite good performance in that case, although R surprised me here as well.Downloaded: datamash-1.0.6.tar.gz and datamash-1.0.6.tar.gz.sig
Then did:
gpg --verify datamash-1.0.6.tar.gz.sig datamash-1.0.6.tar.gz
Which results: gpg: Signature made Tue 29 Jul 2014 03:30:23 PM PDT using RSA key ID 3657B901
gpg: Can't check signature: public key not found
Where can one import that public key, and is it the public key for datamash or gnu?(1) Assaf Gordon <agordon@wi.mit.edu> 4096 bit RSA key 2272BC86, created: 2014-07-09, expires: 2015-07-09
Initial announcement ... http://lists.gnu.org/archive/html/info-gnu/2014-07/msg00007....
I've had a little awk routine that I wrote some years back that does much of this -- it computes (or tabulates) n, sum, min, max, mean, median, standard deviation, and percentiles of the input data series. For generating quick stats, it's quite useful.
I'm looking forward to datamash turning up in my Debian repos.
cat table.txt | datamash transpose
This looks pretty cool. Anyone used it in "real life"?