Useful Unix commands for data science
gregreda.com
gregreda.com
- Column: Create columns / tables from input data
- tr: substitute / delete chars
- join: like a database join, but for text files
- comm: like diff, but you can use it programmatically to choose if a line is in one file, or another, or both.
- paste: put file lines side-by-side
- rs: reshape arrays
- jot: generate random or sequence data
- expand: replace tabs / spacesCheck out the man page for a few snippets: http://www.unix.com/man-page/FreeBSD/1/jot/
It is the older, more flexible uncle of gnu's 'seq' command: http://administratosphere.wordpress.com/2009/01/23/using-bsd...
"Don't pipe a cat".
My test doesn't show a speed improvement, but there are less processes running, and less memory consumed.
bch:~ bch$ jot 999999999 2 99 > data.dat
bch:~ bch$ time cat data.dat | awk '{sum +=$1} END {printf "sum: %d\n", sum}'
sum: 50499999412
real 6m21.111s
user 6m15.506s
sys 0m5.711s
PID COMMAND %CPU TIME #TH #WQ #PORTS #MREGS RPRVT RSHRD RSIZE VPRVT VSIZE PGRP PPID STATE UID FAULTS COW MSGSENT MSGRECV
22342 awk 100.7 05:11.84 1/1 0 17 21 52K 212K 340K 17M 2378M 22341 22306 running 501 311 49 73 36
22341 cat 1.1 00:03.94 1/1 0 17 21 272K 212K 548K 17M 2378M 22341 22306 running 501 268 51 73 36
============== bch:~ bch$ time awk '{sum +=$1} END {printf "sum: %d\n", sum}' ./data.dat
sum: 50499999412
real 6m24.023s
user 6m13.828s
sys 0m2.774s
PID COMMAND %CPU TIME #TH #WQ #PORTS #MREGS RPRVT RSHRD RSIZE VPRVT VSIZE PGRP PPID STATE UID FAULTS COW MSGSENT MSGRECV
22373 awk 100.0 00:30.16 1/1 0 17 21 276K 212K 624K 17M 2378M 22373 22306 running 501 256 46 73 36There are enough variations in ways to do things on Unix that I've sometimes wondered about how easy it would be to identify a user by seeing how they accomplish a common task.
For instance, I noticed at one place I worked that even though everyone used the same set of options when doing a "cpio -p", everyone had their own order they wrote them. Seeing one "cpio -p" command was sufficient to tell which of the half dozen of us had done the command.
I think I'm the only one where I work who uses "sed Nq" instead of "head -N", so that would fingerprint me.
So one day I telnetted into a Solaris machine and immediately typed "lsl" before doing anything else. A short while later a colleague came to my cube. He had been snooping the hme1 interface and saw me login. He didn't need to trace the IP because he knew it was me when he saw 3 telnet packets with "l" "s" "l" in them.
I recommend "The AWK Programming Language" by Aho, Kernighan, and Weinberger, though it seems to be listed for a hilariously high price on Amazon at the moment. Maybe try to pick up a used copy.
By the way: what people need to understand is that in order to use Awk, efficently, you'll either use associative arrays, or structure your script like a sed script, otherwise it will be slow. The interesting thing about both of those, is the regex algorithm Thompson NFA, that is from what I hear around 7 times faster than PCRE that is used in Perl, PHP, Python and Ruby?
http://linuxgazette.net/67/nazario.html
i still use a buttload of awk for data science type uses.
I concur with this recommendation. "The AWK Programming Language", at little over 100 pages, is a classic of programming language instruction. The book jumps right into use cases, it does not waste one's time. This book should be required reading for anyone contemplating writing a handbook on any programming language; my CS bookshelf would be several feet thinner and several times more informative.
It's the first result for me.
http://www.cs.princeton.edu/courses/archive/spr08/cos333/awk...
It deals with things he forgets or needs to remind himself of.
If you're interested in his other personal tutorials, they are here:
http://www.cs.princeton.edu/courses/archive/spr08/cos333/tut...
When Perl was created, one of its advertised goal was to avoid all the time lost trying to work around the limitations of awk, sed and shell.
Use it any place in a pipeline to see a progress meter on stderr. Very handy when grepping through a bunch of big log files looking for stuff. Here is a quick strawman example:
pv /data/*.log.gz | zgrep -c 'hello world'
241MiB 0:00:15 [15.8MiB/s] [==> ] 2% ETA 0:12:12 progress -zf /data/*.log.gz grep -c 'hello world'
progress -f /data/*.log.gz zgrep -c 'hello world'
The second form will show the progress of the decompression process.You can also adjust buffer size, set the length for the time estimate (otherwise we have to fstat the input), and display progress to stderr instead of stdout.
http://svnweb.freebsd.org/base/vendor/tnftp/dist/src/progressbar.c
http://ftp.netbsd.org/pub/NetBSD/NetBSD-release-6/src/usr.bin/{Makefile,progress.c}
Not sure if you prefer binary installs or whether you compile your installs yourself... but I'm sure you could get this to compile on FreeBSD with a little work.One thing I find odd that you have to drop to a full language (awk, perl etc) to sum a column of numbers. Am I missing a utility?
echo "1\n2\n3\n" | sum # should print 6 with hyphothetical sum command
I suppose more generally you could have a 'fold initial op' and: echo "1\n2\n3\n4\n" | fold 0 + # should print 10
echo "1\n2\n3\n4\n" | fold 1 \* # should print 24
But I guess at that point you're close enough to using awk/perk/whatever anyway. Which probably answers my question. echo "1\n2\n3\n+\n+\np\n" | dc -
Or, a little more legibly: > dc
1
2
3
+
+
p
6
[Edited to fix bug]Not sure how to automate this to sum 1000 values without needing to explicitly insert 999 + signs, though. Haven't explored dc in depth myself yet. There's probably some way to do it with a macro or something, but it may not be pretty.
echo -e "1\n2\n3\n4" | paste -sd+ | bc
kind of cheating though! :-)
$ alias sum="xargs | tr ' ' '+' | bc"
$ echo -e "1\n2\n3\n" | sum
6 alias sum='xargs -I{} sh -c "head -c {} < /dev/zero" | wc -c'Another useful one is "hist" which is sort | uniq -c | sort -n -r.
Does everyone else edit command history, stacking up 'grep -v xxxx' in the pipeline to remove noise?
If I'm working on a new pipeline, my normal workflow is something like:
head file # See some representative lines
head file | grep goodstuff
head file | grep good stuff | grep -v badstuff
head file | grep ... | grep ... | sed -e 's/cut out/bits/' -e 's/i dont/want/'
head file | grep ... | grep ... | sed -e 's/cut out/bits/' -e 's/i dont/want/' | awk '{print $3}' # get a col
head file | grep ... | grep ... | sed -e 's/cut out/bits/' -e 's/i dont/want/' | awk '{print $3}' | sort | uniq -c | sort -nr # histogram as parent
Then I edit the 'head' into a 'cat' and handle the whole file. Basically all done with bash history editing (I'm a 'set -o vi' person for vi keybindings in bash, emacs is fine too :-) grep -v "+http" access_log | cut -d \" -f 4 | cut -d \? -f 1 | sed 's/\/$//' | grep -v kmjn.org | sort | uniq -c | sort -nris the same as
> cut -f3 -d' '
cut is amazing for what it does. and most people know only the subset of awk that effectively _is_ cut anyway :D.
paste -sd+|bc echo "1\n2\n3\n" | tr '\n' + | bcInstead this is a set of basic examples of bog-standard tools that every newbie *nix user should be already familiar with: cat, awk, head, tail, wc, grep, sed, sort, uniq
The key word is should... you might be surprised how many "not newbie" nix users are not aware of those commands or how using them in this fashion. Specially awk.
Thanks!
What most users probably don't realize is that the redirection can be anywhere on the line, not just at the beginning. Putting an input redirection at the beginning of the command can make the data flow clearer: from the input file, through the command, to stdout:
< data.csv awk -F "|" '{ sum += $4 } END { printf "%.2f\n", sum }'
(This only works for simple commands; you can't do `< file if blah; then foo; else bar; fi`)True, however, people pointing out UUOC are in fact pointing out that you should not be building a pipeline at all. If you want to apply an awk / sed / wc / whatever command to a file, then you should just do that instead of piping it through a extraneous command.
Sure, as people always mention, in your actual workflow you might have a cat or grep already, and are building a pipeline incrementally; there's no reason to remove previous stuff to be "pure" or whatever. But if you're giving a canonical example, there's no reason to add unneeded commands.
cat data.csv | awk -F "|" '{ sum += $4 } END { printf "%.2f\n", sum }'
From that command, it's unclear whether the sum will be accurate, it depends on the inputs and on the precision of awk. See (D.3 Floating-Point Number Caveats): http://www.delorie.com/gnu/docs/gawk/gawk_260.htmlI don't see how that's more true for awk than it is for any other programming language. Awk uses double precision floating point for all numeric values, which isn't a horrible choice for a catch-all numeric type.
One can simply do awk 'cmds' file.
And if the extra cat is actually making a measurable difference, maybe that's a good signal that it's time to rewrite it in C.
$ cat data.txt | awk '{ print $2+$4,$0 }'|sort|sed '/^0/d'
can be written as $ <data.txt awk '{ print $2+$4,$0 }'|sort|sed '/^0/d'Get-Content .\data.csv | %{[int]$total+=$_.Split('|')[3]; } ; Write-Host "$total"
Powershell is a skill I don't have yet which carries over to ... precisely one declining technical dinosaur (with a penchant for expiring its skillsets).
The Linux toolbox is a set of skills I embarked on learning over a quarter-century ago, most of which goes back another decade or further (the 'k' in 'awk' comes from Brian Kernighan, one of Unix's creators). And while some old utilities are retired and new ones replace them (telnet / rsh for ssh, sccs/rcs for git), much of the core has remained surprisingly stable over time.
The main difference between MinGW and Cygwin appears to be how Windows-native they are considered, which for my own purposes has been an entirely irrelevant distinction, though if you're building applications based off of the tools might matter to you.
#!/bin/sh
printf %s\\n "$@" | awk -F'\t' '
FILE == "/dev/stdin" {
needle[$0] = 1
next
}
needle[$4] {
print $NF
}
' /dev/stdin /etc/pki/CA/index.txt
I find myself using this idiom (feeding data from a file to awk and selecting it with data from standard input) again and again. It's a great way to scale shell scripts to take multiple arguments while avoiding opening the same file N times, or doing clunky things with awk's -v flag. grep -A n -B n
is more easily written: grep -C n
If you think "C" for "context" this is easier to remember too.http://en.wikibooks.org/wiki/Ad_Hoc_Data_Analysis_From_The_U...
Apparently there's a need for it, though, or it wouldn't exist.
This entire HN thread is a perfect example of why we built Manta. Lots of engineers/scientists/sysadmins/... already know how to (elegantly) process data using Unix and augmenting with scripts. Manta isn't about always needing to work on a 10TB dataset (you can), but about it being always available, and stored ready to go. I know we can't live without it for running our own systems -- all logs in the entire Joyent fleet are rotated and archived in Manta, and we can perform both recurring/automated and ad-hoc analysis on the dataset, without worrying about storage shares, or ETL'ing from cold storage to compute, etc. And you can sample as little as much or as much as you want. At least to us (and I've run several large distributed systems in my career), that has tremendous value, and we believe it does for others as well. And that's just one use case (log processing).
Like I said, disclaimers/bias/etc.
m
Cloud services are amazing in a lot of ways, but so far I've found them much more heavyweight for the use-case of running ad-hoc jobs from the Unix command line. You don't really want to write Hadoop code for exploratory data analysis, and even managing a little fleet of bashreduce+EC2 instances that get spun up and down on demand is error-prone and tedious, turning me more into the cluster administrator rather than a user, which is what I'd rather be. Admittedly it's possible that could be abstracted out better in the case where you don't mind latency: I often don't mind if my jobs queue up for a few minutes, which would mean a tool could spin up EC2 instances behind the scenes and then tear them down without me noticing. But I haven't found anything that does that transparently yet, and Manta looks like a more direct implementation of the "illusion of running on an N-core machine for arbitrary N" idea that seems in the same cost ballpark. Definitely going to do some experimentation here, to see if 2010s technology will enable me to keep using a 1970s-era data-processing workflow.
BashReduce is a pretty cool application of many of these utilities.
http://www.slideshare.net/strands/strands-presentation-at-re...
All you need... is the sum of all values in one particular column."
In that case, if speed was paramount, I'd use Kona or kdb. Unquestionably, k is the best tool for that particular job.
Where can I see your experimental design? I'd like to try to replicate your results.
Thankfully you can also write a Perl one liner. Which most of the times is far powerful than awk.
Make all your commands 3x faster:
export LC_ALL=C
Actually use the 32 CPUs you paid for:
sort --parallel=32 ...
xargs -P32 ... export LC_ALL=C
would "make all your commands 3x faster"?With the C locale, text is more or less treated as plain bytes.
[1] http://dtrace.org/blogs/brendan/2011/12/08/2000x-performance...
gnu parallel FTW!
That's part of why Solaris continues to use them in favour of GNU alternatives (although the GNU alternatives are available easily in /usr/gnu/bin).