I know it's not exactly an unknown command, but I didn't know about "sort" until last week.
It's freaking fast and convenient, sorts hadoop reduce results like a champ.
It's freaking fast and convenient, sorts hadoop reduce results like a champ.
# Deflection Col-Force Beam-Force
0.000 0 0
0.001 104 51
0.002 202 101
0.003 298 148
0.0031 104 149
0.004 289 201
0.0041 291 209
0.005 104 250
0.010 311 260
0.020 104 240
You can chain commands together to quickly get ball park figures. I do this all the time when reviewing log data to get simple counts and look for anomalies. $ cat data.txt | grep -v '#' | awk '{print $2}' | sort | uniq -c
1 0
4 104
1 202
1 289
1 291
1 298
1 311
Let me break this down a little. I cat data.txt, which just prints the contents to the console, I use grep to remove the # header, then use awk to print the second data column, the sort the values for my next procedure, and lastly I use uniq -c, which takes the input and counts all the unique values. So you can see that "4 104" means that 4 instances of 104 were found. I use this often to look at httpd logs for IP addresses or to review sudo commands, etc. Very handy!ps. there are optimizations that could be done here, but that what is great about unix and "one function" utilities. There are many ways to the same destination!
cat data.txt | grep -v '#'
you can do: grep -v '#' data.txt
or you can simplify further by moving the grepping into awk, so that the full command becomes awk '!/#/{print $2}' data.txt | sort -n | uniq -c
sort -n is good to use for sorting numbers. You could probably pull the sort | uniq -c into awk somehow, but I have already demonstrated the full extent of my awk knowledge. One of these days I will get around to actually learning it. sort -n -f 2 <data.txt | awk '!/#/ ... awk '!/#/ {a[$2]++}END{for (i in a) print a[i],i | "sort -nk1"}' data.txt
1 0
1 202
1 289
1 291
1 298
1 311
4 104The funny part is that one of the early uses of Google's map/reduce approach was probably to distribute the job of sorting search results among many servers. And here you are sorting Hadoop results with something as small and simple as sort(1).