The fact that most OSs have a POSIX compliant version of it in the base system also make it very valuable, not just for the people crunching data but also to sysadmin/devops people.
The fact that most OSs have a POSIX compliant version of it in the base system also make it very valuable, not just for the people crunching data but also to sysadmin/devops people.
My 99% use case is extracting particular fields from some input stream.
awk '{print $1,$5}'
That's a tremendously common problem and awk is the best tool for the job.I once started diving deeper into awk and there is an impressive amount you can do with it, but I hit a point pretty quickly where it would be cleaner to write the script in a more traditional language where my successor won't be cursing my name for writing something complex in an obscure language.
cut -d $delim -f $field1,$field2,...However, for ad-hoc command-line usage, I still prefer awk, because (1) the default delimiter is white space, which means I don't have to worry about if what I'm looking at might contain tabs intermingled with spaces (or how many spaces) and (2) it makes it much easier to incrementally increase the complexity of my ad-hoc operation, such as if I decide I want to take a sum or average of one of those columns, after all.
It's a narrow tool. Other than the inherent advantage this brings in clarity/simplicity in a script, there's nothing else to it.
Consider a script that originally had now awk but just a
field=$(somecmd | cut -d: -f3)
but now has been modified to have awk elsewhere. Did the above just become less simple or clear? If so, is it worth changing it to the below for clarity? Does it matter if the awk predates the cut, instead? field=$(somecmd | awk -F: '{print $2}')
I say "no" to both, although I do recognize the argument for just using the same, consistent tool everywhere, instead.Can be done with tr, without writing a utility (if by "collapse" you mean what I think you do):
This file t:
$ cat t
the quick brown fox jumped
over
the lazy dog
containing many combinations of spaces, tabs and newlines (whitespace) can be changed to this file t2:$ cat t2
the
quick
brown
fox
jumped
over
the
lazy
dog
by this tr command:
tr -cs "[a-zA-Z]" "\012" < t > t2
That also makes the output more amenable to further processing, including common tasks like finding the frequencies of the words in the input, as mentioned in the "More shell, less egg" post mentioned in this post:
The Bentley-Knuth problem and solutions:
https://jugad2.blogspot.com/2012/07/the-bentley-knuth-proble...
(which I somehow managed to read as an awk invocation).
Re: python - I believe if using things like \w, \d or \s you should be Unicode safe.
Come to think of it, I seem to recall gnu tools should also have Unicode aware pattern/character classes, eg:
https://www.gnu.org/software/gawk/manual/html_node/Bracket-E...
> For example, before the POSIX standard, you had to write /[A-Za-z0-9]/ to match alphanumeric characters. If your character set had other alphabetic characters in it, this would not match them. With the POSIX character classes, you can write /[[:alnum:]]/ to match the alphabetic and numeric characters in your character set.
https://docs.python.org/3/library/re.html#re-syntax
[ed:
Apparently gnu tr is still out in the cold re:unicode;
https://www.gnu.org/software/coreutils/manual/html_node/tr-i...
> Currently tr fully supports only single-byte characters. Eventually it will support multibyte characters; when it does, the -C option will cause it to complement the set of characters, whereas -c will cause it to complement the set of values. This distinction will matter only when some values are not characters, and this is possible only in locales using multibyte encodings when the input contains encoding errors.
]
- outputting my selected fields in an order other than their order in the input lines: you can't `cut -f2,1`, for example.
- use a multi-character or regex field separator: -F'<space><star><comma><space><star>' [edit: rewritten in words to dodge HN's formatting] is a real winner when dealing with hand-written CSV.
- use a record separator other than \n: thanks, Windows, for all those times I've written ORS="\r\n" :P - you can even use a regex record separator, which not even perl supports; I don't need that often, but it's a godsend when I do.
- select based on criteria other than "number of fields from the start": I recently needed the last field from lines of varying field-count, and awk '{print $NF}' was just the thing.
But if you don't need anything like that, cut is a lot more succinct - especially if you need to split on tabs only, which it does by default, while awk defaults to splitting on spaces too.
You can shell out if necessary:
cmd = "whois " $url
while ( cmd | getline ) {
# stuff
}
close(cmd)
Also: coprocesses.Another common usage for awk in my case is to do floating operations (*, +, eq, gt, etc) in a shell script. It avoids installing bc, which is not always present by default.
I used it for other things, like for example comparing software versions in shell scripts.
The most advanced usage of awk in my case was to do some statistic analysis (average, confidence interval, standard deviation) of network simulations.
$ cat foo
foo bar baz
foo qux
$ < foo cut -d' ' -f 2-
bar baz
qux
As far as I know you have to loop through the fields, which is (haha) awkward.Unix was mature even then. There is NO WAY you would have convinced me that 30 years later I'd still be using most of that, still gluing things together with the same shell tools. Or that most of the books would still have a place on the shelf (albeit sometimes in 2nd or 3rd editions), and turn out to be classics.
Ignoring big iron, we'd gone from Apple IIs to Amigas and workstations in just a few years, and hardware was changing as rapidly as the current fashion for JS frameworks does today.
Yet here we are, awk, grep, sed and the other tools still feel like the right, concise answer. So why not learn them, they might greatly surprise you by their staying power.
You need to be cautious however, gnu awk can be somewhat inconsistent at times:
try the following on various distros (and also an *BSD) for example:
awk 'BEGIN {print "9.1" >= "8"}';
awk 'BEGIN {print "9.1" > "8"}';
awk 'BEGIN {print "9.1" < "8"}';
awk 'BEGIN {print "9.1" <= "8"}';
on OpenBDS, you get no output or a syntax error, on Debian, you get the correct result for < and <= in all cases, never for > (it's interpreted as a redirection), and >= only works with newer Debian versions (sid and maybe stretch).
I know that the fix is simple: just put some parenthesis (awk 'BEGIN {print ("9.1" <= "8")}'), but I was quite amused to discover that a few days ago.
And if you really want that in a repo, you can stick that in ansible trivially.
Sometimes instead I just wrap a line of awk in %x() in Ruby to save a few chunks of code.
Most programmers don't understand why I do this though. :(
Hundred of lines of python by two awk lines ? Must be very bad python programmers.
Awk is a compiled language. Your Awk script is compiled once and applied to every line of your file at C-like speeds. It is way faster than Python.
If you learn to use Awk well, you will start doing things with data that you wouldn't have had the patience to do in an interpreted language.
I say this with no disrespect for Python, it's just you should know the tradeoffs of the tools you use. It does not matter if the individual operations in Python are implemented in C. It will still be much slower at looping over lines of a file than a compiled language designed decades ago for that exact purpose.
I have a lot of Python file in one dir:
$ find . -iname "*.py" | wc -l
10429
Finding them and cating them all takes about 0.3 secs: $ time find . -iname "*.py" | xargs cat {} > /tmp/cat.out
...
real 0m0.344s
user 0m0.140s
sys 0m0.175s
So to have something simple that takes a bit of time, I tried to get all lines starting with "print", and output the first thing after that.I'm really bad at awk, so I don't know if there is a better way. I went for the most obvious thing for me:
$ time find . -iname "*.py" | xargs cat {} | awk '/^print/{print($2)}' > /tmp/awk.out
...
real 0m1.111s
user 0m1.165s
sys 0m0.368s
Now, with Python, it's definitely not as easy to type. You have to get a script like. awk wins the expressivity metrics for this use case: import sys
for x in sys.stdin:
if x.startswith('print'):
try:
print(x.split()[1])
except IndexError:
print('') # to match awk behavior
But as for performance, I don't get the huge boost in perfs you are talking about: $ time find . -iname "*.py" | xargs cat {} | python /tmp/test.py > /tmp/python.out
...
real 0m0.762s
user 0m0.862s
sys 0m0.347s
I do get the same output though: $ cmp /tmp/python.out /tmp/awk.out && echo "yes"
yesHere's what I observed: I had a simple text-processing tool I needed that I call "countmerge", which just merges adjacent lines with the same key and adds up their corresponding values. I needed to run it on a lot of large files.
I first wrote it in Python, where it was a significant bottleneck compared to the steps that came before it (split, sort, uniq -c). Eventually I rewrote it in Rust [1], and it was at least 5 times faster, at the expense of a fair amount more low-level code. But then rewriting it in awk [2] turned out to be as fast as what I wrote in Rust, possibly inconclusively faster.
[1] https://github.com/rspeer/countmerge
[2] https://gist.github.com/rspeer/60c87dca1ab550326f8bd6d086452...
>>> with open('data.txt', 'w') as f:
... for l in string.ascii_uppercase:
... for x in range(0, random.randint(1, 100000)):
... f.write('Key {}\t{}\n'.format(l, random.randint(0, 100)))
With this script: import sys
old_key = total = 0
for line in sys.stdin:
key, value = line.split('\t')
if old_key != key:
old_key = key
total = 0
print(key, value, end="")
total += int(value)
I get: $ <data.txt time python3 test.py
Key A 2
Key B 87
Key C 58
Key D 64
Key E 29
Key F 25
Key G 2
Key H 74
Key I 17
Key J 37
Key K 97
Key L 77
Key M 19
Key N 74
Key O 33
Key P 61
Key Q 67
Key R 23
Key S 4
Key T 70
Key U 25
Key V 15
Key W 35
Key X 17
Key Y 31
Key Z 18
1.03user 0.01system 0:01.05elapsed 99%CPU (0avgtext+0avgdata 9564maxresident)k
0inputs+0outputs (0major+1100minor)pagefaults 0swaps
But I can't manage to get the awk version working. It only prints one line on Ubuntu 16.04: $ <data.txt awk -f ./countmerge.awk
Key 0
So I can't check it.