General question: when would one choose to use an awk script over something more general purpose such as a python or ruby script? To me it would make sense to use the latter in most cases.
General question: when would one choose to use an awk script over something more general purpose such as a python or ruby script? To me it would make sense to use the latter in most cases.
AWK is old. 1977 old. Later versions that appeared (nawk and gawk are the most common) helped make it a smoother language, but it's still a pain. There are definite features you will be missing in an awk script:
- Any useful data structure slightly more complex than associative arrays. Try multi dimensional arrays, it's actually fun to do. Once.
- Any useful programming construct to help manage with complexity of scripts longer than a few hundred lines. No classes, no variable scoping, no namespaces in general. Not to mention an extremely permissive compiler.
- Any easy way to deal with the environment (other than text). Try sending http requests in awk, it can be a pain.
In 2015, if you need to write a script you should almost always prefer python/ruby over awk.
Now if you're asking whether you should learn awk? It comes in handy. There are a lot of awk scripts in the wild, and you may need to read them or edit them one day. Also, awk has a fun way of parsing input, which makes for very enjoyable one liners. Learning awk (and some complementary utilities like sed or find) turned me into a "oh, it can be done in a quick one-liner" kind of guy. Definitely recommended.
Yes, being able to casually toss off an awk (or sed) one-liner is a very convenient skill to have.
It's nice that the mentality comes out of using certain languages though.
sed -e 'NR == 123425938039 { print; exit; }' < file
Fast as hell and takes no memory to speak of. Google 'awesome awk' for goodies.[0]: Yes, this happens.
sed -n 123425938039p file?
To be fair, this one won't exit early.1. It is lightweight -- there is no need to load shared libraries, modules, plugins, etc., and it uses less memory than the more general purpose scripting languages do.
2. It comes with BusyBox, so if you have a BusyBox PXE boot environment you can make the AWK scripts work under it, and also work independently of any host OS that gets booted, as long as it's Unix-like. Perl, Python and Ruby are not in Busybox, and their versions can vary a lot across different OS distributions and cluster configurations. AWK avoids this dependency because it's stable and self-contained in one executable file.
3. AWK automatically splits records into fields based on a settable FS, so it makes it very easy to parse things like /proc/stat and /proc/loadavg into $1 $2 $3 ... fields. The code looks nicer and is more compact because of the automatic field variables.
4. The default RS is newline, and AWK is designed to perform actions on records. Most of the scripts I write perform an action upon receiving a newline on their standard input as a trigger. So most of the scripts tend to be of the "BEGIN { } { }" variety -- initialization in the BEGIN section, and then an unconditional action to fire at every newline, such as collecting /proc statistics and reporting them up the cluster hierarchy. AWK is naturally suited for this REPL behavior without needing any boilerplate code.
$ ls -lR | awk '{ print $3, $4 }' | sort -u > user_group_list
and got the list I wanted. (Yes, it's sloppy but I don't care, as it's just a one-off.)
$ find -print0 | xargs -0 stat -c "%U %G" | sort -u > user_group_list
find -printf "%u %g\n"|sort|uniq -cYou have almost answered your own question.
In situations where you are dealing with munging data which has a structure that is implicitly handled by awk (one or more files, consisting of regularly-delimited records, which break into regularly-delimited fields), it is very difficult to beat awk for succinctness.
The second thing is this: Awk has been around for decades and is part of the POSIX standard. A shell script that uses awk commands can be full POSIX compliant and work on any POSIX-like system with minimal changes, without the installation of third-party software.
So, with Awk you can do a lot of things that would otherwise require something like Python. In many cases you can do them with clearer code that has less clutter, and your solution is POSIX, to boot.
The down side of Awk is that it sacrifices reliability to pander to succinctness. Awk does not detect undefined variables; any variable you mention becomes defined. It has loose arithmetic: this is so you can increment a nonexistent variable or array element by one, and it behaves as if it had been defined with zero. Awk only has local variables in functions, and they are modeled as extra parameters (for which you could pass values, but you don't, unless perpetrating a hack).
Basically, whereas you can do actual software engineering in Python and Ruby, you'd be crazy to do it in Awk, even if the lack of libraries and datatypews such isn't an impediment against doing it in Awk.
20-25 years ago I wrote many more things in Awk, up to a Lisp interpreter and a parser generator; but Python/Ruby/etc. have replaced it.
When the others are not around. Sometimes you need to work on weird old or stripped-down boxes that don't have python or even perl - but awk is (almost) always there. Like vi.