Learn to use Awk with hundreds of examples
github.com
github.com
* http://blog.jpalardy.com/posts/why-learn-awk/
* http://blog.jpalardy.com/posts/awk-tutorial-part-1/
* http://blog.jpalardy.com/posts/awk-tutorial-part-2/
* http://blog.jpalardy.com/posts/awk-tutorial-part-3/has some minor issues though
> modern (i.e. Perl) regular expressions
nope, supports only ERE.. doesn't have non-greedy, lookarounds, etc
> $ cat netflix.tsv | awk '{printf "%s %15s %.1f\n", $1, $6, $5}' | sed 1d
could have just added NR>1 condition..
> Alternatively, awk '{print $2}' netflix.tsv would have given us the same result. For this tutorial, I use cat to visually separate the input data from the AWK program itself. This also emphasizes that AWK can treat any input and not just existing files.
awk -F":" 'BEGIN{total=0}{if($3>240)total+=$3}END{print total}' /etc/passwd
Me too. For example, number of unique IP requests in a log file in milliseconds ;)
$ awk '{ print $1 } ' caddy_log | sort | uniq | wc -l
But rarely compose my own from scratch. It's mostly copy paste. And store in admin bin for future use.
$ cat duplicates.txt
abc 7 4
food toy ****
abc 7 4
test toy 123
good toy ****
$ awk '!seen[$2]++' duplicates.txt
abc 7 4
food toy ****
$ awk '!seen[$2]++{cnt++} END{print +cnt}' duplicates.txt
2I'll add a note, thanks :)
awk '{ip[$1]=1}; END{ print length(ip) }'
Use the hash tables directly, assuming you've got the memory. That should be slightly faster as it avoids the sort.also, {ip[$1]} is enough...
;-)
perl -e 'while(<>) { $tot += (split)[2]; }; print "$tot\n"'
It's funny to see that in perl's decline, the stuff it did well is being forgotten and resurrected via tools it at one point had mostly replaced. perl -F: -lane '$total+=$F[2] if $F[2]>240; END{print $total}' /etc/passwd perl -F: -anE '$tot+=$F[2]}{say$tot' perl -F: -anE '$_=$F[2];$t+=$_if$_>240}{say$t'
I wanted to be clever by doing `$t+=$_*($_>240)`, but that's actually one byte larger. :(If my codebase wasn't already in perl I wouldn't necessarily be able to benefit from a large library of my own functions relevant to the problem, but at least I would still have CPAN to fall back on. I'm not sure awk will ever be able to compete with that.
awk -F: '$3>240{total+=$3} END{print +total}' /etc/passwdBEGIN{total=0}
can be skipped (at least in the awk's I've used, not sure if some more recent / strict awk differs), because awk initializes the variable total to 0.
I remember this because I do this all the time to quickly get the (non-recursive) sum of the sizes of the files in a directory - it's pretty much muscle memory from a while now:
ls -l | awk '{ s += $5 } END { print s/1024 " KB" }'
For recursive size, one can use ls -lR or the du command with various options according to need.
So in other words, we can take out the BEGIN block, but then we must remember to change print total to print total + 0.
Also, using uninitialized variables is basically a code golfing stupidity that will bite you in any halfway complicated program.
GNU Awk has a useful --lint argument which spots uses of uninitialized variables. If you make it habit to write code that way, if you then use --lint for finding a bug, you have to deal with false positives.
Nonsense. Not if you know what you are doing, and used it in a known way, which is what I did. The code I wrote works. I tested it on Linux before posting it. Also, such a usage (skipping the initializer) is mentioned (IIRC) in the classic Kernighan & Pike book "The Unix Programming Environment" (still a great resource, though not updated for modern Unix/Linux features), which is where I learned it from, years ago (and hence why I qualified my statement by saying it may not work in more strict or modern awk versions). Fine to talk about other variations but it does not mean that my variation is wrong.
Don't try to read my mind. My intention was not code golfing. Was just sharing some fun info. It's not a big deal to keep the initializer either, I'm quite aware of that.
Your intention can be understood as the promotion of code golfing, as evidenced by these words:
The fragment BEGIN{total=0} can be skipped
by which you're clearly encouraging that other coder to make their code shorter by removing an initialization that works fine.
Your statement (above) can be "understood" as not understanding my prior statement(s), including the one in which I said "Don't try to read my mind" (w.r.t. intention, because you cannot - it is mine (mind), not yours). If you cannot grok that after a second explanation, I have nothing further to say. Good day.
Smile when you say that, stranger!
http://www.thisdayinquotes.com/2011/05/when-you-call-me-that...
#!/bin/sh
x=0;while true;do read a;
test ${#a} -gt 0||exec echo $x
a=${a#*:\*:};n=${a%%:*};
test $n -le 240||x=$((x+n));
done < /etc/passwd #!/bin/sh
IFS=:;x=0;while true;do read a b c d;
test ${#a} -gt 0||exec echo $x
test $c -lt 240||x=$((x+c));
done < /etc/passwdAttribute value minimization would be one limitation, since that regex is for XML, but it's a more robust approach than writing naive "<x>.*</x>"-style regexes.
...isn't that how all tokenizers for all (confusingly-named) context-free languages work?
they are explained here: http://www.catonmat.net/series/awk-one-liners-explained
which I mention it in further reading section
Does awk really provide that more value over sed while being easier or faster to use than a fully-fledged scripting language (thinking of perl, python, etc).?
(and yes, one may argue that awk IS a scripting language, I'm not disputing that, just asking)
sed works on a line by line basis.
awk can work on a whole file. Subsequent line operations can depend on the state of previous lines.
Each has its own operating domain and you have to decide which tool is the best one for the task you have in mind.
What is the definitive tool to process text? Perl? Haskell? Some Lisp dialect?
Definitive? Being snarky, the one you have already installed and are familiar with. Like most I use Awk for one-liners, Perl if I need a little more or better regexes in a one- or two-liner. For the last several years I've been using TXR[1] if it gets complex. Lately I've been doing more fiddling with JSON than text and I'm using Ruby/pry and jq[2].
https://gist.github.com/rlonstein/90d53fdeea31d2137737
about a matter related to the hash bang line in the script.
TXR has a nice little hack (that apparently I invented) to implement the intent of "#!/usr/bin/env txr args ..." on systems where the hash bang mechanism supports only one argument after the interpreter name.
In the interest of fair and accurate disclosure, I earned my bread for 3.5 years debugging Perl code for a living and I've also had formal education in Perl programming at the university. I would never want to do that again.
I also learned recently that GNU Awk has networking support[5]. I have no idea why!
[1S] sed '1d'
[1A] awk 'NR!=1 {print}'
[2S] sed 's/foo/bar/'
[2A] awk '{sub(/foo/, "bar")}'
[3S] sed -n '/start_regex/,/end_regex/p'
[3A] awk '/start_regex/,/end_regex/ {print}'
[4] awk '$1=="foo" {print $5}'
[5] https://www.gnu.org/software/gawk/manual/gawkinet/gawkinet.h...
can be simplified to:
awk '/start_regex/,/end_regex/'
because in awk, if no action is given, the default action is to print the lines (that match the pattern). And if the pattern is omitted but the action is given, it means do the action on all lines of the input.
Edited to change:
print the line (that matches the pattern)
to
print the lines (that match the pattern)
Couldn't agree more!
A great example of this was using (surprised it wasn't mentioned) sed with the -i & 's///g' operators while "cleaning" hundreds (seriously) of HTML/PHP files from injected content at a shared hosting provider.
Similarly
sed 15q
will print only the 1st 15 lines of the input and then terminate. E.g.:
sed 15q file
or
some_command | sed 15q
So, when put in a shell script and then called (with filename arg or using stdin):
sed $1q
is like a specific use of the head command [1]; it prints the first n ($1) lines of the standard input or of the filename argument given - where the value of $1 comes from the first command-line argument passed to the script.
[1] In fact on earlier Unix versions I worked on (which did not have the head command (IIRC), I used to use this sed command in a script called head - similar to tail.
And I also had a script called body :) to complement head and tail, with the appropriate invocation of sed. It takes two command-line arguments ($1 and $2) and prints (only) the lines in that line number range, from the input.
>Too complicated
please try out examples given and let me know if it helps you to understand the syntax better
awk '{ print $1; }'
Other than that... not really. Maybe the advantage would be ubiquity, if you really, really want to avoid Perl.
How about: cut -f 1 -d ' '
You don't even need the -d flag if the you happen to be able to use the default delimiter, like in your example.
apples 1
bananas 2while read a _; do echo $a; done
Personally, I would suggest AWK for one to two liners when doing some one time data transform tasks. Anything more complicated and I would suggest a more "fully-fledged scripting language"
* for field processing, most multiple line processing, use of logic operators, arithmetic, control structures etc, I prefer awk or perl
* this repo is aimed at command line text processing tools, most awk examples given are single line, a few are 2-3 lines. Personally I prefer Python for larger programs
----
See also: https://unix.stackexchange.com/questions/303044/when-to-use-...
According to my potentially miscalibrated gut, referencing a "fully-fledged" scripting language in a shell script or at the command line is an indication that you should probably just be working in that environment in the first place.
There will be exceptions, it depends chiefly on the problem being solved, but overall I prefer my shell scripts to reference utilities with very specific purposes. It feels UNIXier that way.
I myself have implemented XML SOAP command line client, a backup solution, a SAN UUID management application and an automated Oracle RAC SAN storage migration solution, a configuration management, and an Oracle database creation / management applications in AWK.
Usually I develop a thin getopts shell wrapper around an AWK core. Works every time, the executables are on the order of a few KB (the largest so far, the XML SOAP client is 24.5 KB) and they all run like a bandit. Memory requirements are miniscule. Dependencies are minimal: the only external dependency so far in my software has been the xsltproc binary from the libxslt package.
AWK is easier to use than Python or Perl, and is much faster than either of those. Typical code density ratio of Python versus AWK is 10:1, sometimes more. This means that if you have a 650 line Python program, you can implement the same functionality in about 280 lines of AWK, and the program will be far simpler. I've once collapsed a 280+ line Python program into a simple 15 lines of code in AWK.
AWK is an extremely versatile, powerful programming language.
For even more speed, AWKA can be used to transpile AWK source into C and then it will call an optimizing C compiler to compile it into a binary executable. Typical speedup is on the order of 100%, so if your AWK program ran in 12 seconds, it'll now finish in six.
How does this work? I am not saying it can't be done, but the main benefit of Awk seems to be quick one-liners, which are possible because you get "records" (splitting on whitespace) and lines (splitting on newline) and looping for free. But for larger programs, this easily translates to Python; just call readlines(), loop over it, call split() on each line. I would think that at this point, Awk doesn't have much of an advantage anymore... but apparently your experiences are different. What are some Awk constructs that would take a lot more code in Python?
seq 1 30 | awk '
$0 % 3 == 0 { printf("Fizz"); replaced = 1 }
$0 % 5 == 0 { printf("Buzz"); replaced = 1 }
replaced { replaced = 0; printf("\n"); next }
{ print }'
Note that the awk script is far more general than the typical interview question, which specifies the numbers to be iterated in order. The awk script works on any sequence of numbers.Go ahead, write the Python script that behaves exactly as this AWK program does. It will likely be 4x as long, and that's because the number of different patterns and actions to take is quite low. More complex (and hence more situated and less easy-to-understand) use cases will benefit even more from AWK's defaults.
Moreover the pattern expressions are not constrained to simple tests: https://www.gnu.org/software/gawk/manual/html_node/Pattern-O...
They can match ranges, regular expressions, or indeed any AWK expression. They can use the variables managed by the AWK interpreter: https://www.gnu.org/software/gawk/manual/html_node/Auto_002d... (NR and NF are commonly used).
Actions one-way or two-way communicate with coprocesses with minimal ceremony: https://www.gnu.org/software/gawk/manual/html_node/Two_002dw...
All of those mechanisms can be done in a Python script, but they add up to a lot of boilerplate and mindless yet error-prone translation to the standard library or Python looping and conditional logic.
All of which are built in functions...
> To behave like the AWK script it also has to catch an exception and continue when the input cannot be parsed as an integer.
Not quite sure what behavior you're referring to here. When I tested your script, it happily treated "xy" as divisible by 15.
> Go ahead, write the Python script that behaves exactly as this AWK program does.
import fileinput
for line in fileinput.input():
replaced = False
if int(line) % 3 == 0: print("Fizz", end=''); replaced = True
if int(line) % 5 == 0: print("Buzz", end=''); replaced = True
if replaced: print();
else: print(line, end='')
> Moreover the pattern expressions are not constrained to simple testsAnd none of these, except maybe for the range operator, are particularly challenging for python.
2. How would you modify it so it parsed a tab-delimited file and did FizzBuzz on the third column? With awk it is a simple matter of setting FS="\t" and changing $0 to $3?
3. How would you modify it so instead of being output unmodified, rows with $3 that are neither fizz nor buzz output the result of a subprocess called with the second column's contents?
Now you might say that this is all goalpost-moving, but that's the point. AWK is more flexible and less cluttered in situations where the goalposts tend to get moved, but where the basic text processing paradigm stays the same.
def intish(str):
try:
return int(str)
except:
return 0
Can python's default be reproduced as easily in awk?2. You'd insert field = line.split('\t') at the beginning of the loop and then refer to field[2]
3. os.popen or subprocess.run
I buy the "less cluttered" argument when the problem matches awk's defaults. I vehemently disagree with the "more flexible" argument. A problem perfectly suited to awk can easily turn to a poor fit with the addition of a single, seemingly innocuous requirement (e.g. in your subprocess example, log the standard error of your subprocess into a separate file).
!/^[0-9]+$/ {
print "invalid input: " $0 > "/dev/stderr"
exit 1
}
at the beginning of the script. None of the other actions need to be changed; but with your implementation, all of the calls to "int" need to be changed to "intish".I've got the following script (I stopped playing games with line breaks):
#!/usr/bin/env gawk -f
BEGIN {
FS = "|"
}
$2 % 3 == 0 {
printf("Fizz")
replaced = 1
}
$2 % 5 == 0 {
printf("Buzz")
replaced = 1
}
replaced {
replaced = 0
printf("\n")
next
}
{
system("cal " $2 " 2018 2> errors.txt")
}
Which can produce the following output: $ ./script.awk <<EOF
> thing1|0
> thing2|3
> thing3|7
> thing4|13
> EOF
FizzBuzz
Fizz
July 2018
Su Mo Tu We Th Fr Sa
1 2 3 4 5 6 7
8 9 10 11 12 13 14
15 16 17 18 19 20 21
22 23 24 25 26 27 28
29 30 31
$ cat errors.txt
cal: 13 is neither a month number (1..12) nor a name
- What does the equivalent program in Python look like?- How many characters does it have with respect to the number of characters in the awk script? (259 with shebang).
- How many characters would need to change to split by "," instead? (1 for awk). (You can achieve this in Python, but you'll end up spending characters on a utility function.)
- How many characters would need to be added to print "INVALID: " and then the input value for lines with non-numeric values in the second column, then skip to the next line? (55 for awk)
Character adds/changes are the best proxy for "flexibility" I could think of that doesn't go far afield into static code analysis.
I love Python and don't think awk is a good solution for extremely large or complex programs; however, it seems obvious to me that it is significantly more flexible than Python in every line-oriented text-processing task. The combination of opinionated assumptions, built-in functions and automatically-set variables, and the pattern-action approach to code organization, all add up to a powerful tool that's still worth using in order to keep tasks from becoming large or complex in the first place.
Also record splitting and processing is highly configurable in AWK with RS, ORS, OFS and one gets it for free without having to write extra code. And don’t forget that Python needs about 25,000 files just to fire up, while AWK is a single 169 KB executable (on Solaris / illumos / SmartOS). Makes a huge difference come application deployment time.
"Easier to use" - maybe so, on the particular subset of problems that awk was designed for. However, the ease of use upside is limited - awk constructs map pretty much 1:1 onto Python/Perl constructs that are not particularly complicated. Conversely, there is a vast set of problems that are still straightforward to solve in Python/Perl and would be rather awkward in awk.
"much faster" - the comparisons I've seen (and done) usually had awk and perl5 roughly at parity.
"code density ratio 10:1" - I call BS on that one. Sure, with the benefit of hindsight, it's sometimes possible to vastly simplify a script, but that has little to do with the languages involved. There is no awk solution that cannot be expressed in about 2x the lines of Python code (and that 2x is mostly because idiomatic awk puts conditions and code on one line, while Python puts them on two lines).
OTOH I'm not sure I'd recommend bothering to learn it. Python is more verbose in Awk's domain, but not by so much as makes a huge difference, except at the scale of one-liners. (Or a-few-liners, at least.)
Another reason to learn it: the AWK book (by Aho, Kernighan, and Weinberger) is a great very short intro to the spirit of Unix-style coding. You could think of learning Awk as just the price of admission to that intro, paid along the way.
I wrote plenty of Awk in the 90s -- https://github.com/darius/awklisp isn't very representative but it was fun.
(It's a little like functional programming being rediscovered when Lisp has been around since the dawn of time).
* I think Perl got carried away when it added objects and folks started writing large programs with it (although I have written some large scripts for biologists doing genetic studies -- which is interesting popular use case == there are google groups and O'Reilly books focused on this use case).
Also, Larry Wall really humped the shark with Perl 6.
The syntax was simpler than ed, and getting combination of grep, uniq, cut etc correct.
I call it "Excel of the command line"
GAWK: Effective AWK Programming
I added references to it throughout the chapter
Has analogs for all salient POSIX Awk features and most GNU Awk extensions. (Of course, not semantic cruft like the weak type system, or uninitialized variables serving as zero in arithmetic.)
Plus:
* You can embed (awk ...) expressions anywhere, including other (awk ...) expressions.
* You can capture a delimited continuation (awk ...) and yield out of there.
* It supports richer range expressions than Awk. Range expressions combine with other range expressions unlike in Awk, so that you can express a range which spans from one range to another. Also, there are variations of the operator to exclude either endpoint of the range: rng, -rng, rng- and -rng-.
* You can "awk" over a list of strings, possibly an infinitely lazy one.
1> (awk (:inputs '("a" "b") '("c" "d"))
(t (prn nr fnr rec)))
1 1 a
2 2 b
3 1 c
4 2 d
nil
* It has a return value: whatever the last :end returns, or else nil: 1> (awk (:end 42) (:end 43))
[Ctrl-D]
43
Build a list from the first fields of /etc/passwd: 1> (build
(awk (:inputs "/etc/passwd")
(:set fs ":")
(t (add [f 0]))))
("root" "daemon" "bin" "sys" "sync" "games" "man" "lp" "mail"
"news" "uucp" "proxy" "www-data" "backup" "list" "irc" "gnats"
"nobody" "libuuid" "syslog" "messagebus" "avahi-autoipd" "avahi"
"usbmux" "gdm" "speech-dispatcher" "kernoops" "pulse" "rtkit"
"hplip" "saned" "kaz" "vboxadd" "sshd" "oprofile" "ntp" "lightdm"
"colord~" "whoopsie" "postfix")
Type conversion of fields (which are just strings) is achieved by an elegant operator fconv which takes a condensed notation such as (fconv i : r : xz) which means convert the first field to integer as a decimal integer, the last field as a hexadecimal integer and the fields in between as reals. The xz means that if the last field is invalid, it gets converted to zero rather than nil. These letters are just the names of lexical functions available in the awk scope, rather than built-in fconv behaviors.http://web.archive.org/web/20000829071436/http://inferno.bel...
AWK, Perl, Tcl, Scheme, C, Java, Limbo, Visual Basic
What if k scripting language was included in those experiments?
k3:
1. "Basic Loop Test"
\t 1000000(1+)/0
2. "Ackermann's Function Test" \t {:[~x;y+1;~y;_f[x-1;1];_f[x-1;_f[x;y-1]]]}[3;7]
3. "Indexed Array Test" \t x(x;|x:!200000)
4. "String Test" \t f:{(x>#:){(i _ x),(1+i:_.5*#x)#x:,/("123";x;"456";x;"789")}/y};do[10;f[500000;"abcdef"]]
5. "Associative Array Test" \t {+/("0123456789abcdef"16_vs'!x)_lin$!x}40000
6. "File Copy Test" `f 0:(30000 _draw 300)#\:"king "
\t `f 0:0:`f
7. "Word Count Test" \t (#:;+/(+/1<':" "=)';+/#:')@\:0:`f
8. "File Reversal Test" \t `f 0:|0:`f
9. "Sum Test" `f 0:100000#,"-123.456"
\t +/0.0$0:`f
Source: http://web.archive.org/web/20010501041644/http://www.kx.com:...