Awk in 20 Minutes (2015)
ferd.ca
ferd.ca
Why? awk is small enough that it fits in my head, or at least the bits I need every couple of months do. And if I forget, it only takes 20 minutes to put them back in. perl, by contrast, is far too large to fit in my head and comes with an ecosystem that is much larger still.
I know I can use perl for what I use awk for, and when I've raised this point before, people have been quick to explain how to process input by lines, and conditionally do something. For basic stuff, the fact that that's basically all awk does[0] means there's so many fewer ways to do it wrong. I can't say the same of perl, even when I was more familiar with it.
caveat: awk, nawk, mawk, and gawk don't necessarily share the same set of corner cases, and you may not get error messages that make sense to you when you bump into one with an unfamiliar awk.
[0] I know, not really, and especially for gawk. I've written awk scripts a couple hundred lines in the past. It's true enough for the 20 minute version though.
Edited for clarity.
perl -ne "print $_;"
'-n' will run the expression over each line of the input. This is the Awk-like mode. perl -pe "s/foo/bar/"
'-p' will run the expression over each line of input, and print out the (possibly modified) line. This is the Sed-like mode.Slightly more info available in `man perlrun` or https://perldoc.perl.org/perlrun.html
perl -plane 'my $script'
which iterates over all files given on the command-line (or stdin) and + (p)rints every processed line back out
+ deals with (l)ine endings, in and out
+ (a)utosplits every line into @F
I am aware that -n and -p are mutually exclusive, but as -p overrides -n, it's seems simpler to just keep 'plane' in mind and remove the 'p' if necessary.One small test (with a big file) I executed took 1 minute with Perl and 20 minutes with awk.
And then there are those really complicated formats where awk is just not flexible enough.
Awk is really useful, but it doesn't cover the same problem set as Perl does.
Around 10 years ago I rewrote markdown.pl in awk and it was almost 20 times faster. The speedup came from both a much faster startup time and a much faster (albeit simpler) regexp implementation.
I've been a software engineer since 2005 and worked my way up to being a VP of Engineering currently and never had to use either Perl or awk (or similar). I often read about these tools on Hackernews and I find it quite mystifying as I manage to have written Java, Scala, C#, SQL, and so on for 15 years and happily never needed them.
Is this a certain kind of engineering job that requires searching through text files so often and requiring specialized tools? I've managed my whole career with ctrl-f, and highlight-all matches.
I think it depends a lot on the sort of software one works with. A main use of awk and other unix tools for me is as-hoc data munging, looking through production logs, taking bits of data from different sources and comparing them. If you make eg gui applications and sell them or work as a contractor developing in-house applications for clients, you probably don’t see a lot of production logs like that or process them that way. Any data that you expect to process ought to go into a database and then you can use sql (where joins work much better than the unix join command and you can use the actual structure of the data instead of trying to tease it out with the smallest simplest code you can think of). You might be using a debugger or backtracks or reproduction in test to deal with/investigate production issues rather than starting with logs. On the other hand, many other companies will have big production systems that run on lots of different virtual machines and produce lots of logs and have issues that are hard to reproduce (or maybe the logs are easier), and have lots of data spread about random places such that it may be easier to do the hacky thing for one-off cases or to prototype or whatever rather than Doing things more thoroughly or properly. And this set up will also lead to tools that are designed to fit into the rest of the unix ETL tack by the way they output or input data.
On the second point, I think that even as a die-hard emacs user I wouldn’t recommend that anyone try to use emacs as an ide for something like java (or maybe c++). I’d probably try to use emacs a myself but I think my java-editing experience would be worse than for those people who do it with an IDE. More realistically, I’d try to avoid writing java if at all possible.
I mention DevOps because it is hard to think of a situation where you’re just coding Java and C# all day that truly requires awk-or-similar. But just the other day I had a situation where I copied the contents of a CD-ROM to a web server — drag and drop on my Mac — and ended up with the files all in upper case. The program I was trying to run was failing because it requested everything with lower-case URLs. There were hundreds and hundreds of files in a bunch of directories, so I fixed it with Perl in about 30 seconds.
I’m curious if that’s the sort of problem you never encounter, or if it is, what you reach for.
I've also used awk to get button presses from an input stream on a MIDI controller. For me, I found that the up front cost of learning a few awk commands quickly made it's return on investment.
I'm not proud to admit it, but I've used awk in a couple places where I should have properly used `expect' instead. Except, I don't know expect well and haven't gotten to the point where it was worth the investment in learning it.
Imagine you have to do this for 100 files.
I've used perl and awk extensively for all sorts of things. Log parsing embedded device logs to generate reports was a big one for me, that's where I really learned both languages. And they are superb tools for that task. We also used awk quite a bit for configuration parsing as a sort of intermediary between different processes and scripts on those devices.
As I moved into web development, I've found myself using both much less. I haven't touched perl more than once in the last 6 years -- I was given a script to maintain recently but it only needed an hour or two of my time to add a small feature to. That perl script is part of a legacy build system. I still use awk every couple of months but it's just part of some pipeline on the command line.
I could set up a database, write multiple layers of code ... but really?
ETL and anything dealing with importing large data sets come to mind.
To answer the serious question, it's often enough for a first order solution to any problem where you have line and field separated text. Pulling instances of "something weird" out of log files, for instance, is a great use for awk, especially if the fingerprint of "something weird" is scattered through fields in a line in a way that makes grep cumbersome. Or if it spans a couple of lines, especially if you're dealing with a log from a multi-threaded app where there might be irrelevant lines interleaved with the ones you care about.
I haven't written much production-grade awk, but I've often used it as a tool to understand a problem well enough to write a production-grade solution to a problem or fix bugs in application code.
I say it's useful on servers because log files, dump files, or whatever text format servers end up dumping to disk will tend to grow large, and you'll have many of them per server. If you ever get into the situation where you have to analyze gigabytes of files from 50 different servers without tools like Splunk or its equivalents, it would feel fairly bad to have to download all these files locally to then drive some forensics on them.
This personally happens to me when some Erlang nodes tend to die and leave a crash dump of 700MB to 4GB behind, or on smaller individual servers (say a VPS) where I need to quickly go through logs, looking for a common pattern.
Awk/sed and the ETL stack that comes with every Unix-like OS (aka od, tr, cut, sort) and all that are superb data wrangling tools. Log file parsing, data cleaning, even actual data science at scale can be done with these tools. They're extremely efficient, as they're designed from an era when pretty much all interesting data was comparable to or much larger than memory size. As such, you can do a lot of stuff with them that most people don't imagine is even possible. FWIIW for high end data scientists; I don't consider knowledge of these tools to be optional at all. Anyone who hasn't used them in their career hasn't worked on serious problems, and is probably the kind of educated idiot who will suggest you do a job in a giant Hadoop cluster that you could easily do on one machine.
In case you didn't know: Perl 6 has been renamed to Raku, using the #rakulang tag on social media.
> The first occurrence of a variable name defines it as that sum. Subsequent occurrences become the stored value.
With this quote in mind - why is second `SUM` reference replaced with `-138.95` and not `221.81` (stored value) ?
EDIT: Never mind, now I see it's an exception on line 3.
If you want a more thorough/deep exploration of Awk, I recently gave a talk on it (virtually) at Linux Fest Northwest (LFNW) 2020.
Awk: Hack the planet['s text]! Part 1 (Presentation): https://www.youtube.com/watch?v=43BNFcOdBlY
Awk: Hack the planet['s text]! Part 2 (Exercises):https://www.youtube.com/watch?v=4UGLsRYDfo8
If you want to try your hand at the exercises, they are on github: https://github.com/FreedomBen/awk-hack-the-planet
Let's say we have a function CharCount which takes a character and a line of text and returns the number of occurrences of that character:
function CharCount(ch, line,
n)
{
...
}
The line break in the parameter list is an AWK convention and indicates that n is a "local variable." function CharCount(ch, line, n)That said, there's an amazing amount you can do with it, if you really try. Someone once joked that I should try building an IRC bot in AWK, so I did: https://github.com/Marcus316/rufus
There's no pactical reason for it, but it was fun to play with the idea.
Last weekend I was playing around with some data. At first, I thought 'let's just write a line of awk and be done with it' and so I did. The execution took 20 seconds (about 17 million lines) and everything was fine.
Later that day, I came across another task which seemed too complex for an awk one-liner so I took two lines of R and was surprised when R was done within 5 seconds on the same data set.
I was happy because I found a faster tool than the one I had, but the lesson is, that just because you use a proven tool like awk, doesn't mean there aren't any better tools. Find out what works best for you.
[1]: https://brenocon.com/blog/2009/09/dont-mawk-awk-the-fastest-...
$ awk --version
GNU Awk 5.1.0, API: 3.0 (GNU MPFR 4.0.2, GNU MP 6.2.0)Interesting enough: LANG=C increased the performance of GNU awk a bit, but actually decreased the mawk performance.
I agree with the comment below. If you haven't already, try mawk instead of awk. It is often many times faster than awk (and other solutions).
Disclaimer: I don’t actually write any AWK, but learning it is on my bucket list.
AWK -> calculate the difference between every two lines:
awk -F ',' 'NR!=1{printf "%.0f\n", $1-ll}NR==1{print ""}{ll=$1}'
R -> Calculate the min, median, mean, and max: d<-scan("stdin", quiet=TRUE)
cat(min(d), max(d), median(d), mean(d), sep=" ")
AFAIK, calculating the difference to its previous line for every line should be faster for data that is not sorted. I didn't try to write that one in R though.I am sure there ways to make both things faster, I was just surprised as I didn't expect R to be faster with those two naive implementations.
I have just coined this idea the "gateway drug" approach to learning new tools. We should strive to find those introductions that are small enough that someone can digest without a massive upfront time investment to get them past the front door :)
However, at some point, it makes more sense to write your Awk script in something like Python, and my intuition says that that time is shortly after starting a basic Awk script. Using a real programming language is almost always the way to go with code that will become more complex over time (almost all code) and code you have to share with a team (learning curve).
pattern { action } # valid
pattern { # valid
action
}
pattern # not what you think
{ action } # this action has no pattern
There cannot be a newline between pattern an action. This feature allows either the pattern or the action to be omitted without ambiguity. pattern # pattern with default { print } action
{ action } # unconditional action
pattern { # pattern with action
action
}
Items can be put on a single line without ambiguity using semicolons: pattern ; { action } ; pattern { action }
A pattern can have multiple patterns separated by a comma. That syntax admits optional line separation after the comma separators: pattern,
pattern,
pattern {
action
}
The action fires by a match for any of the patterns.The POSIX standard Awk expression grammar has no comma operator, on the other hand; the comma exists only for separating patterns, and function arguments/parameters, and in the print statement syntax.
Plus, JNIL (just now I learned). The in operator allows comma-separated expressions for testing "multi-dimensional" array membership. Here is a hello, world:
$ awk 'BEGIN { a[1,2,3] = 4 ; print (1,2,3) in a }'
1
(Note that multi-dimensional arrays in Awk are simulated; a string index is generated for that 1,2,3).So to use an example from this nice short article: I might do grep GET log-file | awk blah-blah. Then awk doesn’t need to consider the lines I don’t care about. This is especially useful when iteratively writing the awk script.
Six of one half dozen of the other.
1. Writing scripts for environments that only have Busybox. Technically you can write scripts in ash, but I don’t recommend it for anything beyond a couple lines. It’s missing a lot of the features from Bash that make scripting easier, and it’s easy to get mixed up if you’re used to Bash and write things that don’t work. Awk is the best scripting language available, even if you’re doing things that don’t exactly match what it was designed to do.
2. Snippets that are meant to be copy+pasted from documentation or how-to articles. In that case, it’s often not easy to distribute a separate script file, so a CLI “one-liner” is preferred. You also can’t count on Perl, Python, etc. being available on the user’s system, but awk is pretty universal.
For most other cases, I tend to create a new .py file and write a quick Python script. Even if it’s a little more overhead, it helps keep my Python skills sharp, and often it turns out that what I actually want is a little more complicated than my initial idea anyway.
Let's say the myPgm only takes one file name as a command line parameter, then I can so something like this:
dir *.xyz /s/b | mawk "{print 'myPgm -x '$0}" | cmd
If the file paths have spaces in them, then you have to wrap the name within double quotes. I have found it challenging to output those in mawk -- at least easily any way. In this case, I have a windows version of tr.exeIn this case, I can do something like this:
dir *.xyz /s/b | mawk "{print 'myPgm -x ~'$0'~'}" | tr ~ \042 | cmd
Although somewhat crude, it is effective.F.e for %f in (.doc .txt) do type %f
docker container ls -a -q | mawk "{print 'docker rm '$1}" | cmdDoes anyone have concrete, practical examples of use cases where awk made your life much easier?
For example, recently I needed to print out certain fields of an output, but only for a given subset. So I used Awk to create a simple state machine (enable when I see the start of the subset, disable at the end), and print the fields of interest.
Also I use awk often in scripts together with fzf to interactively switch contexts e.g. in cloud provider CLIs.
Awk has a hash like perl which is very efficient. I use an awk expression to print uniq as they come in counted through the hash insert on new instead of uniq which prints at end.
Awk count unique over 300,000,000 ips was as fast as perl and python and smaller memory footprint
I don't know what LWSP gobbling is, though. Google didn't help.
[1] https://stackoverflow.com/questions/21072713/what-exactly-is...
With the parent example it means that column #2 could have any of number whitespaces between it and column #1.
That's like answering "why use Haskell?" with "Because its monads are monoids in the category of endofunctors"
If argument expansion is required, consider using read. That way the argument is given a name. Use "read a b c < x" instead of "b=$(cat x | awk '{ print $2 }')".
Just remember to pipe to read with care, as the right hand part of a pipe is a subshell in which variables are local. So "read a b c < <(echo 1 2 3)" works, "echo 1 2 3 | read a b c" doesn't.
Another way to expand arguments is to simply define a shell function and use $2. Something like "process_line() { echo $2 }".
When you have to reach for something like awk, your script would probably improve by being mostly awk. In which case most people are probably looking at perl or python anyway. Even if awk is a nice language, the arrival of perl mostly killed it. It is not wrong to say perl was the next version of awk.
whereas, tr will still leave a trailing/leading space
for example:
$ echo ' a b c ' | tr -s ' ' | cut -d ' ' -f2
a
$ echo ' a b c ' | awk '{print $2}'
b
and by default awk splits on space/tab/newlines, whereas in cut example above, you get only space as delimitercut has its uses and will be faster than awk, but it depends on the problem being solved
It possible to use `xargs -L 1` to trim the separators, but then you would also add `findutils` as deps.
I just wanted to point out that keeping the dependencies of scripts in mind when programming is also important.
e.g. a regex or even a shell-style pattern match would be very useful. `cut -d ' +' -f2` would get rid of the need for 'tr' in your example.
Even just allowing >1 char delimiters would be helpful. A text list like 'foo, bar, foobar' is too much for cut, but it would be fine if cut accepted parameters like `cut -d ', '`
https://www.bignerdranch.com/blog/a-crash-course-in-awk/
A tl;dr for the following tl;dr: it's great for quick-and-dirty processing of logs.
The tl;dr is that AWK is simple language that lets you use a line-oriented event driven programming model to process text. This abstraction is simple enough to be readable and maintainable, but powerful enough to parse and process text with line-oriented structure.
The example linked elsewhere is worth studying, because it truly is one of the prettiest pieces of programming I've seen in a long while:
If you don't know anything about AWK, here's all you need to know to grok this:
1. An AWK program is a list of [event] { code } pairs. All pairs are run in sequence on each line of input; if the "event" evaluates to true or is a matching regex, the matching code runs.
2. { code } on its own runs unconditionally.
3. The input is automatically whitespace delimited into fields; $1 refers to field 1, $2 to field 2, etc
4. NF is a special value that yields the number of fields
5. Assigning to fields replaces the text in those fields.
6. ($1+0) != 0 implicitly converts $1 to an integer; if it fails, the value is 0.
There is a lot of implicit loose typing in the language, and usage of undefined variables is idiomatic. Functions aren't easy to use, either. So it's not well-suited for programming in the large, or even the medium. But for the scale of programming seen in that link, it's truly a wonderful and simple power tool.
I'd normally write this sort of thing in python. The awk program is usually smaller than a comparable python program, but for me the main selling point is that awk programs a more "UNIX-pipe-native" than python programs are.
At some point, I'll probably rewrite these programs in a more "serious" programming language, once the file sizes get too big or something, but for now they're working great and are easy to extend.
FWIW, you can get around it. I just took a peek because I couldn't remember how to do it off the top of my head.
It looks like I had to use
gawk -F ',' -v FPAT='([^,]+)|("[^"/]+")'
to get the behavior you're looking for. Seems like I nabbed it from https://www.gnu.org/software/gawk/manual/html_node/Splitting...Agreed that this is the sane point at which to pull out python.
The risk of that seems to be if CSV allows embedded escaped quotes in a quoted string. Does it? I don't know. And CSV is pretty loosely defined. For most people it's probably "whatever Excel emits or ingests".
And I think that's why we're on the same page about pulling out python and using a module where somebody has explored what the corner cases are and dealt with them for us already.
In my experience, each implementation is more or less unique. Even if you allow escaping, different formats permit different methods of escaping and you might have to account for each! shudders
It’s much easier to support that in python than it is in awk with regexp. The awk route will eventually make you a regexp wizard, though, which confers it’s own benefits. :)
* https://blog.jpalardy.com/posts/why-learn-awk/ (also discussed on HN: https://news.ycombinator.com/item?id=22108680)
* https://adamdrake.com/command-line-tools-can-be-235x-faster-...
if it becomes lengthy (and again depends on features required), then Python is likely to have inbuilt/3rd-party libraries to make it easier to write and maintain
if it is a question of constructing command line one-liners and using it as a part of other cli tools, then awk wins easily
for example:
awk -F'\\W+' -v OFS=, '{print $NF, $2}' input.txt
prints last column and second column, where non-word characters form the field separator and comma is used as output field separatorPlus I think semantic indentation is not sane language design.
Plus you can pipe it to another command with ease.
Note: The article says that awk patterns can't capture groups. The standard doesn't provide that functionality, but if you use the widely-available gawk implementation, gawk does have that capability (use "match").
Am I the only one?
cat logfile | awk '/ERROR:/ {counts[$1] = counts[$1] + 1}; END { for (day in counts) print day " : " counts[day]}' | sort
I just needed to know how awk programs are structured, the rest is just simple programming!EDIT: I'm not sure if it's actually correct however...
> cat logfile | awk '/ERROR:/ {counts[$1] = counts[$1] + 1}; END { for (day in counts) print day " : " counts[day]}' | sort
Great first program! a bit less verbose could be
> awk '/ERROR:/ {counts[$1]++}END{...}' logfile
there are also ways of sorting the output but within (g)awk (asort & asorti) but sorting externally as you have is more flexible and engages another core which can be faster on large input
grep ERROR logfile | cut -f 1 -d ' ' | sort | uniq -c
It's not detrimental to performance since an empty cat is a no-op in a pipeline. You can have any number of them. But commands should be written for humans to understand, and inserting no-ops is a distraction to the reader.
In the trivial example, "grep needle haystack" reads better than "cat haystack | grep needle".
awk '{ print $2 }' my-file
And it gave me what I wanted. It was cool. awk '++seen[$0] > 1' filename.txtThen, you can even alias it.
awk 'NR==FNR{a[$0]; next} $0 in a' colors_1.txt colors_2.txt
sort -m <(sort -u file1) <(sort -u file2) | uniq -d # common lines
comm -12 <(sort file1) <(sort file2)
# lines unique to first file
comm -23 <(sort file1) <(sort file2)
# lines unique to second file
comm -13 <(sort file1) <(sort file2)
regarding readability, it is the same with any new tool or programming language, you'd need to be familiar with its syntax and idioms, someone not familiar with command line and sort/uniq commands will find your solution as alien as well