Awk: Power and Promise of a 40 yr old language (2021)
fosslife.org
fosslife.org
The AWK book was one of the fundamental books I used to teach myself some coding. The precision of the language is remarkable; I wonder how much will be different in the new edition.
It starts off with similar background and praises for a few minutes before spending the rest of the hour crash coursing you through awk, then leaving you with a very approachable, digestible, and realistic set of exercises (and optionally, a second video covering his solutions).
I keep a text file with a copy of the questions and my solutions on a public gh repo, so I can quickly refer to it from anywhere when needed.
I am much more powerful on the command line because of it.
[Edit]
[0] https://github.com/FreedomBen/awk-hack-the-planet
Originally encountered @ https://news.ycombinator.com/item?id=25144697
For anyone else considering investing the time: I am extremely satisfied with the 2 hours it took to learn + practice the basics. As far as high-yield learning investments go, I’d already put awk up there with my time spent learning Vim and Git.
https://github.com/docker-library/bashbrew/blob/master/scrip...
https://github.com/docker-library/postgres/blob/master/apply...
https://github.com/docker-library/postgres/blob/master/Docke...
awk '/foo=([0-9]+)/ { print $1 }'
although I suppose the syntax would have to be different since $1 has a meaning already.Yes, gawk has a function that returns capture groups, but it's a bit verbose for one-liners. Instead I switch to Perl:
perl -nE 'if (/foo=([0-9]+)/) { say $1 }'
But I wish I could just use Awk. awk -F= '$2 ~ /[0-9]+/ { print $2 }'
With imaginative choice of FS and RS you can push it very far.Whether other people having to deal with such code will appreciate your imagination is another matter, though.
Edit: I missed the detail where you want to specifically match "foo" as lhs, and anywhere on the line. So the correct condition would be even lengthier : ^ ) You have a valid point. Captures would provide for shorter patterns.
grep "\((\d+)\)" -print "$1" file.txtIt'll do capture groups with search | replace
# remove square brackets that surround digit characters
$ echo '[52] apples [and] [31] mangoes' | rg '\[(\d+)]' -r '$1'
52 apples [and] 31 mangoesAdd `-o` option to get only the digits:
$ echo '[52] apples [and] [31] mangoes' | rg -o '\[(\d+)]' -r '$1'
52
31I cite sources way more often than not, this time I got lazy after dithering over whether to go with the definitive ripgrep source page [2] or a decent looking third party(?) tutorial .. pressed for time I did neither.
[1] https://learnbyexample.github.io/learn_gnugrep_ripgrep/ripgr...
grep -oP '\(\K\d+(?=\))'
The above will give all matches in a line though. You can remove `(?=\))` part if numbers are always enclosed in `()` or if you don't care about the `)`[1] : https://unix.stackexchange.com/questions/536657/how-to-refer...
Not really verbose for this particular example though:
$ echo 'foo=42' | awk 'match($0, /foo=([0-9]+)/, m){print m[1]}'
42
$ echo 'foo=42' | perl -nE 'if (/foo=([0-9]+)/) { say $1 }'
42
# TIMTOWTDI
$ echo 'foo=42' | perl -nE 'say $1 if /foo=([0-9]+)/'
42[1] 3.6 using Dynamic Regexps : https://www.gnu.org/software/gawk/manual/html_node/Computed-...
So I added a feature to the TXR Lisp awk macro. There is now an Awk variable called res which holds the result of the condition. If the condition is a regex, then that has the matching part. The fact that the action is executing tells us that the result of the condition is true; but res gives us the specific true value, like the it in anaphoric if macros.
This made it into release 284.
sed -n "s/SENTINAL//g;s/\(.*\)\(foo=\)\([0-9]*\)\(.*\)/SENTINAL\3/;/SENTINAL/!d;s/SENTINAL//;/./p"
shortened to x=$(echo x|tr x '\34');
sed -n "s/$x//g;s/\(.*\)\(foo=\)\([0-9]*\)\(.*\)/$x\3/;/$x/!d;s/$x//;/./p"
Using flex is another option. Faster than AWK, Perl, sed and similarly ubiqitous. flex -8iCrf -o/dev/stdout << eof|cc -xc -O3 -std=c89 -pedantic -W -Wall -static /dev/stdin
int fileno(FILE *);
#define J BEGIN
#define E ECHO
int x;
%option noyywrap noinput nounput
%s x1 x2
%%
foo=[^\n] yyless(1);J x1;x++;
<x1>[0-9]+ if(x==1)E;J x2;
<x2>\n E;x=0;J 0;
<x2>.
.|\n
%%
int main(){ yylex();exit(0);}
eofAnd I learnt a new thing from this article, even though I have been using awk for decades... Functions. Yep!
I only have a couple of surviving examples of the code from back then, but here they are for the curious:
LJPII.AWK is probably the best example. It made a nicely formatted printout of source code on my HP LaserJet II printer. I wish I had one of the printouts it generated, but they are long gone.
Hmm... I wonder if my Brother printer supports the old LaserJet II control codes? Or maybe there is an emulator online?
The code was written for Thompson Awk (TAWK), so some bits would need to be adapted to modern Awks.
Here's a recent one of mine (albeit as usual embedded in a /bin/sh script) that I should try to functionify!
https://www.earth.org.uk/script/BibTeX-to-HTML.sh
Thanks!
https://news.ycombinator.com/item?id=30343373 (38 comments)
I wish this "AWK-like language" came with performance benchmark comparisons.
Edit: Thank you, benhoyt!
https://www.softwarepreservation.org/projects/poplar/doc/Mor...
The finest was a 3rd party who accidentally added a space in a data feed. This was dutifully sucked out of their SFTP server via a bash script, pre-processed using awk into the standard internal format and then picked up by a cron job which ran a python script to inject it into postgres. The outcome was of course that the columns were offset by 1. This caused a huge asset valuation dip and some market alarm.
A proper parser would have rejected the whole dataset as the data row could not be parsed.
Model inputs in great detail and you can throw out a lot of invalid data, but it takes longer to get the code running. Model the input only very crudely and you're up and running quicker, but more open to broken expectations.
Of course, it's way too easy to go the latter route for most programmers...
edit: especially since unicode support allows for use of native Klingon fonts.
[1] : https://www.gnu.org/software/poke/
[2] : https://kernel-recipes.org/en/2019/talks/gnu-poke-an-extensi...
Freepascal/Delphi user and Smalltalk lover here.
I beg to differ.
They're currently trying to transport it to C#, but its slower going than developing in Delphi, ironically, the thing that makes us move to C# is developer availability.
The multi-decade investment in cobol for critical systems (banking) does not make for quick/easy switch.
And handling it with awk: https://stackoverflow.com/questions/45420535/whats-the-most-...
The way gawk handles functional parameters can simulate records (dotted r.attrib notation)
Gawk @ additon [1] permits namespaces (include,load) and other fun syntax/sematic jit enforced separation.
[1] https://www.gnu.org/software/gawk/manual/html_node/Index.htm...
Practically everything you can do in AWK can be done just as easily and quickly in Perl. And Perl absolutely wins when you need to do that one extra thing that AWK really just can't do.
And I say this as a person who switched over from Perl to Python eons ago.
Whatever things I couldn’t do in sh or Awk, I would do in C.
Seriously, Perl is an okay language for quick and dirty things of a tiny/small size. Yeah, it's not the best language for a large development project, but if you do need to parse /etc/passwd or something, not only it's perfectly good as-is, but you'll certainly find something on CPAN that already does it well.
I can't imagine why would one want to do that kind of thing in C. It's just unnecessarily painful, and you'll spend 90% of the time on doing things that don't solve the actual task you need to be solved.
Yeah, in modern times it's gone way downhill, but that's mostly if you intend to do something big with it. I wouldn't use it to start a new, fully featured CMS. But for sysadmin type stuff as an alternative to sh/awk it's still just as usable as ever.
That said, I have to revise my original comment. I forgot that I was a big fan of Tcl/Tk/Expect back in the day. So it's not like my taste is better than anyone else's :)
You could also use
BEGIN {
FS = “: “
}
Which sets the field separator. This would put the keywords into $2 and remove the need for the gensubs :)Anyway, cool blog. I'm reading it.
Sadly I've learned it & cheatsheeted it for future reference but never find myself reaching for it. Part of it is prefering python over shell scripting maybe - awk fits better in a shell scripting world.
foo | awk '{print $4}'
I know there are easier ways. Heck I even wrote a command called "words" (like "foo|words 4" or "words 1 3 2")... but I forget to use it.Also, this blog post helps with the above. For some reason I like how the author cites technical reviews of the article and email correspondence as references.
And congratulations. I’ve been trying for the past year to learn awk, and while I can pretty reliably split text files and extract columns, I’m pretty far from being able to grok awk.
Why not just say understand and move on?
See the "Defining Fields by Content" topic in the manual, which is based around the FPAT variable, specific to Gawk:
https://www.gnu.org/software/gawk/manual/html_node/Splitting...
TIL