Understanding Awk
earthly.dev
earthly.dev
BEGIN{ item_type = 1; item_name = 2; price = 3; sale = 4; #etc }
Now, in place of $1, you'd say $item_type which significantly improves overall readability of the code.
E.g.:
{
name = $1
dob = $2
grade = $3
# ...
# Do stuff with name / dob / grade, etc.
}
If the data are structured, so that there are multiple record types (typically defined by prefix or some other regex) you can put variable assignments within each block. /^rectype1/ { var1 = $1; var2 = $2, ... }
/^rectype2/ { varA = $1; varB = $2, ... }
I prefer to leave BEGIN blocks for defining constants or tables and such.When I wrote my introduction to JQ someone mentioned JQ was tricky but super-useful like AWK. I nodded along with this, but actually, I had no idea how Awk worked.
So I learned how it worked and wrote this up. It is a bit long, but if you don't know Awk that well, or at all, I think it should get the basics across to you by going step by step through examining the book reviews for The Hunger Games trilogy.
Let me know what you think. And also let me know if you have any interesting Awk one-liners to share.
Neat discussions around that sort of thing at least here: https://news.ycombinator.com/item?id=23427479
echo -e "foo bar baz" | choose -1 -2
vs awks echo -e "foo bar baz" | awk '{ print $2, $3}'
I love the effort people are putting into reinventing the core unix tools.I think I'll stick with Awk for now though.
$ choose
bash: choose: command not found...
However, I feel I really should have taken the time to learn Awk better as it could probably be done there, and simply! (It was a good excuse to tinker with rust, but that's an aside.)
$ awk -F: '{len=length($1);if(len>max){max=len;user=$1}}END{print user,max}' /etc/passwd ls -l | tr -s ' ' | cut -d ' ' -f 5 $ printf " one two three" | tr -s ' ' | cut -d ' ' -f 1
$ printf " one two three" | awk '{print $1}'
one ps ax | sed 's/^\s\+//; s/\s\+/ /g;' | cut -d ' ' -f 4 echo -e '1\t2\t3\t4\t5' | expand -t 1 | cut -d ' ' -f 3 $ printf "a b c d e\n1 2 3 4 5" | perl -lanE 'say "$F[2] $F[4]"'
c e
3 5It turns out though that this is because Perl and later Ruby were inspired by AWK and even support these line by line processing idioms with BEGIN and END sa well.
ruby -n -a -e 'puts "#{$F[0] $F[1]}"'
ruby -ne '
BEGIN { $words = Hash.new(0) }
$_.split(/[^a-zA-Z]+/).each { |word|
$words[word.downcase] += 1 }
END {
...But once I got the idea of aggregating the book review data from amazon I felt I had to see it through.
Since then I've made a point of finding man-pages from other systems whenever the manual for a GNU tool is a bit daunting. It tends to lower the learning threshold quite a lot, honestly.
$ man gawk | wc
1568 13030 94207
$ man -l /usr/share/man/man1/awk.1plan9.gz | wc
214 1579 10956
Not trying to detract from this great guide. Just a general tip :)https://www.cs.princeton.edu/courses/archive/spring19/cos333...
Brian Kernighan has a knack for explaining languages very precisely and elegantly.
https://www.pement.org/awk/awk1line.txt
That author also edited the "USEFUL ONE-LINE SCRIPTS FOR SED" page:
https://www.gnu.org/software/gawk/manual/gawk.pdf
and
http://www.cs.unibo.it/~sacerdot/doc/awk/nawkA4.pdf
enough
https://archive.org/download/pdfy-MgN0H1joIoDVoIC7/The_AWK_P...
The gawk one is useful if you're into some of the gnuism specifics
Having relied heavily on the (unofficial, non-GNU) gawk manpage extensively (it's quite good), I instantly started learning very useful features reading the GNU docs. (I still need to fully internalise those). Yes, the full manual is very much better than the manpage.
(Also recommend The AWK Programming Language mentioned here, though I'd suggest the GNU manual adds to that as well.)
For any complex text processing, it's way better and more robust than having a super long pipeline of a bunch of sed/grep.
Most recently, I used awk in a script that parses /proc/mount to grab the mountpoint of a partition, or print something different if the partition isn't mounted. Doable with a bunch of sed/grep and some shell logic? Definitely. But easier and cleaner in AWK, and equally easy to inline in a shell-script.
Also, happy 1337 karma day :).
Including machine hours of work.
Wasn't there a famous story of replacing a Hadoop cluster with an awk script (which was a couple orders of magnitude faster)?
Oh yes, there was: https://news.ycombinator.com/item?id=17135841
I think parsing logs to find pain areas or potential exploit/exfil is a map reduce job, for instance, and grep or awk can manage that just fine.
That said, if you just want to supplement your knowledge of other shell tools and pull out something that can do some obvious text munging, AWK has always looked attractive for the task to me.
There are two common sources of awk for Windows, for example, that drop one exe to provide the interpreter:
http://unxutils.sourceforge.net/
Perl simply wasn't designed to do that.
Had Perl held to a single binary, it could be inside busybox.
Perl cannot be inside busybox, and that hasn't been possible since 1993.
Has a recent version of GAWK (Gnu AWK).
I think if you install git for Windows, you get AWK as well.
Yes, but you also get perl there too.
Internet always seems the similar, pipes|garbage in| and munging on urban dictionary.....
To anyone processing huge quantities of text and text files, someone very likely had the same problem you faced back in the 1980's and there's a Unix/GNU tool for it already.
I'm a huge Linux/Unix fan but sometimes a rethink really works out. I hope Linux will get something similar. I know Powershell is available for Linux but without an adapted userland there's not much benefit
This is probably not important for embedded but doesn’t a pipeline of small scripts (which could be in awk) give you better threading support?
Xargs, GNU parallel or even make then scale that out really quickly.
- Avoid programming in strings (especially in Bash, where nested quotes are full of pitfalls)
- Avoid magic switches that change behavior (like -F)
- Avoid terse or cryptic variable names (like $NF)
- Avoid terse and magical syntax (sorry Perl, happy to leave you behind me)
- Avoid programs that are hard to read
- Avoid programs that are difficult to debug while writing them
- Avoid programs that ignore types
For these reasons, I prefer to avoid awk for anything except the most trivial of tasks. I think the prevalence of scripting languages and the speed of execution and debugging today has made awk not as necessary as it may have been in the 70s. And as to the first point, I'm aware you can write awk scripts in files, and I feel like if your script has gotten complex enough that you need a file, you're creating something unmaintainable and unreadable that would be better suited in a different programming language.
Edit: I should add this article is great and a good introduction to awk, regardless of my personal taste for the tool.
Looking at my command history, I mostly use awk to extract a field like this:
<something> | awk '{print $3}'
(I know "cut" is supposed to do the same thing, but it was never reliable for me - maybe tabs/spaces?) $ cat /bin/awkmail
#!/bin/gawk -f
BEGIN { smtp="/inet/tcp/0/smtp.yourco.com/25";
ORS="\r\n"; r=ARGV[1]; s=ARGV[2]; sbj=ARGV[3]; # /bin/awkmail to from subj < in
print "helo " ENVIRON["HOSTNAME"] |& smtp;
smtp |& getline j; print j
print "mail from: " s |& smtp; smtp |& getline j; print j
if(match(r, ","))
{
split(r, z, ",")
for(y in z) { print "rcpt to: " z[y] |& smtp; smtp |& getline j; print j }
}
else { print "rcpt to: " r |& smtp; smtp |& getline j; print j }
print "data" |& smtp; smtp |& getline j; print j
print "From: " s |& smtp; ARGV[2] = "" # not a file
print "To: " r |& smtp; ARGV[1] = "" # not a file
if(length(sbj)) { print "Subject: " sbj |& smtp; ARGV[3] = "" } # not a file
print "" |& smtp
while(getline > 0) print |& smtp
print "." |& smtp; smtp |& getline j; print j
print "quit" |& smtp; smtp |& getline j; print j
close(smtp) } # /inet/protocol/local-port/remote-host/remote-portFor example: https://www.gnu.org/software/gawk/manual/gawkinet/html_node/...
for example:
#!/usr/bin/python
import smtplib
from email.mime.text import MIMEText
msg = 'hi'
subj='read this!'
smtp_server='mail.example.com'
smtp_from='me@example.com'
smtp_to='you@example.com'
m = MIMEText(msg)
m['To'] = smtp_to
m['From'] = smtp_from
m['Subject'] = subj
s = smtplib.SMTP(smtp_server)
s.sendmail(smtp_from, [smtp_to], m.as_string())
s.quit()
of course, you seem to think in gawk so if that works
for you that's what you should continue doing!by the way, I hacked this example from another script which attached a logfile:
with open(arg.logfile) as f:
log_contents = f.read()
m = MIMEText(log_contents)
you can also use: from email.mime.image import MIMEImage
from email.mime.text import MIMEText
from email.mime.multipart import MIMEMultipart
and then: m = MIMEMultipart()
m.attach(MIMEText('\n\n%s\n\n'%xkcd_img_title))
m.attach(MIMEImage(xkcd_img))The AWK script doesn't need libraries, so it can actually be useful in places where you have awk but not Python.
Awk will treat it as having two columns (by default), while cut will treat each space as it’s own column.
Awk is also a little nicer for whitespace; cut makes specifying the delimiter (with say “-d\ “) a little more vexing.
There is a lot of wisdom in the things you avoid, however I would ask one question, "How often do you use it?"
For me, the best systems are those that can be wordy and prescriptive but as you get to know them you can use more short hand so they "get out of the way" as it were. A good example of that philosophy is keyboard short cuts. When I'm learning a program I'm happy to pause and sling the mouse around to find the thing I need in the labeled menu stack with an appropriate name which also tells me what the keyboard short cut is for that thing. Then as I get better I can just use the short cut and my workflow gets faster. Once I've internalized the keymap my flow is held up by how fast I can think, not by how fast I can take my hand off the keyboard, move the mouse, click and then put it back on the keyboard.
Awk is one of those things that once you internalize what it can do, you can use it for a lot of stuff, and you can do it quickly.
col1,col2,col3
1,2,3
4,"hello, \"world\"",6
"7 buckets",,9
To get the usual awk experience with this very common file format, exactly the type of thing you want to parse with awk, you first need to install gawk, then use a big FPAT regex that needs to be adjusted for any new CSV variant.
I would love to see awk with "CSV mode", where it intelligently handles formats like this if you just pass a flag. I think awk would do well to differentiate itself with excellent 2d dataset parsing functionality, but at least catchup up to the average scripting language would be great.
I'm half expecting someone to say "just pass -csv it does what you want" and if so I'll be very excited.
https://news.ycombinator.com/item?id=28708145
...but if your files are CSV, there is a CSV extension for gawk
@include "csv"
BEGIN { CSVMODE = 1 }It's funny searches for awk CSV seem to yield a bunch of SO questions where the answers are increasingly cumbersome regexes instead of this extension.
Of course, you can't count of this extension being widely installed, but it's great for my own desktop.
I just end up using Python/Perl but I do have a soft spot for awk so it would be cool if good support was built-in.
XML is somewhere in the middle--I've seen some horrible abuses of CDATA sections way back when--but at least there are accepted ways to prove what's invalid.
awk -f ./ucsv.awk -e '{print $5}'
Also this> 4,"hello, \"world\"",6
Is incorrect per https://tools.ietf.org/html/rfc4180 so you should just fix it with a sed -i 's/\\"/""/g' and then just parse as normal.
In principle:
cat textfile.csv | csvquote | awk -f myprogram.awk | csvquote -u > output.csv
Also works for other text processing tools like cut, sed, sort, etc.- Strings are subtly complex, but strings are not variables. You can assign a string, and later handle it as a variable, and not deal with any of the specifics of string-iness. Likewise, you can take a variable, and later treat it as a string (for loosely or not-typed variables).
- Magic switches are not magic, they are options. Virtually every program takes options. Sometimes they impact a lot of things, sometimes a little. Only the context determines how much is "too much".
- Terse/cryptic variables allow you to write complex expressions in a compact form. This allows you to read more in a small space, making it easier to reason about or form complex expressions. Human languages are flush with these, as is mathematics. But you have to balance the terse, cryptic and magical with guilelessness, or it becomes a mess.
- Terse and magical syntax is, again, a feature, not a bug. Using magical syntax I can do in a few characters what would take me many lines with a traditional language, and as we all know, increased number of lines correlates to bugs, in addition to simply making it harder to grok.
- Types aren't ignored, but they may be very loosely enforced. If you want to write a quick program to get something done, typing is a curse. If you want to write a very thorough program, typing is a blessing. In many cases, loosely or untyped programs actually work better than their typed cousins, because they allow for more unexpected behaviors without failing. Failing early and often may be a modern trend, but... it literally means things fail more, and this is often not desirable.
Caveats:
- Programs that are hard to read do indeed suck, and it takes lots of experience to make some kinds of programs easier to read. But that's not an indictment of the program, it's an indictment of the person who wrote it. We don't indict English when somebody writes a document that's impossible to comprehend.
- Interestingly, some of the more popular languages are the worst to debug. Perl is probably one of the easiest languages to debug, not inconsequently because of how good the interpreter is at suggesting to the user what the actual problem was and almost exactly how to fix it.
And if it's something you don't know and are experimenting with, you can also open ipython and play around with the data (kinda like he's doing throughout the article) and keep your variables between each run without having to write intermediate values to disk or keep piping over and over.
Let alone the huge stdlib and 3rd party lib you have access to if needed.
But I sometimes use Awk scripts to quickly analyse a log file to get information you need or to generate a configuration file or something.
I think of the strict strongly-typed nature of a language like Java as equivalent to the formal language a lawyer might use in a legal document: There shouldn't be any place for ambiguity or misunderstanding.
Programming in Awk, on the other hand, is more like having an informal conversation with an old friend; those cryptic variable names is like the slang words that both of you understand and the magical syntax is like the dialect you share that may sound a bit strange to outsiders.
I think there is a place for both approaches in the broad space of computing.
Awk has a reputation for being hard to read (as noted in stevebmark's comment), but when I was using it actively, I tried to treat it as a serious programming language and write readable programs in it.
Several years ago I tracked down a couple of my old Awk programs from around 1990 and posted them here:
SHANEY.AWK is an implementation of the infamous Mark V. Shaney:
https://www.clear.rice.edu/comp200/09fall/textriff/sci_am_pa...
This was probably the first program that made me really impressed with Awk. People were writing rather complicated Shaney implementations in C, and I thought, "this could be really simple in Awk." And it was!
LJPII.AWK is the Awk program I'm most proud of. This was in the days when we had tiny screens and no multiple monitors and you always printed out your code to read it. In my circles we also fond of inserting "separator lines" between functions, in various formats such as this one:
// - - - - - - - - - - - - - - - - - -
So I wrote LJPII to print source code in "two up" format (two pages side by side in landscape mode) on my LaserJet II. It also converted the separator lines into graphical boxes, and tried to avoid splitting a function across multiple pages. It wasted some paper but made nicely readable printouts.I wish I still had some of my old printouts, but they are long gone. One of these days I will have to see if I can update the code to work with the LaserJet emulation in my Brother printer! (It should mostly work, but I wrote this in the old Thompson Awk for DOS, so there are a couple of non-standard things in it.)
Looking at the code again, it's amusing to see some old Windows Hungarian notation which was popular/notorious back then, for example an "f" prefix for a boolean (flag) value, and "af" prefix for an array of flags.
Hungarian aside, I tried to make this code as readable as I could.
Random fun fact! Someone who used to be an avid Awk programmer is Will Hearst (William Randolph Hearst III). It's been many years since I talked with him, so no idea if he still does any Awk programming.
Actually, AWK is a domain specific programming language. When you start treating AWK as such then you can really gain an appreciation for it. I too treated it as a dumb one liner relegated to ingesting cryptic regexp one liners in shell scripts. After reading the original AWK book it completely changed my outlook on the language. I had no idea you could define functions or perform basic math so one could use it for very basic tabular operations such as spread sheets. AWK can even be used as a standalone language outside of shell scrips by writing a program, insert a shebang on the first line calling awk, and mark the file as executable.
But yes, I agree that the original AWK book is really good. After covering some basics and the language reference, it has some fun projects that you can build with AWK.
Every time I read about Awk, I feel very intrigued, but AFAIK, everything Awk can do, Perl can just as well. Is there a compelling reason to learn Awk if you know Perl beyond pure curiosity?
It is installed everywhere pretty much so it's nice to know. I keep thinking the same since I know Perl.
Well, that's what I meant by "pure curiosity". ;-) Which is a perfectly valid reason to learn a language, of course, possibly the best reason one can have.
But I guess when I need a quick and dirty one-shot script, I'll stick to Perl for the time being. It's not officially standard like Awk, but it is available in the default installation of pretty much every open source Unix-like system I use except FreeBSD where it's one of the first packages I install.
Hypothetically, Perl is also available on systems that are not very Unix-like, such as OS/400 (or "IBM i" as it is called these days, apparently) or VMS. (Hypothetically in the sense that I have never had direct contact with one of these systems and have no expectation I ever will, sadly. ... Well, maybe I should count myself lucky - either way, I will probably never know for sure.)
So don't bother as long as Perl is ubiquitous.
While you can write on one line of Perl anything that you can write in one line of Awk, the Perl line will be longer.
So for the simple tasks that you would do from the command line or in a shell script, you can usually save time when writing for Awk.
For complex tasks, where you need more complex control than the implicit loop of Awk and you also need to define various extra variables besides those pre-defined, Awk is no longer more concise.
Then obviously Perl or other scripting languages become more convenient.
The Awk variant had 350 bytes, while the identical Perl variant had 450 bytes.
The extra length in Perl was due to "$" prefix on all variables, an extra "split()" and an extra "while()", which were not needed in Awk.
For slightly simpler versions of those scripts executed as one-liners, the difference in length was even much more in favor of Awk, because Perl needed extra command-line options to execute the script in the command line, while that is the default behavior of Awk.
I just about never have to refer to the man page. I also found that I can go for months without writing any Awk and then knock out a quick script without having to relearn anything.
I'll concede that there are things for which Awk just isn't powerful enough and perhaps I was never that good with Perl to begin with, so if you're already well versed in Perl YMMV
If anyone is interested in learning more, I built a conference talk to teach awk, and a set of exercises also that has gotten pretty positive feedback:
Presentation: https://youtu.be/43BNFcOdBlY
Exercises (for you to try): https://github.com/FreedomBen/awk-hack-the-planet
Exercises (me solving): https://youtu.be/4UGLsRYDfo8
@include "csv"
BEGIN { CSVMODE = 1 } echo one two three|awk '{print $2}'
Are there other ways to do this. Are they faster. cat > awc
#!/bin/sh
test $# -eq 1||exit
exec tr \\40 \\11|exec cut -f "$1"|exec tr \\11 \\40
^D
echo one two three|awc 2
Test it on a file to see if it is faster than awk. time awk '{print $2}' file
time awc 2 < file x(){ tr \\11 \\40;}
x|cut -f "$1"|x x(){ tr \\40 \\11;}
y(){ tr \\11 \\40;}
x|cut -f "$1"|y set a 1
set b 2
define add (n,m) $n + $m
set result [add a b]
I think it would be simple enough to come up with some Awk pattern/actions to parse the above and execute the commands.Awk: The Power and Promise of a 40-Year-Old Language - https://news.ycombinator.com/item?id=28441887 - Sept 2021 (118 comments)
Awk is the coolest tool you don't know - https://news.ycombinator.com/item?id=27039608 - May 2021 (20 comments)
CGI with Awk on OpenBSD Httpd (2020) - https://news.ycombinator.com/item?id=27037113 - May 2021 (22 comments)
The State of the Awk - https://news.ycombinator.com/item?id=25142867 - Nov 2020 (58 comments)
Awk: `Begin { ` Part 1 - https://news.ycombinator.com/item?id=24940661 - Oct 2020 (106 comments)
Show HN: Awk-JVM – A toy JVM in Awk - https://news.ycombinator.com/item?id=23612910 - June 2020 (27 comments)
Running Awk in parallel to process 256M records - https://news.ycombinator.com/item?id=23394024 - June 2020 (101 comments)
The State of the AWK - https://news.ycombinator.com/item?id=23240800 - May 2020 (86 comments)
Awk in 20 Minutes (2015) - https://news.ycombinator.com/item?id=23048054 - May 2020 (126 comments)
Show HN: An eBook with hundreds of GNU Awk one-liners - https://news.ycombinator.com/item?id=22758217 - April 2020 (48 comments)
Learn Awk by Example (2019) - https://news.ycombinator.com/item?id=22455779 - March 2020 (29 comments)
Awk As A Major Systems Programming Language, Revisited (2018) - https://news.ycombinator.com/item?id=22304017 - Feb 2020 (80 comments)
Why Learn Awk? (2016) - https://news.ycombinator.com/item?id=22108680 - Jan 2020 (235 comments)
Learn Just a Little Awk (2010) - https://news.ycombinator.com/item?id=21101478 - Sept 2019 (69 comments)
Awk by Example - https://news.ycombinator.com/item?id=20308865 - June 2019 (21 comments)
Removing duplicate lines from files keeping the original order with Awk - https://news.ycombinator.com/item?id=20037366 - May 2019 (154 comments)
GNU Awk 5.0 - https://news.ycombinator.com/item?id=19671983 - April 2019 (49 comments)
Learn just a little Awk (2010) - https://news.ycombinator.com/item?id=17322412 - June 2018 (244 comments)
The Awk Programming Language (1988) [pdf] - https://news.ycombinator.com/item?id=17140934 - May 2018 (207 comments)
Learn to use Awk with hundreds of examples - https://news.ycombinator.com/item?id=15549318 - Oct 2017 (116 comments)
Awk for multimedia - https://news.ycombinator.com/item?id=15410259 - Oct 2017 (24 comments)
Awk driven IoT - https://news.ycombinator.com/item?id=14735752 - July 2017 (35 comments)
Skip grep, use awk - https://news.ycombinator.com/item?id=14692233 - July 2017 (130 comments)
Awk vs. Perl (2009) - https://news.ycombinator.com/item?id=14647022 - June 2017 (71 comments)
The Awk Programming Language (1988) [pdf] - https://news.ycombinator.com/item?id=13451454 - Jan 2017 (103 comments)
Show HN: 3D shooter in your terminal using raycasting in Awk - https://news.ycombinator.com/item?id=10896901 - Jan 2016 (55 comments)
Awk in 20 Minutes - https://news.ycombinator.com/item?id=8893302 - Jan 2015 (85 comments)
An Awk Primer - https://news.ycombinator.com/item?id=7961848 - June 2014 (28 comments)
A Crash Course In Awk - https://news.ycombinator.com/item?id=6578960 - Oct 2013 (37 comments)
Why Awk for AI? (1997) - https://news.ycombinator.com/item?id=5725291 - May 2013 (53 comments)
Ask HN: Do people build websites in Awk? - https://news.ycombinator.com/item?id=5041323 - Jan 2013 (12 comments)
Why you should learn just a little Awk - A Tutorial by Example - https://news.ycombinator.com/item?id=2932450 - Aug 2011 (76 comments)
Announcing my first e-book "Awk One-Liners Explained" - https://news.ycombinator.com/item?id=2674284 - June 2011 (24 comments)
AWK-ward Ruby - https://news.ycombinator.com/item?id=2486231 - April 2011 (31 comments)
Music with AWK - https://news.ycombinator.com/item?id=2294909 - March 2011 (15 comments)
Exercise #1: Learning awk Basics - https://news.ycombinator.com/item?id=2210085 - Feb 2011 (20 comments)
Why you should learn at least a little bit of Awk - https://news.ycombinator.com/item?id=1738688 - Sept 2010 (62 comments)
Don't MAWK AWK - the fastest and most elegant big data munging language - https://news.ycombinator.com/item?id=815529 - Sept 2009 (22 comments)
I too am happy to see more Awk material in the world, once I learned a bit about it I started reaching for it more and more.
the path that /usr/bin/env returns is (essentially) a global variable that can change underneath you, right? I mean that just screams "variable that may be changed by others" to me.
I've never understood why /usr/bin/env exists.
The /usr/bin/env trick will work on a wide range of systems, in which even common utilities might have numerous locations: /bin, /sbin, /usr/bin, /usr/bin/local, /opt, or others. If you're writing scripts for portability and ohers, this has value.
That said, /usr/bin/env fails on Android/Termux AFAIU.