The State of the AWK
lwn.net
lwn.net
I am speculating of course, but as someone who evangelizes Awk the most common thing I hear from people as to why they don't use it, is that they just don't know the language all that well. For people that don't know much about Awk, it looks really complex and esoteric.
To address the Awk ignorance, I put together a talk for Linux Fest Northwest last year, and it was so popular that I gave it again (virtually due to COVID-19) this year. The conference was fully remote so the talk is on Youtube.
If you've ever wanted to learn Awk, this will take you from zero to proficient. There are exercises as well for practice:
* Presentation: https://youtu.be/43BNFcOdBlY
* Exercises:
- Source: https://github.com/FreedomBen/awk-hack-the-planet
- My solutions with explanation: https://youtu.be/4UGLsRYDfo8
I agree (I side more with Kernighan than Robbins here). My take on what hindered it from larger-scale use: lack of a big standard library (Kernighan more or less says this) -- Python and Go's stdlibs really helped them for example, and some significant quirks like lack of proper local variables (you can get them with extra arguments, but it's weird).
Just skimmed your AWK-learning talk -- very good stuff. I'm relatively proficient, but keen to try your exercises anyway!
(Some embedded systems often do not have perl, but if they are that limited, they might have a very stripped down version of awk too)
Awk is part of posix so is pre-installed and available nearly everywhere, even on many embedded devices. Busybox for example includes an implementation of Awl. Even my super cheap home router had an awk interpreter on it. Perl is also decreasingly commonly installed. On the latest Fedora for example I had to install it explicitly. Awk is still installed by default.
[1]: now available as https://metacpan.org/pod/App::a2p
The exercises look bizarre - "Questions to answer about our payroll data using awk to analyze" - doesn't make sense. That's what DBs are for, or similar tools. It's just an exercise but still.
Final coffin nail is that having awk on your CV is useless, because everyone wants full-stack cassandra, spark and hadoop, and other stupid things. Awk won't help get you hired (try to find it in https://news.ycombinator.com/item?id=23042618)
As an ordinary person whose day job is in the humanities, learning awk, sed, Emacs Lisp and other such tools has allowed me to quickly solve tasks that would have taken me ages to do manually. These tools have docs that are very accessible to the ordinary computer user (after all, they were conceived as something that clerical staff could be trained in). I am happy these tools are still around and I think they should be more widely known.
Well, since that is pretty much AWK's entire reason for being, of course you don't have a use for it.
> That's what DBs are for, or similar tools.
But if you have a text file with a lot of semi-structured data and just need to answer a few questions, using AWK will almost always be faster than designing a schema and importing into a database first, before answering the question.
Oh, and want to know a really good tool for getting data from a text file into a database? AWK!
> Awk won't help get you hired
AWK just lets you accomplish many tasks quickly and without fuss that can create value for your company, and those are things that can go on your resume. At some point in your career, pointing to specific things you have accomplished becomes more important than the keywords on your resume.
I have never had to do that with unstructured data - ah wait, I lie! I pulled data off the covid wiki page, a quick emacs regexping, a bit of macrology, pasted it into a spreadsheet (libreoffice recognises it's tab or whatever separated) and bingo, you're off.
> Oh, and want to know a really good tool for getting data from a text file into a database? AWK!
For smallish amounts, I use emacs regexps. For data too large for emacs, I'd probably use python.
> and without fuss that can create value for your company, and those are things that can go on your resume.
That's not how it works any more. They look for skills. They should not, they should do as you say, but seems not any more.
> specific things you have accomplished becomes more important than the keywords on your resume
again, no. Believe me, I have a track record of delivering, but all my job agents do is match by skills. And that's what their clients want. It's depressing.
So I disagree but thanks anyway
What's that you said, 'convenience and specialization', ah I see, that thought never occurred to me. May be that's why I reach for AWK over Python.
It's very likely that, if you knew AWK well, you could have completed this task faster.
> Believe me, I have a track record of delivering, but all my job agents do is match by skills. And that's what their clients want.
It's hard to find the few companies that actually prioritize results, but it's worth it.
I see[1]:
Title: ben@bensystem76
Dockerfile
ngingx-help.txt
test.pem
Desktop
part2
drew.txt
....
I is only periodically but seems to be the same terminal tearing in[1]. I saw some in the beginning but now at 22:20 I am so annoyed that I come here to complain! ;-)I checked both 720p60 and 1080p60 - I just pause at 22:20 - but plenty before that as well.
Maybe consider re-encode/re-upload if possible. If not: Epileptics beware!
[1] https://i.postimg.cc/6Q2RvbCn/AWK-screen-tearing-in-Youtube....
Nothing to see here. Move along...
Sure you can code anything with it, but there is always another language that will be better at this other thing, and still decent at what awk does. Awk looks more like a DSL than a generalist language.
So why learn awk to save you a few minutes for the rare cases you do need it? Just learn Python/Ruby/etc, it's a better investment, and it has a better crossplateform story.
If you're working on any kind of Unix, it's a niche that you encounter almost constantly.
> Awk looks more like a DSL than a generalist language.
Yes, it is pretty much the ultimate text processing DSL. So if you need to process a text file a line at a time, in most cases AWK is the optimal solution.
> and it has a better crossplateform story.
I think AWK is available on almost every platform where you can install those languages?
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/a...
I love playing with awk. I think it's really cool. But there's a cost to it.
I had no understanding of awk whatsoever and can now fully grok the following command, which returns the number of different cities in a structured file with a header line, addresses and the city in column 5:
awk 'NR != 1 { print $5 }' payroll.tsv | sort | uniq | wc -l
This is beautiful.
awk -F "\t" 'NR != 1 { if (!cities[$5]++) n++ } ; END { print n }' payroll.tsv1. Writing scripts for environments that only have Busybox. Technically you can write scripts in ash, but I don’t recommend it for anything beyond a couple lines. It’s missing a lot of the features from Bash that make scripting easier, and it’s easy to get mixed up if you’re used to Bash and write things that don’t work. Awk is the best scripting language available, even if you’re doing things that don’t exactly match what it was designed to do.
2. Snippets that are meant to be copy+pasted from documentation or how-to articles. In that case, it’s often not easy to distribute a separate script file, so a CLI “one-liner” is preferred. You also can’t count on Perl, Python, etc. being available on the user’s system, but awk is pretty universal.
For most other cases, I tend to create a new .py file and write a quick Python script. Even if it’s a little more overhead, it helps keep my Python skills sharp, and often it turns out that what I actually want is a little more complicated than my initial idea anyway.
Should I have looked at it more carefully?
I think my dream awk-like tool would look something like:
1. Has an “interactive” mode to see what your script is doing as you write it. Something like a combination of less, fzf, and a Bret Victor style debugger showing matches/values of variables at each line
2. Supports things that aren’t list RS-separated records of of FS-separated fields. Some formats I would like are json, csv, records where the fields are all key=value, and maybe some other formats. Support would mean some way to specify patterns for different formats
3. Extracting marching groups from regex matches.
Compared to the Object Oriented PowerShell where every column is just an object property, strongly typed and everything, the string-based bash programming to me seems absolutely bonkers.
Like... what do you do if some text doesn't fit into the space available?
How do you handle Unicode?
What about simple escaping of names like O'Toole, embedded double quotes, leading or trailing spaces that are meaningful, embedded line feeds, etc...
Eventually you have to use a full parser, not a bunch of regexes. I've found that "eventually" to mean: almost immediately, even for supposedly simple problems.
Even seeming trivial things like correctly splitting up an X.500 name as seen in LDAP or PKI is deceptively difficult. Now, if it's a LDAP name that includes quotes embedded in a CSV... err... I don't even know where to begin.
The only tricky parts are embedded line feeds and tabs, and those can be defeated with trivial escape schema, even something as simple as '\r', '\t', '\\'.
Bash quoting also works. If you master the difference between "${x[@]}" and ${x[*]}, keep in mind order of expand / split all the time, and never forget the right kind of quotes, you can have a robust system which can handle arbitrary strings. But I would not recommend this to anyone.
No reason to use tabs, when there explicit group/record/unit separators
I regularly handle 100MB files with Powershell regex , yeah, I wait few seconds here and there, but nothing except ripgrep handles this fast.
Verbosity is nonsensical arugment for anything as you have alises. Bash is equally verbose if you use long parameter names.
edit: There's also https://github.com/akavel/up which could possibly be used in combination with fzf
Although one must be careful, because awk still can overwrite files and shell out to rm.
There's a handy github repo [1] with libraries, has one for csv. For json, see [2]. But yeah, having official support would be better. Or, use tools like xsv, jq, etc that are built for those formats.
You can use match() function to get capture group contents. If you are using GNU awk, then the syntax is much easier with arrays instead of fiddling with substr and RSTART/RLENGTH. For examples, see my repo [3]
[1] https://github.com/e36freak/awk-libs
[2] https://github.com/step-/JSON.awk
[3] https://github.com/learnbyexample/learn_gnuawk/blob/master/g...
2 & 3. Structural regular expressions awk is something I long for. If someone wants to rewrite awk in Rust or whatever, please add this, that would be a killer feature.
http://doc.cat-v.org/bell_labs/structural_regexps/
EDIT: There is an Awk Language Server! - https://github.com/fwip/awk-language-server
I used it plenty over the years, but with a lot of trial and error. Now, I see it as an event engine and it just flows naturally. It’s so great for generating reports from log files, simple csv, data dumps, etc. it’s sad to see it used as ‘cut’ in most scripts.
It’s like a basin wrench. It has a small scope, but once you use it and grok it, you’d never want to do those things with another tool.
That said, I have no use for features above POSIX. It’s essentially a DSL for me, and turning it into Perl ruins the simplicity. I’ll move on to a general purpose language that others can understand and maintain vs use esoteric extensions.
This is something I've been working on as well. I know how to get any utility task I want done in Python, but it requires a lot more work than being able to pipe commands well in a shell. I've been going out of my way to do things in Bash instead as a way to become to familiar with standard Unix utils. It's a fun little thing to do that ends up resulting in more efficient use down the line.
AWK never caught on as a large-scale programming language because it's an esoteric language for processing text-based streams. Not because it didn't support "namespaces". Don't kid yourselves about "what coulda been..."
Had awk been where it is today twenty years ago it could have been the glue systems language instead of bash for non-interactive usage.
Plus, awk's performance is not great, the way you have to introduce local variables is more than awkward, and there are no structs.
I join csv files that each have a header with
awk '(NR == 1) || (FNR > 1)' *.csv > joined.csv
Note this only works if your csv files don't contain new lines. However if they do, I recommend using https://github.com/dbro/csvquote to circumvent the issue.Yesterday I used awk as a QA tool. I had to subtract a sum of values in the last column of one csv file from another, and I produced a
expr $(tail -n+2 file1.csv | awk -F, '{s+=$(NF)} END {print s}') - $(tail -n+2 file2.csv | awk -F, '{s+=$(NF)} END {print s}')
beauty. This allowed me to quickly check whether my computation was correct. Doing same in pandas would require loading both files into RAM and writing more code.However I avoid writing awk programs that are longer than a few lines. I am not too familiar with the development environment of awk, and I stick to either Python or Go (for speed) where I know how to debug, jump to definition, write unit tests and read documentation.
awk -F, 'NR > 2 {s+=$(NF)} END {print s}'I am blown away by how elegant the result is.
Human meaningful field names are made available when records are processed in your awk expressions. This is an improvement over using a delimited text dump of the record and using regular awk with meaningless $1, $2 variables.
I could be wrong, but I believe the relational operators also recognize common record field types like dates and timestamps that a text dump + regular awk couldn't.
The output of the tool is always a human readable serialization.
It is indeed mind-blowingly useful in countless contexts. Disclaimer: I'm the author
Even during that job, as the complexity of the analysis increased, my supervisor/mentor suggested "Have you heard of Perl? You might find that more convenient." And I did find it more convenient, it was clear it could do everything awk provided as well as awk (even using close to the same syntax if you wanted), plus more.
Which led to my first web job, as a university student circa 1996, writing 'dynamically generated web pages' with Perl cgi, for the university. (At this point I haven't written Perl either in at least 15 years, and don't tend to miss it).
The main reason to use awk as opposed to a more convenient language (other than "our legacy code is written in it and it would be a big investment to port it" -- a couple cases in OP) seems to be that it's present on nearly every system. But isn't also Perl?
That the functioning of awk was part of POXIS was something I just learned from this article, and which surprised me. That POXIS includes specifications of entire languages within it! I guess it makes sense, but wow.
But I guess the extensions and advancements to awk being discussed in OP would not be part of POSIX-awk, so not necessarily to be found on a POXIS system, true? Would people using awk for it's POSIX-reliability avoid using new fangled awk innovations?
Truer words have not been spoken, in my opinion. I've used AWK heavily for decades, and still use it for a wide variety of parsing.
The @include and @load directives are extremely useful for shipping your own customizations, but I prefer the maintenance priorities of the JQ maintainers [1], who understand that powerful builtins are what burn into user's minds, making a tool mentally indispensable.
Here's the readfile extension, for example: https://github.com/gvlx/gawk/blob/master/extension/readfile....
But, you're right...there aren't many extensions, and little activity around adding new ones.
To be honest, what I mostly use it for rearranging fields:
something | awk '{print $1" "$2}'
something | awk '{print "xyz:"$0}'
something | awk '{print "cp "$3" "$4}' | sh
what's a shame is how cut or other utilities make it unnecessarily bothersome to rearrange fields.I actually made a utility once called "words" that does:
ls -l | words 1 3 4-5When strptime is in an add-on library, I can't use it.
When processing event streams, it's very natural to express transformation as FILTER + ACTION rules.
AWK is an embodiment of this idea for the domain of text processing where events are lines (or multi-line records).
DTrace uses the paradigm to process low-level system events (like scheduling events, or (kernel) function calls).
It's a good paradigm, worth using.
However, it's easily replicated in every other language with a loop:
while true {
rec = read()
if( filter1(rec) ) { action1(rec) }
if( filter2(rec) ) { action2(rec) }
...
}
You don't need a special language to do event processing like this.Literally every time I use awk I have to google the syntax because I don't use it often enough for it to persist in long term memory.
If only awk's syntax were a subset of a language I already knew, like python or JS
Not necessarily—sometimes it makes sense to use Datalog instead of Prolog, for example.
E.g. Perl and Ruby "-pe"/"-pie" command line options, and assorted related variations both offers ways to do awk-like single-liners that are just minor syntactic sugar over full scripts. But writing a tiny wrapper like that for your preferred language should be easy enough.
"ruby -an -F: -e ' puts $F[0] '" is equivalent to awk -F: ' { print $1 } ' for example
-n adds an implicit "while gets(); ...; end" loop; gets by default sets $_; -n adds an implicit $F = $_.split($;); -F sets $; ; -e takes what follows as a Ruby expression.
"The State of the AWK" should actually be called "The AWK word"
However I learned to like it and use it often in CLI one-liners (mostly to cut out or reformat specific columns, probably the most common usage).
The handful of minutes it took to learn have repaid that investment thousands of times over.
Yes, there will be better solutions for individual cases. But awk is always installed, always available, and for anybody quickly querying semi-structured text data on the Unix command line it is a godsend.