CLI text processing with GNU awk
learnbyexample.github.io
learnbyexample.github.io
sed -ne 'x' -e '/PREV/ {x; /CURR/ p; x}'
> echo -e "PREV\nCURR\nCURR\nCURR\nPREV\nRED" | sed -ne 'x' -e '/PREV/ {x; /CURR/ p; x}'
CURR
This uses sed's hold buffer. I'll break it down: sed -n
The `-n` tells sed no to print anything out. By default, sed prints out whatever is left when processing. We'll tell it with the `p` command when to do so. sed -ne 'x'
`-e` indicates we are specifying one of the scripts sed will execute. The command `x` switches the current line with whatever is in the hold buffer. We'll do this on every line. sed -ne 'x' -e '/PREV/
The next command will only run on lines that contain `PREV`. But, because we've been putting lines in the hold buffer, we'll only execute on lines after `PREV` when it has been switched out of the hold buffer. sed -ne 'x' -e '/PREV/ { ... }'
The braces indicate all commands should be run when we see this match. sed -ne 'x' -e '/PREV/ { x; ... }'
First, we switch the hold buffer with the line buffer. sed -ne 'x' -e '/PREV/ { x; /CURR/ p; ... }'
Then, we only print out the line if it contains CURR. sed -ne 'x' -e '/PREV/ {x; /CURR/ p; x}'
Finally, we switch them back in case there is overlap in our matches. (Give `echo -e "PREV\nPREVCURR\nCURR\nCURR\nPREV\nRED"` a try with this.)All that said, I'm pretty sure the `awk` script is much simpler and more direct, but I wanted to share how one might accomplish this was sed.
The time I spent learning this probably would've been better spend on awk, but this tutorial[0], was so good and so easy, it taught me nearly everything I know about sed.
Plan9's awk(1)[0] man page provides a precise and concise (a few paragraphs) presentation of the core features of all awk implementations.
Tutorials bring practical knowledge, but often lack complete and self-contained descriptions of those nifty little tools.
It's short and to the point, has good examples, and cuts most of the usual fluff like "what is a variable?". Its base assumptions are: You know how to program, and you're here to learn AWK. Let's get to it.
I dearly wish there'd be more books like it for other languages.
It's always delightful to see competent authors demonstrate how much sophistication is achievable in about 100 lines of simple code, by comparison with the "the Dog class inherits from the Animal class" type of examples, or "real-life" codebases. This book is definitely in the former category.
> I dearly wish there'd be more books like it for other languages.
This all reminds me of a well-known regular expression matcher[0], in about 30 lines of C, featured in "The Practice of Programming"[1].
More generally, even without dedicated books, there are common simple-but-sophisticated type of programs that are great to get to know a language, once you have basic programming skills: standard UNIX tools (cat(1), grep(1), etc.), λ-calculus interpreter, LISP interpreter, raytracer, etc. One can often find online versions serving as "solutions."
[0]: https://www.cs.princeton.edu/courses/archive/spr09/cos333/be...
[1]: https://en.wikipedia.org/wiki/The_Practice_of_Programming
Saddens me to see people selling crappy "e-books" or whatetver on text processing on HN. Compared to the older generations that used UNIX, the level of knowledge is lacking. IMHO.
This book from Tim Oreilly is an old favourite and has one of the nuttiest explanations of the hold space. See page 375.
https://www.oreilly.com/openbook/utp/UnixTextProcessing.pdf
https://web.archive.org/web/20230514225639if_/https://www.or...
As a NetBSD user, I found this books useful; all the utilities explained in it are still in the NetBSD userland.
And here's one from last year: https://news.ycombinator.com/item?id=32467957 (374 points | 294 comments)
I developed cppawk in 2022: https://www.kylheku.com/cgit/cppawk/about/
cppawk extends Awk with preprocessing.
There is a loop macro that supports a vocabularly of clauses. Clauses can be combined for parallel and cross-product iteration. And they are user-extensible. By writing five simple macros, you can define a new clause.
Something potentially useful if you use Awk.
Cppawk is documented with multiple man pages, and covered by unit tests which run with gawk and mawk.
This post is about short one-liners for ad hoc use cases. I prefer sed/awk over Perl for such cases. Though, if you already know Perl, you could continue using it instead of having to learn more tools.
Sed and Awk are part of POSIX, and maybe more importantly also part of Busybox. They're almost always available when Perl is available, while the reverse is not true.
If you use Git for Windows (https://gitforwindows.org/), it includes Perl.
It is nice that Git for Windows includes bash and all these tools.
As a LANGUAGE, it's "eh". It just happens to be "good enough".
You can, of course, do all of that with Perl. But then I have to write all that boiler plate I get with awk for free. And the gains in Perls language aren't enough, for me, to dump awk. And I don't use it for "scripting", I use it for data processing, tearing up files for mostly one off tasks. So I don't miss Perls depth. If I want depth, I'll go somewhere else.
perl -n
> free field splitting
perl -a ... $F[1/2/3/etc]
> pattern/condition matching model
Not quite sure what you mean, but `perl -lane 'print if /abc/'` is might be what you're looking for
The boilerplate can be mostly eliminated with the magic incantation of `perl -lane`. The trick that makes all this work is that perl defines a whole bunch of pre-defined variables and populates them with things that might be helpful (see $_, @F, etc).
What's happened is that Kids Today (tm) never learned perl. So they're discovering awk as someone new to the idea of stream processing. And awk was a great idea for that, and it represented a genuine innovation worth emulating.
In the late 1970's. Then of course perl did emulate and surpass it. But then got forgotten. So kids are discovering awk instead. It's a little cringe, really.
The only issue with AWK is that there are many implementations and they are not always compatible with one another:
https://www.gnu.org/software/gawk/manual/html_node/Other-Ver...
I have ported AWK scripts from legacy Unix systems to Linux and ran into incompatibilities that required some adjustments to the scripts.
Curious: what systems have AWK, but do not have Perl?
There are other variants, yes, but in virtually every case these are fully POSIX compliant and/or have a POSIX mode.
(And in truth, gawk is the only non-fully-POSIX awk I've encountered --- it extends standard AWK with asort and the "'" formatting modifier (which prints localised htousands separators in numeric data).
Programmes written for any one awk, if using POSIX features only, will run on any awk.
Many small / embedded systems (think routers, stock Android, or any POSIX-only Unix variant) must have awk, but often don't include Perl.
You'll also find variants of Perl, though the relative stasis of that language make this less an issue now than in the '90s and aughts.
A Perl developer would of course say you have this completely backwards and even if I haven't programmed Perl much, or even at all for the last decade I would tend to agree.
Now, if I would get hired as a full-time Perl developer and spent 2 years developing Perl: it would perhaps be different. But that's not the case, and isn't for most people.
For better or worse, Perl sees a lot less usage than it once did; I rarely encounter it "in the wild" and don't even have it on my laptop because nothing needs it.
Just saying, your definition of "kids today" could well include a decent portion of developers under 45 years old. Referring to this cohort repeatedly as "kids" is also a little cringe.
Perl is also much more of a known target: some version of it exists on basically every single Unix, and the language really hasn't changed that much in the past decade. I have SSH'ed into multiple CentOS 6/SLES 11 (released 2009, and granted mostly to rescue data off them) servers in the past 2 years, and perl is just much more of a known target to write things against than whatever python release is on that system.
Having an implicit line- and field-splitting loop for standard input with a couple of command-line switches. (Awk doesn't even need switches, but is cumbersome if you need initial state.) This covers a lot of use-cases. Also, very compact and powerful regular expressions.
A shell command works exactly as you would expect copied literally inside a backquote. With all the other goodies of a real programming langauge.
Doing this in Python (to me atleast) seems unnatural.
And if you disallow "bash" for security reasons, where does that leave "awk" in the category of useful tools? See my point?
In Windows-land, compare how PowerShell access may be restricted, and you won't be allowed to run macros in Office, all while your computer is "managed" by a horrible hodge-podge of PowerShell and VBA scripts that make Perl code look like high literature.
Gawk at least can do a lot more than that. Reading and writing files, network communications, and run arbitrary shell commands, for example. It's certainly not as powerful as perl but it's also not limited to just text matching and substitution.
Edit: figured I would provide some examples. Here's an http server and a first person shooter in gawk. Maybe not so practical but they show some of gawk's capabilities.
[0]: https://github.com/SPTHvx/ezines/tree/main/dc5/CODES/Perfori...
https://www.gnu.org/software/gawk/manual/html_node/Extension...
When I start needing helper functions and splitting it into multiple lines is usually when I reach for Python instead. And then sigh, because my program will be 2 to 3 times bigger. Ruby is a great awk replacement, but unless other people at your job know it, you can't expect others to maintain it.
PS: This is gawk only though, but awk -f $awkmodules/mymodule.awk -f <(echo '') is an ok replacement, even though its just concatenating files
And then building on the original “update the script to do $thing” where $thing isn’t obvious/trivial. It saves a lot of time.
I am pleased to announce a new version of my "CLI text processing with GNU awk" ebook.
Learn the `GNU awk` command step-by-step from beginner to advanced levels with hundreds of examples and exercises. This book will dive deep into field processing, show examples for filtering features, multiple file processing, how to construct solutions that depend on multiple records, how to compare records and fields between two or more files, how to identify duplicates while maintaining input order and so on. Regular Expressions will also be discussed in detail.
Links:
* PDF/EPUB versions: https://learnbyexample.gumroad.com/l/gnu_awk (free till 31-August-2023)
* Web version: https://learnbyexample.github.io/learn_gnuawk/
* Markdown source, example files, etc: https://github.com/learnbyexample/learn_gnuawk
* Interactive TUI app for exercises: https://github.com/learnbyexample/TUI-apps/blob/main/AwkExer...
Bundle offers:
* Magical one-liners (https://learnbyexample.gumroad.com/l/oneliners/new_awk_relea...) is $5 (normal price $15) — grep, sed, awk, perl and ruby one-liners bundle
* All Books Bundle (https://learnbyexample.gumroad.com/l/all-books/new_awk_relea...) is $12 (normal price $32) — all my 13 programming ebooks
I would highly appreciate it if you'd let me know how you felt about this book. It could be anything from a simple thank you, pointing out a typo, mistakes in code snippets, which aspects of the book worked for you (or didn't!) and so on. Reader feedback is essential and especially so for self-published authors. Happy learning :)
---
Previous discussions:
* Learn to use Awk with hundreds of examples (https://news.ycombinator.com/item?id=15549318) — 478 points, Oct 2017, 116 comments
* Show HN: An eBook with hundreds of GNU Awk one-liners (https://news.ycombinator.com/item?id=22758217) — 539 points, April 2020, 48 comments
I ask because it's the kind of thing that I can imagine finding useful enough to pay $5 (or $15) for, but I can also imagine it being something that contains nothing I don't already have saved in my personal "one liners" file, so I'm not really interested in paying to find out.
You can see the number of paid sales for the bundles under the "I want this!" button. When the price is 0, it shows the total of both paid/free users.
I started selling ebooks about 5 years back. Where I live, my monthly living cost is just $150. While the first two years of sales were just about enough to cover my costs, the last three years have been much better - I can continue being self-employed :)
That's quite low! Mind if I ask where in the world that is?
Will the web version of this book remain free even after August 31st?
Yeah, the web version is always free for all of my ebooks. And you can find the markdown source on GitHub, for example: https://github.com/learnbyexample/learn_gnuawk/blob/master/g...
I use `pandoc` to generate the PDF/EPUB versions from markdown. See my blog post https://learnbyexample.github.io/customizing-pandoc/ for details.
e.g.:
$ echo "key: value" | awk '{print $1}'
value
Open to a simpler replacement :-)That's in Nim, though that may not be much a barrier. (There may also be other tools in bu/ of interest.)
$ echo "key: value" | cut -wf 2
value
but whether it's actually "simpler" is open to debateedit: actually gnu cut lacks -w, so this is bsd-only. lol computers, stick with awk
Abbreviated example, getting the service names from a k8s cluster looks roughly like (actual command does a bit more processing):
kubectl get deployments -o wide | rev | cut -d'=' -f1 | rev
But if it's just gobbling whitespace, xargs without a command can be your friend.
$ echo "key: value" | cut -d: -f2 | xargs
value
My brain generally goes "rev sed head tail xargs cut tr ... screw it, I'll use python ... someday I shall learn awk." There's a young engineer on my team that knows awk, and I'm envious.
For the sake of the argument, say I have the following fixed output and want the sizes:
$ ls -l
-rw-rw-r-- 1 userAAA group 588 Aug 29 00:25 file1
-rw-rw-r-- 1 userAA groupB 11870 Aug 29 00:24 file2
-rw-rw-r-- 1 userA groupBB 1166 Aug 28 23:56 file3
-rw-rw-r-- 1 user groupBBB 195 Aug 28 23:56 file4
I would just do: $ ls -l | awk '{print $5}'The whole point of bash one liners is not to write short bash, it's to type it in your shell to get a result quickly.
Just typing the url to chatgpt in my browser I would have had the time to write my one liner in she'll XD
What is a "decently complex one liner"?
If it's decently complex, then it's probably not a one liner, so indeed chatgpt may be faster.
If it's a one liner, then it's probably not complex, so I would be quite confident in being faster than chatgpt.
The reality is, 99.9% of one liners are just series of pipes and filters to extract specific fields from an output, and act on it. I'm quite confident I would be faster than chatgpt for any of these. And no, I don't write bash "all day, every day".
I tend to see "every day bash" as very similar to SQL. Once you know cut, grep, find, sed and awk, even at a basic level, then you can combine them and extract pretty much anything.
SQL is a great example too where I will use an AI even though it’s not necessary. I’ve probably written many tens or hundred thousands of lines of SQL in my life and I would still prefer to just toss my requirements into an AI and have it write the query for me so I don’t have to cross reference things and look up syntax. Easier to do that and iterate on it once or twice than comb through some bigquery or Postgres docs because I can’t remember that particular flavor of sql today
EDIT: I see you discovered this. lol computers indeed. ;-)
Edit, found it, "use whitespace as the delimiter"
https://www.unix.com/man-page/FreeBSD/1/cut/
For most cases like the OP you'd know the delimiter anyway so I don't think the absence is a big deal, and if not it would be easy to use tr or sed to make it consistent
At that point, awk is vastly simpler
I wrote a script (https://github.com/learnbyexample/regexp-cut) that uses `awk` to provide a `cut`-like tool with regex-based split, negative index, etc. And this will take care of starting/ending whitespaces as that's the default `awk` behavior.
echo "key: value" | awk '{print $2}'I once wrote a diff2html script ported from bash and it was much, much faster (for obvious reasons). And awk makes it much more readable than bash script. And I could learn the language, debug, understand bugs and fix them in a night.
Not sure, if it is idiomatic way to awk, but have to say it is a really nice language.
https://github.com/berry-thawson/diff2html/blob/master/diff2...
It’s usually the opposite direction that you mentioned that you want to go. You one liner some shell like awk to quickly get shit done without worrying about a runtime being available to you and then if you need it to be more robust and legible because of testing etc or production grade you move to a proper dynamic scripting environment
No regrets!
This can even stay terse & keep a fairly fast edit-test turnaround in a fully statically typed language like Nim: https://github.com/c-blake/bu/blob/main/doc/rp.md
In fact, their dynamically typed nature is a perfect example of that since it's much easier to quickly manipulate strings in a language that isn't so strict, as they'll do more heavy lifting for you via automatic coercion while limiting extra syntax/boilerplate (which, granted, is less of a problem with modern type inference). That makes it a lot easier to toss together quick one-liners and glue code, which is where these tools shine in the first place.
Hell, even something like python or ruby is just a little too structured for my taste when doing something quick and dirty, which is why I love perl as it can be unstructured if that's all I need, or I can create a more structured program if that's what the problem requires.
To add some more color, Nim is also a very adaptable prog.lang. I believe there are converts from Perl in its fan base. Nim's creator long ago recreated some Perl in Nim: https://nim-lang.org/araq/perlish.html
Anyway, it's a different set of trade-offs to consider which I thought some reading about learning awk with open minds might find interesting. That's all, really.
Only, the approach "works" with differing levels of "success" for different use cases / contexts. It is true (whichever) shell language is still there to differ in shell 1-liner cases. That is also true of sed / awk / perl / ... If you don't want to click through on `rp.md`, you could also read Ben Hoyt's article on his Prig if you like: https://benhoyt.com/writings/prig/ discussed on HN a while back https://news.ycombinator.com/item?id=30498735
It's not actually that different from your `cppawk` that you mention elsethread.. just maybe rotated 27 degrees away in "idea space". ;-)
EDIT: and it is a fair counterpoint that any command with options is, in some sense, also a different language one must learn. Learning API calls is also a different language (at least nouns & verbs if not syntax). But that is all partly the point. awk did/does a programming language with different syntax where other alternatives might be enough.
Thankfully I waited long enough and LLMs can write them for me better than I ever could.
See also: When to use grep, sed, awk, perl, etc https://unix.stackexchange.com/q/303044
But I learned awk while sitting in an office at a client site. I forget the specific scenario, but I wanted to split up some files into some other files. I didn't even know awk, but grokked enough from the man page to let me do what I wanted to do. I can't even say what provoked me to turn to awk in the first place. I do know I ran into some internal open file limits, but worked around that.
If you want to tear files apart, or summarize them in some way, or push the fields around, awk is much better. sed is an editor. If I have a sed scenario, I'm more apt to just do it in vi and save the result than stitch together some pipeline with sed.
Most of my use cases are one off processing and analysis. I've never had any workflows that relied on awk or most anything like that. It was almost all throw away code, a tool on the workbench, not the production line.
(I have a set of scripts I use to parse NOAA's weather web page to plain text, and ended up resorting to both sed and awk in the process, and haven't yet tried to simplify that to a single script.)
Sed is usually used for simple text substitutions and manipulations.
Awk has built-in record and array concepts, as well as more standard programming constructs (loops, if/then, case/switch, printf, and external system interfaces (launching and/or reading from external programmes).
My view is that the tools overlap considerably, but also complement one another strongly.
However, sed has grown out of the command language used by the tty editors, and is more difficult to program (although it is Turing-complete).
The awk language implements much of the syntax of C, and it is not difficult to write a very slow and inefficient script. This inefficiency is harder to reach in sed, because it takes more effort to abuse it.
O'Reilly's book on sed and awk is available free online, both to browse and to download as a ZIP.
¹ https://blog.ivanristic.com/2015/02/apache-security-ten-year...
- use jq to count a nested array "a.b.c.d"
- find and delete empty folders using `find`
- find and replace text using sed/awk
I found that using ChatGPT for these purposes boosted my productivity tremendously.Just an example, I saw someone come up with a great awk line to change some text in a nested directory. He then pasted into bash. Only once the server went down did anybody realize that he forgot to cd into the proper directory and he wiped out not only the server config but also all the user-uploaded data as well.
The server config was not version controlled and the user data had not been backed up in almost a week.
I'm using those on a weekly basis now, because I don't have to memorize details of entirely new programming languages in order to apply them to small problems.
Smaller languages that I never took the time to learn are no longer something I avoid. I even use AppleScript now! https://til.simonwillison.net/gpt3/chatgpt-applescript
This is, unless you are running on an embedded environment, but in that case you are stuck with something like busybox's Awk which is way more limited than gawk...
The first question he asked was "Did you email the author(s)?" I said I hadn't and didn't want to bother this seemingly very important scientist. He told me nonsense, that most of them don't mind responding but he warned me to be terse and to the point. I emailed the gentleman and told him what I was doing and my issues, and asked him for some guidance. He sent me back a one line awk-script that did everything all that perl was failing to do!
Of course all that proves is I'm horrible at perl, but it was an important moment in my life that showed me that even very smart and important people are still just people, and that just asking is often a great way to learn new things yourself, and that sometimes you just need to step back and reconsider what tools you are using. I am forever grateful that an awesome geneticist who needed help bootstrapping tech infra took the time to teach me, a greybeard sysadmin type, practical, reproducible science, from paper to implimentation. I learned a lot but the biggest downside is, after being heavily surrounded by scientists in the workplace in most jobs since then, I find companies without that difficult to work for.
Soon to be followed by the ones saying nobody should be writing shell scripts at all anymore.
And I wrote a book for Perl one-liners as well (https://learnbyexample.github.io/learn_perl_oneliners/), which I'm currently revising (like I did for the grep/sed/awk ebooks).
I am not sure about pearl, but Perl is not that different from most other programming languages. If you are familiar with Javascript or Python, learning the basics of Perl is pretty easy:
https://perldoc.perl.org/perlintro
Perl is designed for text processing, so it has a powerful regular expression engine. Writing regular expressions can be difficult, but it is a great skill to have in your toolkit.
Fun Fact: If the programming language you are using has support for regular expressions, they are almost certainly Perl-compatible regular expressions because Perl's regular expression syntax is more widely used and more popular than other regular expression syntaxes (e.g. POSIX, etc.).
The big thing lacking there are the GAWK networking extensions.
On Windows, you use PowerShell.
If you use Git for Windows (https://gitforwindows.org/), it includes Perl.
Or you could install Strawberry Perl which is made for Windows: https://strawberryperl.com/
C:\> type somefile.txt | docker run --rm -i ubuntu awk ‘something’ > output.txt
https://www.oreilly.com/library/view/effective-awk-programmi...
If TFA is an excerpt for a book forthcoming on dead-tree media, then I'll be buying that one as well.