The State of the Awk (2020)
lwn.net
lwn.net
"I wish Awk had capture groups. It would fit in so well with typical Awk one-liners to be able to say:
awk '/foo=([0-9]+)/ { print $1 }'
although I suppose the syntax would have to be different since $1 has a meaning already."I use Awk all the time, and the new additions in the article are pretty nice; but for typical uses of Awk, the feature comex wants would make a tremendous difference in usability.
(That said, I do still have the tabs of perldata, perlobj, perlmod, and perlop open because I want to learn it better.)
This random blog post gives something of the flavor.
https://lifecs.likai.org/2008/10/using-perl-like-awk-and-sed...
The equivalent Perl isn't that far off from the above.
perl -nE '/foo=([0-9]+)/ && say $1'
If you're jonesing for a sed replacement (it even has capture groups) try: https://github.com/chmln/sd
Worth being aware of:
Would you consider trying something other than Perl once it is no longer packaged for you? Because it is an x-language, bereft of life.
For a lot of us, we want to develop scripts (and skills) that are portable across different environments. There are limited hours in a workday, and I get the most value out of learning (and using) tools that I can find on the servers / build-machines / workstations that I have to use every day. Those machines run Ubuntu, Debian, Rocky, and RHEL.
Yes, it's a slow process to get new packages into mainstream distros. That's not a bad thing, because those packages have to be maintained for a very long time. Stability is a virtue here.
There are some ecosystems (I'm looking at you, Javascript) where anything older than a year might as well be abandonware. It's great that there are some fast-moving areas in our industry, and also great that there are slow-moving areas.
Don't make the mistake of assuming something is bad just because it's feature-complete. You might be surprised at how feature-rich something like Awk really is.
If your argument boils down to "Awk and Bash are ugly and outdated", I'd encourage you to think more flexibly about the tools you choose. There's nothing wrong with learning the basics of a widespread tool that you are guaranteed to find anywhere.
> if your argument boils down to "Awk and Bash are ugly and outdated"
Not at all. However, Awk was written almost 50 years ago. The awk book is good and it is very true that the thing is installed most everywhere so if you're already invested in it and it does everything you want, keep using it of course. But it just might be possible to improve on a tool after half a century.
I simply shared new tools, that if you like awk, might really be up your ally. They might not be packaged with your favorite package manger anytime soon but you can grab them with cargo - if that's not portable enough so be it. No harm no foul.
If you find unknown syntax construct in Perl code, it is hard to identify what it is, but if you find unknown function in AWK, you can just google its name.
Also, AWK is mandated by POSIX, so one can assume it is installed, while Perl is optional.
Perldoc perlintro. Have you seen real life Perl outside of oneliners?
Example from that article
$ echo "foo\nbar\nbaz" | ruby -ne 'BEGIN { i = 1 }; puts "#{i} #{$_}"; i += 1'
output 1 foo
2 bar
3 bazsh, awk and sed are fine. They are easy, small, powerful tools that are easy to compile and understand.
Perl, python, nushell, etc. The options listed here are great if you're writing cute snippets on the terminal or hacking together some higher level automation.
These more elaborate tools, however, are terrible if you're trying to be lean in the build/bootstrap process and have a small set of auditable, easy-to-compile tools.
The graph on this page illustrate the bootstraping problem well: https://bootstrappable.org/projects/mes.html
These small 50+ years UNIX tools that have zero build dependencies are around for a reason: they are small and have zero build dependencies.
awk 'match($0, /foo=([0-9]+)/, g) { print g[1] }'
works in gawk (using extended match syntax allowing captured groups in the 3rd parameter array). rg 'foo=([0-9]+)' --replace '$1'
or more succinctly: rg 'foo=([0-9]+)' -r '$1'
Example: $ echo 'quux=123 foo=123 bar=123' | rg 'foo=([0-9]+)' -r '$1'
quux=123 123 bar=123
Named groups work too: $ echo 'quux=123 foo=123 bar=123' | rg 'foo=(?P<digits>[0-9]+)' -or '$digits'
123
You can also replace the whole match by combining --replace with --only-matching: $ echo 'quux=123 foo=123 bar=123' | rg 'foo=([0-9]+)' -or '$1'
123
Of course, I understand capturing groups are useful to have in awk when you're using awk. ripgrep can only handle very simplistic cases. But they tend to be quite common.Nowadays I like the nushell approach to the composition:
echo 'quux=123 foo=123 bar=123' | str replace '.*quux=([0-9]+).*foo=([0-9]+).*' $"$2,$1" | from csv -n | each {|r| $r.column1 + $r.column2}
which of course relies on the same regex library (hattip).end
NOTE1: reading variable INPUT. reads from input. assignment to OUTPUT writes to output. normally assigned to STDIN and SDTOUT, and can be configured.
NOTE2: there is no real while or if. flow is through labels and jumps (considered by some as power but i disagree). above example uses another script to transform while and if into labels and jumps. the point is that patterns composition and success match assignment.
digits = '0123456789'
patt = "foo=" SPAN(digits) . num
while line = INPUT
if line ? patt
OUTPUT = num
endif
endI resorted to pandas, where the CSV import has parameters for the thousands and decimal separators.
Also see AWK HN post from 2021 [1].
https://stackoverflow.com/questions/29642102/how-to-make-awk...
$ cat quoted.csv
"Smith, Bob",42
$ awk -F, '{ print $1 }' quoted.csv
"Smith # you want to print the first field: Smith, Bob
$ goawk -i csv '{ print $1 }' quoted.csv
Smith, Bob # that's better!However one of the not so good Devs I worked with used awk to load a large, deeply hierarchical JSON file. They refused to use a library to parse JSON. It was a many hundreds of line monstrosity.
Luckily when they left we were able to parse it in JQ instead .
* https://news.ycombinator.com/item?id=23240800 (213 points | May 19, 2020 | 86 comments)
* https://news.ycombinator.com/item?id=25142867 (207 points | Nov 18, 2020 | 58 comments)
#!/usr/bin/env gawk -f
BEGIN {
for (i=1;i<=100;i++) {
printf(" %2s", i%(3*5)!=0 ? i%5!=0 ? i%3!=0 ? i : "fizz" : "buzz" : "fizzbuzz\n" )
}
printf("\n")
}