The Awk book’s 60-line version of Make
benhoyt.com
benhoyt.com
I patched GNU Awk to have a @let extension that gives you scoped locals (usable in functions as well as in BEGIN/END blocks):
$ egawk 'BEGIN { x = 3; print x; @let (x = 4, y) { print x } print x }'
3
4
3
@ is used because there is at least one other existing extension which is like that: @include.https://www.kylheku.com/cgit/egawk/about/
This was rejected by the GNU Awk project, though. I was encouraged to make a fork and give it some kind of different name, so I did that.
It's curious because gawk cannot for one second claim something like needing to stick to some legacy standard, not with a straight face.
That said, proving it's value in a fork first seems reasonable.
https://lists.gnu.org/archive/html/bug-gawk/2022-04/msg00025...
There is more around it. I had the idea in two other forms.
Initially I had a @param:<ident> syntax which indicated that the given variable is to be allocated in the parameter space (a local variable frame where function parameters go). This only worked inside functions.
Between that and @let was a @local thing.
The maintainer of GNU Awk is one of the two authors of the "Fork y Code Please", the other being the Bash guy:
https://www.skeeve.com/fork-my-code.html
So ...
> Suggestions for new features that:
> Cannot be accomplished using existing features in a straightforward way > Don't (too badly) break compatibility for existing code
So.. in theory your @let should qualify
GNU Awk has switch, which is an extension and not hidden by @.
Is there something in the awk script that makes it advantageous over a shell script?
Edit: I hadn't read the author's conclusion yet when I posted, he agrees
I consider AWK amazing, but I think it should remain where it excels: for exploratory data analysis and for one-liner data extraction scripts... also known as the APL Zone
In this particular case, we're talking about a "make" replacement, so testing the new implementation can be done by simply running "make all" for the project. If it passes, then the new implementation must be identical to the old one in all the ways that actually matter for the project at hand. In all likelihood, for a simple program like this, fixing one bug will also silently fix others because the new architecture is probably better than the old one.
I think both you and the author just don't like AWK if that's the takeaway. What you're describing is literally 1% of the AWK language -- like you don't have to like it, it's weird in many respects but you're treating AWK like it's jq when it's actually closer to like a Perl-Lite/Bash mix. An AWK focused on just those use-cases would look very different.
One of my favorite resources on AWK: https://www.grymoire.com/Unix/Awk.html
What do you find hard to read about it? If you know what make does, I think it is fairly easy to read, even for those who don’t know awk at all, but do know the Unix shell (to recognize ‘ls -t’) and C (both of which, probably the audience for this book knew, given that the book is from 1988)
> I think for something like this I'd prefer a bash script
But would it be easier to read? I doubt see why it would.
But a bash or ksh script would have been less readable than awk.
bash (or ksh88 or ksh93) is powrful and useful but not readable if you're actually using the powerful useful features.
In bash, a lot of functionality comes in the form of brace expansions and word splitting, basically abusing the command parser to get results there is no actual function for. In awk and any other more normal programming language, those same features come in the form of an explicit function to do that thing.
Right. That's one of the reasons why the man page for bash is so long. IIRC, going way back, even the page for plain sh was long, for the same reason.
No, it wouldn’t have been ksh or any other shell, nor C or Perl, nor anything else but awk, in a book titled “The AWK Programming Language”.
awk is like a hidden miracle of utility just sitting there unused on every machine since the dawn of time.
Normally if you want something to be ultra portable, you write it in sh or ksh, (though by now, bash would be ok, I mean there is bash for xenix), but to get the most out of ksh or bash, you have to use all the available features and tricks that are powerful and useful but NOT readable. 50% of the logic of a given line of code is not spelled out in the keywords but in arcane brace expansion and word splitting rules.
But every system that might have some version of bash or ksh or plain sh, always has awk too, and even the oldest plain not-gnu awk is a real, "normal", more or less straighforward explicit programming language compared to bash. Not all that much more pawerful, but more readable and more straightforward to write. Things are done with functions that take parameters and do things to the parameters, not with special syntax that does magic transformations of variables which you then parlay into various uses.
Everyone uses perl/python/ruby/php/whatever when the project goes beyond bash scope, but they all need to be installed and need to be a particular version, and almost always need some library of modules as well, and every python script breaks every other year or on every other new platform. But awk is already there, even on ancient obscure systems that absolutely can not have the current version of python or ruby and all the gems.
I don't use it for current day to day stuff either, there's too many common things today that it has no knowledge of. I don't want to try to do https transactions or parse xml in awk. I'm just saying it's interesting or somehow notable how generically useful awk is pretty much just like bash or python, installed everywhere already, and almost utterly unused.
I once had to rewrite a bash script into awk[1] that is big enough and it made the program more readable and the total time execution came down from 12 mins to less than 1 second.
I think maybe the original bash script would have written badly, (each util command will invoke it's own process and it has to piped to others instead of using awk which will be running in a single process).
[1] - https://github.com/berry-thawson/diff2html/blob/master/diff2...
Pseudo multi-dimensional associative arrays for representing the dependency graph of make. This part:
for (i = 2; i <= NF; i++)
slist[nm, ++scnt[nm]] = $i
The way awk supports them is hacky and not really a multidimensional array, but still is better than what you would have to do with bash, because of split() and some other language features.It would be much easier with any scripting language though, Perl for example.
For those who have taken this advice, they've always told me later they're really glad they did so, and generally express surprise that this isn't more widely known.
(If you already know awk and sed well, then you mightn't view learning perl in addition worth the effort -- I'm not sure either way. This advice is for people that currently are not strong users of either.)
For grep/sed/awk, you also have to worry about implementation differences (GNU/BSD, gawk/mawk/nawk and so on).
However, in my experience, when I start to feel the need to use something more sophisticated than sh/sed/awk, I tend to shy away from Perl in favor of more "robust" languages. Go often is a good-enough substitute (static typing, single-file deployment, trivial cross-compilation); YMMV.
[0]: https://marc.info/?l=openbsd-misc&m=159041121804486&w=2
Only for 1-2 liners, typically though, the moment something grows beyond that, I don't use Awk any more, so really I only use a tiny sliver of what it can do.
(FWIW, I learned Perl before sed and awk, and when I was using Perl every day, it was easy enough to whip up one-liners and throwaway scripts. However, I find that as I stopped using Perl on a day-to-day basis about 17 years ago, I can't produce Perl without re-learning the language; but I can produce sed and awk a few times per year without any refresher. I suspect that -- for me -- the smallness of each of sed and awk has something to do with it. YMMV, of course.)
Python's behavior seems wrong to me. It shows up in rules with a phony target and phony prerequisites, which by definition share the same age (9999) and mtime (0). For example, it wouldn't delete prog in the following rule:
clean: clean-objs
rm prog
clean-objs:
rm *.o
On the other hand, the AWK version has a subtle bug in that it sets to zero the age of a newly updated target: this is not required in the most common cases (because the target will likely be the first file listed by "ls -t" anyway) and makes it incompatible with GNU make in those rare cases when the commands don't actually touch the target. I know they're rare, but just imagine a rule that uses rsync to replace a file with a copy fetched from a remote site only if a newer version exists on that site. If rsync does not download a new version, there's no need to artificially assume that the file was changed, and propagate "upwards" the need to recompile everything that depends on it.Both bugs are easy to correct, though. That could be left as an exercise for the reader!
It’s always been a “thing” with me of not liking to put everything into BEGIN. Kind of a “if I’m doing that, why am I using awk” thing.
Just how I approach problems with awk.
In short, the debate might be something like: What does the computer user prefer more: (a) writing one-liners or (b) writing lengthy programs. Not everyone will have the same answer. Knuth might prefer (b). McIllroy might prefer (a).
Assuming one reading this blog post knew nothing about programming languages, it seems to imply Python is not well-suited for one-liners, or at least not comparable to AWK in that context. Perhaps the interpreter startup time might have something to do with the failure to consider Python for one-liners.
Consider the following AWK one-liner which, for every input line that starts with a letter, prints the line number and the line's second field:
awk '/^[A-Za-z]/ { print NR, $2 }'
The equivalent Python program has a ton more boilerplate: import statements, explicit input reading and field splitting, and more verbose regex matching: import re
import fileinput
inp = fileinput.input(encoding='utf-8')
for line in inp:
if re.match(r'[A-Za-z]', line):
fields = line.split()
print(inp.lineno(), fields[1])"awks and pythons"
the line, lineno and fields would be predifined, and I guess re, os, shutil, pathlib and sys are pre imported. maybe the whole stdlib acts as if it's preimported, while only being imported lazyly
here it would be something like
```
if re.match(r'[A-Za-z]', line): fields = line.split() print(inp.lineno(), fields[1])
```
so
```
cat makefile | pyawk 'if re.match(r"[A-Za-z]", line): print(lineno, fields[1])'
```
I don't see a way out of multiple if statements requiring multiple lines though, otherwise you would have to introduce brackets to Python lol
pawk '/^[A-Za-z]/ (n, f[1])'
By the way, triple backticks don't work on HN. You have to indent by 2 spaces to get a code block. cat makefile | pyawk 'if re.match(r"[A-Za-z]", line): print(lineno, fields[1])'
in a Python one-liner without PAWK by abusing list comprehensions: python -c 'import fileinput, re; [print(re.split(r"\s+", line)[0], fileinput.lineno()) for line in fileinput.input() if re.match(r"[A-Za-z]", line)]' makefile
Edit: Removed a list of other awk replacements to post in a separate comment (https://news.ycombinator.com/item?id=37465164).You can actually get pretty far depending upon boundaries with the always implicit command-option language (when launched from the shell language, anyway). For example, Ben's example can be adapted to:
rp -m^\[A-Za-z\] 'echo nr," ",s[1]'
which is only 5 more characters and only 3 more key downs (less SHIFT-ing) than the space-optimized version of his `awk`. { key downs are, of course, just a start to a deep rabbit hole on HCI ergonometrics ending in heatmaps, finger reach/strain/keyboard layouts, left-right hand switching dynamics, etc., but they seem the most portable idea. }Nim is not Python - it is actually a bit more concise while also being statically typed and can be compiled to code which runs as fast as the best C/C++ (at more expense than one usually wants for 1-liner interactive iteration, though unless you need to test on very large data). That said, I find it roughly "as easy" to enter `rp` commands as `awk`.
If doing this in Python tickles your fancy, Ben actually has an interesting on these ideas: https://benhoyt.com/writings/prig/ you might also find interesting.
EDIT: and while I was typing in a sibling @networked mentions a bunch more examples, but I think my comment here remains non-redundant. I'm not sure even one of those examples has some simple `-m` for auto-match mode (although many would say a grep pre-filter is enough for this).
- Common Lisp: https://github.com/sharplispers/clawk
- Haskell: https://github.com/gelisam/hawk
- Racket: https://gitlab.com/xgqt/racket-rawk
- Tcl: https://wiki.tcl-lang.org/page/owh+%2D+a+fileless+tclsh (disclosure: the page links to my fork)
One use for an awk replacement is emitting more structured data. I have used my fork of owh a few times to emit JSON after awk-style parsing. I know GNU Awk can generate JSON with https://www.gnu.org/software/gawk/manual/html_node/gawkextli..., but I haven't tried it.
grep ^[A-Za-z]|cols 2
You just lose that row number in the original input coordinates feature of Ben's example which could probably be recovered with `grep -n` & `cols -d' :'`, etc., etc. In exchange, you can say `cols 2:5` to get a block of columns trivially. And then, of course, once you have any oft-repeated atom you can save it in a tiny script/etc.A lot of these choices come down to atom discovery & how willing/facile someone is juggling/remembering syntax/sub-languages. In my experience, willingness tracks facility and both are highly variable distributions over the human population.
ruby -nae 'print $.," ",$F[1],"\n" if $_ =~ /^[A-Za-z]/'
-n wraps an implicit "while gets; ... ;end" around the code; "-a" adds an implicit "$F = $_.split" at the start of the loop; "-" takes an expression from the command line; $_ contains the result of the `gets`; $. contains the line number of the last line read.Alternatively:
ruby -ne '$_.match(/^[A-Za-z]+(.*)/) { puts "#{$.}#{$1}" }'
`match` sets $1, $2 etc to the corresponding capture group, and calls the block if successful.The scaffolding would be easy to provide w/Python too, but the extra Awk/Perl-isms to make it convenient is another matter (and while I use them occasionally for one-liners, I will get shouty if I find $1 etc. in production code...).
Even the Ruby differences are sufficient extra noise that I still reach for awk for simple stuff like that.
Here is how I would do that task, assuming (a) I had to do it more than once and (b) I could choose any software. On the computer I'm using, the statically-linked, stripped binary is 50k versus a dynamically-linked gawk which is 623k. This solution is faster than AWK, Python, Go, etc. and uses much less CPU and memory. This is quick and dirty, written in a few minutes. I am not a paid programmer. I'm the so-called average user. I'm not compensated for writing programs.
usage: a.out <-- minimal typing
NB. There is a two space indent added to each line. One must remove exactly two spaces from each line or there will be error messages and this will not compile.
#!/bin/sh
flex -8Crf <<eof
int fileno (FILE*);
int x,y,n=1;
%option noyywrap noinput nounput
%%
^[A-Za-z][^\n]+ {
printf("%d ",n);
for(x=0;x<yyleng;x++){if(yytext[x]==32)y++;
if(y==1)putc(yytext[x],yyout);
}
putchar(10);y=0;
}
\n n++;
.
%%
int main(){ yylex();exit(0);}
eof
cc -O3 -std=c89 -W -Wall -pedantic -pipe lex.yy.c -staticThis (both in awk and python code) seems useless, as the return value from update() is not used anyhere. Am I missing something obvious?
if (update(ARGV[1]) == 0)
print ARGV[1] " is up to date"ie awk '/test/' -> '{ if($0~/test/){print $0} }'