Awk Technical Notes
maximullaris.com
maximullaris.com
As the stuff we were doing in that project got more complex, at some point someone suggested to teenage me "You might want to look at Perl for this now," and then I moved to that. (with the Camel O'Reilly book, of course!)
Haven't touched either one in years now.
Learning new things can be much more overwhelming for me now, I don't know how much is me vs environment. But I am nostalgic for those days where I'd sit down with a print book, and within hours have a grasp of the fundamentals, or within days feel like I had basic fundamental conceptual understanding of the whole dang thing (not of every possible feature, but of the conceptual framework, the big picture).
Whereas if I pick up an Awk book that OP referred, it's likely I could still use it to learn, you can't really do the same with most of modern tech stack.
Or, rather, you can read the two first chapter easily in one sitting. Chapter one gives a brief overview and examples. Chapter two describes the whole language, every function and every variable! The rest of the book is just more examples. I really love this style!
The problem was that the Garmin GPS data for a bike ride I had just completed had split into multiple rides. I used AWK to stitch together the data into one file. I also did some basic linear interpolation to fill in missing data points.
The GPS data is formatted as XML and I was able to parse it fairly robustly using AWK.
Depending on how the xml is structured, it can be possible to just pattern match on the tags if you have something simple to do.
awk is great for it because IRC (or at least the subset that the bot cares about) is relatively easy to parse, and shelling out the shell script that does the actual code evaluation and prints the result back is also fairly straightforward. Someone else used to have such a bot before but they had written it in Rust with a bajillion dependencies; if I had done that I would've had to update dependencies and redeploy it every other week. In contrast I deployed my awk version once and then basically haven't touched it in years.
The following caught my attention in the bash wrapper:
coproc GAWK {
gawk ...
}
<&"${GAWK[0]}" openssl s_client -connect "$IRC_SERVER" -quiet >&"${GAWK[1]}"
This is a cool way to make awk talk over a socket that isn't specific to Gawk. For sockets without TLS you can replace openssl(1) with nc(1). I'll keep it in mind.Edit: You can also use http://www.dest-unreach.org/socat/ and https://nmap.org/ncat/ with awk:
socat "OPENSSL:$host:$port" 'EXEC:awk ...'
socat "TCP:$host:$port" 'EXEC:awk ...'
ncat --exec '/usr/bin/awk ...' --ssl "$host:$port"
ncat -e '/usr/bin/awk ...' "$host:$port"The main difficulties were making the command-line interface and testing. You can't have flags that begin with a dash in portable AWK without a shell wrapper, and I didn't want one. I settled on manually parsing key=value options, which I don't think are bad, just nonstandard. They look like this:
humsize format=%6.1f%1s 'zero= empty'
There is no standard way to test AWK code. For testing I wrote a shell script that checks the program's outputs with grep: https://gitlab.com/dbohdan/humsize/-/blob/122aaed8d65dc8c285.... Don't do this; your tests should give the user (you) better feedback. You may think your program doesn't need anything but a couple of trivial tests that won't ever change; it is a pain when you inevitably are proven wrong. I should have instead had a directory with reference outputs and diffed against them to see what went wrong (my own example: https://github.com/dbohdan/initool/blob/72f65d3fde245ff8660c...).To ensure I didn't introduce portability issues, I set up testing against different awks in GitLab CI.
image: debian:bullseye-slim
before_script:
- apt update
- apt install -y busybox gawk mawk original-awk
- ln -s "$(which busybox)" awk
- busybox wget -O goawk.tar.gz https://github.com/benhoyt/goawk/releases/download/v1.21.0/goawk_v1.21.0_linux_amd64.tar.gz
- tar xzvf goawk.tar.gz
test:
script:
- AWK=false ./test || true
- AWK=./awk ./test
- AWK=gawk ./test
- AWK=./goawk ./test
- AWK=mawk ./test
- AWK=original-awk ./test
Edit: Rephrased and added a nicer shell test example.(Python isn't as nice for one-liner text processing, both because of the lack of Awk heritage--so no built-in regex syntax--and because of the indentation-based syntax requiring newlines for most things.)
grep|ripgrep > awk|sed > most scripting languages > shell
gawk has regexps.
We use Ruby a bit at work. Most coworkers hate it, internal customers scoff at it, and no one's interested in mastering it, using it properly, or considering even fundamental software engineering principles. Tech debt piles up and no one wants to touch it because there's no performance review KPI credit for it.
(h/t Tony Hoare)
The absence of a GC is nice for an embedded language, but I don't think that should be the only criteria. Unless you needed an embedded language that processes text one line at a time awk is probably not a good fit.
function NUMBER( res) {
return (tryParse1("-", res) || 1) &&
(tryParse1("0", res) || tryParse1("123456789", res) && (tryParseDigits(res)||1)) &&
(tryParse1(".", res) ? tryParseDigits(res) : 1) &&
(tryParse1("eE", res) ? (tryParse1("-+",res)||1) && tryParseDigits(res) : 1) &&
asm("number") && asm(res[0])
}
why put yourself through this, when you can just do something like this instead: package parse
import "strconv"
func parse_float(s string) (float64, error) {
return strconv.ParseFloat(s, 64)
}
func parse_int(s string) (int64, error) {
return strconv.ParseInt(s, 10, 64)
} bash> awk -v v="80.1%" 'BEGIN{print v+0.1}'
80.2
gawk has `strtonum`. But yes, parsing in awk generally looks like a pain. With plain positive/negative ints though, not so hard: echo "123456" | awk '{
if ($0 ~ /^-?[0-9]+$/) {
num = 0
sign = 1
start = 1
if (substr($0, 1, 1) == "-") {
sign = -1
start = 2
}
for (i = start; i <= length($0); i++) {
digit = substr($0, i, 1)
num = num * 10 + digit
}
num = sign * num
print "The integer is:", num
} else {
print "Invalid input string:", $0
}
}'Funnily the actual Go JSON decoder code ends up doing something similar during scanning:
https://github.com/golang/go/blob/master/src/encoding/json/d...
...
> The most substantial consequence is that it’s forbidden to return an array from a function, you can return only a scalar value.
This doesn't make sense to me. Does someone understand what it means?
In e.g. C++ a function can return an array without any GC or refcounting, by "moving" the array into the caller's stack.
a[1] = “hello”
a[“world”] = 2
That means that - unlike C arrays - Awk arrays are not a simple, addressable byte range, but a complex data structure with lots of pointers.I suppose you could come up with a way to serialise the array and pop it on the stack but that would be a lot of work, and for the kind of things I use Awk for, the arrays would often be huge.
Maybe the language is simpler without it, and that can be a good reason to avoid it. But I don't buy that it has anything to do with GC.
a[1] = a
Even if we only allowed it on return values: function f(x) {
return x
}
a[1] = f(a)