Fascination of Awk
maximullaris.com
maximullaris.com
Take this function the author wrote for converting a list of integers (or strings) into floats
def ints2float(integerlist):
for n in range(0,len(integerlist)):
integerlist[n]=float(integerlist[n])
return integerlist
Using `range(0,len(integerlist))` immediately betrays how the author doesn't understand python. The first arg in `range` is entirely redundant. Mutating the input list like this is also just bad design. If someone has used python for longer than a month, you'd write this with just `[float(i) for i in integerlist]`.Further down in the function `format_captured` you see this attempt at obfuscation:
freqs=ints2float(filter(None,captured[n].split(' '))[2:5])
Why bother with a `filter`? Who hurt you? freqs = ints2float(captured[n].split(' ')[2:5])
That said, the author's implementation in awk does look pretty clean. I'm just peeved that they straw-manned the other language. floats = [float(v) for v in integerList]
or less idiomatic floats = map(float, integerList) list(map(float, integerList))
for them to be equivalent.(It doesn't matter in a lot of cases, but there are enough edges where it does. Json serialization for one)
> I can't think of many more elegant ways to convert a list of ints to floats, in any language, than `[float(i) for i in integerlist]`.
I think something like `integerlist.map(float)` is at least a contender.
What about `float(integerlist)`
That code was so bad I felt I had to step in too, I used chatGPT to simplify it a bit but it also introduced some errors, so I found what appears to be an input file to test it on [1]. The only difference with the awk program is that it uses spaces while the original python program used tabs.
#!/usr/bin/env python3
import sys
freq, fc, ir = [], [], []
with open(sys.argv[1]) as f:
for line in f.readlines():
words = line.split()
if "Frequencies" in line:
freq.extend(words[2:])
elif "Frc consts" in line:
fc.extend(words[3:])
elif "IR Inten" in line:
ir.extend(words[3:])
for i in range(len(freq)):
print(f"{freq[i]}\t{fc[i]}\t{ir[i]}")
[1] https://dornshuld.chemistry.msstate.edu/comp-chem/first-gaus...Lmaoooo
* Can I solve the problem using one-liners (grep, sed, awk, perl, sort, etc)? Or perhaps from within Vim?
* Can I glue together one-liners with minimal control flow as a Bash script?
* If not, go for Python
---
Discussion for https://blog.jpalardy.com/posts/why-learn-awk/ mentioned at the end of the article: https://news.ycombinator.com/item?id=22108680 (420 points | Jan 21, 2020 | 235 comments)
- can I use ripgrep, fdfind, fzf and choose to do this?
- can I ask chatgpt to do this ?
If I need a script to generate some output (a regular use case that seems to come up for me is random numbers with some property), I tend to use python.
I also use python if I need to do more complicated aggregations or something on tabular data that pandas is better at. Though it's fun to try with `join` and awk sometimes (parsing csv can get tricky).
If I need to plot something I tend to use jupyter notebook but it's way more satisfying to use gnuplot, which I mention because it fits naturally into workflows that use shell tools like awk.
You should install xsv.
Like unix or linux? One person's "bucket of old shit" is another person's reliable swiss army knife that is always there when you need it.
My (admittedly very lame) anology: you're hiring a ninja assassin and have two candidates. One tells you about all the special swords and staffs and and smoke bombs he carries, the other says I just need my hands. Who do you hire?
And by the way, I would prefer the first ninja if the second one is handless.
But I absolutely know what you are talking about. I often SSH into new machines, that I may not have root access on, that may not even have internet access, or be a distro with a package manager (e.g. some switch running a custom distribution).
In those situations it is a huge advantage to know the tool, rather than try to do some gymnastics to get your tools onto the box or the data off the box.
These "old" and "broken" tools are installed as defaults on all these systems for a reason.
This is enormously portable, and does not require any new software installations for Android users who have root.
Requiring xsv would reduce availability to a fraction of where it can be run now.
#!/bin/sh
find /data \
-name WifiConfigStore.xml \
-print0 |
xargs -0 awk '
/"SSID/ { s = 1 }
/PreShared/ { p = 1 }
s || p {
gsub(/[<][^>]+[>]/, "")
sub(/^[&]quot;/, "")
sub(/[&]quot;$/, "")
gsub(/[&]quot;/, "\"")
gsub(/[&]amp;/, "\\&")
gsub(/[&]lt;/, "<")
gsub(/[&]gt;/, ">")
}
s { s = 0; printf "%-32.32s ", $0 }
p { p = 0; print }
' | sort -f
Note that the -print0 null processing is not POSIX. This is a reasonable compromise of standards compliance, as it does not reduce the base of available users.I did try to do this first with arrarys, but awk segfaulted.
> parsing csv can get tricky
Yes I'd say awk is unsuitable for csv files that can contain text fields with arbitrary text, that includes commas, newlines, quotes etc.
It gets hackier and becomes the wrong tool as the number of edge cases to handle increases.
import html
print(html.unescape(foo))
And the best part - you don't need to debug/update the (g)sub list every time you stumble upon new weird &whatever; too. And there are a lot of those out there:https://www.reddit.com/r/LineageOS/comments/gzm7to/wifi_pass...
The code as written will convert
&lt;
to <
When it should (presumably) be <git://bitreich.org/xml2tsv
Then parsing it with AWK should be easy.
Compared to straight Python (i.e. not Numpy/pandas), it can also be surprisingly fast[1]. I experienced this personally on a trivial problem: take a TSV of (I,J) pairs, histogram I/J into a given number of bins. I can’t remember the exact figures now, but it went like this: on a five-gig file, pure Python is annoyingly slow, GNU awk is okayish but I still have to wait, mawk is fast enough that I don’t wait for it, C is of course still at least an order of magnitude faster but at that point it doesn’t matter.
[1] https://brenocon.com/blog/2009/09/dont-mawk-awk-the-fastest-...; note that the original author of mawk has made a release since then, at https://github.com/mikebrennan000/mawk-2, that doesn’t have the crashes I encountered with 64-bit builds of Dickey’s fork
I mean, sometimes for small edits sed it's better, and awk for some tabular based files by using xml2tsv or lots of TSV related tools.
But for medium sized projects, Perl it's the obvious tool against something that requieres something similar to awk/sed but more complex data parsing.
That, and shell, sed and awk are standard tools which are easy to learn, every unix user should learn basics of (even a java developer), and this won't change anytime soon.
However, this can't be said of Perl - it is a powerful tool, but it became culturally obsolete and deprecated, and most unix users in the present and the future won't be bothered to learn it, when learning more modern languages like Python is a better investment.
The problem was that every time someone else was asked to add any feature to it, they freaked out at the language choice and I had to get on a plane.
Pipes are great because they enable you to trivially send data between programs, and they're terrible for the same reason. While the execution time on modern computers for the average data size isn't noticeable, on larger datasets or repeated execution, it absolutely is. If you don't have to pipe, don't.
I most recently did that an hour ago, and a few hours ago, pretty much every day.
awk '! x[$0]++' foo
It is just so damn good at the things it is good at, but nobody learns to use it, because learning Awk is inefficient in the grand scheme of things. There are better things you can learn.
There is one place where Awk has undeniable superiority—and that is its use in environments where bureaucratic rules prohibit the distribution of programs / code (Perl, Python), but where Awk is permitted.
It's absolutely efficient to learn, because there isn't much to learn.
I mean, I like Awk, but I’ve also been programming for a long time and I’ve amortized the cost of learning Awk.
CLI world is full of bespoke text interfaces. AWK is the tool for dealing with those programmatically.
pdf / csv / excel export from my three webbanks, a bit of pdftotext, or soffice conversion just to pipe to awk to augment it and render properly formatted spreadsheet
Awk + bash could easily recreate most existing code in a couple of lines
I think the confusing factor with awk is that it allows you to leave out variuos levels of structure in the really simple scripts, meaning that the same scripts you see around will look quite different.
E.g. all the following would be the same (looking for the string "something" in column 1, and printing only those lines):
'$1 == "somestring"'
'$1 == "something" { print }'
'($1 == "something")'
'($1 == "something") { print }'
... to give a small example.
At least this confused me a lot in the beginning.
awk '{print $3}'
is what cut should have been all along. Specifically, this will give you the third space delimited field, where multiple spaces are coalesced. cut -d ' ' -f 3
will get you whatever is between the 3rd and 4th space.[1]: https://brendaneich.com/2010/07/a-brief-history-of-javascrip...
$ query_something | awk 'generate commands' | sh
For larger programs, I wrote and use ngetopt.awk: https://github.com/joepvd/ngetopt.awk. This is a loadable library for gawk that lets you add option parsing for programs.
It’s well suited to iterative composition of the commands: I’ll write the query/find part, and (with ctrl P) add the awk manipulations, and then pipe to sh.
If it doesn’t have side effects you can pass through “head” before “sh” to check syntax on a subset.
ls | awk '{if (length($1)==7) {print "cat " $1 }}' | sh
it is something you really aren't supposed to do because bad inputs could be executed by the shell. Personally the control structures for bash never stick in my mind because they are so unlike conventional programming languages (and I only write shell scripts sporadically) so I have to look them up in the info pages each time. I could do something like the above with xargs but same thing, I find it painful to look at the man page for xargs.When I show this trick to younger people it seems most of them aren't familiar with awk at all.
For me the shell is mostly displaced by "single file Python" where I stick to the standard library and don't pip anything, for simple scripting it can be a little more code than bash but there is no cliff where things get more difficult and I code Python almost every day so I know where to find everything in the manual that isn't on my fingertips.
Though you’d have to be confident that running it twice is going to give the same results. If it’s remote data that could change then weird/bad/nasty things could happen.
For anything non trivial, best to separate those steps and generate a temp script to execute.
I run stuff like You mentioned (piping to a shell) and also system() frequently. It depends on many factors which one I'll choose.
(FWIW, I'm also quite decent in many shell flavors on many Unix/Linux variants, so that is another determinant)
Eg, the Busybox ash(1) that I frequently work with does not support arrays, but its awk(1) does...
Once I wrote a 300-line Awk script to install a kernel driver. It would scan a hardware bus, and ask the user questions before loading the driver onto the system. Lots of fun!
*Especially keeping in mind that these people wrote things like this for fun.
https://github.com/crossbowerbt/awk-webserver
I know, just because you can doesn't mean you should.
This is an AWK script that serves HTTP requests by listening on port 8080.
It defines several functions:
1. `send` function takes in status code, status message, content, content type, and content length and sends an HTTP response with the provided information.
2. `cf` function checks if a path contains `..` and returns 0 if it does, otherwise returns 1.
3. `mt` function determines the MIME type of a file using the `file` command.
The script sets the record separator RS and output record separator ORS to \r\n which is the line ending used in HTTP.
It enters an infinite loop listening for incoming connections on port 8080. When a connection is established, it reads the HTTP request line by line using `getline` and processes the `GET` request. If no path is provided, the script serves `index.html`. It checks if the requested path is safe using the `cf` function, and if it is a file, reads the file and sends an HTTP response using the `send` function. If the path is not valid or the file is not found, it sends a "404 Not Found" response.
gawk '@load "filefuncs"
@load "readfile"
function send(s, e, d, t, b) {
print "HTTP/1.0 " s " " e |& S
print "Content-Length: " b |& S
print "Content-Type: " t |& S
print d |& S
close(S)
}
function cf(x) {
split(x, y, "/")
for (z in y) {
print "FOUND " y[z]
if (y[z] == "..") {
return 0
}
}
return 1
}
function mt(f) {
c = "file -b --mime-type " f
r = ""
while ((c | getline z) > 0) {
r = r z
}
close(c)
return r
}
BEGIN {
# Change to the specified directory
if (ARGV[1] != "") {
if (chdir(ARGV[1])) {
print "Failed to chdir to " ARGV[1]
exit
}
ARGC = 1
}
# Set the record separator and output record separator
RS = ORS = "\r\n"
# Listen for incoming connections
while (1) {
S = "/inet/tcp/8080/0/0"
while ((S |& getline l) > 0) {
split(l, f, " ")
if (f[1] == "GET") {
p = substr(f[2], 2)
}
if (p == "") {
p = "index.html"
}
stat(p, s)
if (cf(p) && s["type"] == "file") {
m = mt(p)
o = readfile(p)
send(200, "OK", o, m, s["size"])
break
}
n = "<html>Not Found</html>"
send(404, "Not Found", n, "text/html" RS, length(n))
break
}
}
}'
I added some comments to explain what each section of the code does. Let me know if you have any questions!You can also pretty print gawk source with -o[filename], using - as the filename to print to stdout, so you can just run the oneliner version, but lead with "gawk -o- " and it will print the pretty version.
Edit: Also, the performance is terrible. Gawk's listen sockets suck because you get no granular control over listen/accept. The socat based one I replied to is probably much better.