Awk As A Major Systems Programming Language, Revisited (2018)
skeeve.com
skeeve.com
Was the increased productivity due to Awk or was it due to Perl?
Also Perl has better support for programming-in-the-large than awk, with modules, lexical scoping and CPAN.
Booking.com, IMDB and I believe the Amazon frontend are all written in Perl.
perl -anE 'say($F[0]) if /error/' big.log >/dev/null 0.85s user 0.02s system 99% cpu 0.867 total
mawk '/error/ {print $1}' big.log > /dev/null 0.21s user 0.03s system 99% cpu 0.246 totalhttps://github.com/mikebrennan000/mawk-2/blob/master/about-m...
Plus if you're going to talk about older builds of mawk then you can't really ignore older versions of Perl as well (which was originally released in 1987). Otherwise you're not making a fair comparison.
I should add, I have absolutely nothing against mawk. It just wasn't something available on any of the POSIX systems I used in the 90s. tbh even now it's an optional install but at least it is an easier install than it was in the 90s.
Local variables unfortunately can only be obtained in the form of extra parameters (which are not passed by the caller).
function foo(x, y, # params
z, w) # locals
{
}
There is a convention to separate the two by some obvious whitespace.There are no block scoped-locals: all locals have to go into the parameter list. Initial values cannot be specified.
(Speaking of which, there are hacks for simulating the feature of optional parameters with defaulted values.)
In GNU Awk, from the following experiment, the scope appears lexical:
function bar()
{
return x
}
function foo(x)
{
x = 3
return bar();
}
BEGIN { x = 42; print foo(); }
The output is 42, which means that the "x = 3" assignment to the local variable x in foo does not affect the access to the free variable x in bar, as it would under dynamic scope.My recollection is there was a lot of disdain for scripting languages because they could not match the speed of C or other low level languages. Today, computer speed is so many orders of magnitude faster it’s hard to believe the speed of scripting languages was ever an issue.
John Ousterhout’s paper on programmer productivity gains comes to mind:
https://web.stanford.edu/~ouster/cgi-bin/papers/scripting.pd...
The good old days.
Then I realized most all the awk/sed stuff looked very similar to Perl, which I already knew, and I ended up just becoming very good at Perl 1-liners.
For these jobs even an "import re" in Python is too much typing, compared to ubiquitous regular expressions in sed or awk. Awk is great at handling lines of delimited fields. Again in Python you would be writing your own boilerplate to read lines from files, splitting these lines, converting strings to numbers etc.
Once things get complicated though you should probably push whatever transformations you need to do to a database.
The other day I set out to use BASH to mass rename, and I had such trouble backslash-escaping the dot characters and whitespaces to prevent sed from interpreting them specially. Eventually I gave up and searched for the pythonic way to rename them, and it was as simple as a string replacement followed by a call to os.rename() inside a for-loop. It was a breath of fresh air to escape “command line Kung Fu” and fearing the thrashing from shell globbing.
To be fair, I got my start in using sed and awk as powertools in the BASH command pipeline, but I don’t miss them compared to a language with strong data types and simple built-in methods for handling complex manipulation. Python is built in to basically every Linux distro that has BASH, and for the sake of simple transformations, it offers a lot of succinct methods that work on either 2.7 or 3.x with no external packages.
But if you're dealing with changing stuff in text files greater than 25 gigs I'd always recommend using the unix cmd line tool set for simple modifications if performance is a key concern. It works the same everywhere and can rip through big files. But I'm a Vim guy and I live in the shell so I'm surely quite biased lol.
There are plenty of pretty easy one liners though that can rename files depending on if you can define a pattern correctly though.
echo 'TV Show Name 2014.L337.H4XXX1080p-420.mkv' | awk -F . -v q="'" '{print q$0q" "q$1".mkv"q }' | xargs mv
find . -type f -name '*.mkv' -printf '%f' | awk -F . -v q="'" '{print q$0q" "q$1".mkv"q}' | xargs mv
python -c 'import sys; f=open(sys.argv[1]);print(len(f.readlines()))' .zshrc
But I wouldn't recommend python as a good one-liner language. So the question remains, what is awk particularly good at?
awk -F, '{ print $2 }'Now remember, this was popular when perl wasn't even something you could depend upon on all systems.
These days there's no reason you should even do anything terribly complicated in a nasty oneliner. Just save a nice function somewhere and call it with a concise alias.
Processing regular (i.e. machine-, not human-generated) tabular text files line by line. If you already know Python there are probably few reasons to learn awk.
One of those could be working in constrained environments where you have awk but no Python. Awk is part of POSIX and part of Busybox, so there are almost no environments which have a working Python installed but no Awk, while the reverse is common enough (e.g. embedded systems). Apple plans to remove Python, Perl and Ruby from the default MacOS install in the next version, but probably not Awk (POSIX).
(And I hadn't realized Apple was planning on getting rid of Python, Perl, and Ruby from default installs! Wow.)
Then there's flexibility. With bash you have a massive ecosystem of command line tools, including the ability to in-line code for tools like sed, awk and perl right there in your script. Python has a great standard library, but often the best way to do something in Python is using a third party module, so now you need to curate a module library across all your boxes. Also suppose your script needs to do things other than just process text, suppose I want to ssh onto a box and run a command? Is Paramiko installed, no? So now I need to invoke ssh in a sub-shell which is a PITA in Python compared to bash. Yes it's easier in later 3.x versions but now we're back to version soup again.
We have bash scripts running in our environment that are probably 20 years old, and if I write a new one now it will still work fine in 20 years time. Often it will be a quarter of the length of an equivalent Python script and much easier to understand and reason about.
Then again, there are many cases where a Python script makes a huge amount more sense, often run as part of a bash script.
For now I'm teaching myself CL tools, and I feel if I need to pipe more than 5 times in one liners I'd rather use Python.
I know I like it, but I'd have to be working with it alot more for the basics to sink in.
And the thing is that in almost all cases Python or (ugh) bash can get the same job done so I rarely pick up Awk.
Mawk on debian is super fast and can determine statistics like messages/sec, uniq ip addresses for a particular user, etc from 10GB log files very quickly.
Several of our "grunt work" servers run macOS. So you're either left with a seriously outdated set of utilities, or need to write wrappers to first attempt to use gX versions first, and then fallback to standard names.
The options I'm trying to use aren't GNU specific, and can be found across the BSDs. But Apple is old, so you can't fully count on POSIX.
Strange thing to say since macOS is officially certified UNIX(r) POSIX, and has been for many years:
* https://www.opengroup.org/openbrand/register/brand3653.htm
[0] https://www.opengroup.org/csq/repository/noreferences=1&RID=...
I have to look at the manual to remember how to write loops in bash, so I write an awk script that writes a bash script and pipes to 'bash'. People will tell you this is a bad idea because you get character escaping risks as with SQL injection - they are right, but it is so much fun.
I have done this with three layers of code generation.
I do this on occasion, it is fun~
The only instance where awk has utility for me is when the program is short enough to be explicitly specified at the command line, for example:
$3 > 10 { print $6 - $5 }
for that it is awesome, you don't have to look inside another program to figure exactly what is happening, it is explicit etc. it is also super fast, much-much faster than splitting with a typical scripting languagefor anything more complicated than that, it offers very few benefits (I'd be hard pressed to name any) and significant limitations.
Thus IMO the problem with awk is that it does too little and offers to little room to grow.
But it could be that I was just never shown one.
Bash is a glue language, while Awk a scripting one, so it rather makes sense to compare Awk with Python/Ruby/similar.
Granted, this is quite a constraint and perhaps there aren't many problems that fit this model, but when you do encounter one of those problems it's good to have the right tool for the job.
The man pages are nice but I didn't have the patience to start reading every thing to just do simple stuff like replace regex pattern with a content of a file located at the path generated from a capture group of that regex and some other stuff.
Comparatively, perl was also easier to find stuff for in their docs. I ended up using that for some places.
Not quite sure what you mean but it does sound like awk was the wrong tool for the job there. For the sort of templating I'm thinking of shell scripts or m4 would have been a better tool. Taking some structured data and piping it to one of those is where awk shines (that and pattern matching).
I routinely use awk when cut is obsessing about IFS awk does LWSP compression so you can get ' one two. three' to match properly when cut thinks field 1 (oh god, code which counts from 1 not zero..) is a ' ' space awk '{print $1}' just works.
I used awk to compile a list of unique IP addresses seen over GB inputs, 350m+ unique IPs. it was within scale both for memory footprint and speed of python and perl, for hash constructs. Basically, Brian coded it efficiently, all the perl claims of maximal hash efficiency did not add much in terms of speed OR size outcome.
I choose to code in python3, but I use awk for one liners. Its great. I avoid gawk-isms. I don't see the need.
At some point I thought that had the ideal use case for awk (a git --graph filter) and spent an evening desperately putting it together because, as other commenters mentioned, it's hard to find good documentation and examples online. Sure, I have a fast and mostly-working filter now, but the code is also hard to understand or even debug. On the other hand, the examples linked in the article are actually a lot more readable than I expected, so maybe it's something to consider for small but frequently-used log parsing scripts.
Sometimes people observe I'm using three tools with pipes to do one job, and I freely admit grep <pat> file | awk '{print $2}' | sed -e 's/this/that/g' is probably stupid, but I do think of these atoms as tools for the job. Grep aside, sed and awk should be fully interchangeable for many pipe jobs, and when not BEGIN{} ... END{} you could do the whole thing in awk or sed simply. If it has pre- and post- states, Awk is ideal. But.. the mind does what the fingers remember.
Pipes are cheap.
And even then, things can fall apart. Python2 made it, but now it's going away. (Yes, there is a Python3 often present, but neither is a substitute for the other.)
Great tool though