Why you should learn at least a little bit of Awk
gregable.com
gregable.com
Years later I ran into Brian Kernighan at a conference and told him the story, ending it with "and that's when I knew she was the woman for me." He looked at me like I was nuts.
This is 100% true. A coworker of mine implemented an elevation-bitmap-to-3d-model conversion tool in 160 lines of Awk. It ran faster than our "good" Matlab tool by a factor of 10.
Awk (or Perl) doubles the usefulness of Unix. Most of the common commands in Unix are query commands. When you need to start manipulating queried data, Awk is where the rubber meets the road. Piping data through the shell stops being read-only, and becomes interactive.
Could you give a bit more details there? I don't have any experience with matlab, but I tend to think of awk as fast to write code in (and start up), though not particularly fast in execution. (Roughly on par with Python, i.e., usually good enough.)
Also: "I have since found large datasets where mawk is buggy and gives the wrong result. nawk seems safe." makes me uneasy, as does the fact that it was unmaintained for a while.
Still, the One True Awk still has my favorite opening line in its "b.c" source file:
/* lasciate ogne speranza, voi ch'intrate. */Interestingly enough, windows powershell structures its cmdlets in the same way. Makes lots of sense for stream processing as you said.
If the unit of input in this kind of stream processing system doesn't match the problem domain exactly, things get very difficult very quickly.
Regular expressions are my favourite secret weapon; So many problems are made simple by regular expressions and so few people (outside of IT) know of them.
I was curious enough that I bought and read it just at the end of summer. It really is excellent. Highly, highly recommended.
Ieursalimschy's _Programming in Lua_ ("PiL") was written in a similar style. I recommend it quite highly, too. Great language, great programming book.
Also, the PSD, SMM, and USD books (_4.4BSD Programmer's Supplementary Documents_, etc.) are dry, but also have excellent introductions to several classic Unix tools. They're included as documentation in some BSD installations, and should be easy to find otherwise. The intros to lex and yacc are particularly good.
It's incredibly handy, yet the language is small enough that you can learn most of it in an evening, with just a bit longer if you don't know regular expressions.
I do tend to use Perl for things where speed matters, though, especially with large amounts of data going through a regex--- Perl's regex engine seems considerably faster than any awk (or especially sed) I've tested, at least on a few examples I've ported in the past. I was surprised once to get an 8x speedup by porting a 3-line sed script to a 3-line perl script (it was basically doing s/ABC/A\nC/g on a multigigabyte file). I've heard mawk can be speed-competitive with Perl, though.
It's based on PEGs, a different formalism than regular expressions. PEGs are more expressive - they're able to handle balanced, recursive structures, for example. LPEG is a nice middle ground between regular expressions and a full LALR(1) parser.
awk "{print $0}"
does not work. Awk programs need single quotes to prevent bash expansion.
awk 1I can recommend this text file of awk one-liners:
http://www.pement.org/awk/awk1line.txt
And for completeness, here's one for sed:
O yeah and let's not forget Perl.
In other words, awk is unbeatable for stream crunching. (That's the point of being domain specific, by the way.)
I can say the same for ruby and python (and perl).
From personal experience, as an awk script/program becomes more important - it will evolve with more requirements and it will start to be clunky. It just isn't practical to stick with it since you'll eventually need the features/libraries that the other languages have. Given the choices we have today, why even start with awk?
On the performance side, you can always just use Lua if that's really important.
I write a lot of little awk scripts, but if they grow past ~5 lines, they usually get rewritten in Lua. (Perhaps eventually with inner loops in C.) Still, Awk is simple and useful enough that it's still worth knowing.
Doesn't every language have regular expressions built in now? Again I still fail to see the point of writing it in Awk when you can write something small and fast in a more powerful and modern language.
It's a higher-level approach than typical scripting languages, and that's why it can be so concise - the model makes a lot of unpacking and looping implicit. It's a DSL for stream-processing problems which are easy phrased as "count these", "transform this into that", etc.
Are you familiar with Prolog? It uses a similar approach, but can match on whole trees (and other complex, nested data structures), not just a list of $N string/numeric tokens. Also, it supports backtracking - at any point, if it reaches a dead end, it can back up arbitrarily and try a different approach. Sometimes slow, but very handy for prototyping.
I agree that using another language than awk makes sense after a few lines, but it's still a sweet spot for 1-5ish line programs. Since awk itself is small enough that a two page cheat sheet is sufficient, it's worth keeping around. Perl (for example) has many nooks and crannies I forget about if I don't use it frequently.
The parent post mentions Prolog, which is a good example, but there are several others worth trying that frequently come up on HN; Scala, Haskell, F#, and Ocaml spring to mind.
I can't speak for Scala, but the PM in Haskell and OCaml is a bit different since it's informed by the static typing. When patterns have variant types (i.e., x is either Foo, Bar, or Baz * int), it also checks for complete coverage. Same general concept, different flavor. Also very useful.
I mentioned Prolog in particular because its emphasis on unification and backtracking make it the most pattern-matching-centric programming language I've seen. Where other languages have pattern matching, it almost is pattern matching.
Also, there are well-known ways to compile pattern specifications into efficient decision trees, so while it's a very expressive abstraction, it's not necessarily an expensive one. If they're being constructed at runtime (as they are in my Lua library), you can generally get a big improvement by just indexing on the patterns' first fields and doing linear search thereafter.
Again, my point is that you can do the same thing in Ruby, Python, Perl, or Lua just as easily and fast as you can in Awk. Awk used to have a nice niche years back. It pretty much lost that niche the second Perl got popular. It's even more irrevant now that Ruby and Python are even easier and faster to build stuff with quickly
About the only real pragmatic reason I can think of for learning Awk is to migrate existing Awk scripts, that started as nice useful one liners but eventually evolved into spaghetti, to Python / Ruby
| awk '{print $2}' | ...
| python -c 'import sys
for line in sys.stdin:
try:
print line.split()[1]
except IndexError:
print' | ... | perl -nae 'print $F[1], "\n"'
| ruby -nae 'puts $F[1]'
Your local python oil vendor may have a one liner for that language as well.http://zork.net/~nick/loyhargil/if/if.awk
For comparison, here are all the published examples of this exercise in a variety of systems:
http://www.firthworks.com/roger/cloak/
I won't say it's the best tool for this job, but I feel that the awkishness provides a certain elegance to some aspects.
I'd bet that people get shit done 10x+ times faster in awk/lua/python/ruby/lisp/whatever until having to work with nasty C++-specific libraries dominates, though. (C is friendlier that way.)