The Awk Programming Language (1988) [pdf]
archive.org
archive.org
> But the real reason to learn awk is to have an excuse to read the superb book The AWK Programming Language by its authors Aho, Kernighan, and Weinberger. You would think, from the name, that it simply teaches you awk. Actually, that is just the beginning. Launching into the vast array of problems that can be tackled once one is using a concise scripting language that makes string manipulation easy — and awk was one of the first — it proceeds to teach the reader how to implement a database, a parser, an interpreter, and (if memory serves me) a compiler for a small project-specific computer language! If only they had also programmed an example operating system using awk, the book would have been a fairly complete survey introduction to computer science!
To bytecode; I wanted to use the awk-based compiler as the initial bootstrap stage for a self-hosted compiler. Disturbingly, it worked fine. Disappointingly, it was actually faster than the self-hosted version. But it's so not the right language to write compilers in. Not having actual datastructures was a problem. But it was a surprisingly clean 1.5kloc or so. awk's still my go-to language for tiny, one-shot programming and text processing tasks.
http://cowlark.com/mercat (near the bottom)
(...oh god, I wrote that in 1997?)
But no, there's always one. :-)
I agree. This idea doesn't receive enough attention. If you pick your constraints you can make a particular envelope of uses easy and ones you don't care about hard.
AWK's choice to be a per line processor, with optional sections for processing before all lines and after all lines is self-limiting but it defines a useful envelope of use.
Not sure what that would mean. I think the tool was designed to be a user's programming language. I liken to think that `awk` was the Excel + VBScript of its days.
Now, you've posted a compiler/interpreter in awk for a C-like language that could allow easier porting of a C compiler's source. Hmmm. The license would have to be BSD so the BSD's could use it, too. Or pieces of it in my own solution.
I have a feeling whatever comes out of this won't make it into next edition of Beautiful Code. ;)
Let me also say that if you actually want to use this for anything you're crazy. I wrote it when I was... younger... and when I had no idea what I was doing. The only thing it's useful for these days is looking at and laughing at.
AFAIK they're both ubiquitous, though you might need a particular awk like gawk for library functions, depending on what you need to do. Nowadays I'm way more likely to use Python, though of course it's a much bigger dependency.
Sorry about the injury, and good luck -- I'd like to hear how it goes.
Thanks for sharing awklisp. Nice reading for a Sunday morning.
Here is the result: https://github.com/dbohdan/all-caps-basic. It targets C and uses libgc along with antirez's sds for strings. The multi-pass design with each pass consuming and producing text is intended to make the intermediate results easy to inspect, making the compiler a kind of working model. The passes are also meant to be replaceable, so you could theoretically replace the C dependencies with something else or generate native code directly in Awk. You can see some code examples in test/. Unfortunately, the compiler is very incomplete. I mean to come back to it at least to publish an implementation of arrays.
Which meant that it was perfectly allowable for it to be hacky and non-future proof, which it was.
Here's part of the code which read local variables definitions (in C-like syntax):
function do_local( \
nt, st, n, t){
nt = readtype()
st = tokens
ensuretoken(readtoken(), token_word)
n = tokens
outofscope(n, 1)
ldb[n] = "var"
ldb[n, "type"] = st
ldb[n, "sp"] = sp - spmark
emit("# local " n " of type " st " at " (sp-spmark) "\n")
tokens is the value of the current token, ldb is the symbol table; you can see how I'm faking a structure using an associative array indexed by keyword.There's nothing actually very wrong with this code, but there's no type safety, barely any checking for undefined variables, no checking for mistyped structure field names, no proper data types at all, in fact... awk wouldn't scale for making the compiler much more complicated than it currently is. But it did hit the ideal sweet spot for getting something working relatively quickly for a one-shot job. It's still really good at that.
* Cannot have an array as an input to a function
* Cannot return an array from a function
* Meta-programming or pointers are only (barely) available in gawk
* There is an `@include` statement for `gawk` that is not part of POSIX, and there is no name spacing involved.
* Functions names can only exist in the global name space
There are some reasons somebody felt an urge to create perl... Still loving awk, and using it every day for text processing jobs.
Larry Wall (creator of Perl) says something pretty close to that here[1]:
"I was too lazy to do it in awk because it would have been hard to get awk to jump through the hoops I was wanting it to jump through. I was too impatient to wait for awk to finish because it was so slow. And finally, I had the hubris to think I could do better."
http://cahighways.org/wordpress/?p=8019
http://ieeexplore.ieee.org/document/213253/
BLACKER's heavy lifting in security was done by high-assurance kernel called GEMSOS:
http://www.cse.psu.edu/~trj1/cse443-s12/docs/ch6.pdf
It was a classified work for TRW whose details took a long time to get released. Might be why he rarely mentioned the BLACKER project in its origins. Possibly trying to obfuscate it a bit to avoid breaking laws.
This is neat! Question for you regarding the formatting/syntax used-
Why the slash-newlines in function declarations?
E.g.
function scope(name, \
s) {It found a couple bugs because the test suite is quite comprehensive. I think it's somewhat interesting that 5000 or so lines of C code polished over 20 years still has memory bugs.
I didn't fix the bugs, but anyone should feel free to clone it and maybe get some karma points from Kernighan. Maybe he will make a 2017 release. He is fairly responsive to email from what I can tell :)
Typical might be the word you're looking for there. The CompSci people doing static analysis for C programs often apply them to popular FOSS. They find new errors about every time.
But the C code is invalid. In many circumstances, a one byte buffer overrun is exceedingly likely to "work", but ASAN flags those as bugs with 100% reliability.
If I recall, it also had eiter a use-after-free or double free. The former can obviously cause problems but may not, not sure about the latter.
Anyhow, I asked Brian if we could base it off the one true awk and he tarred up ~bwk/awk and sent it to me.
I love that guy, the culture of the Bell Labs people and the people that worked with them is great.
I've stolen a bunch of awk ideas over the years. BitKeeper (first DSCM) has a programming "language" for digging info out of the repository. For example, this:
http://www.mcvoy.com/lm/bkdocs/dspec-changes-json-v.txt
prints out the repo history as a json stream. One of my guys said that it couldn't be done, heh, it could be :)
Everyone should learn some awk, it's so handy.
ls -lR /path/to/dir | awk ' { s += $5 } END { print s / 1024 " K" } '
$5 is the 5th field of the output, which is the file size field in the case of ls output. The code inside the first set of braces runs once for every line of input (which comes from standard input, so from the ls command, in this case), and the code inside the second set of braces runs at the end of the input, calculating and printing the desired result of the total of all file sizes for files found by ls, in kilobytes. It can easily be changed to output the total in bytes or megabytes by dropping the '/ 1024' or adding another one after the first. Variable s is initialized to 0 by default at the start.
You can get similar info with "du -hs /path/to/dir" but the ls plus awk pipeline lends itself to more customization, such as adding conditions for the type or owner of the file, etc.
find /path/to/dir -type f -printf "%s\n" | awk ' { s += $0 } END { print s " bytes" } '
The find command has much powerful file filtering capabilities than that of the ls command and works better with weird characters in filenames.
There is also the -print0 option to find to handle filenames with newlines in them.
-print0 may be non-POSIX and a GNU extension.
POSIX has -print, but interestingly, in some Unixes I have seen that not using -print still prints the filenames found, by default.
If no expression is present, -print shall be used as the expression. Otherwise, if the given expression does not contain any of the primaries -exec, -ok, or -print, the given expression shall be effectively replaced by: ( given_expression ) -print
http://pubs.opengroup.org/onlinepubs/9699919799/utilities/fi...
You are correct regarding the GNU option -B, which would not be relevant here, and regarding the fact that neither of these options are in POSIX. Thanks for pointing out those limitations.
http://www.drdobbs.com/tools/examining-the-tawk-compiler/184...
I ended up writing a couple of command-line email utility programs with it that I sold, for a while.
"What happened" was industry-wide standardization on dynamic-linking of prerequisite (library) code (IOW, this code stays in separate DLL files typically stored in system-global locations), leading to the need for "installer" software whose purpose (I presume; I've entirely avoided dealing with that stuff) is to ensure that all prerequisite dynamic libraries are upgraded to the minimum version needed by the SW being installed, replaces any old version(s) of the program with the new, and modifies the Windows Registry in various and sundry ways (can you say "system-global variables run amok"?). The solution which I prefer is to build static-linked .EXEs (binaries) instead of dynamic-linked. Convincing toolchains to do this is a small exercise for the reader. OBTW: I think go (golang) static-links by default.
I stopped using TAWK compiler when I discovered Lua (5.1; IMHO a substantially better language than TAWK (this is not a criticism of TAWK)). I even went so far as to commission a "Lua Compiler" for Win32 which behaved almost identically to the TAWK compiler; I used this with great success for a few years. Unfortunately it was an internal tool which I lost access to when I departed that employer.
P.S. IIRC Borland Delphi also builds static-linked EXEs by default. I wrote one Delphi 2 (Win NT 4.0 era) program whose source code I've kinda lost track of which still runs fine on Win10 x64. TAWK and Lua are more productive languages than Delphi/Pascal, and it's trivially easy to add your own C library functions to Lua (for improved performance or added functionality), so I gravitated toward an overall preference for Lua.
And freepascal, nim. Perhaps ocaml (but might require some magic to generate a standalone exe? I belive unison is available as just an exe file?).
Ed: and rust?
I wrote a simple command line statistics tool that uses awk to calculate sum, stddev, and more. https://github.com/numcommand
It's short, clear, and concise. It's useful and helps you solve real problems with AWK. Who could ask for anything more?
The fact that one of the authors is Brian Kernighan is partly why, IMO. Just a few days back, I commented here in reply to someone about the quality of his K&P and K&R books (both of which I've used for trainings), on Unix and C respectively.
The number one example for me is counting by string in a csv file:
>> awk -F',' '{a[$1] +=1} END {for(v in a) print v,a[v]}'
Not that this is particularly difficult stuff, it's just a bit exhausting to find myself typing that over and over again. I'd love a more concise alternative to this.
Also, 'sort | uniq -c' is not a viable alternative for very large files.
A useful fact that I've seen some people didn't know: Unix metacharacters such as star and question mark (for filename matching, $ (in various uses such as $ star, $#, $!, etc.) - are all expanded by the shell, not by individual commands, so use of metacharacters is actually available to all commands and scripts that are run at the shell prompt - not just to selected or built-in ones.
Contrast that with DOS (at least in earlier versions) which had the problem that some commands supported wild-card characters such as for filename matching, but others did not. You could write your own logic for that using OS API calls (Int 21H etc., IIRC), or calls named like FindFirst and FindNext)but it was not built-in and freely available.
Edited for formatting.
It's really a low bar to clear. In fact, it would be far more surprising if a language such as awk with counters, conditional statements, the ability to jump to statements (i.e., loops), and the ability to change memory were not Turing complete.
Edit: clarity
Bought this book 2nd hand online. This book on one day costs $150, and on the next $2. The first bit has been an awesome read, never got to read much more. Tend to read much more from $READER. Sure this PDF will get me going again!
You'll never use it in industry, but a few weeks with Prolog will bend your mind in just the right ways and teach you more about how you can model computation differently than six months with a Lisp. It's also cool as shit and really fun to program in. Prolog is a language you can learn just for the sheer joy in expanding your notions of what programming is, or at least could be.
$> python2.7 pdfid.py The_AWK_Programming_Language.pdf
PDFiD 0.2.1 The_AWK_Programming_Language.pdf
PDF Header: %PDF-1.6
..
/Page 0
/Encrypt 0
/ObjStm 7
/JS 0
/JavaScript 0
/AA 1
/OpenAction 0
/AcroForm 0
/JBIG2Decode 222
...
It has /AA which is an automatic load action, and it has a lot of objects which could contain javascript, would need closer scrutiny I think.Wouldn't a tech pdf of a popular book that is impossible to obtain legally in digital form be an excellent vector to deliver malware to tech users with probably lots of stored credentials to resources?
GAWK: Effective AWK Programming
http://www.nongnu.org/txr/txr-manpage.html#N-000264BC
It has direct counterparts to all POSIX features, plus a number of extensions similar to ones found in Gawk, as well as some of its own: for instance, range expressions which freely combine with other expressions (including other range expressions), and range expressions which exclude either or both endpoints.
Given, AWK's problem space is very small, but still...
In fact, you can even read indices of an array without declaring it. If you do, it's auto-declared as having 11 elements (with indices 0 to 10), filled with zeroes.
This was still supported as late as VB for DOS. I wouldn't be surprised if VB6 also had this behavior.
$ perl -e '++$i && print $i;$foo[2]=99;$foo[1]=88;for (@foo) {print}'
18899crunchgen, a compiled C program, has to call AWK.
Anyone out there do AWK-less builds?
Why did I need to learn a little AWK?
Because I could work out how crunched binaries were built without knowing some AWK.
Best thing about AWK IMO is the C-like syntax.
For anyone learning C and AWK concurrently, this kills two birds with one stone.
Loading the file into Excel took literally minutes as Excel tried to parse every field. It bogged down a 16GB RAM machine.
Using awk and uniq, the total run time of getting a solution , including reading the many MB of files and generating a summary into another file, was about 6 seconds.