The Awk Programming Language
awk.dev
awk.dev
Gawk and awk will soon have a new "--csv" option that enables proper CSV input mode (parsing files with quoted and multiline fields per the CSV RFC). I'm really glad Arnold Robbins added a robust "--csv" implementation to Gawk, too, because that's really the most-heavily used version of AWK nowadays. I've already got CSV support in my own GoAWK implementation, and I'll be adding "--csv" to make it compatible.
I'm really glad this new updated version is coming out!
Awesome!!!! Super excited to see this!
To be really useful as a format it would just need for text editors to: -display something distinct for the field separator (some editors do this) -treat the record separator character like a carriage return (not aware of any editors that do this)
This made me think of WordPerfect's "reveal codes" functionality. :)
(Word's "Reveal Formatting" is supposedly similar.)
Which would be trivial too.
* https://miller.readthedocs.io/en/6.8.0/file-formats/#csvtsva...
I have programs that handle it.
* https://jdebp.uk/Softwares/nosh/guide/commands/console-flat-...
It will be nice to have a awk cookbook for CSV. In terms of CSV maniupulation and querying there is only a limited number of operations and I think there is potential to standardize those operation using AWK.
I guess for the people that are still using nawk, you can set up an AWK envvar so you can { awk -f $AWKU/ucsv.awk -f <(echo '{print NR, $1}') }
https://github.com/Nomarian/Awk-Batteries/blob/master/Units/...
It's not really a one-liner, neither something big, but one can take that as an example regarding that awk is really not just for one-liners.
Meanwhile having `--csv` support is really nice. I'd also like to see things like a builtin `length` function to be standard.
[1]: https://github.com/nvm-sh/nvm/ [2]: https://github.com/nvm-sh/nvm/pull/2827/ [3]: https://github.com/nvm-sh/nvm/blob/9a769630d7/nvm.sh#L1703-L...
The language has some quirks. To declare temporary variables, it's common practice to add extra arguments to functions that won't be used. And traversal of associative arrays is implementation-dependent. I'm not sure what the situation is regarding locale and UTF-8 support.
EDIT: Looks like Brian Kernighan added Unicode support last year.[1]
[0] https://github.com/siraben/awk-vm/blob/master/vm.awk
[1] https://github.com/onetrueawk/awk/commit/9ebe940cf3c652b0e37...
Not really. Later on the book just ran out of line-matching examples to go through and started doing regular programming instead :P. When I actually write AWK code I rely on line-matching and using a variable to handle state.
> Hey you should read the AWK book, it even says how to write a VM!
> Why would I ever want to use AWK for that?
> Well, the input is a text file with one space-delimited instruction per line.
> Hmm... You have a point.
Busybox has their own independent AWK implementation.
https://busybox.net/ https://frippery.org/busybox/
Also see the first edition of the AWK manual online here:
If we're counting minimalist implementations, there's micropython that even runs on microcontrollers that cost a less than 2 dollars.
I’m a big fan of Forth, but it’s not available everywhere. One has to adapt.
It does not mandate Perl or Python.
<https://pubs.opengroup.org/onlinepubs/9699919799/utilities/a...>
<https://pubs.opengroup.org/onlinepubs/9699919799/utilities/c...>
There are many systems which lack Perl or Python, but include awk.
You might be carrying an Android device at the moment --- if you drop to its default userland, that provides a bunch of utilities, including awk, via Busybox. But not, so far as I'm aware, either Perl or Python.
(You can of course install Termux which will then give you both Perl and Python, along with Node.js, ruby, and a whole slew of other scripting and compiled languages. But so long as we're considering stock installs, it's sed and awk.)
edit: this prompted me to write up a little note showing how: https://notes.billmill.org/visualization/graphs/gnuplot/A_ba...
(The nice feature of feedgnuplot of course is that you can _also_ render the plots to images, which youplot can't)
But you could feed them through `textimg`[1] to generate PNGs.
[1] 27623 14272 22218 21267 19037 989 27116 32405 23261 27104 7793 9432 7776 28832 13521 10783 29261 32193 30367 20358 22611 2023 19607 9844 3516 6510 16533 8378 22986 17043 14628 13392 22799 23847 29212 23690 17779 17059 28211 26180 32061 22740 7911 12018 4508 9801 9578 15350 9554 15517 11112 405 22054 2743 26609 7843 713 10975 2830 1126
[2] http://rjp-hosted-files.s3.amazonaws.com/sparkline-demo.png
I know exactly enough to be dangerous and have meant to deep dive for almost a decade.
1) succeed, and regret the messiness of the solution
or
2) fail, and find a non-awk way to handle it.
I really tried to like awk, but its portability hasn't been enough of a feature to raise it above other scripting languages for me. Especially if I'm going to end up in an editor
"Dark corners are basically fractal - no matter how much you illuminate, there is always a smaller but darker one." - - Brian Kernighan (quoted in the GNU Awk book)
I see awk as a DSL to be honest. Yes, it can be used as a general purpose language, but that quickly becomes, as you say, awkward :D
Like many DSLs, it is simple, fast and lightweight as long as it is used for it's intended purpose. Once you start using it for something else, these advantages evaporate pretty quickly, because then you have to essentially work around the DSL design to get it to do what you want.
Aside from that I've mostly been using it for quick statistics [2], but it quickly moves into perl territory...
1: https://github.com/9001/asm/blob/hovudstraum/etc/bin/beeps#L...
0: https://github.com/bentxt/microperl-standalone
1: Original article from 2000 by the author Simon Cozens: https://www.foo.be/docs/tpj/issues/vol5_3/tpj0503-0003.html
It depends on scale.
If you have some quick parsing to do, then awk will get you started quickly, but as you expand your experimentation on what you want to extract/manipulate, it may not be easy to add onto the awk beginnings of your "one liner".
But if you start with awk-like† syntax but invoking it with Perl, then if you find you have to expand, Perl has more elbow room.
The intention is not to 'go big', which those other languages may be better at, but to more easily 'start small'.
† IIRC, Larry Wall wanted a utility that had awk/(s)ed-like syntax for text manipulation, just 'with more'.
#!/usr/bin/perl
while (<>) {
# various processing here
# $ARGV is set to either "-" for piped input, or the current filename
# $_ is the data of the current line
}
That (<>) construct accepts data from stdin, redirection or file(s) named as arguments and iterates over the data. There's lots of things like that throughout the language.It can be very useful and they are pretty robust. I often found Perl scripts running for years and years without issues at different companies.
My main issue with Perl-scripts is that they often are not "readable" by anybody but the original creator. Which of course left the company. (not a fault of Perl itself tough)
But your millage may vary and any script can be made (un)readable.
Anyone writing Perl scripts like this should not be trusted with any programming language.
Perl scripts are no less readable than bash scripts or Awk scripts. This is because so much of Perl was written to do the same work as bash, awk, sed, and the other related Unix text processing command line programs, but all under one roof.
Don't believe me? Take a look for yourself:
Most programming languages can be obfuscated. That does not mean people write code in those programming languages like that:
Javascript: view-source:https://www.google.com/
The truth is that insulting Perl is considered stylish by some, so many people do despite knowing little to nothing about Perl and having never used it.
However, if you want Perl to be hilariously unreadable, why not write it in Latin:
https://metacpan.org/dist/Lingua-Romana-Perligata/view/lib/L...
Or Klingon:
fn print_d(t: &'static impl Display) {- Want to cut through and move loam, compost, sandy, and compacted soil? You're gonna want a rounded shovel.
- Want to break up rocky, clay soil? A pick mattock will penetrate deep, breaking up soil, shattering smaller rocks, and is used as a lever to uproot. A tiller is a faster method but disturbs the soil more.
- Want to dig a narrow, deep hole? An augur will quickly break up rocks and soil in a shaft and move them upwards.
What do you use the Perl tool for?
- Quickly and efficiently open files, read line by line, analyze text, and perform any kind of operation you can think of, with complex data structures, objects and modular code, using very few lines of code.
- Executing external commands with a shell, returning their output, and making complex yet short programs easily with arguments to the interpreter from a command line.
One day while avoiding working on something important, I spent half a day learning Perl in order to implement something related to a build tool that was being used in the important thing I was avoiding.
I was blown away. It's a really delightful language. Its big downfall is that it makes it feel good to do something "clever."
Perl is a joy to write, and a devil to read. I liked it, and wish I had started my career earlier so I could have enjoyed Perl in its heyday.
I have similar feelings about Ruby.
In fact, Perl remains remarkably robust if you stack clever tricks on top of each other.
1. You can write totally unreadable perl. It is probably the single worst language in this regard most programmers will run into. Be careful to make your code readable.
2. Keep your amount of perl small. 200-300 lines is a good bit of it.
So for quick bang it out scripts that want to parse text etc... perl is great. For writing a major application, not so much.
That stuff is just awkward and painful in Python by comparison.
... {
print $0 | "command"
}
"command" is executed once, and the pipe is kept open until closed explicitly by close("command"), at which point the next invocation will execute it again. The command string itself acts as a key for the pipe file descriptor.And of course, no mention of awk is complete without the "uniq" implementation, which beats the coreutils uniq in every way possible (by supporting arbitrary expressions as keys and not requiring sorted input):
!a[$0]++I was expecting the book, but the page itself says "This page is a placeholder for material related to the second edition of The AWK Programming Language."
It's fine if this is a placeholder page (and an awesome excuse to read talk about AWK here on HN :) ) but I want to be sure that I'm not missing the book itself.
~
One of my first big projects at my first job fresh out of college was using sed & awk to semi-automate the transformation of semi-unstructured data into a database.
IIRC I couldn't completely automate because it contained author names, from global naming conventions. (parsing names correctly is deceptively complex) They had somewhat arbitrary #'s of initials ranging from 0-3.
Again, IIRC, I could easily accommodate 0 or 1 initial (followed by \.) but trying for more would make the regex I was using too greedy and pull in part of the article abstract. These were scientific books and journals.
So I scripted a sed & awk program to detect the possibility of > 1 initials and when that occured, I'd pipe the record into nano for a quick review where I manually inserted the correct \. characters for the initials.
It was decades of back-catalogue publications for digitization so I sat there for days, listening to music on an original 1st gen iPod, waiting for my duct-taped kludge of a program to pipe one of thousands of records into a nano session every few minutes. This was on an Apple G4 workstation running OS X, where I earned my real bash scripting chops. It was an awful hack by today's standards, but at the time, accomplishing what was expected to be a 1-year long project in ~1 month, it was seen as nearly miraculous.
>I used awk until I learned Python (long ago). For me, awk was yet another example of the "worse is better" approach to things so common in unix. For example, if you make a syntax error, you might get a message like "glob: exec error," rather than an informative message. "Worse is better" is probably a good strategy in business and for getting things done, but still, mediocrity and the sense of entitlement that so often goes with carelessness, sickens me.
[0] https://news.ycombinator.com/item?id=13457265
Long live the Unix Hater's Handbook! (Unix is fine, and so are the criticisms herein. Some of these criticisms have been eclipsed by ongoing development.) https://en.wikipedia.org/wiki/The_UNIX-HATERS_Handbook
Awk's claim to fame in my world is that it's cognitive activation energy for anyone who has taken the 3-4 hours to learn the language from start to finish (and that's the awesome thing about the language - it really is about 3 hours of concentrated attention) - is essentially nil. You see a bunch of ugly not really structured text 500 MB files that you can't pull into pandas, or easily parse into python dicts? No problem - awk will tear through them for you and get the information you want in < 60 seconds, including the time you took to write your (almostl always single line) of code.
That's Awk's sweet spot.
I'm not complaining that someone banged out awk (speaking figuratively) on a Friday afternoon to do something and not have to stay after work. Excellent! My complaint is that the failure to address technical debt has negatively affected the productivity of millions, if not tens of millions, of people, often working under pressure, for DECADES.
It's benefited from extraordinarily enlightened stewardship, kept it's minimalism and strengths, and will finally get a key enhancement (UTF-8 support).
The first edition manual is probably the greatest example I've ever seen of technical writing as well.
For many python users, it’s the only language they know. Often, they see programming in python, as part of their “identity” - so they’re overly invested in it, to the detriment of other wonderful languages, like awk.
I used to code perl myself, back in the day - but I came to appreciate the simplicity of awk, and now it’s one of my favourites. I no longer code perl, as a consequence, as I believe awk to be far more elegant! I wouldn’t have done so, if I was overly invested in being a “perl programmer”.
Well, the fact is that I have to write such parsers. That's very sad, but has no chance of being fixed. So it's good to know Awk.
I think Erik Naggum had this exact criticism of Perl.
See https://hn.algolia.com/?q=The+AWK+Programming+Language for discussion on the first edition
Didn't know there was a list of `awk` implementations: https://www.gnu.org/software/gawk/manual/html_node/Other-Ver...
And: Brian Kernighan adds Unicode support to Awk https://news.ycombinator.com/item?id=32534173
> "foo","bar,baz"
$ echo '"foo","bar,baz"' | awk -v FPAT='"[^"]*"|[^,]*' '{print $1}'
"foo"
$ echo '"foo","bar,baz"' | awk -v FPAT='"[^"]*"|[^,]*' '{print $2}'
"bar,baz"
For a more robust solution, see https://stackoverflow.com/q/45420535 or use other tools like https://github.com/BurntSushi/xsvecho '"foo","bar,baz","boo"' | awk -F"\",\"" '{print $1}' "foo
echo '"foo","bar,baz","boo"' | awk -F"\",\"" '{print $2}' bar,baz
echo '"foo","bar,baz","boo"' | awk -F"\",\"" '{print $3}' boo"
Realizing that I have to strip the quotes that remain.
Edit. formatting.
EDit, again, from your link, the following is more terse and too my taste (still needs strips):
awk -v FPAT='("[^"]*")+'
I mention this specifically, here, because of the CSV point. Marcel handles CSV, e.g. "read --csv foobar.csv" reads the foobar.csv file, parses the input (getting quotes and commas correct), and yields a stream of Python tuples, splitting each line of the CSV into the elements of the output tuples.
Marcel also supports JSON input, translating JSON structures into Python equivalents. (The "What's New" section of marcel's README has more information on JSON support, which was just added.)
# This function takes a line i.e. $0, and treats it as a line of CSV, breakin
# it into individual fields, and storing them in the passed in field array. It
# returns the number of fields found, 0 if none found. It takes account of CSV
# quoting, and also commas within CSV quoted fields, but doesn't remove them
# from the parsed field.
# use in code like:
# number_of_fields = parse_csv_line($0, csv_fields)
# csv_fields[2] # get second parsed field in $0
function parse_csv_line(line, field, _field_count) {
_field_count = 0
# Treat each line as a CSV line and break it up into individual fields
while (match(line, /(\"([^\"]|\"\")+\")|([^,\"\n]+)/)) {
field[++_field_count] = substr(line, RSTART, RLENGTH)
line = substr(line, RSTART+RLENGTH+1, length(line))
}
return _field_count
}
It's not perfect but gets the job done most of the time and works across all awk implementations. mlr --icsv --otsv cat examplefile
* https://miller.readthedocs.io/en/latest/10min/My other problem is that I want to accomplish things, not learn a tool, and it generally takes me a bit longer than it should to decide to actually learn something and not just hack at it.
Is it still worth it to be "the awk guy" at work?
(based on my experience where people who could've benefited from awk for a one-liner dependably reach for sheets/excel rather than something like python or perl)
I picked up this little book from my University library once, and it was a fantastic read.
1. It is a lot faster than awk/perl/grep/sed combos
2. Way a lot readable and maintainable
3. More powerful than awk with it's string functionalities
4. Same availability as awk in OSs since last decade