Learn just a little Awk (2010)
gregable.com
gregable.com
In an attempt to fix the bug, I opened mirexpress's code. And all my confidence in my programming ability vanished when I saw its innards. I understand that the code may have been written by scientists who had no experience in programming, but I have never been so utterly _disoriented_ by bad code. Anyways, after hacking away at the mess for about 3-4 hours, I realized that this was a fool's errand and thought I'll just phone it in the next day saying I couldn't do it. I went to sleep thinking that it was already late and I'd get late for work the next day.
- 5 minutes later -
I woke up with a start, recalling this nifty tool called awk. I had last used it maybe 3 years ago, and before that only in college. But I could see how awk could do some of the things which mirexpress was claiming to do. So I fire up my computer, write an awk script - 2 lines only! TWO FUCKING LINES! And it runs like a charm - eats away at megabytes of sample data and gives me results I can show. So then like any rational person, I spent the remaining hours re-discovering awk and forgot to sleep. Pissed away the whole next day (and some part of the day after that too!) :-D
It's really fascinating that this nifty little tools invented DECADES ago are still going strong, and there's been no _evolutionary_ leap in areas where tools like awk/grep/sed excel at.
Don't write one-liners if they're not that simple. Write ten-liners! Add comments! Commit them to version control! Awk can be a nice, readable little language if you're not trying to win at code golf all the time.
Awk is a pretty terse language, and that gives you plenty of room to put in whitespace and comments and still have something short and sweet. I think that the idea of writing code to be nicely readable by your collaborators, or your future self, might not have been around in the early days of UNIX.
("But it's just a one-off thing I'm never going to do again, why should I save it to a file?", you may ask. Is what you do important? Then you'll probably have to do something like it again.)
It’s an invaluable asset and saves me tons of time on a regular basis. I’d rather search my OWN hard drive for an example of my OWN code for how to - for instance - use a CTE to recursively populate a date dimension table, than to search Google and see someone else’s code, to refresh my memory.
It’s why my Programmer’s Compendium [0] may only ever be useful to me. If you try looking at some of the more complete pages there, you might think they’re useless, but they’re, in fact, all that I need to write down.
However, mine is public, and there’s not much harm in these things being public, because perhaps it can be useful to someone else one day. I would encourage people to share knowledge by default.
So while I also think making it public is a good idea, I would also say never be under pressure to make it accessible to others!
What’s good for learning is almost never good for reference.
[0] https://qasimk.gitbooks.io/programmers-compendium/content/
Maybe. But stuff like The Perl Cookbook helped me a lot in the past.
[1] The chapter on iterators and generators is excellent, and full of useful ideas and code snippets you can reuse and build upon. I think a large part of that chapter was written by Raymond Hettinger, who also has designed and implemented much of those features in Python, IIRC. This book is written by many contributors, for the different sections and recipes.
[2] Interestingly, the Cookbook for Python 3 takes a somewhat different approach. It still has recipes, of course, but they are presented without too much discussion. The authors expect you to (and say so explicitly up front) use more of your time and thinking to figure out how they work (and to read the relevant Python and external library docs for the background information needed), rather than giving detailed explanations (not that the explanations in the previous edition are very detailed, but they are there). David Beazley and Brian Jones are the main authors of this one, IIRC.
Both models have their merits, IMO. In fact, overall, I think the approach of the Python 3 Cookbook may be better, at least for experienced programmers, because it makes you think and do more on your own (using the book as a base), from which you grow more as a programmer.
I use it to save all sorts of clever snippets on the command line. Helps me revise some nifty commands periodically and thereby grow expertise in them with time.
Furthermore, a lot of it contains stuff that I'm doing for work at any given moment. Some queries are designed to run against databases at work. I don't know what kind of hot water I could get in for letting that code loose on the internet, but I'd rather not find out.
I may consider making some sort of "best of" project and putting some stuff out there, though. It would be fun pick 5-7 concepts that I've struggled with, and spend some time writing about them and making them presentable for the public.
It's not worth it. As soon as you go over about 3 lines, you're better off switching to Ruby, Python or even Perl.
I also recall a more recent interview where when asked about his language preferences, he admitted he does very little programming anymore. When the interviewer pressed for what language he would use now (sigh), he said he would just use Python.
By designing a terse language, they made people want to do things with it that are short and sweet, preventing it from feature-creeping into a full-featured language where you'd write thousands of lines of spaghetti. Is that the idea?
Either you won’t understand it later, or you’re just wasting effort.
I wonder if the problem is twofold: 1) a lack of education compounded by 2) the rapid evolution of computer systems.
Unix is a rare beast in that not just its philosophy but even its components survive to this day and remain relevant[1]. People in unrelated fields rebuild tools that could just as well be assembled using Unix's basic components, but they're just not aware of them. And why would they look for these antiquated tools? They've been trained to reasonably expect old tools to have been replaced by newer, better, more featureful ones.
[1] AWK was created before I was born, and I'm among the more senior engineers in my 20+ team.
From what I've seen, amid all the research, failed experiments, constant fight for funding (in some cases) and "life" in general, people have much less time to learn these critical things like programming / tools so useful to them. They will do it at some point, but probably not till their life depended on it. On the other hand, if life depended on me recognizing those amino acids, we'd all be :-)
Compounded by the fact that in science once you have the result the code is probably just an artifact so there's no real reason to refactor it
Could you explain what you mean by this?
It's easier to break the mould if you're not bound by other people's software, but starts to look awfully sciency if you explore too far : D So it's useful to keep wiping the slate to not be bogged down by previous experiments
pd, processing, chuck, cm/clm, or just old-school mod-tracking are some good ways to get going
It's a shame to see so much bloated, over abstracted development these days with more emphasis on delivery than doing one thing well.
https://en.wikipedia.org/wiki/Demoscene
After seeing what the demo scene is able to pull off in very little code makes me hope that the art of fundamentals & single responsibility principals like in Unix/Linux resonates with younger developers the ability of truly understanding a problem scope before implementing a solution. Above all else have fun with it, be curious on even a low level how your program functions at an os & hardware levels. Things like what is the von Neumann bottleneck? Or even those building boolean logic gates in minecraft is inspiring to want to know "how does it work? And why?"
[0] http://people.fas.harvard.edu/~lib113/reference/unix/writing...
But that article was archived from the IBM dW site some time ago, after being there for some years. However, I wrote to IBM and got the PDF of the article, and put it on my Bitbucket account, with the C code for the utility.
This post on my blog describes the article:
https://jugad2.blogspot.com/2014/09/my-ibm-developerworks-ar...
And here is the Bitbucket project for selpg, the utility used as a case study, with the C code and the article text (as PDF):
https://bitbucket.org/vasudevram/selpg
The article and all the source files are here:
https://bitbucket.org/vasudevram/selpg/src
It may be of use to people who want to progress beyond using Unix / Linux command-line utilities (in C, but the principles and techniques can be adapted to other languages like Python, Ruby, etc.), to writing such utilities themselves, along with integrating them into shell scripts and pipelines.
https://github.com/RetroBSD/retrobsd/tree/master/src/cmd/awk
One of these days, when the mood strikes, I'll upload the matlab script I inherited from my Ph.D supervisor.
Scientific programming is, by and large, an abysmal state of affairs.
I will never forget my gut wrenching feeling when I opened up one of the main provisioning scripts (in Python) to read the first two lines:
True = 0
False = -1
The rest of it was an unmaintainable disaster that may have worked. I learned so much that summer, having been forced to learn it all myself. True = 0
False = -1
<<Could you please explain what is so egregious about those two lines for those people that may not see it? Asking for a friend.
This script is changing it to be backwards, and python is flexible enough to allow that. But I would expect this would break most python libraries, because it's such a core language feature that's being completely restructured.
[1] https://docs.python.org/3/whatsnew/3.0.html#changed-syntax
(Python has a search path: it looks first on the current module, then on the built-in __builtins__ module, which is where the built-in constants and functions reside.)
Is there a valid use case for changing the numeric values of True and False?
Python 2 has _many_ quirks like this one. Python 3 was not just "it's unicode".
No! Actually I take it back this _is_ definitely still a terrible idea and will lead to certain madness for all involved. :)
However maybe the professor needed 'true/false' to represent something to be used in his formulas so this made sense to him. But who knows.
>>> True = 0
>>> False = -1
>>> if False: print('oh no')
False is now truthy and True is falsey. Note this only works in python2.https://codesearch.debian.net/search?q=%28%3Fm%29%5E%5Cs%2AT...
For the curious, the workflow builds ML models to predict binding to unwanted proteins for new drug compounds, based on drug molecule-protein target interaction data, and is implemented with out Go-based SciPipe workflow library [0].
The rawdata is a 18GB tsv file (ExcapeDB [1]), and AWK really really shines for this.
I sometimes wish we had an SQL-like language for expressing all these computations, because we ended up doing some pretty complex join-like filtering and stuff. But from my tests with SQLite, I would not get a query answer until after like ~5 minutes with SQLite, so I gave up. With AWK, I can do a "head -n 10" on the result of the AWK operation (or on the input data file), and immediately verify that it works as expected. This seems to be one of the strongest points of the unix tool philosophy: Ability to check partial results in no time, and so iterate quicker on developing long "pipelines" of chained commands.
It also fits perfectly as a way to build scipipe workflow components, since having the component code in a separate language (AWK), rather than as inline Go, which is also possible in SciPipe, it is super-easy to send this component of to the HPC resource manager (given that we have a shared parallel filesystem): Just prepend the SLURM [2] salloc command with its parameter, before the command.
you might be interested in datajoint:
Learning the basics of awk, sed, etc is almost a requirement.
Most of my time is spent on the terminal and without these I'd shoot myself.
Haha! I've been there!
Awk is amazing and when paired with sed it exponentially increases what one can achieve on the command line.
I tend to one line most of my things in bash and tend to write a script if I have a multiple use case.
Such is the folly in bioinformatics. Most of the stuff that we use tend to be one off. :/
If you are interested in learning Awk, I highly recommend "The AWK Programming Language" by Aho, Kernighan, and Weinberger. It's about the same size as the original "The C Programming Language" and is equally well-written. Previously on HN: https://news.ycombinator.com/item?id=13451454
I'm a big fan of small utilities :) - as I sometimes say in my email sig; but more importantly, I'm a big fan of Kernighan et al, where by "et al" I mean the others from the core early Unix days, such as Dennis Ritchie, Rob Pike, Ken Thompson and many unnamed others, from whom I (and tons of others) learned about the Unix command-line (tools), the shell (scripting), and the Unix philosophy [1].
Had written this just a few weeks ago on HN, in the thread titled "Technical Writing: Learning from Kernighan", but worth repeating here in the context of this thread:
https://news.ycombinator.com/item?id=17163276
It's a list of his books. I guess many may not know of some of them - I know I didn't.
[1]:
The Unix Philosophy in One Lesson:
http://www.catb.org/esr/writings/taoup/html/ch01s07.html
Attitude Matters Too:
"If someone has already solved a problem once, don't let pride or politics suck you into solving it a second time rather than re-using."
I'd strongly argue it's overzealous. As much as I agree "reinventing the wheel" is dangerous, tempting, and can quickly spiral to yak shaving, but Unix itself, and all the good it brought, is a prime example of "solving [a problem] a second time" after Multics.
In other words, I'd restate it in Sage Speak™: "Don't do this. Except when you need to." ;P Or, just want to have fun :P
Not sure about that. I mean, I know it came after Multics and was inspired by it (due to some of the early Unix people having worked on Multics), including that the name was originally Unics (I've heard, as a word play on Multics, because it was originally written by one person or was originally a single-user OS, maybe), but I am not so sure that all the good it brought was from Multics. Likely Unix brought some new stuff too. Others who know better may be able to say more.
>In other words, I'd restate it in Sage Speak™: "Don't do this. Except when you need to." ;P Or, just want to have fun :P
Good one. A bit Zennish :) Check out one of ESR's other compilations, the Unix Koans of Master Foo, if not seen already ...
http://www.catb.org/esr/writings/unixkoans/introduction.html
1. A simple interpreter for an awk-like language called qawk. qawk is like awk except that it allows for querying by field name rather than field number. For instance, it allows doing
{ print $country, $population, $capital }
instead of the more cryptic { print $1, $3, $5 }
2. An awk program that takes another awk program (in their example, a sorting algorithm) and outputs a version of that program modified to include profiling statements and an END section that outputs the results of those profiling statements to some file; then, another awk program that reads the data in that file and inserts that data back into the original awk program, thereby approximating where the hotspots are.There's a lot more in the book besides these, but to me these are the coolest programs because they are the awk-iest, by which I mean that they loop over lines of input, split the fields of those lines, and then manipulate the fields. Some of the programs in the book don't do this; instead, they consist of a single large BEGIN block with typical for-loops, arrays, etc. Used in this way, awk is just yet another dynamic language.
Am I right that qawk was included as a program in the text? Did they ever follow up with further uses?
BEGIN { readrel("relfile") }
/./ { doquery($0) }
where- relfile is a file containing the field attributes used in various database files,
- readrel is a function that parses the relfile and stores the fields in a dictionary, and
- doquery is a function that takes a qawk query, converts it to an awk query by replacing the field names with their corresponding field numbers, and then executes the awk command.
The whole thing runs about 60 lines.
Also, Awk isn't great for making reports, which is why Perl 5 to this day has an awkward report creation system[1] that looks like some COBOL refugee instead of idiomatic perl code.
If you're of a certain age and read that and grimace as you immediately understand why this would happen, it does rather put into perspective the misery of dealing with, say, Webpack configuration.
Did email clients not handle "dot stuffing" back then? That is, if a line begins with a single dot, the client would automatically insert another dot right before it. Then, at the receiving end, the client would remove the extra dot at the beginning of the line.
That seems ridiculous; where is it substantiated?
When Wall posted Perl to comp.sources.unix for the first time, he wrote "If you have a problem that would ordinarily use sed or awk or sh, but it exceeds their capabilities or must run a little faster, and you don't want to write the silly thing in C, then perl may be for you."
Or rather, not Larry Wall, but the apparent newsgroup moderator added that text, lifting it from the Perl manual page.
Thus he was pitching it as something that performs faster than awk and sed, with a greater range of capabilities.
> News was maintained in separate files on a master machine, with lots of cross references between files. Larry's first thought was "Let's use awk." Unfortunately, awk couldn't handle opening and closing of multiple files based on information in the files. Larry didn't want to have to code a special-purpose tool. As a result, a new language was born.
So that's why it's the Practical Extraction and Reporting Language. He wanted to extract data from files and generate reports.
Why?
I could be wrong though. Awk is one of those tools like vi where you can use it for years and still be discovering new features.
Good point. The BEGIN and END patterns do work for global headers and footers, totals, etc., but not for per-page stuff. You can do it yourself with some extra awk code, but yes, you have to write it. Not difficult, though. I guess it was not designed to be a Crystal Reports-like reporting tool, with report bands and what not.
>I could be wrong though. Awk is one of those tools like vi where you can use it for years and still be discovering new features.
Agreed :) Not only new features, even new uses for existing features, because, although it is a sort of DSL or little language (but a programmable one), the area it is applicable to, pattern matching and data processing of many kinds, is vast.
As many others pointed out in this thread, the fact that the reading of input is built-in to it (whether from standard input or files given as command-line arguments), saves you a bit of boiler-plate code each time (cumulatively) you write an awk program using that feature. So does the pattern-action model, with those two defaults for missing pattern or action (match all lines, or print). And again as others have said, Perl, Ruby, etc. have that feature too (the first one).
That must be where the name comes from, right?
https://www.gnu.org/software/gawk/manual/gawk.html
It's a pretty good book which teaches you both techniques and the nitty-gritty. I recommend this.
it is available from internet archive, so i guess a legit copy.
I think Kernighan also stated somewhere that awk was (also) designed as a helper-tool to learn C.
All in all, I think it's a great first language, even if you're initially more compelled by Lisp family languages. If you have time to learn only one language, then awk is not a bad choice; It'll open many doors in the Unix/KISS-world.
I can also state that "The Awk programming language" is, among other things, an excellent introduction to computer science or "the way programmers think" in general. A remarkably well written book for general audience.
As a result, I’m WAY less likely to reinvent the wheel. I love to argue down tech proposals at work with something along the lines of “Linux already does that for you”. This sort of understanding can easily save man-years of effort on moderately complex undertakings.
As I’m writing this I again wonder if some sort of professional licensing is needed in software engineering. A bioinformatics PhD has absolutely no way of assessing potential CS consultants, and picking the right one could save so much money, time, and effort.
The fact that most OSs have a POSIX compliant version of it in the base system also make it very valuable, not just for the people crunching data but also to sysadmin/devops people.
My 99% use case is extracting particular fields from some input stream.
awk '{print $1,$5}'
That's a tremendously common problem and awk is the best tool for the job.I once started diving deeper into awk and there is an impressive amount you can do with it, but I hit a point pretty quickly where it would be cleaner to write the script in a more traditional language where my successor won't be cursing my name for writing something complex in an obscure language.
$ cat foo
foo bar baz
foo qux
$ < foo cut -d' ' -f 2-
bar baz
qux
As far as I know you have to loop through the fields, which is (haha) awkward. cut -d $delim -f $field1,$field2,...- outputting my selected fields in an order other than their order in the input lines: you can't `cut -f2,1`, for example.
- use a multi-character or regex field separator: -F'<space><star><comma><space><star>' [edit: rewritten in words to dodge HN's formatting] is a real winner when dealing with hand-written CSV.
- use a record separator other than \n: thanks, Windows, for all those times I've written ORS="\r\n" :P - you can even use a regex record separator, which not even perl supports; I don't need that often, but it's a godsend when I do.
- select based on criteria other than "number of fields from the start": I recently needed the last field from lines of varying field-count, and awk '{print $NF}' was just the thing.
But if you don't need anything like that, cut is a lot more succinct - especially if you need to split on tabs only, which it does by default, while awk defaults to splitting on spaces too.
It's a narrow tool. Other than the inherent advantage this brings in clarity/simplicity in a script, there's nothing else to it.
Consider a script that originally had now awk but just a
field=$(somecmd | cut -d: -f3)
but now has been modified to have awk elsewhere. Did the above just become less simple or clear? If so, is it worth changing it to the below for clarity? Does it matter if the awk predates the cut, instead? field=$(somecmd | awk -F: '{print $2}')
I say "no" to both, although I do recognize the argument for just using the same, consistent tool everywhere, instead.However, for ad-hoc command-line usage, I still prefer awk, because (1) the default delimiter is white space, which means I don't have to worry about if what I'm looking at might contain tabs intermingled with spaces (or how many spaces) and (2) it makes it much easier to incrementally increase the complexity of my ad-hoc operation, such as if I decide I want to take a sum or average of one of those columns, after all.
You can shell out if necessary:
cmd = "whois " $url
while ( cmd | getline ) {
# stuff
}
close(cmd)
Also: coprocesses.Can be done with tr, without writing a utility (if by "collapse" you mean what I think you do):
This file t:
$ cat t
the quick brown fox jumped
over
the lazy dog
containing many combinations of spaces, tabs and newlines (whitespace) can be changed to this file t2:$ cat t2
the
quick
brown
fox
jumped
over
the
lazy
dog
by this tr command:
tr -cs "[a-zA-Z]" "\012" < t > t2
That also makes the output more amenable to further processing, including common tasks like finding the frequencies of the words in the input, as mentioned in the "More shell, less egg" post mentioned in this post:
The Bentley-Knuth problem and solutions:
https://jugad2.blogspot.com/2012/07/the-bentley-knuth-proble...
(which I somehow managed to read as an awk invocation).
Re: python - I believe if using things like \w, \d or \s you should be Unicode safe.
Come to think of it, I seem to recall gnu tools should also have Unicode aware pattern/character classes, eg:
https://www.gnu.org/software/gawk/manual/html_node/Bracket-E...
> For example, before the POSIX standard, you had to write /[A-Za-z0-9]/ to match alphanumeric characters. If your character set had other alphabetic characters in it, this would not match them. With the POSIX character classes, you can write /[[:alnum:]]/ to match the alphabetic and numeric characters in your character set.
https://docs.python.org/3/library/re.html#re-syntax
[ed:
Apparently gnu tr is still out in the cold re:unicode;
https://www.gnu.org/software/coreutils/manual/html_node/tr-i...
> Currently tr fully supports only single-byte characters. Eventually it will support multibyte characters; when it does, the -C option will cause it to complement the set of characters, whereas -c will cause it to complement the set of values. This distinction will matter only when some values are not characters, and this is possible only in locales using multibyte encodings when the input contains encoding errors.
]
Another common usage for awk in my case is to do floating operations (*, +, eq, gt, etc) in a shell script. It avoids installing bc, which is not always present by default.
I used it for other things, like for example comparing software versions in shell scripts.
The most advanced usage of awk in my case was to do some statistic analysis (average, confidence interval, standard deviation) of network simulations.
Unix was mature even then. There is NO WAY you would have convinced me that 30 years later I'd still be using most of that, still gluing things together with the same shell tools. Or that most of the books would still have a place on the shelf (albeit sometimes in 2nd or 3rd editions), and turn out to be classics.
Ignoring big iron, we'd gone from Apple IIs to Amigas and workstations in just a few years, and hardware was changing as rapidly as the current fashion for JS frameworks does today.
Yet here we are, awk, grep, sed and the other tools still feel like the right, concise answer. So why not learn them, they might greatly surprise you by their staying power.
You need to be cautious however, gnu awk can be somewhat inconsistent at times:
try the following on various distros (and also an *BSD) for example:
awk 'BEGIN {print "9.1" >= "8"}';
awk 'BEGIN {print "9.1" > "8"}';
awk 'BEGIN {print "9.1" < "8"}';
awk 'BEGIN {print "9.1" <= "8"}';
on OpenBDS, you get no output or a syntax error, on Debian, you get the correct result for < and <= in all cases, never for > (it's interpreted as a redirection), and >= only works with newer Debian versions (sid and maybe stretch).
I know that the fix is simple: just put some parenthesis (awk 'BEGIN {print ("9.1" <= "8")}'), but I was quite amused to discover that a few days ago.
And if you really want that in a repo, you can stick that in ansible trivially.
Sometimes instead I just wrap a line of awk in %x() in Ruby to save a few chunks of code.
Most programmers don't understand why I do this though. :(
Hundred of lines of python by two awk lines ? Must be very bad python programmers.
Awk is a compiled language. Your Awk script is compiled once and applied to every line of your file at C-like speeds. It is way faster than Python.
If you learn to use Awk well, you will start doing things with data that you wouldn't have had the patience to do in an interpreted language.
I say this with no disrespect for Python, it's just you should know the tradeoffs of the tools you use. It does not matter if the individual operations in Python are implemented in C. It will still be much slower at looping over lines of a file than a compiled language designed decades ago for that exact purpose.
I have a lot of Python file in one dir:
$ find . -iname "*.py" | wc -l
10429
Finding them and cating them all takes about 0.3 secs: $ time find . -iname "*.py" | xargs cat {} > /tmp/cat.out
...
real 0m0.344s
user 0m0.140s
sys 0m0.175s
So to have something simple that takes a bit of time, I tried to get all lines starting with "print", and output the first thing after that.I'm really bad at awk, so I don't know if there is a better way. I went for the most obvious thing for me:
$ time find . -iname "*.py" | xargs cat {} | awk '/^print/{print($2)}' > /tmp/awk.out
...
real 0m1.111s
user 0m1.165s
sys 0m0.368s
Now, with Python, it's definitely not as easy to type. You have to get a script like. awk wins the expressivity metrics for this use case: import sys
for x in sys.stdin:
if x.startswith('print'):
try:
print(x.split()[1])
except IndexError:
print('') # to match awk behavior
But as for performance, I don't get the huge boost in perfs you are talking about: $ time find . -iname "*.py" | xargs cat {} | python /tmp/test.py > /tmp/python.out
...
real 0m0.762s
user 0m0.862s
sys 0m0.347s
I do get the same output though: $ cmp /tmp/python.out /tmp/awk.out && echo "yes"
yesHere's what I observed: I had a simple text-processing tool I needed that I call "countmerge", which just merges adjacent lines with the same key and adds up their corresponding values. I needed to run it on a lot of large files.
I first wrote it in Python, where it was a significant bottleneck compared to the steps that came before it (split, sort, uniq -c). Eventually I rewrote it in Rust [1], and it was at least 5 times faster, at the expense of a fair amount more low-level code. But then rewriting it in awk [2] turned out to be as fast as what I wrote in Rust, possibly inconclusively faster.
[1] https://github.com/rspeer/countmerge
[2] https://gist.github.com/rspeer/60c87dca1ab550326f8bd6d086452...
>>> with open('data.txt', 'w') as f:
... for l in string.ascii_uppercase:
... for x in range(0, random.randint(1, 100000)):
... f.write('Key {}\t{}\n'.format(l, random.randint(0, 100)))
With this script: import sys
old_key = total = 0
for line in sys.stdin:
key, value = line.split('\t')
if old_key != key:
old_key = key
total = 0
print(key, value, end="")
total += int(value)
I get: $ <data.txt time python3 test.py
Key A 2
Key B 87
Key C 58
Key D 64
Key E 29
Key F 25
Key G 2
Key H 74
Key I 17
Key J 37
Key K 97
Key L 77
Key M 19
Key N 74
Key O 33
Key P 61
Key Q 67
Key R 23
Key S 4
Key T 70
Key U 25
Key V 15
Key W 35
Key X 17
Key Y 31
Key Z 18
1.03user 0.01system 0:01.05elapsed 99%CPU (0avgtext+0avgdata 9564maxresident)k
0inputs+0outputs (0major+1100minor)pagefaults 0swaps
But I can't manage to get the awk version working. It only prints one line on Ubuntu 16.04: $ <data.txt awk -f ./countmerge.awk
Key 0
So I can't check it. awk '/^func Test/{p=1}; p; /^}/{p=0}'
EXPLANATIONFirst of all, this is 3 separate "commands", separated by ';'. In order:
/^func Test/ {p=1}
— if line matches regexp '^func Test' (i.e., starts with "func Test"), then set variable p to 1 (a.k.a. "True"). p
— equivalent to any of: p { print }
p { print $0 }
{ if (p) { print $0 } }
meaning: if variable p is true-ish (in case of this script, if p==1), then print current line (if action is not specified after a condition, then it's by default {print}). /^}/ {p=0}
— you may have guessed already now: stop printing after encountering end of function (line starting with '}'). $ sed -n '/^func Test/,/^}/p' file
Which means: match 'func Test' at the beginning of a line and print (p command at the end) until you find a line beginnning with '}'.This is because the print command (p) accepts 2 addresses to delimit a range. I used regexes as addresses but, for instance, can use line numbers also:
$ sed -n '10,20p' file
This prints file's line from 10 to 20.By default, sed prints the pattern space (modifications to each line) at the end of the script. The -n I used is to avoid that.
That said, there's one extra advantage with the awk script, that by rearranging the expressions appropriately, I can choose whether to include or exclude the the first and last line in the output (i.e. 'FIRST; p; LAST' vs. 'FIRST; LAST; p' or 'p; FIRST; LAST', etc.) Is this also possible with sed? :)
Plenty of people that would rather drop to Perl or Python once awk gets too cumbersome for the command line.
Still, I'm not certain if the original comment is complaining about the tendency toward one-liners by actually writing less code (something the article alludes to, "only the inner part of the for loop") or merely cramming what would otherwise be readable as a multi-line, well-indented program into a single line.
I tried answering a few awk questions being verbose and explaining the steps. But I noticed that most are not asking to learn but to get a quick fix. Soon after that came the second realization: several of the answers did not understand what they where doing.
My conclusion: We are living in a cut'n'paste world were actual understanding is less valued.
Naturally code reuse is nice. But I doubt that many users are able to construct their magic awk snippets from the ground up. Hence the prolific terse magic snippets.
Interesting... to me the value (er, a value) is that you can write
/PATTERN/ { print $3 }
instead of import re, sys
for line in sys.stdin:
if re.search('PATTERN', line):
print(line.split()[2])
And another value is that it's more likely to be preinstalled than Python.I really like the concept of Wolfram Mathematica Notebooks as you can essentially embed everything into a reproducible analysis (what they call a computational essay). The free and open source equivalent is Jupyter notebooks.
awk '!_[$0]++'Don't know why it gives me joy to write those things even when I know it would be far better to write the same in three nested for loops (which are less error prone, more maintainble and gives more flexibility like easy to "break").
Secondly and this is a personal preference but the code looks more readable when the for statements are on different lines telling you which array is being iterated compared to the nested functions where the output of one function is piped as input to another.
Just observe the kind of kindergarten antics some in the _sec world will get up to...
I like the baked in logic for dealing with record/field formatted data. One thing that seems to hang people up is the lack of a concatenation operator. You just put things next to each other. It looks a little weird. Having associative arrays is a nice plus-up from shell programming.
awk '{if ($(NF-2) == "200") {print $0}}' logs.txt
as awk '$(NF-2) == "200" { print $0 }' logs.txt
IMO, this style scans easier for most of the grep-ish use cases.Using "200" in quotes makes sense if you're looking for an exact string match; this will fail if the datum is 0200, or 200.0.
Ok, time for someone that doesn't regularly use awk to say something. I understand that awk is great. The language being terse is nice if you use it regularly, but otherwise it's very easy to forget, and it's strange to look at it the first time. 90% of the programmers probably don't regularly need the kind of functionality awk provides, and the few times they do, they can create a simple script in a language they know better.
Awk can be a super useful tool, but I think it's reasonable that most programmers don't use it, and the language is not designed to be used by everyone.
In my opinion, a more universal alternative to this would be something like a web application that allows visual programming and translates the operations to awk, and possibly other languages too (the visual language could encode only a subset of awk, only common operations). You would learn awk yourself if you use it regularly, and you would know that you can simply use an intuitive interface to solve those formatting problems you need to handle from time to time otherwise.
(The twitter thread that follows also has some great one-liners!)
https://gist.github.com/jaysoffian/e41ca479d70e60efe59fded93...
Backstory: circa 2000, at an early cloud company that no longer exists, we used Solaris boxes and provisioned them using JumpStart. They booted to a minimal state (we called them "embros") using DHCP. Later, our datacenter ops folks would have to switch their networking from DHCP to static as part of moving them to another network where a DHCP server didn't exist. So I wrote them this menu driven program they could run on console to help with the static configuration. (There were no high-level languages on the box save for shell, sed and awk. Even then, I apparently I had to use /usr/xpg4/bin/awk.)
For some tasks it's amazingly effective, and it's usually installed even if Python isn't.
As a shell, the tab completion, parameter completion, long names, makes it easier to discover, easier to understand, and easier to remember.
Then it's a high level language too so if there's something that needs scripting, it doesn't mean a complete change from shell to Python/Ruby/Perl, it stays PowerShell.
Then it's a .Net language too, so if there's something getting a bit big for it, it doesn't mean a complete change to C#/Java instead, it means a small change to PowerShell with .Net methods, then maybe PowerShell with a C# core (like Python with a C module, but still much easier to create).
Of course there's a bit of XKCD "fix having too many things by adding another thing" going on.
But, the fact that it covers the common shell tools with all their different syntaxes reasonably well, and it can be tuned to approach the speed of C# as well, makes it useful for a whole lot of situations, despite having a pretty huge syntax and list of warts, it still seems to come out well.
For quick filtering it's really great, e.g. you can parse simple (pretty-printed) XML tags via `-F[<>] '{print $2}'` (for a quick glance on the data - of course not a good idea in production).
To my surprise, it was powerful enough to write a somewhat limited implementation of Snake: https://github.com/johshoff/snawk
awk 'BEGIN { ORS=" "} { print $0 }' $2 | awk --re-interval -v pat="$1" '
{
cut_content = substr($0, 1, match($0, "==== Refs"))
orig_content = substr($0, 1, match($0, "==== Refs"))
idx = 0
regex = "(^|[^a-zA-Z]{1})" pat "([^a-zA-Z]{1}|$)"
while (match(cut_content, regex)) {
if (RLENGTH > 0) {
prot = substr(cut_content, RSTART-1, RLENGTH+1)
gsub(/[:punct: ]$/, "", prot)
gsub(/^[:punct: ]/, "", prot)
print prot, "\t", substr(orig_content, idx+RSTART-400, RLENGTH+800)
idx += (RSTART + RLENGTH -1)
cut_content = substr(cut_content, RSTART+RLENGTH)
}
}
}'I tried writing a script that would extract entries from logs (e.g. if every entry started "====CRITICAL ERROR LOG 2018/06/15====", I wanted it to start there and grab every line below it until it hits the next log entry starting with the "=====" header), but I gave up. Your script might do something similar, if I can figure out what it's looking for.
Another option might be to use multi-line records (https://www.gnu.org/software/gawk/manual/html_node/Multiple-...). First pre-process the file to add an empty line before each header, then use blank lines as the record separator and newlines for the field separator.
/^=====/ { inentry = !inentry; next }
inentry
Possibly without 'next' if you want the marker lines too.Intent: Given a folder with multiple text files of academic papers from PubMed, remove the contents after the '====Refs' header. Then in each file find occurrences of any term in a keywords file and extract it along with 400 chars before and after it. Print this output as a tsv.
input: $1) A text file containing keywords of interest that need to be searched in the files separated by `|`.
$2) File containing paragraphs of lots of text with headers starting with '===='. There isn't a formatting requirement really. Just that, in my case the files happened to be so.
# Note: Awk string commands have indexes that start with 1 (not 0).
# The first Awk command just removes the newlines in file $2 and replaces them
# with space. `pat` contains the contents of the file $1: 'keyword1|kewword2|...'.
# --re-interval allows use of, well, regex intervals.
awk 'BEGIN { ORS=" "} { print $0 }' $2 | awk --re-interval -v pat="$1" '
{
# both variables are initialized with full-content of text file
# with the part after '==== Refs' removed.
cut_content = substr($0, 1, match($0, "==== Refs"))
orig_content = substr($0, 1, match($0, "==== Refs"))
# index (could be better named as offset, but life is too short to refactor).
idx = 0
# match the keywords, even if it appears between non-alphabet chars.
# e.g -keyword1. is a match.
regex = "(^|[^a-zA-Z]{1})" pat "([^a-zA-Z]{1}|$)"
# loop until matches are found.
while (match(cut_content, regex)) {
# RLENGTH is the length of the matched keyword.
if (RLENGTH > 0) {
# set the matched keyword to the variable `prot`.
# RSTART is the character position of the matched keyword in the file.
# This is the first occurrence of it in the file when read from the start.
prot = substr(cut_content, RSTART-1, RLENGTH+1)
# clean-up punctuation and/or space from the beginning and end of `prot`.
gsub(/[:punct: ]$/, "", prot)
gsub(/^[:punct: ]/, "", prot)
# print the `prot` and 400 chars before and after it using idx
# as an offset. Separate the results by tab char.
print prot, "\t", substr(orig_content, idx+RSTART-400, RLENGTH+800)
# update offset to character position right after the current matched keyword
idx += (RSTART + RLENGTH -1)
# cut-out all the text up until the position of the
# last matched keyword (including it) and on the remaining text
# repeat the above steps in the loop. This is so that the search
# continues throughout the file and doesn't keep finding the same
# keyword over and over again. See `while` condition.
cut_content = substr(cut_content, RSTART+RLENGTH)
}
}
}'
Edit: Formatting. HN doesn't wrap text in code-block.My big new power move for awk is -F. You can use any regex as a field separator! Mind blown!
perl -lane 'print "$F[1] $F[2]"'
> awk '$4>3 { x+=$3; print $2, $3, x}' perl -lane '$F[3] > 3 && {$x += $F[2], print "$F[1] $F[2] $x"}'
Don't know about perl6, though.Edit: Better syntax per comments.
Scalar value @F[0] better written as $F[0] at -e line 1.
Scalar value @F[1] better written as $F[1] at -e line 1.
In Perl 5 the array elements of @a are $a[0] $a[1] etc. However @F[0] works too because a "slice" is returned.
Want to throw a few examples my way?
Disclaimer: never learned perl
awk '
BEGIN { print "0" }
$4 > 3 { print $3"+p" }
' | dc
[0] https://www.gnu.org/software/bc/manual/dc-1.05/html_mono/dc....Sure, the gnu version is more powerful, but the plan9/original awk implementation handles most cases I run into, and I can now look up the gnu extras when I need them.
df -m | awk '{p+=$3}; END {print p}'
I bought a book titled "Awk One-Liners Explained" on a whim a while back. And to this day, I consider it to be umong the most "useful" $6 I've ever spent in terms of productivity.Awk has great abilities, but when I want to use it I always have to spend a lot of time reading.
I've always wanted to come up with something in Python that combined the awesomeness of awk with my ability to pick up and hack something together in Python, but I've never been able to come up with the right semantics.
One thing I think could be really useful is the regex lines range, where you give it two regexes and that block of code gets executed on every line in that range.
It can be really helpful in mundane data/logs analysis, for example.
If I have to look up something 2-3 times in a month it's usually committed to memory. Any less and it's probably better forgotten anyway, haha.
#!/usr/bin/python3
import re, sys, json, other, useful, packages
def printf(s):
print(s.format(**globals()))
for NR, line in enumerate(sys.stdin):
match = eval(sys.argv[1])
if match:
eval(sys.argv[2])
Then you can use this in command line as somecommand | pyawk.py "re.search(r'^(pattern)\s(otherpattern)', line)" "printf('{NR}) {match[1]}')"
I deliberately didn't pass pattern straight into re.search but rather provides the option to do other forms of parsing/matching in that argument, for example json parsing. I just spent 2 minutes writing this script so it's obviously not perfect and you probably won't be able to use it for json out of the box but give it a few more iterations and it will most likely work. $ echo -e $'line one\nline two' | ruby -ne 'puts $_.split(/\W+/)[1]'
one
tworuby inherited a lot of structure and ideas from perl which it inherited from awk and that thought makes me irrationally happy
As a big fan and advocate of a low level language (C) and a high level language (Python), as a combo, being a very powerful paradigm, I still can't believe some of the things the, shall we call it, shell piping workflow, let's you accomplish in one line and in a matter of minutes.
Here's what you do.
Say you have a text log file. It's "semi structured" (like most log files) in that you can extract quite a bit of information using things like grep/sed/awk, but still not fully parse it (without a very complex parser; we don't want to get into that). What's the first thing you do with the log file? You view it:
view logfile
(view is just opening the file in vim in read-only mode, if you don't have it, you can use "vim -R"). You inspect the file and then quit. Try this instead: cat logfile | view -
(don't forget the hyphen at the end). Again, inspect and quit (e.g., using :q). You might say, what a round-about/inefficient way to view a logfile. But. Now you can do something in the middle: cat logfile | do something | view -
inspect and quit. Or more than one thing cat logfile | do something | do something else | view -
and you keep adding the piped commands until you're satisfied with your output. Once you're satisfied, you can dump the output into another text file as a "report" by replacing the last "| view -" with "> reportfile". cat logfile | many | processing | commands | later > reportfile
and you can do this multiple times to generate multiple report files of various kinds. And you can concatente some of those report files cat file1 file2
or put some columns of some of the report side by side paste -sd' ' <(cat file1 | pick column) <(cat file2 | pick column)
The possibilities are endless.To give one example from my shell history. Some times I have a dozens of pdf files open in my linux desktop (using evince pdf viewer) and I wanna restart the computer but don't want to lose track of which pdf files were open. There may be automated ways of doing this but let's say there aren't any. I start with ps ax:
ps ax | view -
A long log. I want to pick only lines that have evince: ps ax | grep evince | view -
Now I included the 'grep evince' line too, which I don't want: ps ax | grep evince | grep -v grep | view -
Good. But I don't care about all the columns of ps log, except for the last one (the one that shows the full path of the file). That's column 6: ps ax | grep evince | grep -v grep | awk '{ print $6 }' | view -
Looks good. Generate a report file from this: ps ax | grep evince | grep -v grep | awk '{ print $6 }' > openpdfs_YYYYMMDD.txt
Then I close all my pdfs, and reboot. And I don't know grep, awk except for very basic things like what I already did. But I know similar tidbits about many other commands. If I have to perform regex search only I use grep, but for search and replace I use sed. Sometimes cut is more handy for column selection than awk. If I don't know the command but I know what I need to do, I just google it, and 99% of the time, I can find a stackoverflow post where a linux based one-liner is mentioned which is pretty much a drop-in for my piping workflow.Finally, if I want to traverse through each line of the entry and process one by one, I use the while loop. For example, if instead of dumping the file paths of the pdf I wanted to format it a little bit using directory name and file name, I'll do this:
ps ax | grep evince | grep -v grep | akw '{ print $6 }' | while read -r dfnam; do echo 'Directory:' $(dirname ${dfnam}) 'File:' $(basename ${dfnam}); done | view -
And again, view or dump.Try this, and you would realize that the possibilities are endless.
Very useful stuff, and as you say, the possibilities are endless.
I had done something sort of analogous, an experimental Python tool called pipe_controller, which is not about piping commands to each other in the traditional Unix sense; rather it is about "piping" the output of one function to a second one, and the output of the second to a third one, and so on, as many as one needs, under the control of a for loop. The net effect is like normal composition of function calls, like f(g(h(x))), but some other interesting effects can be achieved by doing it with a for loop, and changing some things at run time:
After first creating pipe_controller (which is simple, really), I played around with using it in a few different ways, and found that it can be used for at least a couple of interesting things:
- running a "pipe" (of those functions) incrementally (something like the technique you showed), and saving / viewing the output of each intermediate stage;
- swapping components of the "pipe" at run time, under program control, which again can lead to some interesting use cases.
I blogged a small series of posts about pipe_controller and such uses of it. Here are the two last or so posts, and the previous posts can be reached by following links in those posts:
Swapping pipe components at runtime with pipe_controller:
https://jugad2.blogspot.com/2012/10/swapping-pipe-components...
Using PipeController to run a pipe incrementally:
https://jugad2.blogspot.com/2012/09/using-pipecontroller-to-...
The pipe_controller code is here:
I have thought of this myself too, both in terms of python and in terms of scheme. For example, for your python function composition example f(g(h(x))), shell piping would look like this:
cat x | h | g | f | view -
while in scheme it would be like python, just a little different:
(f (g (h x)))
Essentially the point is, by using existing syntax and language facilities, can we mimic the seamless shell piping workflow (e.g., with no annoying nesting).
I think one issue that python or scheme will have is that, when something goes wrong, debugging the problem would still be very annoying in python/scheme, whereas in shell piping, you simply remove some tail-end processing commands to view and earlier output, then fix your issue, then reintroduce the tail-end commands that you removed, possibility with some modifications. So "view -" acts as a debugger, not just for visual inspection of your report.
Anyway this is an interesting area. Keep up the good work. I would just like to mention that this is related, in fact it is, part of a much larger paradigm of programming called dataflow programming, sometimes called stream processing. And also has connection with reactive programming. (which if you lookup in wikipedia [1] falls into a larger paradigm called declarative programming).
Thanks for the encouragement and links. Will check them out.
Generate PDF from a Python-controlled Unix pipeline:
https://jugad2.blogspot.com/2016/01/generate-pdf-from-python...
Code for recent post about PDF from a Python pipeline:
https://jugad2.blogspot.com/2016/01/code-for-recent-post-abo...
> If you are the type of person interested in Awk, you are probably the type of person I'd like to see working with me at Google. If you send me your resume (____@gmail.com)
Awk is a great tool. Glad to see this.
The task was very simple in Python, but would be absolutely trivial and fun in AWK if it a) had first class unicode support, including conversion b) had a support for quoted fields. I think it needs a new variable, like QC (Quote Character) of FD (Field Delimiter).
I would put money on a kickstarter of another crowdfunding initiative to modernize AWK. I don't mean by slapping a Python or other programming language on it, but by fishing long-standing issues with it. I think a for loop like in Python, Rust and VimL would be a better fit to an otherwise simple language (only C-style fors are available in BEGIN/END).
I miss in awk mainly a structural regular expressions mode. With it I would not miss lack of structures or that functions like gsub do not return a string instead of assigning it to a variable.
This stands out like a sore thumb in a language famous for its brevity and frictionless syntax. There should be a better way.
https://www.gnu.org/software/gawk/manual/html_node/Scanning-...