Why Learn Awk? (2016)
blog.jpalardy.com
blog.jpalardy.com
I use awk because I like to visually refine my output incrementally. By combining awk with multiple other basic unix commands and pipes, I can get the data that I want out of the data I have. I'm not writing unit tests or perfect code, I'm using rough tools to do a quick one-off job.
For instance, "mail server x is getting '81126 delayed delivery' from google messages in the logs, find out who is sending those messages".
# get all the lines with the 81126 message. Get the queue IDs, exclude duplicates, save them in a file.
cat maillog.txt | grep 81126 | awk '{print $6}' | sort | uniq | cut -d':' -f1 > queue-ids.txt
# Grep for entries in that file, get the from addresses, exclude duplicates.
cat maillog.txt | grep -F -f queue-ids.txt | grep 'from=<' | awk '{print $7}' | cut -d'<' -f2 | cut -d'>' -f1 | sort | uniq
Each of those 2 one-liners was built up pipe-by-pipe, looking at the output, finding what I needed. It's not pretty, it's not elegant, but it works. I'm sure there's a million ways that a thousand different languages could do this more elegantly, but it's what I know, and it works for me.
I completely disagree.
... | grep foo | awk ‘{print $6}’ | ...
becomes
... | awk ‘/foo/{print $6}’ | ...
If you start working this into your awk habits you’ll find delightful little edge cases that you can handle with other expressions before the block (you can, for example, match specific fields).
awk FS=: '{print $1}' instead of cut -d: -f1
awk FS="<" '{print $2}' instead of cut -d'<' -f2
awk FS=">" '{print $1}' instead of cut -d'>' -f1echo test,123 | awk -F, '{print $1}'
awk 'BEGIN {FS=":"};{print $1}'
One benefit of the FS variable over -F, at least in original awk, is that by using FS the delimiter can be more than one character. I guess that's why I remember FS before I remember -F. More flexible. $ echo 'Sample123string42with777numbers' | awk -F'[0-9]+' '{print $2}'
string awk -v FS="\t"EXAMPLES Print and sort the login names of all users:
BEGIN { FS = ":" }
{ print $1 | "sort" }
The above is from the GAWK manpage. FWIW, the first example under EXAMPLES uses FS not -F.There is nothing wrong with using FS instead of -F.
FS is not used on the command line and doing so is asking for trouble. FS is a built-in variable and as such is treated specially.
cat > 1.awk << eof
{ print $ARGC }
eof
echo|nawk -f 1.awk FS=":"
echo|gawk -f 1.awk FS=":"
echo|nawk -f 1.awk -v FS=":"
echo|gawk -f 1.awk -v FS=":" echo "" | awk '{print Bla;}' Bla="Hello."In awk, I couldn't find how to do this. I tried /\bfoo\b/ and /\<foo\>/ but neither worked. I don't know why and don't care enough which brings me to my major awk irritation ...
It doesn't use extended or perl REs, which makes it quite different to ruby, perl, python, java. Now, according to the man page it does; at least on OSX (man re_format) but as mentioned it didn't work for me.
Details
$ echo fish | awk '/\bfish\b/'
gets nothing, vs $ echo fish | perl -ne '/\bfish\b/ && print'
fishawk(1) does not support word-boundary metacharacters https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=171725
$ printf "fishstick\nfish\ngoldfish\n" | awk '/\<fish\>/'
fishGNU awk also supports \y which is same as \b as well as \B for opposite (same as GNU grep/sed)
Intererstingly, there's a difference between the three types of word anchors:
$ # \b matches both start and end of word boundaries
$ # 1st and 3rd line have space as second character
$ echo 'I have 12, he has 2!' | grep -o '\b..\b'
I
12
,
he
2
$ # \< and \> strictly match only start and end word boundaries respectively
$ echo 'I have 12, he has 2!' | grep -o '\<..\>'
12
he
$ # -w ensures there are no word characters around the matching text
$ # same as: grep -oP '(?<!\w)..(?!\w)'
$ echo 'I have 12, he has 2!' | grep -ow '..'
12
he
2!There's no point in awk if perl etc are ubiquitous and more consistent.
As the article points out, other languages will have a lot more ceremony around opening and closing the file, breaking the input into fields, initializing variables, etc.
A little bit of knowledge here goes a LONG way.
This is something I need to bring up to my coworkers, I should write some sort of basic guide to unix tools for them.
* Projection (Π): awk and cut for simple cases
* Selection (σ): grep for simple cases, otherwise sed & awk
* Rename (ρ): sed
* Set operators: join, comm...
[1] https://lagunita.stanford.edu/courses/DB/RA/SelfPaced/course...
Disclaimer: I was one of the authors
[1] https://www.thomasrebele.org/publications/2018_report_bashlo... [2] https://www.thomasrebele.org/projects/bashlog/datalog
That's like 90% of my use of awk right there. I don't know of any easier equivalent of "awk '{ print $2 }'" for what it does.
Posted partially so the Internet Correction Squad can froth at the mouth and set me straight, because I'd love to be showed to be wrong here.
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/c...
But, just as we can't say "awk offers the -b flag to process characters as bytes", we can't really say that cut offers any extensions not defined in the standard.
An implementation could, sure. I'd prefer that it didn't, writing conformant shell scripts is hard enough.
https://pubs.opengroup.org/onlinepubs/9699919799/utilities/c...
Tools should improve, and standards should eventually catch up & pick up the good parts.
People who need to work with legacy systems can just opt not to use those extensions, but one day they too will benefit. Others benefit immediately.
When I'm looking to do something more complicated, I'd rather look to tools from outside the POSIX standard. The good ones give me less fuss, because they can typically be installed on any POSIX-compliant OS, which is NBD, whereas extensions to POSIX tools tend to create actual portability hassles.
Select characters 3-6 (inclusive) in the second field:
$ echo Testing LengthyString | awk '{print substr($2,3,4)}'
> ngth
If you want to select columns from the entire line, then:
$ echo Testing LengthyString | awk '{print substr($0,3,4)}'
> stin
Is that what you meant?
Or you can use loops, e.g.,
echo $(seq 100) | awk '{ for (i = 2; i <= 7; i++) { print $i; }; }' cut -d, -f10-30
(selects from field 10 to 30)Not saying this can't be written in awk with more code, but we were talking about field selection ergonomics.
It doesn't detract from the point at hand - which is perfectly valid - but it's worth noting that there's a confusion here with regards the terminology: "fields" vs "columns". I thought they were referring to "columns of characters" whereas the added explanations[0] are about "columns of fields". That makes a difference.
But as I say, yes, I agree that to select a range of columns of fields, especially several fields from each line, is definitely better with cut.
Repo has a few other shortcuts, too:
Obviously, other people have different experiences which is why I quality it so. (I only narrowly missed it at the beginning, but I started in webdev, and we never quite took direct feeds from the mainframes.) But I don't think it's too crazy to say UNIX shell, to the extent it can be pinned down at all, works most natively with variable-length line-by-line content.
SAP comes to mind. I think it does support various different formats, but for reason or another fixed-width seemed to be some kind of default value (that's what I usually got when I asked for SAP feed at least, but that was years ago).
Admittedly I have encountered fixed-width text formats in the wild. But the last such occasion was about 15 years ago. (It was for interacting with a credit card processor to issue reward cards.)
and my favourite....... - source code.
My first intro to awk was using it to process COBOL code to find redundant copy libs, consolidate code, and generally cleanup code (very very crude linting). And it was brilliant. Fast, logical, readable, reliable - was everything i needed.
It is also an eminently suitable language for teaching programming because it introduces the basic concept of SETUP - LOOP - END . which is exactly the same as one will find in most business systems, you find it in arduino sketches, hell you even find it in a browser which is basically just a whole universe of stuff sitting atop a very fast loop that looks for events.
AWK fan for sure - my heirachy of languages these days would be cmd line where there is specific command doing all i need, AWK for anything needing to slice and dice records that dont contain blobs, python for complete programs, and python+nuitka or lazarus or C# when need speed and/or cross platform.
Does `cut -f2` not work? My complaint with cut is that you can't reorder columns (e.g. `cut -f3,2` )
Awk is really great for general text munging, more than just column extraction, highly recommend picking up some of the basics
Edit to agree with the commenters below me: If the file isn't delimited with a single character, then cut alone won't cut it. You need to use awk or preprocess with sed in that case. Sorry, didn't realize that's what the parent comment might be getting at.
$ echo ' 1 2 3' | cut -f2
1 2 3
$ echo ' 1 2 3' | cut -f2 -d' '
$ echo ' 1 2 3' | awk '{print $2}'
2
"-f [...] Output fields are separated by a single occurrence of the field delimiter character." FS = "[0-9]"I don't think there is, because cut separates fields strictly on one instance of the delimiter. Which sometimes works out, but usually doesn't.
Most of the time, you have to process the input through sed or tr in order to make it suitable for cut.
The most frustrating and asinine part of cut is its behaviour when it has a single field: it keeps printing the input as-is instead of just going off and selecting nothing, or printing a warning, or anything which would bloody well hint let alone tell you something's wrong.
Just try it: `ls -l | cut -f 1` and `ls -l | cut -f 13,25-67` show exactly the same thing, which is `ls -l`.
cut is a personal hell of mine, every time I try to use it I waste my time and end up frustrated. And now I'm realising that really cut is the one utility which should get rewritten with a working UI. exa and fd and their friends are cool, but I'd guess none of them has wasted as much time as cut.
or
echo ' 1 2 3' | tr -s ' ' '\t' | cut -b 2- | cut -f2
In my experience, most column based output uses variable number of spaces for alignment purposes. Tabs can work for alignment, but they break when you need more than 8 spaces for alignment.
Most utilities don't use a tab character as separator, and that's what cut operates to by default. Can't cut on whitespace in general, which is what's actually useful, and what awk does.
Only way to get cut to work is to add a tr inbetween, which is a waste of time when awk just does the right thing out of the box.
Agree in general. Only exception I'd make to this is when you're selecting a range of columns, as someone else mentioned elsewhere in the thread. I typically find (for example) `| sed -e 's/ \+/\t/g' | cut -f 1,3-10,14-17` to be both easier to type and easier to debug than typing out all the columns explicitly in an awk statement.
| tr -s ' ' '\t'As others have pointed out, no. It should! (Said the guy sitting comfortably in front of his supercomputer cluster in 2020. No, I don't do HPC or anything; everything's a supercomputer by the time that cut was written's standards.) But it doesn't. Going out on a limb, it's just too old. Cut comes from a world of fixed-length fields. Arguably it's not really a "unix" tool in that sense.
"highly recommend picking up some of the basics"
I have, that's the other 10%. I've done non-trivial things with it... well... non-trivial by "what I've typed on the shell" standards, not non-trivial by "this is a program" standards.
I'm not sure if you refer spefically to cut, but Perl has something similar and approximaly terse:
> echo 'a b c' | perl -lane 'print $F[1]'
Also, Perl can slice arrays, which is something that I really miss in Awk.
Every commercial embedded Linux product that I've seen uses Busybox (or maybe Toybox) to provide the coreutils. If awk is available on a system like that, it's almost certainly Busybox awk.
And Busybox awk is fine for a lot of things. But it's definitely different than GNU awk, and it's not 100% compatible in all cases.
I can't edit my OP but I'm already downvoted so that will suffice.
GPLv3 (which was written in 2007) has much tougher restrictions. It's the license for most of the GNU packages now, and GPLv3 packages are impractical to include in any firmware that also comes with secret sauce. So most of us in the embedded space have ditched the GNU tools in our production firmware (even if they're still used to _compile_ that firmware).
It should be noted that GPLv2 actually had a (much weaker) version of the same restriction:
> For an executable work, complete source code means all the source code for all modules it contains, plus any associated interface definition files, plus the scripts used to control compilation and installation of the executable. [emphasis added]
(Scripts in this context doesn't mean "shell scripts" necessarily, but more like instructions -- like the script of a play.)
So it's not really a surprise that (when the above clause was found to not solve the problem of unmodifiable GPL code) the restrictions were expanded. The GPLv3 also has a bunch of other improvements (such as better ways of resolving accidental infringement, and a patents clause to make it compatible with Apache-2.0).
The reason why I said it's impractical to include GPLv3 code in a system that also has secret sauce (maybe a special control loop or some audio plugins) is more about sauce protection.
If somebody has access to replace GPLv3 components with their own versions, then they effectively have the ability to read from the filesystem, and to control what's in it (at least partially).
So if I had parts that I wanted to keep secret and/or unmodifiable (maybe downloaded items from an app store), I'd have to find some way to separate out anything that's GPLv3 (and also probably constrain the GPLv3 binaries with cgroups to protect against the binaries being replaced with something nefarious). Or I'd have to avoid GPLv3 code in my product. Not because it requires me to release non-GPL code, but more because it requires me to provide write access to the filesystem.
And I guess that maybe GPLv3 is working as intended there. Not my place to judge if the license restrictions are good or bad. But it does mean that GPLv3 code can't easily be shipped on products that also have files which the developer wants to keep as a trade secret (or files that are pirateable). With the end result that most GNU packages become off-limits to a lot of embedded systems developers.
The point of the license condition is that once a device has been sold the new owner should have as much control as the developer to change that specific device. If no one can update it then the condition is fulfilled.
never heard that before
I wouldn't call it bloat, but yes it is much bigger. At the time you had C (really fast, but cumbersome) and Awk/Bash (good prototyping tools, but not good for large codebases). Perl was the perfect answer to something that is fairly fast, relatively easy to develop in, and easier to write full-sized codebases
Perl is the language, and perl is the implementation. Spelling it with ALL CAPS announces that someone knows little about the language.
Regarding slim/embedded distros, it depends on the use cases, and the definition of "slim". It's hard to make broad statements on their prevalence, and regardless, I've never stated that one should use Perl instead and/or that it's "better"; only stated that the option it gives is a valid one.
echo a b c | perl -pale '$_=$F[1]'But it's nice to have awk for the slightly more complicated cases, up until it's easier to use Python or another language.
cdd : like cd but can take a file path
- : cd - (like 1-level undo for cd)
.. : cd ..
... : cd ../.. (and so on up to '.......')
nd : like cd but enter the dir too
udd : go up (cd ..) and then rmdir
xd : search for arg in parent dirs, then go as deep as possible on the other branch (like super-lazy cd, actually calls a python script).
ai : edit alias file and then source it
Also I set a bindkey (F5 I think) that types out | awk '{print}' followed by the left key twice so that the cursor is in the correct spot to start awk'in ;D
# Bind F5 to type out awk boilerplate and position the cursor
bindkey -c -s "^[[[E" " | awk '{print }'^[[D^[[D"
Edit: better formatting (and at work now so pasted the bindkey)
You can submit a new feature request/patch to GNU coreutils' cut, but they'll probably just tell you to use awk.
Edit: Nevermind, it's already a rejected feature request: https://lists.gnu.org/archive/html/bug-coreutils/2009-09/msg... (from https://www.gnu.org/software/coreutils/rejected_requests.htm...)
It's a sharp and not-entirely-welcome change from the 80s and 90s.
Here's another that would be great but will never be added: I want bash's "wait" to take an integer argument, causing it to wait until only that number (at most) of background processes are still running. That would make it almost trivial to write small shell scripts that could easily utilize my available CPU cores.
I use shell because I have to, not because I like it. I dread maintaining shell scripts which have a bunch of awk and sed in them.
The Unix ideal of small single-purpose tools and text processing is separable from these old warhorses.
I never understood the whole religion around unit tests. Integration tests are often far easier to write and far more valuable.
Like you said, unit tests are really nice when testing for a known, expected output.
Unit tests that are effectively testing mocks and crazy stubs because your method has side effects? Not for me.
You need both unit tests and integration tests.
I have written a tool (also a script) myself that allows you to write unit tests, manage your scripts in git, load scripts in other scripts, etc. Maybe I will post a Show HN in the coming weeks, but at the moment I would like to round up some edges before posting it.
In my experience, the biggest problem is that there are many different runtime environments out there that differ in detail and make it hard to write scripts that run everywhere. But programmers not applying all their skills (e.g. writing tests) to build scripts are also part of the problem.
Even better if I can write such utilities in a more modern programming language like Go or Rust, if the organization has personnel with expertise in such languages.
As an example, you have some data you need to fix as input to some program. you incrementally try to fix it with perl:
1) run program with data, observe errors, infer what needs fixing ->
2) write perl to fix data, modify data ->
3) repeat from (1) until no errors
You have no expectation you'll ever need to fix data corrupted in the same way ever again.
Also rectang: "Even better if I can write such utilities in a more modern programming language like Go or Rust..."
Thank you for demonstrating how Linux decays due to lack of experience and NIH syndrome.
For me, that's THE reason to use it. It's terseness is what allows it to be efficient enough to be used primarily interactively.
So for example, selecting columns is easy, and can be done by name.
Here's a useful little snippet demonstrating converting and processing the CSV output of the legacy "whoami" Windows command. It lists the groups a user is a member of without poking domain controllers using LDAP. It always works, including cross-forest scenarios.
WHOAMI /GROUPS /FO CSV `
| ConvertFrom-Csv `
| Where-Object Attributes -NotLike '*used for deny only*' `
| Select-Object -ExpandProperty 'Group Name'
It looks like a lot of typing, but everything tab-completes and the end-result is human readable. I find that people that prefer terseness over verbose syntax are selfish. They simply don't care about the future maintainers of their scripts.PowerShell can also natively parse JSON and XML into object hierarchies and then select from them. That's difficult in UNIX. The Invoke-RestMethod command is particularly brilliant, because instead of a stream of bytes it returns an object that can be manipulated directly.
$p = Invoke-RestMethod 'https://ipapi.co/8.8.8.8/json/'
echo $p.org
echo $p.country_name
PowerShell Core is available on Linux and is great for ad-hoc tabular data processing! Give it a go...I'm a pretty advocate for how it works in these use cases now.
Generally I still use Bash\Sed\AWK for the one off instances when ever I need something done there and then. But if I'm writing something that's going to be involved in my CI\CD with a high chance of needing to be maintained. I'll nearly always write it in PowerShell.
And anything more complex than that I typically just use python.
And some language decisions are just asinine, e.g. how function parameters and variable scope works, fall through by default although you almost never want that, etc..
But hey, you have socket support! Sounds to me like things have developed in the wrong direction.
And of course no one on your team will know awk.
I found the idea of rule-based programming interesting, but the way it interacts with delimiters and sections (switching rules when encountering certain lines) doesn't work well in practice.
I also found the performance to be very disappoinging when I compared it to C and python equivalents.
Awk is there for a reason: to be small. That's why the O'Reilly press book is called "Sed & Awk", because they were orignally written to work together in the early days of unix dating back to the late 70's. Sed (1974) & Awk (1977) are in the DNA of unix, Python is something totally different.
The only reason I could've seen to use awk was to throw code together more quickly in a DSL.
However this is much less the case than I had hoped. For the one liners there are usually specialized tools like fex that are easier to use and faster (for batch processing).
When I compared my C/python/awk programs the difference was msec/sec/minutes. As soon as I use such a program repeatedly it starts to hurt my productivity. And the development time is not orders of magnitude slower in non-awk languages.
Python is absolutely not available everywhere one can find Awk. I've never seen a system with Python but not Awk, but have seen many systems with Awk but not Python (excluding the BSDs, where Python is never in base, anyhow).
Actually, not many years ago I used to claim that I never saw a Linux system with Bash that lacked Perl, but had seen systems with Perl that lacked Bash. (And forget about Python.) This was because most embedded distros use an Ash derivative, often because they used BusyBox for core utilities or a simple Debian install. Perl might not have been default installed, either, but invariably got pulled in as a core dependency for anything sophisticated. Anyhow, the upshot was that you'd be more portable, even within the realm of Linux, with a Perl script than a Bash-reliant shell script. Times have changed, but only in roughly the past 5 years or so. (Nonetheless, IME Perl is still slightly more reliable than Python, but variance is greater, which I guess is a consequence of Docker.)
One thing to keep in mind regarding utility performance is locale support. Most shell utilities rely on libc for locale support, such as I/O translation. Last time I measured, circa 2015, setting LC_ALL=C resulted in significantly improved (2x or better, I forget but am being conservative) standard I/O throughput on glibc systems.[1] I never investigated the reasons. glibc's locale code is a nightmare[2], and that's more than enough explanation for me.
Heavy scripting languages like Perl, Python, Ruby, etc, do most of their locale work internally and, apparently, more efficiently. If you don't care about locale, or are just curious, then set LC_ALL=C in the environment and test again. I set LC_ALL=C in the preamble of all my shell scripts. It makes them faster and, more importantly, has sidestepped countless bugs and gotchas.
For the things I do, and I imagine for the vast majority of things people write shell scripts for, you don't need locale support, or even UTF-8 support. And even if you do care, the rules for how UTF-8 changes the semantics of the environment are complex enough that it's preferable to refactor things so you don't have to care, or can isolate the parts that need to care to a few utility invocations. In practice, system locale work has gone hand-in-hand with making libc and shell utilities 8-bit clean in the C/POSIX locale, which is what most people care about even when they care about locale.
[1] The consequence was that my hexdump implementation, http://25thandclement.com/~william/projects/hexdump.c.html, was significantly faster than the wrapper typically available on Linux systems. My implementation did the transformations from a tiny, non-JIT'd virtual machine, while the wrapper, which only supports a small subset of options, did the transformation in pure C code. My code was still faster even compared to LC_ALL=C, which implied glibc's locale architecture has non-negligible costs.
[2] To be fair, it's a nightmare partly because they've had strong locale support for many years, and the implementation has been mostly backward compatible. At least, "strong" and "backward compatible" relative to the BSDs. Solaris is arguably better on both accounts, though I've never looked at their source code. Solaris' implementation was fast, whatever it was doing. musl libc has the benefit of starting last, so they only support the C and UTF-8 locales, and in most places in libc UTF-8 support simply means being 8-bit clean, so zero or perhaps even negative cost.
http://widgetsandshit.com/teddziuba/2010/10/taco-bell-progra...
Someone ought to write - Zen and the art of Unix tools usage.
I don't know enough about the 'real way' or the 'taco bell way', but interested to know --- is this doable the way Ted describes in the article via xargs and wget?
- sed/awk to extract URLs, one by line
- xargs and wget to download each page from the previous output
"This is the opposite of a trend of nonsense called DevOps, where system administrators start writing unit tests and other things to help the developers warm up to them - Taco Bell Programming is about developers knowing enough about Ops (and Unix in general) so that they don't overthink things, and arrive at simple, scalable solutions"
It's not possible for developers to know enough about Ops, just as it's not possible for Ops to know enough about development, because they are different jobs. Moreover, devs are doomed to create terrible solutions because of their job, and Ops are doomed to create kludgy hacks for those terrible solutions because of their job. DevOps is just an attempt to get them to talk to each other frequently so that horrible shit doesn't happen as frequently.
Also, the real Taco Bell programming is actually to use only wget, no xargs. It takes a whole lot of basically every option Wget has, and a very reliable machine with a lot of RAM, but you can crawl millions of pages with just that tool. xargs and find make it worse because you don't get the implicit caching features of the one giant threaded process, so you waste tons of time and disk space re-getting the same pages, re-looking up the same hostnames, etc. (And that's Ops knowledge...)
The Zen of Unix is to try to move towards not using the computer at all. One-liners are part of that path, but so is minimizing the one-liner. http://www.catb.org/~esr/writings/unix-koans/ten-thousand.ht...
> so you waste tons of time and disk space re-getting the same pages, re-looking up the same hostnames, etc. (And that's Ops knowledge...)
If it's my job to write a web scraper, it's absolutely my job to think about/solve this problem.
Is this a new trend thing?
$ perl --help
...
-F/pattern/ split() pattern for -a switch (//'s are optional)
-l[octal] enable line ending processing, specifies line terminator
-a autosplit mode with -n or -p (splits $_ into @F)
-n assume "while (<>) { ... }" loop around program
-p assume loop like -n but print line also, like sed
-e program one line of program (several -e's allowed, omit programfile)
Example. List file name and size: ls -l | perl -ae 'print "@F[8..$#F], $F[4]"'Awk example:
ls -l | awk '{print $9, $5}' or ls -lh | awk '{print $9, $5}'
Seems a whole lot simpler. To me. I find if you have to write exhaustive shell scripts then maybe you can look for something more verbose like Perl, I guess.
> The print statement shall write the value of each expression argument onto the indicated output stream separated by the current output field separator (see variable OFS above), and terminated by the output record separator (see variable ORS above).
The default value for OFS is <space> and for ORS, <newline>.
No, lack of commas in output and broken filenames with spaces.
ls -l | awk '{print $9 "\t" $5}'
That is about as much as i'm willing to do for this.
Even though there's frequently value in adding whitespace to programs, many of them are just fine as one liners :)
e.g. this gets the title of a webpage for you: ``` $ perl -Mojo -E 'say g("mojolicious.org")->dom->at("title")->text' ```
https://ferd.ca/awk-in-20-minutes.html
Also, a handy trick is to combine awk and cut. For example I had a log line that had a variable amount of columns just in one field, but immediately after the field was a comma. I cut based on the comma:
cut -d, -f1,2
and then awk'd the last column:
cut -d, -f1,2 | awk '{ print $2" "$5" "$NF }'
So, sometimes awk and cut can help each other.
[0] https://gregable.com/2010/09/why-you-should-know-just-little...
I found that reading these -- sort-of half-assed structured data, but with page-chunked artefacts and idiosyncrasies -- was difficult on a line-by-line basis, and thought idly "this would be a lot easier if I could process by page instead".
Text was laid out in columns, and the amount of indenting (the whitespace between columns) was significant. So preserving this somehow would be Very Useful.
Suddenly those pesky '^L' formfeeds were an asset, not a liability. Let's treat the formfeed ("\f") as a record delimiter, and the newline ("\n") as a field delimiter. We can parse out the actual columns based on witespace, for each line:
BEGIN { RS="\f"; FS="\n" }
{
pageno = NR
lines = NF
for( line=1; line<=lines; line++ ) {
ncols = split( $line, columns, " {2,}", gaps )
}
}
This gives me:- The running tally of pages.
- Each line of the page as an individual record.
- Via the split() function, an array of columns separated by two or more spaces, which are saved as an array of gaps so I have the whitespace to play with.
Edge cases and fiddling ensue, but that's the essential bit of the code there.
Since the lines are an array, I can roll back and forth through the page (basically being able to read forward and backwards through the text record), testing values, finding out where column boundaries are, etc., and then output a page's worth of content, transposing to a single-column format, with appropriate whitespacing, when done.
In testing and debugging the output (working off of 20+ documents of 100s to ~1,000 pages), a lot of test cases, scaffolding, diagnostics, etc., have been created and removed to make sure the Right Things are happening. Easy with awk.
https://www.gnu.org/software/gawk/manual/html_node/Multiple-...
As an example, if I want a sorted list of all open files under the home directories on CentOs I can do this:
lsof | awk '{ print $10 }' | grep ^/home/ | sort | uniq
You can drop another command as such:
lsof |awk '($10 ~ /^\/home\//) {print $10}' |sort -u #! /bin/bash
# lsof-tree: list open files in a given directory tree (default /home)
NAME=9 # set to 10 for CentOS
BASE="${1:-/home}"
lsof | awk -v NAME=$NAME '{print $NAME}' | grep "^$BASE" | sort -uI guess it depends on how often you need that particular pipeline. Every day? Sure, make a script. Every few months? Nah, I won't remember it anyway,or probably I remember that I've made a script like that but then I have to start searching my bin directory in the end using more time than just writing the pipeline in the first place.
Why you use awk?
[0]: http://rc3.org/2014/08/28/surprisingly-perl-outperforms-sed-...
Remembering which Perl command-line arguments simulate awk’s line-by-line processing is harder than just remembering awk.
I think if you know Perl really well and can remember the command line arguments - particularly -E, -n, -I and -p - then it’s a good swap in substitute for grep, sed, awk, cut, bash, etc when whatever 5 min task you’re working on gets a tiny bit more complex.
Similarly a decent version perl 5 seems to be installed everywhere by default.
I’m curious to know if anyone would say the same about python or any other programs? I’m not particularly strong in short python scripting.
I do, however, use it for JSON pretty printing in a pipeline: python -mjson.tool IIRC.
Because AWK is not suited for CSV. Please prove me wrong!
I had to parse 9million lines. Some of which contain "quoted records", others, same column, are unquoted. Some contain comma's, in the fields, most don't. CSV is like that: more like a guideline than actual sense.
Two hours of googling and hacking later, I gave up and rewrote the importer in Ruby, in under 5 minutes.
Lesson learned: I'll stay clear of AWK, when I know a oneliner of Ruby (or Python) can solve it just as well. Because I know for certain the latter can deal with all the edgecases that will certainly pop up.
(It’s Python, and you can use it as a library as well)
Awk would chew through that no problem.
> Some of which contain "quoted records", others, same column, are unquoted.
In which case, there is the FPAT variable which can be used to define what a field is. FPAT="\"[^\"]\"|[^,]", which means "stuff between quotes, or things that are not commas", would probably have worked for you. (EDIT: Looks like formatting has gotten hold of my FPAT and I don't know how to stop it... hopefully it is still clear where asterisks should be)
> Some contain comma's, in the fields, most don't. CSV is like that: more like a guideline than actual sense.
Well, I would say that's absolutely false. You can't just put the delimiter wherever you fancy and call it a well-formed file. Quoting exists for the unfortunate cases your data includes the delimiting character (ideally the author would have the sense to use a more suitable character, like a tab).
This is just a retort to prevent your post from dissuading readers from awk, which is a fantastic tool. If you actually sit for half and hour and learn it rather than google to cobble together code that works, it is wonderful. I also don't think it is valid to base your judgement of a tool on what was apparently garbage data.
But if you want to be in a world where people only deal with well specified files like RFC 4180 (for some definition of well specified), your quick field pattern doesn’t conform. It doesn’t handle escaped double quotes or quoted line breaks. If you’re using your quick awk command to transform an RFC 4180 file into another RFC 4180 file you’ve just puked out the sort of garbage you were railing against.
While awk is a great tool if you’re dealing with a csv format with a predictable specification, and probably could be made to bend to the GP will with a little more knowledge, it gets trickier if you’re dealing with handling some of the garbage that comes up in the real world. What’s worse is the programming model leads you down the path of never validating your assumptions and silently failing.
I love awk for interactive sessions when I can manually sanity check the output. But if I’m writing something mildly complex that has to work in a batch on input I’ve never seen, I too would reach for ruby.
The lesson I took wasn't that awk sucks, though. The lesson was that CSV is not trivial, and should not be parsed with regex or string matching. It's a standard with variants, and rolling in a library will pay dividends, especially if you're parsing a wide variety of different dialects of CSV.
A related lesson I took is that once your awk script grows beyond a certain level, graduate it up to a real language. I love awk, but it excels at small scale text munging. It's not suited to anything more involved than that. If translating an awk program is a major task, then the program was already too big to begin with.
But if your file is correct csv, and you use gawk, this does the trick: https://www.gnu.org/software/gawk/manual/html_node/Splitting...
I usually load CSV data into PostgreSQL to do anything with it; mostly wrote this Awk library for fun. So I'm not going to argue that Awk is the best language for doing this kind of thing, but it is possible.
So use a dedicated tool or library or you will run into trouble one day.
For one off jobs, I open it in Excel or similar and save as tab or pipe delimited text. This usually plays much more nicely with command line utilities, assuming it didn't mangle any numbers.
Contact information is hn handle @ yahoo.com.
[1] - https://www.commandlinefu.com/commands/matching/awk/YXdr/sor...
[1] https://github.com/learnbyexample/Command-line-text-processi...
Here's the thing: these arguments all too commonly focus on subjective notions of "simplicity", and toy examples divorced from actual common practise, and or solid comparable benchmarks.
Show me a range of practical examples, for each competing env (awk, sed, perl, python, ruby, maybe bash).
Include:
- time it takes to teach a total novice (the time it takes to learn whatever is needed that example, not the entire language)
- how easy it is to recall said knowledge at a later date
- how fast example is multiplied by how much you are likely to use it = actual time saved in terms of execution. for small, fast examples the difference is irrelevant, a 10x speedup that is 0.1s vs 0.01s is meaningless.
- how extendable an example is. Hence the original example should include a series of extensions to the original task, to demonstrate how flexible / composable they are: e.g. task 1) count lines in a file; task 2) count lines in a file, then add 42 to it;
I suspect awk falls behind in practicality vs perl (which can do simple one-liners, but also more complicated constructs), but perhaps has a hidden virtue wrt speed in more expensive tasks, ala https://news.ycombinator.com/item?id=17135841 or https://news.ycombinator.com/item?id=20293579
And you have immediate access to so many useful modules (csv, json, xml) and can easily extract code fragments into functions.
You can also execute shell-like commands with subprocess.check_output() without ever worrying again about escaping strings or accidentally splitting them at spaces or whatever.
Clever one liners are difficult to comprehend. It's better to break them up to a few variable assignments with descriptive, long names without abbreviations.
The only problem is that it was written by people using ed on PDP computers and that kind of shows. The primary logic is "filter out lines X and apply transform Y" is completely natural to someone using an editor like ed, but is fairly foreign to modern computer users. Most people aren't going to take the time to learn an obscure commandline tool these days, especially since it comes from the Dennis Ritchie school of "errors are like angry housewives, you know what you did" debugging.
I think your missing the point of awk. The O'Reilly sed and awk book has some complex examples, but when I look at my own usage they are all toy examples within a much larger scope. It's more like a special DSL extension for my shell than something I'd pick to build the entire solution, so a comparison to perl, python and ruby don't really make sense, they are general purpose languages but awk just has a couple of features that make it a very specialized yet useful one.
As an example, a have a system for importing and parsing log files that mostly done from a shell script, awk is used in two parts. The first is to transform a structured and easy to read file (records '\n\n' separated) into a csv easier to consume for bash, there's probably quite a few options to do this from tr to bash and it's done inline. The second is to filter the results down to what I need, so I have scripts like:
#!/usr/bin/awk -f
/some common error I don't care about/ { next } #skip line
/other common error/ { next }
/Error/ { print $0 } #this prints error lines, alternatively:
/Error/ && !errors[gensub($1", "", "g", $0)]++ { print $0 } #print each error once
{next} #skip everything else
Apart from one single line which wasn't in the original that's something you could teach a total novice in minutes, the /pattern/{action} syntax is about as simple as programming can be. Execution speed could probably be improved with a specific program but I suspect the bottleneck would be the spinning disk anyway, I run this over hundreds of MB every few minutes and it's not a problem, when I run it manually it's near instantaneous, I spend longer waiting for the desktop calculator to open up these days.perl is general purpose, but that doesn't mean it can't be used for one-liners.
in everyday life however, many small problems are a bit too much for the shell but too little for Perl/Python/whatever.
awk fits very nice in there.
I live in this future and it's beautiful.
Steps to programming without regular expressions:
1) find a PEG library for your language of choice
There is no step 2.
It is a recursive decent parser with the tiny tweak that productions are ordered (not a set) and short circuit.
A 90's language you did not have to imagine that saved you from regex in this way was[is] REBOL with it built in `parse` dsl.
examples here http://www.rebol.org/search.r?find=parse&form=yes
[0]https://en.wikipedia.org/wiki/Parsing_expression_grammar [1]http://www.rebol.com/
For those not familiar with PEG, like me, here: https://en.wikipedia.org/wiki/Parsing_expression_grammar
I found that fairly abstract, a Python example provided more concrete examples: https://github.com/erikrose/parsimonious
https://softwareengineering.stackexchange.com/questions/1949...
For a huge number of simple tasks, awk is available and sufficient. It's largely a subset of Perl, so yes, there's some skills overlap, but there are times where knowing awk is the right tool and the available tool will pay off.
(Use the right tool for the job.)
But it was the ability to string together these types of command line tools that made it possible.
I believe this was intentional by the authors.
In the early days of UNIX, I think more of its users knew C. Today, it is probably a much smaller number. However, learning AWK today, IMO, can help someone who also intends to learn C.
Best of all worlds. I wish there would be a way to get back the time I spent with debugging edge cases of different regex implementations on various OS's.
Maybe I’m dumb but I’ve never come up with a separator regex that is quite right.
Is this not what you mean?
Edit: This comment might be helpful to you: https://news.ycombinator.com/item?id=22110036
It looks like FPAT from your linked article is for gawk. Gawk is great, but it's not everywhere. Still - it's good to know. Thanks!
FS = ","
or: awk -F ',' <program>
If you're working with CSV data that has quoted strings with embedded commas, FPAT is your friend: FPAT = "([^,]+)|(\"[^\"]+\")"
See: https://www.gnu.org/software/gawk/manual/gawk.html#Splitting...brew install gawk
good to go
:)
1. Uses shell-only commands (echo, for) - most robust; but things like basename/dirname and regex's vary by shell (sh, bash, zsh, ksh)
2. Uses /bin - might run into missing a binary but not likely, still robust and allows a richer set of tools (e.g., uname, chmod and admin-ish things live in /bin)
3. Uses /usr/bin - runs risk if missing packages, likely not very robust (packages drop things in here, like gzip, yacc, gcc)
4. Uses /usr/local/bin or /opt/local/bin - definitely requires package installs, least robust