awk '{print $1","$2}' | sed '1i count,word'
when you can just add a BEGIN block: awk 'BEGIN { print "count,word" } { print $1","$2 }' awk '{print $1","$2}' | sed '1i count,word'
when you can just add a BEGIN block: awk 'BEGIN { print "count,word" } { print $1","$2 }'https://ia802309.us.archive.org/25/items/pdfy-MgN0H1joIoDVoI...
There's plenty of examples and exercises, plus an entire chapter dedicated to regular expressions.
https://www.gnu.org/software/gawk/manual/gawk.html
The manpage alone is quite useful, though lacks some useful details.
The awk FAQ provides general background, including diffeerences between implementations:
https://m.youtube.com/watch?v=43BNFcOdBlY
In an afternoon, you can learn enough to be more than dangerous.
The presenter drafted up a set of (easy, yet practical) companion exercises and a pt2 video with his own answers. I cannot recommend this talk enough.
Although I hadn't the opportunity to put the knowledge to use and have since forgotten it (i don't work in tech), i find comfort in knowing i can reacquire the power of awk even sooner than the already-short first time around.
awk < file '
{
$0=tolower($0); # lowercase conversion
gsub( "[^a-z ]", "", $0); # remove non-alpha + space chars
for(i=1;i<=NF;i++0) {
# increment word count if word length > 2
if(length($i)>2) count[$i]++
}
};
END{
# report
for(word in count) printf( "%-20s %i\n", word, count[word])
}
' | sort -k2nr -k1
This omits the header (can be trivially added). The external sort can be internalised in gawk using asort(), or by printing to sort via a command: cmd="print -2kr -k1"
for<loop> printf( <args> ) | cmd
close(cmd)This one did. Never left it, as it happens.
C-x C-e FTFW.
(I expanded, indented, and commented the example for HN.)
However the point often missed when people see awk piped into other CLI tools is that code is intentionally optimised for one time writing rather than maximising awk's usefulness.
It's the same with the GPs comment too. In that example the awk code was longer than the awk|sed equivalent.
Don't get me wrong, I am a big fan of awk myself. But sometimes people get so hung up on better usage of awk that they lose sight of the point behind the command line example.
real 0m0.176s
user 0m0.060s
sys 0m0.060s
On an early-2015 Android tablet running Termux.I've thrown multimillion row datasets at awk (usually gawk, occasionally mawk, nawk on OSX, and, hell, busybox on occasion) without any practical performance issues. I'm virtually always writing for one-off or project-based analyssys, not live web-scale realtime processing. A second or even ten won't be missed.
I'm also aware that building a pipeline out of grep / cut / sed / tr / sort / uniq / awk is often conceptually nearer at hand. It almost always mirrors how I start exploring some dataset.
But a quick translation to straight awk gives cleaner code, more power, easier conversion to a script, and access to a small library of awk-based utilities I've acumulated.
All with far-more-than-adequate performance.
We've spent far more time discussing this than coding, let alone running, it.
500 records isn't a large dataset. Not even close. Large would be orders of millions to billions. And yes, I have had to munge datasets that large on many occasions.
> But a quick translation to straight awk gives cleaner code, more power, easier conversion to a script, and access to a small library of awk-based utilities I've acumulated.
- cleaner: only if you find awk readable. Plenty of people don't. Plenty of people find Go or Python more readable.
- more powerful: again depends. You wouldn't have multithreading builtin like pipelines would. And Python and Go are undoubtedly more powerful than awk. I'm not knocking awk here, just being pragmatic.
- easer conversion to a script: at which point you might as well skip awk entirely and jump straight to Go or Python (or any other programming language)
- and access to a small library of awk-based utilities I've accumulated: that only benefits you. If you're having to sell the benefits of awk to someone then odd are they don't have that small library already to hand ;)
> We've spent far more time discussing this than coding, let alone running, it.
Some of the larger datasets I've had to process have definitely taken longer to munge than my reply here has taken to type :)
Disclaimer: I've honestly not got a problem with awk, I used to use it heavily 20 years ago. But these days its value is diminishing and a lot of the awk evangelists seem to miss the point of why awk isn't well represented in blog posts any more. It's both more verbose than pipelining to coreutils and less powerful than a programming language -- it's that weird middle ground that doesn't provide much value to most people aside those who are already invested into the awk language. Some might see that as a loss but personally I see that as demonstrating the strength of all the other tools we have at our disposal these days.
All I'm doing is citing a few reasons why someone might prefer a terser pipeline but I hadn't realised this wasn't supposed to be an objective conversation and since I have no interest in engaging in pointless language fanboyism I'm just going to leave you to it.
a quick translation to straight awk gives cleaner code, more power, easier conversion to a script, and access to a small library of awk-based utilities I've acumulated.
https://news.ycombinator.com/item?id=24822697
You're not discussing what I've actually written.
Good day.
I use to be very comfortable using awk/sed/perl/sort/uniq/tr/tail/head from the CLI for the sort of data cleaning this article is talking about. However, over the past year I've found I use VisiData https://github.com/saulpw/visidata for interactive work.
If I need to clean up the data first, I'll use mlr or jq as input to Visidata. If my data is too dirty for mlr, then I'll use Unix toolbox tools mentioned as input to mlr, jq or VisiData.
VisiData provides some ability to script, but when possible I prefer to have the shell do the scripting with all the tools mentioned as input to Visidata.