A Crash Course In Awk
blog.bignerdranch.com
blog.bignerdranch.com
$0.02
If there's one argument against using perl in place of these other tools it's simply the cognitive overhead of learning the extra stuff perl brings to the table on top of bash, sed and awk. On the other hand, picking up non-OO old-school perl if you're already a proficient shell scripter should be a day's work ... the important thing is knowing when to switch tools.
The advantage of learning perl is that you can replace both utilities in your piped command line and only have to remember one syntax. And it integrates just as well as sed and awk do in any pipeline. But quite often, it will obviate the need for piping through other utilities.
Perl is much closer to sed and awk than any of the other scripting languages you mention. As noted in other posts, you can even automatically convert awk scripts to perl. And it doesn't take much effort to convert sed scripts to perl where you can leverage more powerful pattern matching and procedural facilities to boot.
If you already know sed and awk, by all means keep using them, they're fine. But if you're new to the Unix command line, you'll get the most bang for your buck by learning perl instead of awk and sed.
import json
f=open("bus-stops.json")
j=json.load(f)
for a in j:
print a['no'],",",a['lat'],",",a['lng'],",",a['name']
A simple "import json", "help(json)" at the Python command line, and 2 minutes later I was done.
Also - I'm able to understand my code the next day - something I was never able to do with Perl, but for some reason I can with Python.I probably spend 90% of my time in sed/awk, and 10% of my time in Python. Haven't touched Perl in 10 years - not because it isn't an awesome language (it really is) - it's just that I have room in my head for one full blown language at a time, and Python replaced Perl for me.
perl -naE 'say $F[0], $F[1]'
perl -nae '$tot+=$F[1]; $c++;END{print $tot/$c}'
perl -ape 's/^.client="(.)",dev.*/\1/' foo.txt
And finally: use JSON;
use IO::All;
use feature 'say';
$f = io('bus-stops.json')->all;
$j = decode_json($f);
$,=',';
for (@$j) {
say @$_{'no','lat','lng','name'};
}
I still think newbies would get much more value out of learning perl basics than spending any time on sed/awk intricacies. But to each his own.I'd argue with most others, Perl is the single most useful command line tool. Not the only one, of course. But afaik, you can't e.g. load a JSON lib in awk as part of a pipeline. (I deserialize dumped data structures multiple times a week in [ad hoc testing with] pipelined cmds.)
Imho, if you know Ruby or PHP as the back of your hand, don't learn another scripting language for command line use. Learn some completely different language for some other use, instead.
gawk -vFPAT='[^,]*|"[^"]*"'
http://stackoverflow.com/questions/4205431/parse-a-csv-using...Are there multiple variants of coding '"' in CSV fields? I don't know -- but some people who do know are those who write the CSV libs I use!
Edit: And as your link notes, it fails for embedded \n:s. Imnsho, awk needs csv (and json, etc) builtin, preferable as a plugin architecture. But then, why not just use the Perl superset?
Hierarchical formats like JSON are a little different, because they don't fit the awk model very well. You could add functions to work with JSON, but working with it this way wouldn't be very awk-like. You're better off preprocessing the JSON into records with another tool to make it more awk-friendly, or simply using another language altogether.
Join the dar... cough, Perl side, we have cookies. :-) We have CSV parsers and everything else, all the way up to e.g. good web libraries and the best OO among the scripting languages (Moose, ~ like the Common Lisp OO environment; more or less std for new Perl projects today.)
And there is more! You can reuse most everything you know from awk! Write: perldoc perlrun
Check for -n, -p, -i, -E flags. And, as many have noted, there is a2p.
http://perldoc.perl.org/5.16.2/perlrun.html
http://perldoc.perl.org/5.16.2/a2p.html
But the main reason is that we have fun. An insane programming language which throw all this "minimal mathematical notation" stuff out the window with some linguist inspirations, but still works wonderfully (do insist on keeping to the coding standards in your group. Seriously. At a minimum -- lie and say that you do that, when people interview for a job at your place. :-) )
[0] https://github.com/stedolan/jq [1] https://github.com/onyxfish/csvkit
You can use semi-colons instead of newlines, and colons don't require a newline.
python -c 'for c in "abc": print c; print c'
Imports and actual pipelines are a little more tedious, it can be done, but python isn't as straightforward as perl or shell-ish tools for pipelines.(Last I looked -- the answer was no. It removes a use case for ideology, sigh.)
print "hi" if True else "bye"
I eventually got pretty good at doing both functional and OO-style programming (each where appropriate, and straight-up imperative scripts in many other places) in POSIX-compliant shell.
And that's kind of why I like awk as much as I do, in spite of its limitations. There's so little to the language, it's of practical use, and as a precursor to tools like perl, python, and ruby, it's historically interesting, too. And I know a lot of people who like to learn neat, useful things who don't know anything about awk.
Thanks for reading, btw - I think perl is another older tool that could stand for some evangelizing. I don't write perl, though, so I'm not the guy to do it.
That said, the loop concept and using next is also very valuable piece of information. So thanks to both!
This produces same output as your awk script (except it's sorted alphabetically):
#!/usr/bin/perl -F, -an "$@"
$wins{$F[0]} += $F[6];
$losses{$F[0]} += $F[7];
END {
print "manager,total_wins,total_losses\n"
for (sort keys %wins) {
print "$_,$wins{$_},$losses{$_}\n";
}
}
EDIT: should also mention that you can automatically convert awk scripts to perl with the a2p utility which should already be installed along with perl. #!/usr/local/bin/gawk -E
BEGIN{
FS=","
}
{
total_wins[$1]+=$7;
total_losses[$1]+=$8;
}
END{
print "manager,total_wins,total_losses"
n = asorti(total_wins, managers)
for (i=1; i<=n; i++){
m = managers[i]
print m "," total_wins[m] "," total_losses[m]
}
}
(I recommend using gawk over mawk or nawk. You might need to install it over the awk that came with your OS.) $ seq 1 100 | awk '!($1%3){printf "Fizz";p=1}!($1%5){printf "Buzz";p=1}!p{printf $1}{printf "\n";p=0}' !($1%3){$2="Fizz"}!($1%5){$2=$2 "Buzz"}!$2{$2=$1}{print $2}