Sculpting text with regex, grep, sed, awk, emacs and vim (2012)
matt.might.net
matt.might.net
jq -c .'select(.server_name == "slow_server") | .end_time - .start_time' < my_log_file
where your log file might look like '{"server": "slow_server", "timings": {"end_time": 1406611619.90, "start_time": 1406611619.10}}'
to get your web request timings.Because it's line-oriented, it also works seamlessly with other tools, so you can pipe the output to, say, sort, to find the slowest requests.
A good follow-up to read, from the same person, is his article on relational shell programming: http://matt.might.net/articles/sql-in-the-shell/
The comment on jq (which I'd never seen) had me thinking about the relational shell programming again.
One could implement a remarkably robust relational DB at the shell with jq.
$^(11+)(\1)+$
see OPs linked article http://zmievski.org/2010/08/the-prime-that-wasnt for detailshttp://doc.cat-v.org/bell_labs/structural_regexps/
Many people have tried to generalize unix pipes and homogeneous data, few have succeeded.
For even longer ones I just started using perl with /x, so you can uses insignificant whitespace and comments.
One very useful feature in pcregrep is outputting the matched subpattern only. For example if you do:
echo 'abcdefg' | pcregrep -o2 'a(bc)d(ef)g'
It will output only second matched subpattern.If you want a proper language to support your scripting goals, then if you go right up to something like Ocaml or Haskell you'll skip all the pointless stringly-typed problems of perl/ruby/python.
Haskell and OCaml and Java and C++ are about equally badly suited for the job. No, they don't make a good scripting languages. And they don't even want to. Why would anyone try to write shell scripts with them is really beyond me.
These are tools that are not going away tomorrow just because something better exists.
Languages like Awk and especially tools (you could say "language" because it's strictly true but, come on) like sed have built in safeguards against writing code that stretches for more than a certain number of lines. That safeguard is that it's a really awful experience to actually do that. As a result these scripts tend to be short and to the point.
Perl does not have this safeguard.
perl -ne 'print $1 if /foo="(.*?)"/'
awk '/foo=".*"/ { ??? }'
You can do it with gawk, but it's ugly: gawk 'match($0, /foo="(.*?)"/, a) { print a[1] }'
Another is manipulating hexadecimal numbers, which is also a gawk extension.Now that Python has eclipsed Perl's popularity, it's just a matter of time before you start seeing the same level of quality issues in Python. The untrained, non-programmers will be creating write-only scripts in the new language soon enough.
It was just a few months ago that I debugged some Python scripts for a QA department at a smallish company. This code was the equivalent of any nightmare that I've seen in Perl. Not only was it all very "un-Pythonic", it didn't use classes, it hardly used subroutines, and it was equal parts of commented out tries along with the "working" code. The gem was a script that wrote another Python script and executed it (written because the author only knew how to initialize multidimensional arrays, but didn't know how to build them on-the-fly).
(And yes, there were popular scripting languages before Perl. I remember arguing the superiority of Bourne shell scripts of C-shell scripts...)
Perhaps we can retire the notion of "write-only" Perl -- all languages of sufficient complexity provide the means to obfuscate.
I thought that we were supposed to stop using classes in python.