Pipes and Filters
blog.petersobot.com
blog.petersobot.com
• Though the first pipeline is didactic, it can be done entirely within awk:
awk '
BEGIN { l=0 }
/purple/ {
if(length($1) >= l) { word = $1; l = length($1) }
}
END { print word }' < /usr/share/dict/words
• Named pipes are neat, but you can also use subshells and additional FDs (I am in no way arguing this is more clear): (
(
(
echo out
echo err >&2
) | while read out; do echo "out: $out"; done >&3
) 2>&1 | while read err; do echo "err: $err"; done
) 3>&1
• Bash has "set -o pipefail" for cases where you want any process in the pipeline that exits non-zero to cause the entire pipeline to exit non-zero.http://www.gnu.org/software/bash/manual/html_node/Pipelines....
"If pipefail is enabled, the pipeline’s return status is the value of the last (rightmost) command to exit with a non-zero status, or zero if all commands exit successfully"
http://www.ivarch.com/programs/pv.shtml
> pv - Pipe Viewer - is a terminal-based tool for monitoring the progress of data through a pipeline. It can be inserted into any normal pipeline between two processes to give a visual indication of how quickly data is passing through, how long it has taken, how near to completion it is, and an estimate of how long it will be until completion.
Example: the raspberry pi has pretty slow SD performance and the USB bus can get hogged. If you record audio and want to encode + write it to SD you can easily get buffer overruns. Easily solved by a 10sec buffer between arecord and flac in my case.
arecord -D hw:1,0 -v --fatal-errors --buffer-size=192000 -f dat -t raw | dd bs=480000 | flac --endian=little --channels=2 --bps=16 --sample-rate=48000 --sign=signed -o /mnt/usbstick/`date '+%s'`.flac -
Gotta love Linux :)It highlights a similar pipeline-oriented architecture and eventually ends up being sort of mindblowing.
There are lots of ways to tighten up the example, if needed.
Considering the tiny overhead of an additional cat process, UUOC these days feels like nitpicking.
head bigfile.1.txt | grep | awk | stuff
and refine things, and when output looks right, it's a simple "Ctrl+A Meta+D cat RET" to run it on the full output. Or vice versa if I suddenly want to go back to testing part of bigfile (or exchange the cat for "grep something").If I want to change that to "< bigfile.1.txt", I have to "Ctrl+A Meta+D < Meta+F Meta+F Meta+F Ctrl+D Ctrl+D" – the extra keypresses are to delete the first "|" symbol. And if I suddenly want to change it back to head or grep, I have to reinsert the | (also I often by habit do Meta+D instead of Ctrl+D at the beginning of the line, which doesn't work as intended if the first token is "<" instead of "cat").
Those useless cats are quite handy when doing a lot of shell work.
Also, that's a Useless Use of head, since grep has the option "-m10"
"(If we move grep to run immediately after cat, and before putting data into Redis, this operation runs more than 1,200 times faster.)"
Which does indicate that in some cases the UUOC is justified (thou in this case the cat remains)
...But, yeah, exactly. :)
Various searches on the subject revealed plenty of people noting how meaningless Internet points are, leading to an additional meta-rub: not only did I fall for the hover trick, I also was childish enough to google for a svbtle kudos undo. sigh
I don’t mind the useless use of `cat` as it can enhance readability for some people. However, I would suggest replacing the Bash while loop with a for loop:
ls *.flac |
while read song
do
flac -d "$song" --stdout |
lame -V2 - "$song".mp3
done
for song in *.flac
do
flac -d "$song" --stdout |
lame -V2 - "$song".mp3
done
Using Bash’s file globbing avoids problems with enumerating the output of `ls` [1]. It also avoids an unnecessary Bash sub-shell being spawned for the while loop that follows the first pipe. More importantly, I think it’s a lot more readable while still demonstrating how pipes can be efficiently used to process any amount of FLAC files.Here's a video about it: https://www.youtube.com/watch?v=f_0QlhYlS8g
Also, I've been looking at Factor, but I have a hard time getting into the mindset of the paradigm, since I'm used to Python (although I'm very much used to using pipes in the terminal). Are there any types of programs that you prefer to write in Factor, and others you prefer to write in -- for example -- Python?
_____________
< unimpurpled >
-------------
\
\
Part way through my second viewing of the article, I thought, "what is 'unimpurpled'". Wiktionary didn't know. Google doesn't return useful results for it, even. M-W, finally, clued me in: it's an obsolete term, with an "un" prefix, for the verb "empurple", which means to make purple[1].[1] And a few similar things. https://en.wiktionary.org/wiki/empurple
x |> a|> b |> c|>s->d(s,y)|>e|>...
in julia instead of e(d(c(b(a(x))),y)) or (e (d (c (b (a x)) y))
...or whatever is your flavour. I reckon it is impossible to make a serious case against that readability gain.[1] julialang.org
f (g (h x)) == f $ g $ h $ x
== f . g . h $ x
\x -> f (g (h x)) == f . g . h
The "compose" combinator, (.), is especially pertinent for making pipelines. Idiomatic Haskell code uses it constantly---especially for its natural mechanism of eliminating "points" like that `x` above. These are usually better described by their type than any variable name given (especially since the variable name cannot be machine checked for meaningfulness, unlike the type) and so are best eliminated.In many libraries there is also a reverse apply function defined, often as &
f (g (h x)) == x & h & g & f
== x & f . g . h
which is more popular when using other operator chains to describe functions as in lens.I prefer to use command in this scenario.
$ command time
You might be interested in the classic McIlroy-Knuth dialogue:
http://www.leancrew.com/all-this/2011/12/more-shell-less-egg...
The Unix "Software Tools" philosophy -- small kits composable into something big -- is deeply shared by functional programming.
I don't think one begot the other. It was just obviously the natural thing to do during that era.
Yes, that was a good article. The comments on the post were interesting too. Saw it a while ago, and for fun, wrote two quick solutions for the problem, in Python and shell:
http://jugad2.blogspot.in/2012/07/the-bentley-knuth-problem-...
A reader, Veky, wrote an interesting comment on my post too.
Basically, the example can be shortened to the following:
grep purple /usr/share/dict words | # Find words containing 'purple' in the system's dictionary
awk '{print length($1), $1}' | # Count the letters in each word
sort -n | # Sort lines ("${length} ${word}")
tail -n 1 | # Take the last line of the input
cut -d " " -f 2 | # Take the second part of each line
cowsay -f tux # Put the resulting word into Tux's mouth