For the Love of Pipes
blog.jessfraz.com
blog.jessfraz.com
The Unix philosophy is documented by Doug McIlroy as:
Make each program do one thing well. To do a new job, build afresh rather than complicate old programs by adding new “features”.
Expect the output of every program to become the input to another, as yet unknown, program. Don’t clutter output with extraneous information. Avoid stringently columnar or binary input formats. Don’t insist on interactive input.
Design and build software, even operating systems, to be tried early, ideally within weeks. Don’t hesitate to throw away the clumsy parts and rebuild them.
Use tools in preference to unskilled help to lighten a programming task, even if you have to detour to build the tools and expect to throw some of them out after you’ve finished using them.
I really like the last two, if you can do them in development then you are then you have a great dev cultureBut of course that's probably not the author's intended context.
> Make each program do one thing well. To do a new job, build afresh rather than complicate old programs by adding new “features”.
> Expect the output of every program to become the input to another, as yet unknown, program. Don’t clutter output with extraneous information. Avoid stringently columnar or binary input formats. Don’t insist on interactive input.
> Design and build software, even operating systems, to be tried early, ideally within weeks. Don’t hesitate to throw away the clumsy parts and rebuild them.
> Use tools in preference to unskilled help to lighten a programming task, even if you have to detour to build the tools and expect to throw some of them out after you’ve finished using them.
I'm not too convinced on the wrapping of code, though.
However you're definitely right about dropping the max-width property.
What would solve most of the problems is HN actually implementing markdown instead of the current half-assed crap.
- Slack product development
(that was a joke but is likely the answer to your question)
But I think the very limited formatting is just fine anyway. For the above comment as an example, I agree the code formatting looks awful, especially on mobile. But the version with >'s is ok, and I don't think proper bullet points or a quote bar would have improved it dramatically.
[0]: Actually not Markdown but a subset but it's not important.
It's way more restful purveying a page of uniformly restrained text.
HN already has shitty italics (shitty in that it commonly matches and eats things you don't want to be italicised e.g. multiplications, pointers, … in part though not only because HN doesn't have inline code). "bold" can just be styled as italics, or it can be styled as a medium or semibold. It's not an issue, and even less worth it given how absolute garbage the current markup situation is.
Just give me the triple-tilde code block syntax please!
Meh. It does literal code blocks, they work fine.
That's pretty much the only markup feature which does, which is impressively bad given HN only has two markup feature: literal code blocks and emphasis.
It's not like they're going to add code coloration or anything.
And while fenced code blocks are slightly more convenient (no need to indent), pasting a snippet in a text editor and indenting it is hardly a difficult task.
TaoUP has a longer discussion[1] of the Unix philosophy, which includes Rob Pike's and Ken Thompson's comments on the philosophy.
[1] http://www.catb.org/esr/writings/taoup/html/ch01s06.html
"Those who don't understand Unix are condemned to reinvent it, poorly." (Henry Spencer)
Could you recommend some "original sources" to learn from, instead? Ideally in book form?
I don't know if it's possible to have impartially "fair" discussion of editors. Skimming now, I can see how vi lovers would hate some characterizations there. But it does try to learn interesting lessons from them.
It does NOT simply equate "Emacs has UNIX nature" so you can't just prove something like "TAOUP mentions Emacs, Emacs is GNU, Gnu is Not Unix => TAOUP is not UNIX, QED" ;-)
http://www.catb.org/esr/writings/taoup/html/ch13s02.html
bias disclaimers: I learnt most of what I know of unix from within Emacs, which I still use ~20 years later. I learnt more from Info pages than man pages (AIX had pretty bad man pages). I suspect you have a different picture of unix than I. And I now know better than arguing which editor is better ;)
But I found TAOUP articulated ideas I only learnt through osmosis. I'm looking forward to reading a better articulation if you know one.
Because I'd consider the latter to be far worse than the former.
Powershell pipes are an extension over Unix pipes. Rather than just being able to pipe a stream of bytes, powershell can pipe a stream of objects.
It makes working with pipes so much fun. In Unix you have to cut, awk and do all sorts of parsing to get some field out of `ls`. In poweshell, ls outputs stream of file objects and you can get the field you want by piping to `Get-Item` or sum the file sizes, or filter only directories. It’s very expressive once you’re manipulating streams of objects with properties.
I'm guessing you've mentioned using `ls` as a simple, first-thing-that-comes-to-mind example, which is cool. I just wanted to point out that if a person is piping ls's output, there are probably other, far better alternatives, such as (per the link below) `find` and globs:
https://unix.stackexchange.com/a/247973
That has been my experience, at least.
Or are the resulting chain of pipes so short lived, it doesn't matter?
I’m not saying that this makes powershell pipes better/worse, just that this problem isn’t unique. Microsoft tends to be reasonably committed to backwards compatibility but I don’t know the answer to the question
> get some field out of `ls`
If you're parsing the output of ls, you're doing it wrong. Filenames can contain characters such as newlines, etc, and so you generally want to use globbing/other shell builtins, or utils like find with null as the separator instead of newline.
I have heard this several times, but either I do not understand it or I disagree. Do you mean parsing the output of the ls program? Parsing ls output is not wrong, the program produces a text stream that is easy and useful to parse. There's nothing to be ashamed when doing it, even when you can do it in a different, even shorter way. I do grep over ls output daily, and I find it much more convenient than writing wildcards.
For given paths, stat command essentially lets us directly access their inode structs, and invocations are nicely concise. The find util then lets us select files based on inode fields.
Both tools do take a bit of learning, but considerably less than grep and regexs. Anyway, I've personally found find and stat to be really nice and ergonomic after the initial learning period.
Unix pipes can pipe objects, binary data, text, json encoded data, etc.
The problem is it adds a lot of complication and simply doesn't offer much in practice, so text is still widely used and binary data in a few specific cases.
Disclaimer: I do cat-piping myself quite a bit out of habit, so I'm not trying to look down at the author or anything like that! :)
I had recently built a set of tools used primarily via pipes: (tool-a | tool-b | tool-c) and it looks clearer when I mock (for testing) one command (cat results | tool-b | tool-c) instead of re-flowing it just to avoid cat and use direct files.
<file command | command | command
is perfectly fine.Does anyone have a better way to do this kind of thing?
[1] https://core.tcl.tk/expect/index [2] https://pexpect.readthedocs.io/en/stable/
i.e.
prefer `apt-get install -y` over `yes | apt-get install foo`
Instead, shell script should be optimized for readability and portability and I think it is much easier to understand something like 'read | change >write' than 'change <read >write'. So I like to write pipelines like this:
cat foo.txt \
| grep '^x' \
| sed 's/a/b/g' \
| awk '{print $2}' \
| wc -l >bar.txt
It might be not the most efficient processing method, but I think it is quite readable.For those who disagree with me: You might find the pure-bash-bible [1] valuable. While I admire their passion for shell scripts, I think they are optimizing to the wrong end. I would be more a fan of something along the lines of 'readable-POSIX-shell-bible' ;-)
In any case, it's probably just a matter of personal taste.
It is instantly plainly obvious to me what each step of their shell script is doing.
While I can absolutely understand what your shell script does after parsing it, it's meaning doesn't leap out at me in the same way.
I would describe the prior shell script as more quickly readable than the one that you've listed.
So, perhaps it's not a question of one being more silly than the other—perhaps the author just has different priorities from you?
That should be gsub, shouldn't it? (sub only replaces the first occurrence)
grep '^x' < input | sed 's/foo/bar/g'
to be very readable, as the flow is still visually apparent based on punctuation. cat input | grep '^x' | sed 's/foo/bar/g'
Is far more readable, in my opinion. In addition, it makes it trivial to change the input from a file to any kind of process.I'm STRONGLY in favor of using "cat" for input. That "useless use of cat" article is pretty dumb, IMHO.
OK, to save some of my face, this will work:
grep 'foo' <(input) | sed 's/baz/bar/g'
... at least in zsh and probably bash. input | grep foo | sed ... diff <(prog1) <(prog2)
and get a sensible result.And sometimes programs just refuse to read from stdin but do just fine with an unseekable file on the command line. True, you do have this:
input | recalcitrant_program /dev/stdin
... but it's a bit of a tossup as to which one's more readable at this point. They're both relying on advanced shell functionality.> diff <(prog1) <(prog2)
> and get a sensible result.
That is called process substitution and is exactly the kind of use case that it's designed for. So yes, process substitution does make sense there.
> input | recalcitrant_program /dev/stdin
> ... but it's a bit of a tossup as to which one's more readable at this point. They're both relying on advanced shell functionality.
There's no tossup at all. Process substitution is easily more readable than your second example because you're honouring the normal syntax of that particular command's parameters rather than kludging around it's lack of STDIN support.
Also I wouldn't say either example is using advanced shell functionalities either. Process substitution (your first example) is a pretty easy thing to learn and your second example is just using regular anonymous pipes (/dev/stdin isn't a shell function, it's a proper pseudo-device like /dev/random and /dev/null) thus the only thing the shell is doing is the same pipe described in this threads article (with UNIX / Linux then doing the clever stuff outside of the shell).
cat input | grep '^x' | sed 's/foo/bar/g'
→ sed '/^x/s/foo/bar/g' <input sed -n '/^x/s/foo/bar/gp' <input
This may be an inadvertent argument for the ‘connect simpler tools’ philosophy.Now back to the topic of "cat", which is a great example of why shell scripts are minefields.
Replace "foo.txt" with a user supplied variable, let's call it "$F". It becomes cat $F | blah_blah... I mean cat "$F" | blah_blah, first trap, but everyone knows that.
Now, if F='-n', second trap. What you think is a file will be considered an option and cat will wait for user input, like when no file is given. Ok, so you need to do cat -- "$F" | blah_blah.
That should be OK in every case now, but remember that "cat" is just another executable, or maybe a builtin. For some reason, on your system "cat --" may not work, or some asshat may have added "." in your PATH and you may be in a directory with a file named "cat". Or maybe some alias that decides to add color.
There are other things to consider, like your locale that may mess up you output with comas instead of decimal points and unicode characters. For that reason, you need to be very careful every time you call a command and even more so if you pipe the output.
For that reason, I avoid using "cat" in scripts. It is an extra command call and all the associated headaches I can do without.
You're not wrong, but I think it's worth pointing out that's a trap that comes up any time you exec another program, whether it's from shell or python. I can't reasonably expect `subprocess.run(["cat", somevar])` to work if `somevar = "-n"`.
(Now, obviously, I'm not going to "cat" from python, but I might "kubectl" or something else that requires care around the arguments)
I think that you forgot to edit the "I mean" to "echo $F" :)
I periodically think it would be a good idea to organize a language around.
For example, you can write this pipeline as:
grep '^x' foo.txt \
| sed 's/a/b/g' \
| awk '{print $2}' \
| wc -l > bar.txt
This is by no means scientific, but I've got a LaTeX document open right now. A quick `time` says: $ time grep 'what' AoC.tex
real 0m0.045s
user 0m0.000s
sys 0m0.000s
$ time cat AoC.tex | grep what
real 0m0.092s
user 0m0.000s
sys 0m0.047s
Anecdotally, I've witnessed small pipelines that absolutely make sense totally thrash a system because of inappropriate uses of `cat`. When you `cat` a file, the OS must (1) `fork` and `exec`, (2) copy the file to `cat`'s memory, (3) copy the contents of `cat`'s memory to the pipe, and (4) copy the contents of the pipe to `grep`'s memory. That's a whole lot of copying for large files -- especially when the first command grep in the sequence usually performs some major kind of reduction on the input data!That said, I suspect the example would be much faster if you didn't use the pipeline, because a single tool could do it all (I'm leaving in the substitution and column print that are actually unused in the result):
awk '/^x/{gsub("a","b");print $2; count++}END{print NR}' foo.txtDoesn't gsub(/a/, "b") do the same thing as s/a/b/g?
Of course, the example @arendtio uses is correct, because they obviously care about such things.
(Similarly, the first thing I used to do on Windows was set my prompt to [$p] because many years ago I also accidentally nuked a part of Visual Studio when I copied and pasted a command line that was prefixed with "C:\...>". Whoops.)
$ less foo | bar
Is similar too: $ bar < foo
Except that less is typically more clever than that and might be more like: $ zcat foo | bar
Depending on the file type of foo.It's an abstraction that held up quite well, but its starting to show its age.
However, how can you ensure the output type of one program matches the input type of another?
A program emits one type, and the other program accepts another.
Something will be needed to transform one type into another. Imagine doing that on the command line.
Powershell's probably the closest to success, because it could be pushed out unilaterally. Without that I'm not sure it would have gotten very far, not because it's bad, but again because nobody else seems to have gotten very far....
Some of that discussion has not aged well :)
>The receiving and sending processes must use a stream of bytes. Any object more complex than a byte cannot be sent until the object is first transmuted into a string of bytes that the receiving end knows how to reassemble. This means that you can’t send an object and the code for the class definition necessary to implement the object. You can’t send pointers into another process’s address space. You can’t send file handles or tcp connections or permissions to access particular files or resources.
Thank goodness.
Thankfully switching shells is as painless as switching text editors.
So, somewhere between, "That wasn't as bad as I feared," and, "Sweet Jesus, what fresh new hell have I found myself in"?
Im working on a library for this kind of intra-pipeline negitiation. It's all drawing-board stuff right now but I coobbled together a proof of concept:
https://unix.stackexchange.com/a/495338/146169
Do you think this is a reasonable way to achieve the magic that some users want in their pipelines? Or are ancient Unix gods going to smite me for tampering with the functional consistency of tools by making their behavior different in different contexts?
I was imagining an algorithm where each pipeline-aware utility can derive port numbers to use to talk/listen to its neighbors. I may be able to use http content negotiation wholesale in that context.
This reminds me of how busted Linux is for not having a SO_PEERCRED. You can actually get that information by walking /proc/net/tcp or using AF_NETLINK sockets and inet_diag, but there is a race condition such that this isn't 100% reliable. SO_PEERCRED would [have to] be.
However the problem I face is how do you pass that data type information over a pipeline from tools that exist outside of my shell? It's all well and good having builtins that all follow that convention but what if someone else wants to write a tool?
My first thought was to use network sockets, but then you break piping over SSH, eg:
local-command | ssh user@host "| remote-command"
My next thought was maybe this data should be in-lined - a bit like how ANSI escape sequences are in-lined and the terminals don't render them as printable characters. Maybe something like the following as a prefix to STDIN? <null>$SHELL<null>
But then you have the problem of tainting your data if any tools are sent that prefix in error.I also wondered if setting environmental variables might work but that also wouldn't be reliable for SSH connections.
So as you can see, I'm yet to think up a robust way of achieving this goal. However in the case of builtin tools and shell scripts, I've got it working for the most part. A few bugs here and there but it's not a small project I've taken on.
If you fancy comparing notes on this further, I'm happy to oblige. I'm still hopeful we can find a suitable workaround to the problems described above.
I was hoping to stick with bash or zsh, and just write processes that somehow communicate out of band, but I think we're still up against the same problem.
One idea I had was that there's a service running elsewhere which maintains this directed graph (nodes = types, edges = programs which take the type of their "from" node and return the type of their "two" node). When a pipeline is executed, each stage pauses until type matches are confirmed--and if there is a mismatch then some path finding algorithm is used to find the missing hops.
So the user can leave out otherwise necessary steps, and as long as there is only one path through the type graph which connects them, then the missing step can be "inserted". In the case of multiple paths, the error message can be quite friendly.
This means keeping your context small enough, and your types diverse enough, that the type graph isn't too heavily connected. (Maybe you'd have to swap out contexts to keep the noise down.) But if you have a layer that's modifying things before execution anyway, then perhaps you can have it notice the ssh call and modify it to set up a listener. Something like:
User Types:
local-command | ssh user@host "remote-command"
Shell runs: local-command | ssh user@host "pull_metadata_from -r <caller's ip> | remote-command"
Where pull_metadata_from phones home to get the metadata, then passes along the data stream untouched.Also, If you're writing the shell anyway then you can have the pipeline run each process in a subshell where vars like TYPE_REGISTRY_IP and METADATA_INBOUND_PORT are defined. If they're using the network to type-negotiate locally, then why not also use the network to type-negotiate through an ssh tunnel?
This idea is, of course, over-engineered as hell. But then again this whole pursuit is.
Yeah we have different starting points but very much similar problems.
tbh idea behind my shell wasn't originally to address typed pipelines, that was just something that evolved from it quite by accident.
Anyhow, your suggestion of overwriting / aliasing `ssh` is genius. Though I'm thinking rather than tunnelling a TCP connection, I could just spawn an instance of my shell on the remote server and then do everything through normal pipelines as I now control both ends of the pipe. It's arguably got less proverbial moving parts compared to a TCP listener (which might then require a central data type daemon et al) and I'd need my software running on the remote server for the data types to work anyway.
There is obviously a fair security concern some people might have about that but if we're open and honest about that and offer an "opt in/out" where opting out would disable support for piped types over SSH then I can't see people having an issue with it.
Coincidentally I used to do something similar in a previous job where I had a pretty feature rich .bashrc and no Puppet. So `ssh` was overwritten with a bash function to copy my .bashrc onto the remote box before starting the remote shell.
> This idea is, of course, over-engineered as hell. But then again this whole pursuit is.
Haha so true!
Thanks for your help. You may have just solved a problem I've been grappling with for over a year.
You're right that it is in fact possible for a command to find the preceding and following commands using /proc, and figure out what content types they produce / want, and do something sensible. But there won't always be just one way to convert between content types...
Me? I don't care for this kind of magic, except as a challenge! But others might like it. You might need to make a library out of this because when you have something like curl(1) as a data source, you need to know what Content-Type it is producing, and when you can know explicitly rather than having to taste the data, that's a plus. Dealing with curl(1) as a sink and somehow telling it what the content type is would be nice as well.
I notice that function composition notation; that is, the latter half of:
> f(g(x)) = (f o g)(x)
resembles bash pipeline syntax to a certain degree. The 'o' symbol can be taken to mean "following". If we introduce new notation where '|' means "followed by" then we can flip the whole thing around and get:
> f(g(x)) = (f o g)(x) = echo 'x' | g | f
I want to write some set of mathematically interesting functions so that they're incredibly friendly (like, they'll find and fix type mismatch errors where possible, and fail in very friendly ways when not). And then use the resulting environment to teach a course that would be a simultaneous intro into both category theory and UNIX.
All that to say--I agree about finding the magic a little distasteful, but if I play my cards right my students will only realize there was magic in play after they've taken the bait. At first it will all seem so easy...
At some point you just want a Haskell shell (there is one). Or a jq shell (there is something like it too).
As to the pipe symbol as function composition: yes, that's quite right.
> This means that you can’t send an object and the code for the class definition necessary to implement the object.
To some degree you can and I do just this with my own shell I've written. You just have to ensure that both ends of the pipe understands what is being sent (eg is it JSON, text, binary data, etc)? Even with typed terminals (such as Powershell), you still need both ends of the pipe to understand what to expect to some extent.
Having this whole thing happen automatically with a class definition is a little optimistic though. Not least of all because not every tool would be suited for every data format (eg a text processor wouldn't be able to do much with a GIF even if it has a class definition).
> You can’t send pointers into another process’s address space.
Good job too. That seems a very easy path for exploit. Thankfully these days it's less of an issue though because copying memory is comparatively quick and cheap compared to when that handbook was written.
> You can’t send file handles
Actually that's exactly how piping works as technically the standard streams are just files. So you could launch a program with STDIN being a different file from the previous processes STDOUT.
> or tcp connections
You can if you pass it as a UNIX socket (where you define a network connection as a file).
> or permissions to access particular files or resources.
This is a little ambiguous. For example you can pass strings that are credentials. However you cannot alter the running state of another program via it's pipeline (aside what files it has access to). To be honest I prefer the `sudo` type approach but I don't know how much of that is because it's better and how much of that is because it's what I am used to.
> Actually that's exactly how piping works
Also SCM_RIGHTS, which exists exactly for this purpose (see cmsg(3), unix(7) or https://blog.cloudflare.com/know-your-scm_rights/ for a gentler introduction and application).
That's been around since BSD 4.3, which predates the Hater's Handbook 1ed by 4 years or so.
MacOs wasn't always layered on unix, and the unix-haters' handbook predates the switch to the unix-based MacOs X.
Not to put too fine a point on it, but they found religion. Unlike Classic (and early versions of Windows for that matter), there was more to be gained by ceding some control to the broader community. Microsoft has gotten better (PowerShell - adapting UNIX tools to Windows, and later WSL, where they went all in)
Still, for Apple it meant they had to serve two masters for a while - old school Classic enthusiasts and UNIX nerds. Reading the back catalog of John Siracusa's (one of my personal nerd heroes) old macOS reviews gives you some sense of just how weird this transition was.
find . -name '*.el' | while read el; do [ -f "${el}c" ] || echo $el; done
Of course this uses the dreaded pipes and doesn't support the extremely common filenames with a newline in them, so let's do it without them find . -name '*.el' -exec bash -c 'el=$0; elc="${el}c"; [ -f "$elc" ] || echo "$el"' '{}' ';'The following works very nicely for me:
* find . -name '*.el' -exec file {}c ';' 2>&1 | grep cannot(havent tried above, not sure I recommend that you do)
This makes /dev/null a valid source file in many languages, C included.
You could write a executable that accepts piped input and throws it away.
When would it exit though? Would it exit successfully at the end of the input stream? That sounds sensible.
That would be behaving like an executable wouldn't it?
A process that attempts to read from a closed pipe receives SIGPIPE. The default disposition for SIGPIPE is to terminate the program (similar to SIGTERM or SIGINT). So yeah, assuming that the previous program in the pipeline closes its stdout at some point (either explicitly, or implicitly by just exiting), then our program would die of SIGPIPE the when it tries to read() from stdin and the pipe's buffer has been depleted.
However, our program could also set SIGPIPE to ignored and ignore the EPIPE errors that read() would return in that case. In that case, it could run indefinitely. But at this point, you're way past normal behavior.
Cf http://trillian.mit.edu/~jc/humor/ATT_Copyright_true.html or https://twitter.com/rob_pike/status/966896123548872705
https://bugs.debian.org/919341
However, executing /dev/null doesn't seem to work on Linux:
$ sudo chmod 777 /dev/null
$ /dev/null
bash: /dev/null: Permission deniedYou could direct, or redirect the flow to /dev/null. Or pipe to /dev/null. Or redirect the pipe to /dev/null?
So from a metaphor point of view either would fit.
Although of course you don't use the pipe construct to direct to a file. Which would suggest piping is wrong?
And then on the third hand, we all know what it means so what's the problem.
So I would say theres war, famine and injustice in the world. Don't worry about posix shell semantics. :)
Sometimes they work great -- being able to dump from MySQL into gzip sending across the wire via ssh into gunzip and into my local MySQL without ever touching a file feels nothing short of magic... although the command/incantation to do so took quite a while to finally get right.
But far too often they inexplicably fail. For example, I had an issue last year where piping curl to bunzip would just inexplicably stop after about 1GB, but it was at a different exact spot every time (between 1GB and 1.5GB). No error message, no exit, my network connection is fine, just an infinite timeout. (While curl by itself worked flawlessly every time.)
And I've got another 10 stories like this (I do a lot of data processing). Any given combination of pipe tools, there's a kind of random chance they'll actually work in the end or not. And even more frustrating, they'll often work on your local machine but not on your server, or vice-versa. And I'm just running basic commodity macOS locally and out-of-the-box Ubuntu on my servers.
I don't know why, but many times I've had to rewrite a piped command as streams in a Python script to get it to work reliably.
Because why would it? If the UNIX philosophy is to separate out tools and pipe them, then the UNIX philosophy should be to pipe through gzip and gunzip, not for ssh to provide its own redundant compression option, right?
While this may be your experience, the mechanism of FIFO pipes in Unix (which is filehandles and buffers, basically), is an old one that is both elegant and robust; it doesn't "randomly" fail due to unreliability of the core algorithm or components. In 20 years, I never had an init script or bash command fail due to the pipe(3) call itself being unreliable.
If you misunderstand detailed behavior of the commands you are stitching together--or details of how you're transiting the network in case of an ssh remote command--then yes, things may go wrong. Especially if you are creating Hail Mary one-liners, which become unwieldy.
One issue I did used to have (before I discovered ‘-o pipefail’[1]) was the annoyance that if an earlier command in a pipeline failed, all the other commands in the pipeline still ran albeit with no data or garbage data being piped to them.
[1] https://stackoverflow.com/questions/1550933/catching-error-c...
It essentially hijacks pipe's input and output into browser where you can play with the Ramda command. Then you just close browser tab and Ramda CLI applies your changed code in the pipe, resuming its operation.
Now I'm thinking all kinds ways I use pipe that I could "tee" through a browser app. I can use browser for interactive JSON manipulation, visualization and all around playing. I'm now looking for ways to generalize Ramda CLI's approach. Pipes, Unix files and HTTP don't seem directly compatible, but the promise is there. Unix tee command doesn't "pause" the pipe, but probably one could just introduce pause/resume output passthrough command into the pipe after it. Then web server tool can send the tee'd file to browser and catch output from there.
https://github.com/mpw/MPWFoundation/blob/master/Documentati...
One aspect is that the coordinating entity hooks up the pipeline and then gets out of the way, the pieces communicate amongst themselves, unlike FP simulations, which tend to have to come back to the coordinator.
This is very useful in "scripted-components" settings where you use a flexible/dynamic/slow scripting language to orchestrate fixed/fast components, without the slowness of the scripting language getting in the way. See sh :-)
Another aspect is error handling. Since results are actively passed on to the next filter, the error case is simply to not pass anything. Therefore the "happy path" simply doesn't have to deal with error cases at all, and you can deal with errors separately.
In call/return architectures (so: mostly everything), you have to return something, even in the error case. So we have nil, Maybe, Either, tuples or exceptions to get us out of Dodge. None of these is particularly good.
And of course | is such a perfect combinator because it is so sparse. It is obvious what each end does, all the components are forced to be uniform and at least syntactically composable/compatible.
Yay pipes.
* vipe (part of https://joeyh.name/code/moreutils/ - lets you edit text part way through a complex series of piped commands)
* pv (http://www.ivarch.com/programs/pv.shtml - lets you visualise the flow of data through a pipe)
* pee: tee standard input to pipes (`pee "some-command" "another-command"`)
* sponge: soak up standard input and write to a file
Though in zsh and bash you can create pee using tee: `tee >(some-command) >(another-command) >/dev/(null`
https://github.com/ghemawat/stream
Edit: Jeff Dean, not James Dean
I would have loved to see some awesome pipe examples though.
A simple virus scanner in one line of pipe:
https://everythingsysadmin.com/2004/10/whos-infected.html
And a bunch of pipe tricks that are oh so wrong but oh so useful:
And many programming languages have libraries for something similar:. iterators in rust / c++, streams in java/c#, thinks like reactive
Iterators don't really fully capture what a pipe is though? Theres no parallelism.
And streams don't have the conceptual simplicity of a pipe?
Pipes are concurrent, not necessarily parallel. Iterators are concurrent, and can be parallel (https://docs.rs/rayon/0.6.0/rayon/par_iter/index.html).
f |> g(1)
would be equivalent to g(f, 1) [1, 2, 3].map(n => n + 1).join(',').length
Is basically like this shell command: seq 3 | awk '{ print $1 + 1 }' | tr '\n' , | wc -c
(the shell version gives 6 instead of 5 because of a trailing newline, but close enough)The Python:
length(','.join(map(lambda n: n+1, range(1, 4)))
is a bit closer, but the order's now reversed, and then jumbled by the map/lambda. (Though I suppose arguably awk does that too.)Nim has an interesting synthesis where a.f(b) is only another way to spell f(a, b), which (I think) matches the usual behavior of |> while still allowing familiar-looking method-style syntax. These are equivalent:
[1, 2, 3].map(proc (n: int): int = n + 1).map(proc (n: int): string = $n).join(",").len
len(join(map(map([1, 2, 3], proc (n: int): int = n + 1), proc (n: int): string = $n), ","))
The difference is purely cosmetic, but readability matters. It's easier to read from left to right than to have to jump around.1:3 |> x->x+1 |> x->join(x,",") |> length
1:3 |> x -> map(y -> y+1, x) |> x -> join(x, ",") |> length
All those anonymous functions seem a bit distracting, though https://github.com/JuliaLang/julia/pull/24990 could help with that. (1:3 .|> x->x+1) |> x->join(x,",") |> length1:3 |> x->x.+1 |> x->join(x,",") |> length
Note the dot in x.+1, that tells + to operate on each element of the array (x), and not the array itself.
list.select { |x| x.foo > 10 }.map { |x| x.bar }...
Etc.I wouldn't make the same argument about the IO monad, which I think more in terms of a functional program which evaluates to an imperative program. But most monads are not like the IO monad, in my experience at least.
Forgive me if I'm misreading this syntax, but to me this looks like plain old function composition: a call to `select` (I assume that's like Haskell's `filter`?) composed with a call to `map`. No monad in sight.
As I mentioned, monads are more about collapsing structure. In the case of lists this could be done with `concat` (which is the list implementation of monad's `join`) or `concatMap` (which is the list implementation of monad's `bind` AKA `>>=`).
And monads are not "more about collapsing structure". They are just a design pattern that follows a handful of laws. It seems like you're mistaking their usefulness in Haskell for what they are. A lot of other languages have monads either baked in or an element of the design of libraries. Expand your mind out of the Haskell box :)
> list.select { |x| x.foo > 10 }.map { |x| x.bar }
Where's `bind` or `join` in that example?
Thanks for clarifying; I've read a bunch of Ruby but never written it before ;)
From a quick Google I see that "select" and "map" do work as I thought:
https://ruby-doc.org/core-2.2.0/Array.html#method-i-select
https://ruby-doc.org/core-2.2.0/Array.html#method-i-map
So we have a value called "list", we're calling its "select" method/function and then calling the "map" method/function of that result. That's just function composition; no monads in sight!
To clarify, we can rewrite your example in the following way:
list.select { |x| x.foo > 10 }.map { |x| x.bar }
# Define the anonymous functions/blocks elsewhere, for clarity
list.select(checkFoo).map(getBar)
# Turn methods into standalone functions
map(select(list, checkFoo), getBar)
# Swap argument positions
map(getBar, select(checkFoo, list))
# Curry "map" and "select"
map(getBar)(select(checkFoo)(list))
# Pull out definitions, for clarity
mapper = map(getBar)
selector = select(checkFoo)
mapper(selector(list))
This is function composition, which we could write: go = compose(mapper, selector)
go(list)
The above argument is based solely on the structure of the code: it's function composition, regardless of whether we're using "map" and "select", or "plus" and "multiply", or any other functions.To understand why "map" and "select" don't need monads, see below.
> the list could be an eager iterator, an actual list, a lazy iterator, a Maybe (though it would be clumsy in Ruby), etc.
Yes, that's because all of those things are functors (so we can "map" them) and collections (so we can "select" AKA filter them).
The interface for monad requires a "wrap" method (AKA "return"), which takes a single value and 'wraps it up' (e.g. for lists we return a single-element list). It also requires either a "bind" method ("concatMap" for lists) or, my preference, a "join" method ("concat" for lists).
I can show that your example doesn't involve any monads by defining another type which is not a monad, yet will still work with your example.
I'll call this type a "TaggedList", and it's a pair containing a single value of one type and a list of values of another type. We can implement "map" and "select" by applying them to the list; the single value just gets passed along unchanged. This obeys the functor laws (I encourage you to check this!), and whilst I don't know of any "select laws" I think we can say it behaves in a reasonable way.
In Haskell we'd write something like this (although Haskell uses different names, like "fmap" and "filter"):
data TaggedList t1 t2 = T t1 [t2]
instance Functor (TaggedList t1) where
map f (T x ys) = T x (map f ys)
instance Collection (TaggedList t1) where
select f (T x ys) = T x (select f ys)
In Ruby we'd write something like: class TaggedList
def initialize(x, ys)
@x = x
@ys = ys
end
def map(f)
TaggedList.new(@x, @ys.map(f))
end
def select(f)
TaggedList.new(@x, @ys.select(f))
end
end
This type will work for your example, e.g. (in pseudo-Ruby, since I'm not so familiar with it): myTaggedList = TaggedList.new("hello", [{foo: 1, bar: true}, {foo: 20, bar: false}])
result = myTaggedList.select { |x| x.foo > 10 }.map { |x| x.bar }
# This check will return true
result == TaggedList.new("hello", [false])
Yet "TaggedList" cannot be a monad! The reason is simple: there's no way for the "wrap" function (AKA "return") to know which value to pick for "@x"!We could write a function which took two arguments, used one for "@x" and wrapped the other in a list for "@ys", but that's not what the monad interface requires.
Since Ruby's dynamically typed (AKA "unityped") we could write a function which picked a default value for "@x", like "nil"; yet that would break the monad laws. Specifically:
bind(m, wrap) == m
If "wrap" used a default value like "nil", then "bind(m, wrap)" would replace the "@x" value in "m" with "nil", and this would break the equation in almost all cases (i.e. except when "m" already contained "nil").It actually somewhat changes the way you write code, because it enables chaining of calls.
It's worth noting there's nothing preventing this being done before the pipe operator using function calls.
x |> f |> g is by definition the same as (g (f x)).
In non-performance-sensitive code, I've found that what would be quite a complicated monolithic function in an imperative language often ends up as a composition of more modular functions piped together. As others have mentioned, there are similarities with the method chaining style in OO languages.
Also, I believe Clojure has piping in the form of the -> thread-first macro.
[0] https://caml.inria.fr/pub/docs/manual-ocaml/libref/Pervasive...
So lazy collections / iterators, and HoFs working on those.
The pipe operator itself is mostly a way to denote the composition in reading order (left to right instead of right to left / inside to outside), which is convenient for readability but not exactly world-breaking.
I only have experience using RxJS, but it's incredibly powerful.
julia> 1:5 |> x->x.^2 |> x->2x
5-element Array{Int64,1}:
2
8
18
32
50This example has a nested hash map where we try to get the "You got me!" string.
We can either use `:a` (keyword) as a function to get the value. Then we have to nest the function calls a bit unnaturally.
Or we can use the thread-first macro `->`, which is basically a unix pipe.
user=> (def res {:a {:b {:c "You got me!"}}})
#'user/res
user=> res
{:a {:b {:c "You got me!"}}}
user=> (:c (:b (:a res)))
"You got me!"
user=> (-> res :a :b :c)
"You got me!"
Thinking about it, Clojure advocates having small functions (similar to unix's "small programs / do one thing well") that you compose together to build bigger things.https://github.com/linpengcheng/PurefunctionPipelineDataflow
Nowadays, most of the user-facing desktop programs have GUIs, so the 'pipe' operator that composes programs is the user himself. Users compose programs by saving files from one program and opening them in another. The data being 'piped' through such program composition is sort-of typed, with the file types (PNG, TXT, etc) being the types and the loading modules of the programs being 'runtime typecheckers' that reject files with invalid format.
On the first sight, GUIs prevent program composition by requiring the user to serve as the 'pipe'. However, if GUIs were reflections / manifestations of some rich typed data (expressible in some really powerful type system, such as that of Idris), one could imagine the possibility of directly composing the programs together, bypassing the GUI or file-saving stages.
the pipes in your typical functional language (`|>`) is not a form of function composition, like
```
f >> g === x -> g(f(x))
```
but function application, like
```
f x |> g === g(f(x))
x |> f |> g // also works, has the same meaning
f |> g // just doesn't work, sorry :(
```
What is a "typical functional language" in this case? I don't think I've come across this `|>` notation, or anything explicitly referred to as a "pipe", in the functional languages I tend to use (Haskell, Scheme, StandardML, Idris, Coq, Agda, ...); other than the Haskell "pipes" library, which I think is more elaborate than what you're talking about.
A small additional bit of information: this style is called "tacit" (or "point-free") programming. See https://en.wikipedia.org/wiki/Tacit_programming
(Unix pipes are even explicitly mentioned in the articles as an example)
The shell have the weird(?) behavior of ONE input and TWO outputs (stdout, stderr).
Also, can redirect in both directions. I think a language to be alike pipes, it need each function to be alike:
fun open(...) -> Result(Ok,Err)
and have the option of not only chain the OK side but the ERR: open("file.txt") |> print !!> raise |> print
exist something like this??? cat /etc/passwd | grep root | awk -F: '{print $3}'
ruby -e 'puts File.read("/etc/passwd").lines.select { |line| line.match(/root/) }.first.split(":")[2]'
A little more verbose, but the idea is the same.https://alvinalexander.com/scala/fp-book/how-functional-prog...
Edit: Also "column" to format your output into a table.
Edit: s/files/fields/
I have seen plenty pipe use in bash scripts, especially for build and ETL purposes. Other languages, have ways to do pipes as well (e.g. Python https://docs.python.org/2/library/subprocess.html#replacing-...) but I have seen much less use of it.
It appears to me, that for more complex applications one rather opts for TCP-/UDP-/UNIX-Domain sockets for IPC.
- Has anyone here tried to plumb together large applications with pipes?
- Was it successful?
- Which problems did you run into?
Some functional programming styles are pipe-like in the sense that data-flow is unidirectional:
Foo(Bar(Baz(Bif(x))))
is analagous to: cat x | Bif| Baz |Bar| Foo
Obviously the order of evaluation will depend on the semantics of the language used; most eager languages will fully evaluate each step before the next. (Actually this is one issue with Unix pipes; the flow-control semantics are tied to the concept of blocking I/O using a fixed-size buffer)The idea of dataflow programming[1] is closely related to pipes and has existed for a long time, but it has mostly remained a niche, at least outside of hardware-design languages
I got the whole thing working from training to prediction in about 3 weeks. What I love about Unix shell commands is that you simply can't abstract beyond the input/output paradigm. You aren't going to create classes, types classes, tests, etc. It's not possible or not worth it.
I'd like to see more devs use this approach, because it's a really nice way to get a project going in order to poke holes in it or see a general structure. I consider it a sketchpad of sorts.
If you write them cleanly they don’t suck and crucially for me, bash today works basically the same way as it did 10 years ago and likely in 10 years, that static nature is a big damn win.
I sometimes wish language vendors would just say ‘this language is complete, all future work will be bug fixes and libraries’ a static target for anything would be nice.
Elixir did say that recently except for one last major change which moved it straight up my list of things to look at in future.
- My mail setup uses NMH, everything is automated. I can mime-encode a directory and send the resulting mail in a breeze.
- GF's photos from IG are being backed up with a python script and crontab. Non IG ones are geotagged too with a script. I just fire up some cli GPS tools if we hike some mountain route, and gpscorrelate runs on the GPX file.
- Music is almost everything chiptunes, I felt interesent on any mainstream music since 2003-4. I mirror a site with wget and it's done. If they offered rsync...
- Hell, even my podcasts are being fetch via cron(8).
- My setup is CWM/cli based, except for mpv, emulators, links+ and vimb for JS needed sites. Noice is my fm, or the pure shell. find(1) and mpg123/xmp generate my music playlist. Street View is maybe the only service I use on vimb...
The more you automate, the less tasks you need to do. I am starting to avoid even taskwarrior/timew, because I am almost task free as I don't have to track a trivial <5m script, and spt https://github.com/pickfire/spt is everything I need.
Also, now I can't stand any classical desktop, I find bloat on everything.
Transducers are compostable algorithmic transformations that, in a way, generalize map, filter and friends but they can be used anywhere where you transform data. Transducers have to be invoked and handled in a way that pipes do not.
Anyone interested should check out Hickeys talks about them. They are generally a lot more efficient than chaining higher order list processing functions and since they don't build.intermediate results they have a lot better GC performance.
Streams in most decent languages closely adhere to this idea.
I especially like how node does it, in my opinion one of the best things in node. Where you can simply create cli programs that have backpressure the same as you would work with binary/file streams, while also supporting object streams.
process.stdin.pipe(byline()).pipe(through2(transformFunction)).pipe(process.stdout)One criticism would be that ggplot2 uses the "+" to add more graph features, whereas the rest of tidyverse uses "%>%" as its pipe, when ideally ggplot2 would also use it. One of my most common errors with ggplot2 is not utilizing the + or the %>% in the right places.
Of course, Hadley admitted it was because he wrote ggplot2 before adopting pipes into his packages.
This is, in fact, a Useless Use of Cat [1]. POSIX shells have the < operator for directing a single file to stdin:
figlet <file
[1] http://porkmail.org/era/unix/award.htmlLike always, simple stuff will be simple (http://write.flossmanuals.net/pure-data/wireless-connections...) and complicated stuff will be complicated (https://ni.i.lithium.com/t5/image/serverpage/image-id/96294i...).
No silver bullet guys, sorry. If you take out the complexity of the actual blocks to have multiple small blocks then you just put that complexity at another layer in the system. Same for microservices, same for actor programming, same for JS callback hell...
That is not actually simple because the data is flowing across two completely different message passing paradigms. Many users of Max/MSP and Pd don't understand the rules for such dataflow, even though it is deterministic and laid out in the manual IIRC.
The "silver bullet" in Max/MSP would be to only use the DSP message passing paradigm. There, all objects are guaranteed to receive their input before they compute their output.
However, that would make a special case out of GUI building/looping/branching. For a visual language designed to accommodate non-programmers, the ease of handling larger amounts of complexity with impunity would not be worth the cost of a learning curve that excludes 99% of the userbase.
Instead, Pd and Max/MSP has the objects with thin line connections. They are essentially little Rube Goldberg machines that end up being about as readable. But they can be used to do branching/looping/recursion/GUI building. So users typically end up writing as little DSP as they can get away with then uses thin line spaghetti to fill in the rest. That turns out to be much cheaper than paying a professional programmer to re-implement their prototype at scale.
But that's a design decision in the language, not some natural law that visual programming languages are doomed to generate spaghetti.
Edit: clarification
in one hand, this simplifies the semantics (and it's the approach I've been using in my visual language (https://ossia.io)), but in the other it tanks performances if you have large numbers of nodes... I've worked on Max patches with thousands and thousands of objects - if they were all called in a synchronous way as it's the case for the DSP objects you couldn't have as much ; the message-oriented objects are very useful when you want to react to user input for instance because they will not have to execute nearly as often as the DSP objects, especially if you want a low latency.
I can't remember the name of it, but there's a Pd-based compiler that can take patches and compile them down to a binary that performs perhaps an order of magnitude faster. I can't remember if it was JIT or not. Regardless, there's no conceptual blocker to such a JIT-compiled design. In fact there's a version of [expr] that has such a JIT-compiler backing it-- the user takes a small latency hit at instantiation time, but after that there's a big performance increase.
The main blocker as you probably know is time and money. :)
It's a great learning tool for learning a new programming language as well as the interface between the boundaries of the program are very simple.
For example I wrote this recently https://github.com/djhworld/zipit - I'm fully aware you could probably whip up some awk script to do the same, or chain some existing commands together, or someone else has written the same thing, but I've enjoyed the process of writing it and it's something to throw in the tool box - even if it's just for me!
* find
* cal
* vi
* emacs
* ls
These don't use one of standard input/standard output. (edited) and are not fully pipeable.I don't recall seeing a list of programs--tools, in the original description--that distinguish between pipeable and not-pipeable programs.
Also, none of the corrective cat/grep code in these threads point out that grep in fact takes file names, so "cat foo | grep stuff" is just a silly no-op.
What about standard output? That is, can vi be used in a pipe?
The ! command sends the selection through external command(s) as STDIN and then replaces the selection with STDOUT from the command(s). For example, grep or sort, but can be any command that works with pipes. Buffer is replaced with output (sorted file for example). Undo with U to go back to original data. Redo with R to go forward to transformed data. Command line history is available to add more commmands or correct when you type ! again.
Edit a file. Select block (Visual mode shift-V) and type ! or use the whole file with gg!G command. Type in the commands you need to run.
Vim also reads stdin if you give - as the filename, like “ls -l | vim -“ so you we can use it at the end of a pipe instead of redirecting to a file.
Like I said, I use it as an interactive debugger to assemble pipelines and see the results.
vi / vim filter commands: http://vimdoc.sourceforge.net/htmldoc/change.html#!
pipeline | vim - +"file $mytempfile"
Then if you quit vim without saving $mytempfile won't exist, if you save and quit it will and further processing can be done. I've got a view scripts like this to to things like viewing csv files after being piped through column, after saving they're converted back to csv.- Cal can be parsed (it's shows in the Unix Programming Environment, from 1983)
- Vi is a visual editor, it can be used as a front-end for ed/ex commands for I/O anyway. Kinda like the acme(1) of its day. You have both :w and :r. And, hint: it can input text from external pipes.
- Emacs is not Unix
- ls(1) is not meant to be parsed on files' content, that's the shell globbing for.
If the -l option is specified, the following information
shall be written for files other than character special and
block special files:
"%s %u %s %s %u %s %s\n", <file mode>, <number of links>,
<owner name>, <group name>, <size>, <date and time>,
<pathname>
There's no stat command in POSIX. More practically, the BSD and GNU versions of stat are completely incompatible. If you want to query a file's size or other metadata from a portable shell script, you need to parse the output of ls.i think the unix pipeline concept doesn't quite scale to other domains that try to exchange a different unit of information between the pipeline elements.
I believe it is a huge success in the crowd it was aimed for. (non-trivial size Windows admins)
https://jabustyerman.com/2017/09/29/auditing-your-r-data-pip...
This basically allows you to tag (wrap) intermediate functions in your pipe, and then inspect the objects generated by that function later.
Copying your public key somewhere? `cat ~/.ssh/id_rsa.pub | pbcopy`
Pretty-printing some JSON you copied from a log file? `pbpaste | jq . | pbcopy`
It's one of those little tools I use daily and don't think about much, but it's incredibly useful.
I gave a talk about including a bit about pipes in Angular: https://youtu.be/Gv7IfU78vxw?t=578
a |1 b | c | d, |1 e
aka piping a to both b,c,d and e branches ? a foo
; bar
; baz
would sequentially send "foo", "bar" then "baz" to "a", then would return the result of the last call.The ability to fork iterators (which may be the semantics you're actually interested in) also exists in some languages e.g. Python (itertools.tee) or Rust (some iterators are clonable)
a (foo duh meh) <- classic composition here
; bar
; baz a.(foo.duh.meh)
in e.g. ruby.edit:
a | tee >(b | c | d) | e
Can do with named pipes, but an operator would be nice
(I love pipes too btw)
https://github.com/linpengcheng/PurefunctionPipelineDataflow
Using the input and output characteristics of pure functions, pure functions are used as pipelines. Dataflow is formed by a series of pure functions in series. A dataflow code block as a function, equivalent to an integrated circuit element (or board)。 A complete integrated system is formed by serial or parallel dataflow.
data-flow is current-flow, function is chip, thread macro (->>, -> etc.) is a wire, and the entire system is an integrated circuit that is energized.
For those of us that know who Jess is, the post is a little lacking. I went there to read about some cool things she's doing with pipe, but was disappointed at a post that contained nothing.
But to people who follow her and are rather new to this way of doing things, this could be a good read.
Either way, it is nice to be reminded of the unix philosophy about chaining tools together for junior and senior alike.
Edit: typo fixes
I guess being honest around here is the wrong thing.
And I also didn't upvote the article, because it is indeed quite basic. Though apperantly enough people did find it interesting enough to make it to the frontpage ^^
Having said that, it does seem like the author's posts seem to get upvoted no matter what by a contingent of readers. I also found TFA to be quite basic, and it couldn't have taken very much effort to write (I'm assuming the author is already well familiar with pipes based on her background). The article would have been more interesting with some more technical details or experimental results - or perhaps some novel information that most people aren't aware of.
I think my point was more valid, in this case, given that this isn't some random person trying to contribute something, I'd have been less critical in that case and not commented, but this is Jess Frazelle. Those of us that know the name know she's an extremely skilled engineer, and _that_ is why I was disappointed at the post.