Unix: When pipes get names
itworld.com
itworld.com
In addition to doing cool things like passing open files over a UNIX socket, or credentials (your remote server can verify the UID of whoever opened the file in the other end!) the SOCK_DGRAM UNIX sockets are reliable and accept packets beyond normal UDP limits. This is useful for load balancing (you don't have to proxy your connection but can pass the new accepted socket directly to whowever is going to send data to it).
So if you want to make a simple server yet requesting/replying good size packets you can bind it to AF_LOCAL/SOCK_DGRAM, then the client calls bind() and sendto() while the server calls bind, then recvfrom() (which fills in the address of the caller) and sendto() back.
http://www.normalesup.org/~george/comp/libancillary/ [code]
http://www.lst.de/~okir/blackhats/node121.html [description]
<{ls}
that string would get replaced by a filename, which when opened would be hooked up to the stdout (or stdin for >) of the invoked command.Used as an argument, this allowed you to essentially do none-linear piping in a mostly transparent way. I say mostly because I don't think you could seek.
Every once in a while I find myself wishing there was something like this in bash. Maybe there is...
diff <(ls /path/to/dir/1) <(ls /path/to/dir/2)
in bash/zsh - im not sure if its a per executable thing or actually built into the shell?It's built into the shell in zsh, and it must also be in bash. By way of explaining how I know, I am first going to tell you a little bit about the program called strace, but probably not backwards-talking little people:
http://en.wikipedia.org/wiki/Strace
When invoked like so:
strace diff <(ls /path/to/dir/1) <(ls /path/to/dir/2)
strace positions itself between the running process which it invokes (in this case, diff) and the OS kernel. It then prints for you all of the communications traffic, in the form of system calls and return values, between the kernel and the process. Because it runs after the shell has expanded the command line, it gets to see the command line as the program sees it, as opposed to how it was typed.Now, the output scrolls by pretty fast (there is a -o option to save the output to a file, which I neglected to bother with here), so I saw the end first:
stat("/proc/self/fd/11", {st_mode=S_IFIFO|0600, st_size=0, ...}) = 0
stat("/proc/self/fd/12", {st_mode=S_IFIFO|0600, st_size=0, ...}) = 0
open("/proc/self/fd/11", O_RDONLY) = 5
open("/proc/self/fd/12", O_RDONLY) = 7
and so on. Obviously, diff is getting information from /proc/self/fd/11 and /proc/self/fd/12. Weird. Where do those come from?Scrolling up to the top of my xterm, we have the answer:
execve("/usr/bin/diff", ["diff", "/proc/self/fd/11", "/proc/self/fd/12"], [/* 68 vars */]) = 0
You can look up execve for yourself; the upshot is, its second argument is the argv passed in to the process by the runtime environment, which reflects the command line after shell expansion. Thus, the <(ls foodir) stuff is a shell feature, and not per-command. echo <(:) <(command)
Edit to add: both ksh and zsh support process substitution, too: ksh: <(command)
zsh: either =(command) -- uses a temporary file
or <(command) -- uses a named pipe
[1]: http://tldp.org/LDP/abs/html/process-sub.html ls -l|tee >(wc -l)
Works in bash, but not sh.One minor correction I'd made though, it that you don't need '-l' in your 'ls' parameters as 'ls' automatically outputs one item per line when you're piping it's output, plus the '-l' option adds an additional line.
Though if you did still want to see the file permissions but have a correct line count then the following would crop the 1st line from 'ls -l':
ls -l | sed 1d | tee >(wc -l)
Weirdly though, I thought 'ls' had an option to show a file count, but I can't find it in 'man'.It's akin to bringing up Mozart's writing symphonies at the age of 8 to the proud mother of a somewhat disabled child that has finally mastered the potty at the tender age of 12.
Seriously, you can literally list every single new IPC introduced in *nix in the last 3 decades and counter it with a better design and execution under Plan9. It's simply the result of a proper design as opposed to an afterthought. No real point in bringing it up unless you can constructively use the information to fix Linux.
But the discussion isn't about Linux.
http://man.cat-v.org/plan_9_2nd_ed/3/pipe ls | pee "./cruncher" "./muncher"
will produce output from both ./cruncher and ./muncher in a mixed order. But often it would be more useful to synchronize the output so that all of ./cruncher output comes first and then ./muncher output.You could redirect ./cruncher and ./muncher output to individual named pipes and then pipe a 'cat cruncherpipe muncherpipe' to follow that:
mkfifo cout mout
ls | pee "./cruncher >cout" "./muncher >mout" | cat cout mout
But that requires the manual precreation of the named pipes which isn't very clean. Using <(./cruncher) on the command line would automatically create an output pipe but doesn't allow redirecting the pipeline to each of the subcommands.Any brilliant ideas?
LS=$(ls)
./cruncher <<END
$LS
END
./muncher <<END
$LS
ENDConductor, thanks for posting the link.
mkfifo file; script -f file
Other user simply does cat fileIn one shell:
mplayer -input file=named_pipe
In a different shell:
echo 'pause' > named_pipe
or:
echo 'quit' > named_pipe
So you can control your music player easily. (I bind the echo script to hotkeys.)
Example 2:
mysql LOAD DATA INFILE wants a filename. So I made a named pipe and sent my data there (putting that command in the background since it will hang till the data is read). Then I use that named pipe as the filename for load data.
Example 3:
You want to send your data over ssh, but you also want to send commands over ssh (you can't do both at the same time, except in the command line which might not be enough). You could send the data over a second ssh session - but then you can't do your preliminary commands first.
So you make a named pipe on the server and send the data to it in one ssh session. Then you open a second ssh session, do your preliminary commands, then your data command using the named pipe as a filename.
(There are other ways to do this with renaming file descriptors in a shell, but a named pipe is easier, and much more flexible.)
Example 4:
You want to start a long running process that waits for data. You let it read from the named pipe, and then you can write to the named pipe from any other program as needed.
SSH already has a Control facility for this.
The Control facility is for multiplexing the data streams, but that solves a different problem.
So, I set up a FIFO like described, and then I bound the key '1' in Emacs to echo 'p' to it. Voila. I could seamlessly pause and unpause the video while typing away furiously. The strategy was quite a hack, but it worked very well for my purpose.
This is not to say named pipes do not have their uses. Just that you should think before reaching for them first.
Having said that, I have had occasion to use named pipes in a shell script. Basically, I had two programs and I wanted to determine which lines were in one, but not both, of their outputs. Using named pipes, I was able to do something like:
mkfifo foo
./prog1 > foo &
./prog2 > foo &
cat foo | sort | uniq -u
Effectively, this provides a way to combine the output of multiple commands. I took it as a matter of faith that doing this won't interleave lines. I assume that as long as both programs flush only at line breaks then there is not any problem. sort < foo | uniq -u
not work with named pipes?The advantage to
cat foo | sort | uniq -u
is that it is (marginally) easier to drop in another filter in front of sort. <foo sort | uniq -u
(redirections can be at the beginning too.)Also: sort -u instead of sort|uniq.
{ ./prog1; ./prog2; } | sort | uniq -u
IIRC, they're guaranteed not to interleave write()s less than PIPE_BUF (4k?) but if your program is buffering internally it might not end writes on a newline.
mkfifo split_aa
mkfifo split_ab
mkfifo sorted_aa
mkfifo sorted_ab
split -n2 input_file split_ &
sort split_aa >> sorted_aa &
sort split_ab >> sorted_ab &
sort -m sorted_aa sorted_bb > final_sorted
It's left as an exercise for the reader to determine if this actually saves any time, and if this technique could be used to implement a truly parallel sort. sort --parallel=2 input_fileThis is mostly to illustrate a technique of "parallel processing using fifos and split input" It's almost like a miniature mapreduce running locally. See also "bashreduce" (https://github.com/erikfrey/bashreduce) which I didn't write.
mkfifo /tmp/buf;
cat /tmp/buf | /bin/sh -i 2>&1 | nc -lp 31337 > /tmp/buf
And then, in another terminal: nc 127.0.0.1 31337http://mywiki.wooledge.org/ProcessManagement#I_want_to_proce.... is another neat trick: when a background process ends, make it echo to some pipe, some other "controller" process is continually reading from that pipe. Since the read is blocking, it'll wait patiently for the echo until it does whatever it should do (maybe even based on whatever it read from the pipe). So much less hacky than sleep-polling :-)
I was recently messing around with Arduinos and Raspberry PI's - totally clueless about baud rates and Serial in general but had the bug so kept going, burned a few chips then a few more and finally did something that didn't smell of smoke : )
Then it went ding - the | in Linux was passing data one byte at a time like a Serial connection to an Arduino was. Somehow I've never looked at computers the same way since. Time to go play with some named pipes!
So probably something is using the write function with just one character at a time. The standard wrapper over write (i.e. the stdio functions like puts and printf) can be set to buffer the data first then write it out.
The only difference in behavior I can think of is timing. Generally, when you read a text file, you see what is written at that moment in time (barring race conditions). With named pipes, you read what has been written, and then wait until the program writing closes the pipe. If anything gets written in the meantime, you still see it. This lets you use them for message passing in a way that normal files do not.
The blocking/signalling is still missing though.
But also, by writing out to /tmp on memory-backed systems, the lack of blocking means you grow memory use potentially indefinitely if the reader is slow or delayed for some reason. That will ultimately turn into swapping, which is just disk access again.
There's also no great way to truncate the beginning of a regular file.
Out of curiosity, is there some good reason you'd do this instead of just using a mechanism that's specifically made to solve this problem?
What's wrong with
$ ./producer > my_rambacked_file & $ tail -f my_rambacked_file > ./consumer
Well, okay, I guess this way you don't get a signal that the producer is finished.
http://manpages.courier-mta.org/htmlman7/pipe.7.html
> POSIX.1-2001 says that write(2)s of less than PIPE_BUF bytes must be atomic: the output data is written to the pipe as a contiguous sequence. Writes of more than PIPE_BUF bytes may be nonatomic: the kernel may interleave the data with data written by other processes. POSIX.1-2001 requires PIPE_BUF to be at least 512 bytes. (On Linux, PIPE_BUF is 4096 bytes.) The precise semantics depend on whether the file descriptor is nonblocking (O_NONBLOCK), whether there are multiple writers to the pipe, and on n, the number of bytes to be written.