Stop Piping Cats
ibm.com
ibm.com
cat file | xargs foo | grep bar | sort | wc -l
It just looks nicer than: < file xargs foo | grep bar ...You failed at that?
I'm intrigued because it always occured to me that you were a particularly fastidious person (with all due respect).
I often come accross issues that would seem for all the world to be a pointless artifact of an accepted convention. If this accepted convention is an acquired trait and one learns exclusivly from a limited meduim (irc frustrates me), I find it very curious that they persist as they sometimes act as a difficult barrier for a thorough understanding. This xargs thing, though I know almost nothing about it, seems to parallel similar problems I've encountered and I for one would be interested to hear your take on this (and HN in general).
Or just some food for thought for your next blog post maybe?
The point was, when I used "cat file |", I knew it was going to work, and it did. When I tried to eliminate the cat by using the program's built-in ability to read a file, I had to read several manpages before determining that it was not possible. All because some tutorial says "cat is useless", when it clearly saved me much more time than the extra CPU time it used.
And if you are actually asking; xargs is just a utility to read command-line-arguments from stdin. `echo -e "foo\nbar" | xargs rm' === `rm foo; rm bar' (or depending on the xargs implementation; `rm foo bar'). It kind of reminds me of a "functor map" operation, where stdin is a functor (of command-line arguments), and command-line programs are functions. (I will now mention that xargs also does "join" on the "results" of the "function", which is very ... monad-like. But "Monads are teh awesome and everything is one" is my second-least-favorite Internet meme, so I will spare you. :)
So true.
My requirement for a shell command is that it be reasonably easy to assemble, return something resembling the correct answer, and that it run in a reasonable amount of time.
If your requirements are more strict than those, it's time to write a real program.
I tried to figure out how to make xargs read from a file
instead of stdin, and was unsuccessful.
"xargs -a <filename>" First option described in the man page on Red Hat. diff <{ps} <{sleep 3; ps} diff <(ps) <(sleep 3; ps) (xargs foo | grep bar | sort | wc -l) < file
It makes the pipeline one command, both lexically (verb comes first) and concretely (the pipeline is kicked off in a subshell) verb file
cat file | verb1 | verb2 | verb3
(verb1 | verb2 | verb3) < file
verb123 file
EDIT: Consider the calling conventions, which one of these handles it's arguments differently? Assume that verb123 is an equivalent to the pipeline -- the subshell+stdin construction lends itself to shell aliases: alias verb123="(verb1 | verb2 | verb3) < "Of course, with the usual argument vs stdin conventions, either of the middle two could also be rewritten:
verb1 file | verb2 | verb3
This is probably how I'd write such a chain.The point is a logical improvement anyway (not burying the input argument near the beginning). I'm kind of surprised that the bash folks haven't turned cat into a builtin like they did with time and some of the other coreutils.
It's too bad it's about 30 years too late to stem the tide of shit like cat -v: http://harmful.cat-v.org/cat-v/
Here's cat as a pure sh builtin:
shcat {
for arg in "$@"; do
exec 3<>"$arg"
while read line <&3; do
echo "$line"
done
exec 3>&-
done
}
The shell by it's very nature can't just do exactly one task well, it's a programmable environment for living in. The cancer that's bloating UNIX was the way that the BSD and especially the GNU crews took simple tools and cross-pollinated them randomly with stupid shit. Try running "/bin/true --help" on a GNU system sometime -- there's a damn good reason why "your shell may have its own version of true". % /bin/true --help
Usage: /bin/true [ignored command line arguments]
or: /bin/true OPTION
Exit with a status code indicating success.
--help display this help and exit
--version output version information and exit
NOTE: your shell may have its own version of true, which usually supersedes
the version described here. Please refer to your shell's documentation
for details about the options it supports.
Report bugs to <bug-coreutils@gnu.org>.
I wouldn't necessarily call that 'bloated' unless you feel that any program that uses one bit more than absolutely necessary should be scrapped as 'bloated beyond belief.'> there's a damn good reason why "your shell may have its own version of true".
Because why exactly?
(wc Lines) . sort . (grep "bar") . (xargs "foo") =<< file
But note in this case that all the data "flows in the same direction". wc Lines $ sort $ grep "bar" $ xargs "foo" =<< file Prelude> :info $
($) :: (a -> b) -> a -> b -- Defined in GHC.Base
infixr 0 $
Prelude> :info =<<
(=<<) :: (Monad m) => (a -> m b) -> m a -> m b
-- Defined in Control.Monad
infixr 1 =<<
Prelude> :info .
(.) :: (b -> c) -> (a -> b) -> a -> c -- Defined in GHC.Base
infixr 9 .
Incidentally, the parens in my example are actually unnecessary. Function application is about 10, and (.) is 9.It's too easy to make a mistake and type >
..and that can ruin your whole day.
I find that to be a far more pernicious design error -- they should have made the longer token the destructive one, or used another character in it.
xargs foo < file | grep bar | sort | wc -lThese days, if you hit a performance wall from spawning too many cats, I would think you switch to some scripting language where you have everything in one process. Premature optimization, people...
cat file.txt | sort | uniq | wc -l ## notice uniq
it gives you the count of unique lines. If you omit the sort in that instance it will fold duplicates together a/b/a is three lines, not two.
cat file | xargs foo | grep bar | sort | wc -leg
cat file | less
cat file | grep thing
cat file | grep otherthing
cat file | grep otherthing | cut stuff
instead of less file
grep thing file
grep otherthing file
grep otherthing file | cut stuff cat file | less
less file
The second version is 5 fewer keystrokes. On to the next command: cat file | grep thing
grep thing !$
The second example is the same or fewer keystrokes. In both cases, you have to type "grep thing". In the first, you have to press the up arrow and backspace over "less" (at least three keystrokes), and in the second, you have to type an extra " !$".I'll skip moving to "cat file | grep otherthing" or "grep otherthing !$", and consider the change to get to
cat file | grep otherthing | cut stuff
grep otherthing file | cut stuff
In both cases, you have to type " | cut stuff". If you key in the second example as "!! | cut stuff", that's an extra two keystrokes. If you key in the first as up arrow + "| cut stuff", that's only 1 extra keystroke.In total, my version saves keypresses in the specific example and doesn't seem much worse in general.
~ $ mkdir -p tmp/a/b/c
~ $ mkdir -p project/{lib/ext,bin,src,doc/{html,info,pdf},demo/stat/a} $ echo project/{lib/ext,bin,src,doc/{html,info,pdf},demo/stat/a}
project/lib/ext project/bin project/src project/doc/html
project/doc/info project/doc/pdf project/demo/stat/a
(I added a linefeed to prevent wrapping.)I only drop into a shell at most for 20 minutes a day at the moment, so a lot of the really neat time-saving trick simply don't stick in my head due to disuse...
mv .xinitrc{,.bak}
or, if you have some stuff mounted under another mountpoint, such as when installing gentoo: umount /mnt/gentoo/{proc,dev,boot,}
That empty entry expands to the base path, which can be a nice shortcut.Old fink example:
sudo fink install lib{png,jpeg,ssl,whatever}{,-{dev,shlibs}}
which expands to: sudo fink install libpng libpng-dev libpng-shlibs libjpeg libjpeg-dev libjpeg-shlibs libssl libssl-dev libssl-shlibs libwhatever libwhatever-dev libwhatever-shlibsI am ashamed to admit that I usually say "apt-get install libfoo.*", wait for the downloads to start, hit Control-c, and then cut-n-paste the package names I actually want onto the command-line.
It sounds bad when I type it out, but it's really not the most horrible thing ever. But your way is definitely like 83x better.
A perfect example of the infinite features that can be found in bash manual page. I've been using bash since the early 90's and I've written lots of non-trivial programs in bash, and I know a lot many other people don't know and yet I had somehow managed to miss this pearl.
$ cd
$ ls tmp
- some files lists -
$ mkdir -p tmm/a/b/c
oops (remember your esc key) tail -f access.log | grep --line-buffered "GET /blog/post "In general I far prefer the (it seems to me) more Unixy way of having lots of very simple commands and chaining them together over the use of extra options and arguments.