Disclaimer: I do cat-piping myself quite a bit out of habit, so I'm not trying to look down at the author or anything like that! :)
Disclaimer: I do cat-piping myself quite a bit out of habit, so I'm not trying to look down at the author or anything like that! :)
Instead, shell script should be optimized for readability and portability and I think it is much easier to understand something like 'read | change >write' than 'change <read >write'. So I like to write pipelines like this:
cat foo.txt \
| grep '^x' \
| sed 's/a/b/g' \
| awk '{print $2}' \
| wc -l >bar.txt
It might be not the most efficient processing method, but I think it is quite readable.For those who disagree with me: You might find the pure-bash-bible [1] valuable. While I admire their passion for shell scripts, I think they are optimizing to the wrong end. I would be more a fan of something along the lines of 'readable-POSIX-shell-bible' ;-)
In any case, it's probably just a matter of personal taste.
It is instantly plainly obvious to me what each step of their shell script is doing.
While I can absolutely understand what your shell script does after parsing it, it's meaning doesn't leap out at me in the same way.
I would describe the prior shell script as more quickly readable than the one that you've listed.
So, perhaps it's not a question of one being more silly than the other—perhaps the author just has different priorities from you?
That should be gsub, shouldn't it? (sub only replaces the first occurrence)
grep '^x' < input | sed 's/foo/bar/g'
to be very readable, as the flow is still visually apparent based on punctuation. cat input | grep '^x' | sed 's/foo/bar/g'
Is far more readable, in my opinion. In addition, it makes it trivial to change the input from a file to any kind of process.I'm STRONGLY in favor of using "cat" for input. That "useless use of cat" article is pretty dumb, IMHO.
OK, to save some of my face, this will work:
grep 'foo' <(input) | sed 's/baz/bar/g'
... at least in zsh and probably bash. input | grep foo | sed ... diff <(prog1) <(prog2)
and get a sensible result.And sometimes programs just refuse to read from stdin but do just fine with an unseekable file on the command line. True, you do have this:
input | recalcitrant_program /dev/stdin
... but it's a bit of a tossup as to which one's more readable at this point. They're both relying on advanced shell functionality.> diff <(prog1) <(prog2)
> and get a sensible result.
That is called process substitution and is exactly the kind of use case that it's designed for. So yes, process substitution does make sense there.
> input | recalcitrant_program /dev/stdin
> ... but it's a bit of a tossup as to which one's more readable at this point. They're both relying on advanced shell functionality.
There's no tossup at all. Process substitution is easily more readable than your second example because you're honouring the normal syntax of that particular command's parameters rather than kludging around it's lack of STDIN support.
Also I wouldn't say either example is using advanced shell functionalities either. Process substitution (your first example) is a pretty easy thing to learn and your second example is just using regular anonymous pipes (/dev/stdin isn't a shell function, it's a proper pseudo-device like /dev/random and /dev/null) thus the only thing the shell is doing is the same pipe described in this threads article (with UNIX / Linux then doing the clever stuff outside of the shell).
cat input | grep '^x' | sed 's/foo/bar/g'
→ sed '/^x/s/foo/bar/g' <input sed -n '/^x/s/foo/bar/gp' <input
This may be an inadvertent argument for the ‘connect simpler tools’ philosophy.Now back to the topic of "cat", which is a great example of why shell scripts are minefields.
Replace "foo.txt" with a user supplied variable, let's call it "$F". It becomes cat $F | blah_blah... I mean cat "$F" | blah_blah, first trap, but everyone knows that.
Now, if F='-n', second trap. What you think is a file will be considered an option and cat will wait for user input, like when no file is given. Ok, so you need to do cat -- "$F" | blah_blah.
That should be OK in every case now, but remember that "cat" is just another executable, or maybe a builtin. For some reason, on your system "cat --" may not work, or some asshat may have added "." in your PATH and you may be in a directory with a file named "cat". Or maybe some alias that decides to add color.
There are other things to consider, like your locale that may mess up you output with comas instead of decimal points and unicode characters. For that reason, you need to be very careful every time you call a command and even more so if you pipe the output.
For that reason, I avoid using "cat" in scripts. It is an extra command call and all the associated headaches I can do without.
You're not wrong, but I think it's worth pointing out that's a trap that comes up any time you exec another program, whether it's from shell or python. I can't reasonably expect `subprocess.run(["cat", somevar])` to work if `somevar = "-n"`.
(Now, obviously, I'm not going to "cat" from python, but I might "kubectl" or something else that requires care around the arguments)
I think that you forgot to edit the "I mean" to "echo $F" :)
I periodically think it would be a good idea to organize a language around.
For example, you can write this pipeline as:
grep '^x' foo.txt \
| sed 's/a/b/g' \
| awk '{print $2}' \
| wc -l > bar.txt
This is by no means scientific, but I've got a LaTeX document open right now. A quick `time` says: $ time grep 'what' AoC.tex
real 0m0.045s
user 0m0.000s
sys 0m0.000s
$ time cat AoC.tex | grep what
real 0m0.092s
user 0m0.000s
sys 0m0.047s
Anecdotally, I've witnessed small pipelines that absolutely make sense totally thrash a system because of inappropriate uses of `cat`. When you `cat` a file, the OS must (1) `fork` and `exec`, (2) copy the file to `cat`'s memory, (3) copy the contents of `cat`'s memory to the pipe, and (4) copy the contents of the pipe to `grep`'s memory. That's a whole lot of copying for large files -- especially when the first command grep in the sequence usually performs some major kind of reduction on the input data!That said, I suspect the example would be much faster if you didn't use the pipeline, because a single tool could do it all (I'm leaving in the substitution and column print that are actually unused in the result):
awk '/^x/{gsub("a","b");print $2; count++}END{print NR}' foo.txtDoesn't gsub(/a/, "b") do the same thing as s/a/b/g?
I had recently built a set of tools used primarily via pipes: (tool-a | tool-b | tool-c) and it looks clearer when I mock (for testing) one command (cat results | tool-b | tool-c) instead of re-flowing it just to avoid cat and use direct files.
<file command | command | command
is perfectly fine.Does anyone have a better way to do this kind of thing?
[1] https://core.tcl.tk/expect/index [2] https://pexpect.readthedocs.io/en/stable/
i.e.
prefer `apt-get install -y` over `yes | apt-get install foo`
(Similarly, the first thing I used to do on Windows was set my prompt to [$p] because many years ago I also accidentally nuked a part of Visual Studio when I copied and pasted a command line that was prefixed with "C:\...>". Whoops.)
$ less foo | bar
Is similar too: $ bar < foo
Except that less is typically more clever than that and might be more like: $ zcat foo | bar
Depending on the file type of foo.Of course, the example @arendtio uses is correct, because they obviously care about such things.