Useful Uses of cat
two-wrongs.com
two-wrongs.com
cat filename.txt
Up | grep “thing I want”
Is fewer keystrokes than cat filename.txt
grep “thing I want” filename.txt
Or more likely cat filename.txt
grep filename.txt “thing I want”
grep “thing I want” filename.txt grep “thing I want” !$
Bash (and similar) will replace !$ with the last parameter of the previous command.This is a trick I’ve used lots when wanting to perform a non-piping operation on a command I’ve ‘cat’ed (eg ‘rm -v !$’)
I’d never criticise anyone for “useless” use of ‘cat’ though. If the fork() overhead was really that critical then it wouldn’t be a shell command to begin with.
Press Escape, then the number of the argument from the previous command, confirm with Ctrl+Alt+Y.
Example:
> command arg1 arg2 arg3
Escape, 1, Ctrl+Alt+Y gives you arg1. cat filename.txt
grep "what I want" $_
expands "$_" to "filename.txt" bind '"\e."':yank-last-arg
So [esc] _ does it on systems where I haven't customised my environment. A major drawback though is it doesn't go through history like alt . does.However, when 'set -o vi' is enabled you can easily go through history and edit with familiar vi keystrokes, or press 'fc' and fix the previous command in EDITOR, or 'v' to edit the current command in EDITOR.
I think you meant 'vv' and I wonder if it's not something that's set up by oh-my-zsh ? It's great though
I wasn't aware that bash and zsh did it differently, I assumed they'd both use the same readline - now I'm aware there's more to understand about it :)
I find that zsh's mode is actually better than readline's For example it can handle visual mode which is why the same function is `vv` and not just `v` Zsh can handle text objects too, very useful to be able to ci" for example Fish's can too and is quite good IIRC
Too bad the `v` command does not work in gdb so it seems it's more of a bash thing than a readline thing
Relevant : https://superuser.com/questions/1543120/make-readline-edit-i...
Then I can look at it before hitting return
useful shorthand
> sudo !!
for example.
Again, bash specific.
cat log | not spam1 | not spam2 | not 'spam(3|4)' | .... | less alias -g V='| grep -vE'
alias -g L="| $PAGER"
Then you can do cat log V spam1 V spam2 L
I also like alias -g G='| grep'
alias -g X='| xargs'
The possibilities are endless...Really I’d be more worried about accidental invocations. Aliases are not scoped, so if you’re dealing with one of those programs with non standards option handling and uppercase single character switches…
As for accidental invocations, yeah I agree it feels dangerous, but in a few hundred thousand lines of shell history since I set them, never had a problem.
What I really wish for is some sort of tool that would let me pipeline like that, but also easily examine each step in the chain for sanity. Sort of a workbook for shell.
function expand-alias() {
zle _expand_alias
zle self-insert
}
zle -N expand-alias
bindkey -M main ' ' expand-aliasAnd an accidental typo is not really a big deal, certainly not an "injection possibility".
UUoC criticism, to me, belongs when one sits down to script.
$ cat file
$ grep stuff alt-.
Alternatively, make use off the READNULLCMD mechanism in Zsh: $ < file
translates to $ ${READNULLCMD:-more} < file
Thus you can $ < file
then UP (or ctrl-p which I find more ergonomic) and continue with "grep stuff": $ < file grep stuff
(Redirections can be anywhere in the command.)https://zsh.sourceforge.io/Doc/Release/Redirection.html#Redi...
< file.txt grep pattern
Less keystrokes. < file.txt cat
But then one has to ctrl-w cat. It is a pity that this is not an alternative to cat for a single file: < file.txtcat “filename.txt”
Up | grep “thing”
Up | grep -v “not thing”
Up | grep -v “other thing”
Etc. it’s just easier to build this way even if the initial cat is unnecessary.
> grep “thing I want” filename.txt.
…every time
I use "cat ... | ... | ... " like in TFA and just like many in this thread because it simply makes sense. It's more intuitive. It's easier to read. It requires less braincycles to remember how this or that command wants its parameter passed, etc.
I think the "useless use of cat" movement made its time: it failed. Many of us are never going to give up our use(less|ful) of cat (you decide). So stop wasting your time complaining about it.
"useless use of cat" goes into the latter bucket. Complaining at me about it does not actually improve the code; it's just a nag about a bad habit that, arguably, isn't even a bad habit.
"Useless" uses of cat aren't bad habits during interactive usage, for all the reasons people mention here which I won't rehash.
For scripts, however, the story is different than for one-off commands. For one thing, it's slower due to the extra forks and copying of data across pipes, so there's at least that. For another, it prevents the command from inspecting the other end of its pipe, which can negatively impact usage in some case. (For example, if the program knows its input is from a terminal, it may flush its output on every newline it sees.) Moreover, a bunch of the arguments for the interaction case (like "it's fewer keystrokes" or whatever) don't even apply to the script case in the first place...
The end result here is that you definitely shouldn't assume some habit is just fine with scripting merely because it's fine when you're typing on the terminal, or vice-versa.
For shell scripts, I would argue quite vehemently that the most important goals should be correctness and readability, with performance being a very distant third concern. I'd even be tempted to argue that performance shouldn't be a consideration at all, except of course that argument would be misinterpreted to support some absurd edge case until I'd have to admit that of course performance is a little bit of a concern. But in any case, I can't recall a single example of a cat pipe being the root cause of an unacceptable performance problem in a shell script.
On the readability point, the example that probably irritates me most often is a cat pipe into some commands into a while loop. I much prefer this:
cat file.conf | sed -e 's/pattern/replacement/g' -e 's/reallybigolhonkinpattern/other-replacement/g' | tr... | while read line; do...
to this: sed -e 's/pattern/replacement/g' -e 's/reallybigolhonkinpattern/other-replacement/g' wrongfile.conf | tr... | while read line; do...
or this: sed -e 's/pattern/replacement/g' -e 's/reallybigolhonkinpattern/other-replacement/g' | tr... | while read line; do
stuff...
done < file.conf
...and that's a pretty common pattern where the edge case of reading input from a terminal doesn't apply.So this is the kind of thing that makes me go "shut up shellcheck" instead of "thanks shellcheck!"
But: performance was just one of the problems I cited. I gave you more than that -- one was a correctness reason (which you do care about) and had nothing to do with performance. And, again, incorrect buffering (which can make the script literally unusable in some cases) was just one example. I've seen needless redirection interfere with Ctrl+C handling too, though I don't recall the exact example. Oh, and there's terminal coloring and ANSI escape processing too, which programs detect differently. Point is, being unable to see the end of the pipe can definitely cause an unnecessary mess in some cases.
As for readability - honestly, part of the reason you find it less readable is that you're missing something else. Namely, this:
sed -e 's/pattern/replacement/g' -e 's/reallybigolhonkinpattern/other-replacement/g' wrongfile.conf | tr... | while read line; do...
should really have been: sed -e 's/pattern/replacement/g' -e 's/reallybigolhonkinpattern/other-replacement/g' -- wrongfile.conf | tr... | while read line; do...
which is in fact both more correct (at least when the file name isn't hard-coded, which is the common case in shell scripts) and more readable than your example; you can immediately spot where the file name is. The difference between that and cat "$blah" | sed ... is very minor at that point (and in fact you should be doing cat -- "$blah" as well...); anybody reading a command like sed without a pipe input knows to look for an input argument. The important point regarding readability here is, it's not like the code gets overly tricky if you write it one way vs. another way. It's just a matter of spending 1-2 extra seconds glancing over. So it's very much a minor thing to be prioritizing above everything else. (If the logic became harder to reason about, that'd be a different story, and it'd put more weight on the readability aspect.)Yes, I saw your other points, and I chose this example because it is an example drawn from real-world use where there is zero objective reason to wag a finger about "useless use of cat". Those other points are not relevant in this example, and piping a cat into some other commands into a while loop is pretty typical shellcode. Forcing me to move a filename argument into the middle of a long line for stylistic reasons should be obviously wrong. It is one case where shellcheck is over-reaching and being a nuisance rather than helping me catch errors.
This has been argued better and to death already: https://stackoverflow.com/a/16619430, http://oletange.blogspot.com/2013/10/useless-use-of-cat.html, https://news.ycombinator.com/item?id=23341711, https://news.ycombinator.com/item?id=36116208, https://news.ycombinator.com/item?id=6367319, https://news.ycombinator.com/item?id=1116085, etc.
By all means don't build something where you have cascading effect and need to retest an entire pipeline, but this is _not_ it.
P.S.
And if you really really want to keep it separate, just do "< access.log head -500 etc etc etc" (no I didn't forget a pipe. And yes the "< inputfile" works even if it's in front of what you're calling).
Or just use `cat` and let the pipe separate the different steps. "< access.log head" is nice but it breaks this representation where each step is piped into the next one. Sure, once you’re done fiddling you can rewrite the thing to remove the "cat", but when you are constructing the thing I find it clearer to use cat.
Or just… don’t?
No it doesn't?
< access.log head -n 500 | grep mail | perl -e …
is completely valid, and reads right-to-left as well as the cat version. IMO using stdin is preferable to either solution in TFA. < infile some_cmd > outfile
and real clarifies what's a shell command, what's a redirection, etc. infile > some_cmd > outfile
Then the arrow would be pointing in the right direction. But then it would be unclear whether "infile" is a file or a command. Which is why people use: cat infile | some_cmd > outfile
You can now interpret "cat" as a keyword that specifies that "infile" is a file. <infile some_cmd >outfile
Just like you wouldn't add spaces in the middle of '2>&1' when redirecting stderr to stdout: <infile some_cmd >outfile 2>&1Mostly because if you do it doesn’t work: '2 >&1' is not the same as '2> &1' (invalid syntax) which is not the same as '2>& 1', which …is the same as '2>&1'.
[0]: newline, ‘||’, ‘&&’, ‘&’, ‘;’, ‘;;’, ‘;&’, ‘;;&’, ‘|’, ‘|&’, ‘(’, or ‘)’.
[1]: https://www.gnu.org/software/bash/manual/bash.html#Simple-Co...
This is one of those things where I think it is until it isn't.
I sometimes second-guess myself when I think I might be over-single-responsibilifying. "Well in practice these two things are so trivial that this feels a little silly."
It often turns out to have been a good call in hindsight, especially when working with other people who aren't necessarily thinking about these things at all. If the responsibilities have been sufficiently split up, they're more likely to change only the part that needed to be changed, and less likely to complectify the two things together that really shouldn't've been. Or when I go "oh wow that thing that I thought I overly-abstracted sure composes well with this unexpected new thing!"
Hardcore separation of concerns is just another method of defensive programming.
> < access.log head -500 etc etc etc
It's too bad that the syntax is so different. Why does the first stage not end with "|"? There's space for shell syntax improvements, here. Maybe a 'cat'-like builtin that translates `cat foo | bar` into `bar <foo` so you can have the nice syntax but don't needlessly create processes would leave everyone happy.
Uh... I dunno, but my lizard brain thinks that the whole idea of mediating filesystem operations on storage and IPC mechanisms like pipes is a lot more complicated, magic, and deserving of a single command than merely filtering the data on stdin.
I agree with the article and the logic, and think this historic meme was basically wrong originally. You string up your chain of pipelines with the first element being "where does it come from?" and not merely whatever the first operation happens to be just because that operation allows for some kind of file input or redirection syntactically.
I learned some bash from an old-timer who would write an infinite-loop like this:
while :; do
# loop body here
done
This works because the `:` is a way to set a label, and it implicitly returns 0. It's just a weird wrinkle of the language. So, why not do `while true`? On old systems, `true` was not a builtin and would call `/usr/bin/true`. Writing the loop this way saves a process fork on each iteration.On a modern system, you'd be hard pushed to measure the difference, so it really doesn't matter which style you prefer.
Do you have a source for that? I thought it was just POSIX built-in for true. Like `.` vs. `source`. What's a label in this context anyway?
for(;;){
// loop body here
}Nope. Unix shell doesn't have labels (are you mixing with DOS batch files?).
: is a shell builtin that does nothing. In the bash man page, look for the first entry of the "SHELL BUILTIN COMMANDS" section. https://www.gnu.org/software/bash/manual/html_node/Bourne-Sh...
<input X|Y|Z >output
The point of this syntax is that I can readily replace it with F() { X|Y|Z }
<input F >output < infile x | y | z > outfile
I just didn't like how the filename was so close to the command name ;) </proc/0/environ xargs -0
I also don't tend to want enormous volumes of text in my terminal scrollback so I generally view files or pipe verbose commands to `less`, then when I find what I want to send to the terminal I use the `|` less command to pipe it to `cat`.Or to grab just a few lines for my later reference:
kubectl get po/my-pod -o yaml | less
/* find the lines I'm interest in */
-N
|^sed -n 34,35p[1] https://evalapply.org/posts/shell-aint-a-bad-place-to-fp-par...
cat /dev/sdb > backup.img # make a disk image
cat /dev/sdb > /dev/sdc # clone disk
cat ~/Downloads/* # play Russian roulette with your terminal
cat > file # minimalistic text editor, ^D to exit saving, ^C to exit erasing the file
cat << wq > file # nearly complete emulation of ed
grep -r bongo . | cat # shorter than typing --color=never
cat -v file # cause 20 points of damage to wizards of bell labs
cat file > file # empty a file without removing the file
cat meow meow > meows # duplicate file contents dostuff () {
if [[ $1 = clean ]]; then
grep -v dirt
else
cat
fi | do_other_stuff
}
See also https://www.in-ulm.de/~mascheck/various/uuoc/ (via https://lobste.rs/s/rtvp2u/useful_uses_cat#c_0xpqkr )`Y x | Z` is Verb-Subject-Object.
That's why I prefer using cat "uselessly".
< x Y | Z > w
Y takes input redirected from x, piped into Z, which outputs into w.I.e.
x | Z | tee w | Y
? that's... something else entirely. < x Y | Z > w
where x and w are files, not commands.Something that "cat file | ..." advocates might be overlooking is that a redirection ("<inputfile", ">outputfile", "2>errorfile") can appear anywhere within a simple command, so these:
command -option < file
< file command -option
command < file -option
are all exactly the same -- and of course very similar to cat file | command -option
If the purpose of typing "cat file | command" is to put the input file at the beginning (which does make logical sense), you can achieve the same thing with "< file command". Admittedly, it does look at bit strange if you're not accustomed to it. (It even works with csh and tcsh.) < <(curl http://...) command -option
Because that's what you get if you address the actual argument and still insist on input redirection.Input redirection is inconsistent with every other command to retrieve data. Not only does it not have the same syntax, it's combining two actions into one step of the pipeline.
I would say that < <(command arg) is a useless use of process substitution (UUoPS). You just want command arg |.
The redirection variant doesn't eliminate command and does not move arg out of command's argument position; it's just superfluous syntax.
Just because we want "< file" instead of "cat file |" doesn't imply that we want "< <(command arg)" instead of "command arg |". It's not even the same rewrite pattern at all.
However, if someone wrote:
cat <(curl https://example.com/file) | next
then that now the UUoC pattern "cat file |". We can apply the transformation to eliminate cat: < <(curl https://example.com/file) next
Now in so doing, we have moved the process substitution such that there is an obvious match for the UUoPS pattern. We apply that rewrite rule as well: curl https://example.com/file | next"Get data" is the pipeline step we're talking about. Using "< file" combines it with the first transformation step, instead of keeping it as its own separate step as all other such data sources require.
It isn't implementation-wise, but it is semantically. If anything, < is a useless syntax that should never have been in the shell to begin with. If you want to peephole optimize cat... just do that, just like it replaces tons of other commands with built-ins.
I usually prefer to do "< file command" though. "cat" adds an extra layer of indirection, forcing stream processing and hiding the original file, but that's usually unnecessary. If you really don't trust the program, it is not an adequate solution anyways.
I do think he could have woven in the Parnas source a bit better. The first quote seems to rely a lot on the surrounding context of the essay and its never contextualized. This bit:
> The problem is that these subsets and extensions are not the programs that we would have designed if we had set out to design just that product.
Just feels a bit disjointed in the article. I get the high-level message but thought it could be woven into the blog authors narrative better. Anyways, it sounds like a good article though.
I actually noticed your reply because I came back to this comment to save your blog. Not just for the content but the page style. The eggshile white (or similar) color, font, foot(side)notes. It's nice.
https://en.wikipedia.org/wiki/101_Uses_for_a_Dead_Cat
"It consisted of cartoons depicting the bodies of dead cats being used for various purposes, including anchoring boats, sharpening pencils and holding bottles of wine."
"By December 7, 1981, it had spent 27 weeks on the New York Times Best Seller list."
Things like this:
head -n 500 access.log | grep ...
head -n 500 <access.log | grep ...
Feel like you start with the filename, then go leftwards to the first operation, then start reading rightward again through the pipe. At least in my brain, it feels slightly more awkward.
<access.log head -n 500 | grep ...
... though that's less familiar to many, I'm sure.I'm realizing now another (and potentially stronger influence) is just years of muscle memory starting pipelines with cat.
It's funny, when learning programming, I think Haskell was the language that introduced me to the pattern of having a chain of operators processing a stream to build up a result (and I'd later cover it again in SICP), and I loved how clean it looked compared to imperative code. But I now find it one of the harder to read languages due to it all being prefix, whereas Java/Kotlin/C#/Javascript now all have stream constructs that use method calls, so read left-to-right, source-to-sink
And I'm reminded that I need to give Forth a proper go sometime
[1] https://en.wikipedia.org/wiki/Uniform_Function_Call_Syntax
<access.log sed -n '/mail/p; 500q' | perl -e ...
If perl is processing the file line-by-line then filtering lines by regex and stopping at line X is trivial, and you don't even need sed.I have heard snarky Perl putdowns ad nauseam at work and on HN and may have regretted using it a handful of times but I can say worse or similar for other popular tools, languages...
awk '/mail/ && NR <= 500 {...}' access.log
If you want N matched lines: awk '/mail/ && i < 500 {i++; ...}' access.logbut that absolutely kills modularity / composability
so yes, "cat xyz" plays a source - and can be replaced with another source - without touching all other stuff.
This. I often have to read logs and being able to apply the same oneliner on older logs just by adding a "z" in front of "cat" to read .log.gz instead of .log files is extremely useful.
perl -ne 'last if $. > 500; /mail/ && print' access.log
When you're done with the part after the &&, remove the 'last if $. > 500'For me, the most useful use of (gnu) cat, is cat -A weird.file
It saves my day or solves weird issues (X-files) with files generated by (not so) junior sysadmins, copy/pastes, end of lines, invisible diffs, etc... many times each year.
cat > file.txt # after this command, just type your text
Or to append : cat >> file.txt
Can also work with Here Doc syntax : cat > file.txt << EOF
This is the text to write down
EOFBut probably the former.
Like, R can pipe, and pandas can .pipe() but they both complete the function on the entire data before it pushes a copy of the output to the next step.
Why isn’t this flow of data a primitive outside of shell?
shite_template_common_default_page() {
cat <<EOF |
$(shite_template_common_header)
<main id="main">
$(cat -)
</main>
EOF
shite_template_standard_page_wrapper
}
[1] https://evalapply.org[2] https://github.com/adityaathalye/shite?tab=readme-ov-file#te...
< access.log head -n 500 | grep mail | perl -e
Now you can delete "head -n 500 |".> If we then delete only the head processing step we’re left without a step that transforms the string access.log into the contents of the access log.
By introducing "cat access.log" we have the same problem: if we delete only the cat processing step, we're left without a step that transforms the string access.log into the contents. For the useless cat to have the nice property that you can cleanly delete it from the command line, you need:
< access.log | cat | head -n 500 ...
:)It should be the same as giving your pipeline a file input handle via < access.log… but why take the risk?
$ echo "foo" > access.log # create the access.log with some content
$ < access.log | cat | head -n 5
# no output
This works: $ cat access.log | head -n 5
foo
$ < access.log cat | head -n 5
foo
$ < access.log cat | cat | head -n 5
foo
$ < access.log head -n 5
foo- uses: cat file.txt | grep foo # doesn't know why or that you don't have to
- uses: grep foo file.txt # knows that's religiously the right way, knows maybe one or two options and that "cat" is bad.
- uses (when appropriate): cat file.txt # knows all the ways to do it, pros/cons of each, can make sound judgements and trades in real-time
all three groups are actually mostly doing nothing that wrong.
if someone or shellcheck points out "useless use of cat", and it's directed at the 1st or 2nd category it's helpful. it's either saying "was that a mistake?" or "did you know that you didn't have too...".
if someone points out "useless use of cat" just to be a pedant to throw rocks at an adult who knows full well what they are doing, not as helpful.
all that being said, if you are in that 3rd category, it's not just knowledge but maturity required not just personal preference alone. the author sure seemed to have a chip on their shoulder. i suspect the truth is somewhere in-between.
( head -n40000 filelist_h.txt | grep 'v\|$' >> /dev/null; ) 0.00s user 0.02s system 43% cpu 0.045 total
( tail -n40000 filelist_h.txt | grep 'v\|$' >> /dev/null; ) 0.00s user 0.01s system 22% cpu 0.066 total
( shuf -n40000 filelist_h.txt | grep 'v\|$' >> /dev/null; ) 0.71s user 1.37s system 12% cpu 16.874 total
( cat filelist_h.txt | head -n40000 | grep 'v\|$' >> /dev/null; 0.00s user 0.01s system 101% cpu 0.017 total
( cat filelist_h.txt | tail -n40000 | grep 'v\|$' >> /dev/null; 0.05s user 0.45s system 54% cpu 0.930 total
( cat filelist_h.txt | shuf -n40000 | grep 'v\|$' >> /dev/null; 0.60s user 0.91s system 96% cpu 1.565 total
The results are fairly repeatable:- as expected, head and tail alone are equally quick alone
- shuf is ridiculously slow due to overhead compared to cat | shuf
- cat | shuf is just a bit slower than cat | tail.
- cat | tail is slower than tail.
- cat | head is faster than head (but only because of overhead)
Caveats are that this is WSL2 and the file is 480 MB (5 million lines) in a mounted Windows directory, although that helps magnify that slow I/O can influence how you pipe commands.
Seems like there's a random-access optimization that's severely impacting I/O performance.
I use the extra time waiting for WSL2 to get a coffee.
I wish I could write "useless cat" without people pestering me about it.
https://zsh.sourceforge.io/Doc/Release/Expansion.html#Comman...
You describe a really clever use case for rev. On the surface, my intuition was that reversing, cutting, and reversing again wouldn't work if you had escape sequences, but I suppose that's a moot point with cut!
<x Y | Z w>
Terminated: 15
Which cat died first?
Fascinated to know the actual answer though.
<filename.txt grep mypattern
That syntax works the same as "grep mypattern <filename.txt".>cat access.log | head -n 500 | grep mail | perl -e …
>we find that cat performs two responsibilities:
>1. Printing an error to stderr if the file doesn't exist
>2. Copying a file to stdout
`tr` is the only one I can think of, excluding ones that really can only operate on already-open file descriptors (the `read` builtin, `flock` in certain modes, ...)
What? Who cares, but also, if you do, then try this:
<the_file the_pipeline_here
taking his example: <access.log head -n 500 | grep mail | perl -e ...
<access.log grep mail | perl -e ...
<access.log perl -e ...
Yes, you have a lot of freedom over where you put the redirections. I often write nroff -man >/dev/null blah.1
to check for errors in a man page I'm editing, and you can see that intersperses I/O redirection with command-line words.So if you're really concerned about what edits you might have to make to
head -n 500 access.log | grep mail | perl -e ...
to remove the head(1) and/or the grep(1) and so have to move where the `access.log` goes, well, just write <access.log head -n 500 | grep mail | perl -e ...
and now you don't have to move `access.log`.> The natural solution is a useless use of cat.
The problem with useless uses of cat(1) is mainly that it betrays a misunderstanding of how the shell works, so "that's a useless use of cat" is a way to teach someone something they're missing. (Useless uses of cat are also useless uses of CPU cycles and energy, but the vast majority of the time those will be in the noise, so it's not a huge deal.)
I treat cat as a stream-fier: given a file it creates a stream on stdout
BTW Is it true that the origin of cat name is that in slang to cat means to vomit ? (to vomit a file to stout) ? This is what I always known, but recently I read that cat stand for conCAT
I prefer cat as vomit....
It's also possible to make a much faster cat (I have, considered naming it cheetah),10%+ faster.
What I do like is doing something like:
diff -aui <(xxd binary1) <(xxd binary2)- pest control
- entertainment
- transporting small solar arrays
I definitely need to hear more about this one.
cat like sleeping in sun.
attach solar panel to cat.
cat occasionally moves to remain in sun as it naps.
From this gem of a DEFCON talk: https://youtube.com/watch?v=rJ5jILY1vlw