Therefore every time you use it to spool one file into a pipeline, that is technically an abuse!
Therefore every time you use it to spool one file into a pipeline, that is technically an abuse!
Do they mean the same thing?
As for cat, the utility, I'm afraid we'll never stop seeing people doing
cat file|prog1|prog2
even when it makes absolutely no sense whatsoever.If it did, I might as well do
cat file|cat|prog|cat|prog2|cat -
I mean, why not? It "looks nicer" than prog file|prog2 or
prog < file|prog2
Programmers love their cats.Besides, ken & Co. aren't daft. con would be short for concatenate. :-)
See: http://man.cat-v.org/unix-1st/1/cat http://man.cat-v.org/unix-6th/1/cat
I'm pretty sure based on the timeline of pipe (~1973, roughly V4) that the cat command (~1971 V1) predates pipes.
<file ./command --args
and even ./command <file --args
works fine in Bash and Zsh.For the downvoters: please time how long it takes to do something like `cat $file | awk '{print $1}' ` and `awk <$file '{print $1}'`
cat file | foo
foo <file
assuming foo only reads stdin so `foo file' isn't possible, is that with the latter the shell will open file for reading on file descriptor 0 (stdin) before execing foo and the only cost is the read(2)s that foo does directly from file.With the needless cat we have cat having to read the bytes and then write(2) them whereupon foo reads them as before. So the number of system calls goes from R to R+W+R assuming all reads and writes use the same block size and more byte copying may be required.
$ ls -lh foo
-rw-r--r-- 1 ori ori 954M Sep 5 22:57 foo
$ time cat foo | awk '{print $1}' > /dev/null
real 0m1.631s
user 0m1.452s
sys 0m0.540s
$ time awk <foo '{print $1}' > /dev/null
real 0m1.541s
user 0m1.376s
sys 0m0.160s
This was run from a warm cache, so that the overhead of the extra IO from a pipe would dominate.But if you add up the "user" and "sys" time in the cat example, you see that it took 1.992s of actual cpu-time... Which is actually about a 30% increase in cpu-time spent.
The perf decrease wasn't visible because you have multiple cores parallelizing the extra cpu-time, but it was there.
~/desktop$ du -h c.dat
11G c.dat
~/desktop$ time cat c.dat | awk '{ print $1 }' > /dev/null
real 0m53.997s
user 0m52.930s
sys 0m7.986s
~/desktop$ time < c.dat awk '{ print $1 }' > /dev/null
real 0m53.898s
user 0m51.074s
sys 0m2.807s
cat CPU usage didn't exceed 1.6% at any time. The biggest cost is in redundant copying, so the more actual work you're doing on the data, the less and less it matters.It is pretty much just a matter of principle.
Heh. Be careful with this, though: ^ is the caret (note spelling) according to most sources of information about these things.
Random Fun Geekery Time: Back in the Before-Before, the grapheme in ASCII at the codepoint ^ is now was an up-arrow character, which is why BASIC uses ^ for exponentiation even though FORTRAN, which came first and which early BASIC dialects greatly copied, uses .
These days, ↑ is U+2191, or ↑ in HTML.
http://www.alanwood.net/unicode/arrows.html
That is, two asterisks in a row.
Since everything's a file* in UNIX (and ilk,) it's actually not an abuse.
* Pedantic variations such as "everything's a bytestream" or "everything's a file descriptor" notwithstanding.