Unix Commands I Abuse Every Day
everythingsysadmin.com
everythingsysadmin.com
A lot of people like writing bash for loops, I will try and avoid that as much as possible, xargs -n1 is the bash equivalent of a call to 'map' in a functional language.
For instance, let's say you want to create thumbnails of a bunch of jpegs:
find images -name "*.jpg" | xargs -n1 -IF echo F F | sed -e 's/.jpg$/_thumb.jpg/' | xargs -n2 echo convert -geometry 200x
Additionally, it's fully parallelizable as xargs supports something akin to pmap.
Your trivial example:
find images -name "*.jpg" | parallel convert -geometry 200x {} {.}_thumb.jpg
GNU parallel makes sure output from the commands is the same output as
you would get had you run the commands sequentially. This makes it
possible to use output from GNU parallel as input for other programs. find . -print | xargs ls
`find` dumps the results to stdout, the pipe shuttles `find`'s stdout to `xargs` stdin, `xargs` uses its stdin as arguments for `ls`.(Don't run this in your home directory)
[1] Pedants will realize this is actually equivalent to `ls .* *` as hidden files are found by `find`.
Note also that "-print" is the default command for find, so you can leave it off. Other commands include -print0 (NUL-delimited instead of newline-delimited) and -exec.
find ./ -name '*.log' | xargs rm
Finds all log files and map them to 'rm' commands. e.g. if it finds system.log and rails.log it will run the command `rm system.log rails.log`.xargs will automatically do things like break up very long lists into multiple command calls so that it doesn't exceed the maximum number of arguments a command can have.
Other useful things about xargs are '-P <NUM>' which lets you run the same command in parallel. I use this with curl to do ghetto benchmarks.
The next major flag is `-n <NUM>` which changes the number of arguments per command call. e.g. `-n 1` will run the command per argument passed to xargs.
And the last flag I commonly use is `-I {}`. This sets `{}` as a variable which contains the arguments. (This also forces `-n 1`). This is useful for things like:
find ./ -name '*.log' | xargs -I{} sh -c "if [ -f {}.gz ]; then rm {}; fi" find ./ -name '*.log' | xargs rm
Only do that if you know exactly what '*.log' will expand to (i.e. don't use it in scripts and avoid using it on the command line). This is because the delimiter for xargs is a newline character, but filenames can have a newline character in them. This can lead to unexpected results.Almost everywhere I see xargs used, find ...-exec {} ; will work as well and find ...-exec {} + may work even better.
find ./ -name '*.log' -print0 | xargs -0 rm
Fixes that issue and xargs is far more efficient, it doesn't launch a new process for each line like exec does, but far more importantly, xargs is generalizable to all commands so you only have to learn it once; exec is just an ugly hack on find, you can't generalize it across all commands; xargs is much more unixy.Exec doesn't either if you use "+" as a terminator:
find . $(options) -exec command {} \;
executes one command per match, find . $(options) -exec command {} +
executes a single command with all matches> exec is just an ugly hack on find
Obvious and complete nonsense, -exec is both an important predicate of find and a useful and efficient tool.
-print0/xargs -0, now that is a hack.
find -name '*.log' -delete find ... -print0 | xargs -0 ...
or: ... | xargs -d'\n' ... find . -iname '*.pdf' -exec pdfgrep -i keyword {} +vim $(grep -lr foo | xargs)
and doing what I need to do on a file by file basis. Otherwise, for renaming functions and the like, I do a lot of:
find . -name foo_fn exec sed -i s/foo_fn/bar_fn/g '{}' \;
I generally love abusing bash. Just today I was asked about how to rename a bunch of files, specifically containing spaces, and came up with either of these two options:
find -name foo_bar -exec cp "'{}'"{,.bak} \;
and
for file in $(find -name foo_bar); do cp "$file"{,.bak}; done
Ultimately, the great thing is, if you learn CTRL-R, you can always search for these types of commands and modify them as necessary for the particular task at hand and not necessarily remember them. One I use all the time, to push git branches upstream is the following:
CTRL-R --set-
which gives me:
git push -f --set-upstream origin `git rev-parse --abbrev-ref HEAD`
This is entirely unique in my history, a common part of my workflow, and trivially searchable.
I also enjoy being able to perform something along the lines of:
vim $(bundle show foo-gem)
vim $(grep -lr foo | xargs)
is doing? Assuming a missing directory after the `foo' why can't it just be vim $(grep -lr foo .)
I don't see that xargs's default behaviour is adding anything.For
find . -name foo_fn exec sed -i s/foo_fn/bar_fn/g '{}' \;
you may find -exec's + of use. The above has the fork/execve overhead for each file found. find -name '*.[ch]' -exec sed -i 's/\<foo\>/bar/g' {} +Thanks for the suggestion about -exec +; I will have to remember it in the future.
Furthermore, instead of the pipe to sed and extra xargs, it might be clearer and simpler to do something like:
zargs -n 1 **/*.jpg -- make-thumb
Where "make-thumb" is a short script (or even a zsh function, if you care about saving a fork for each input file) containing: convert -geometry 200x $1 ${1%%.jpg}_thumb.jpg
But, in real life, instead of writing such a script or function, what I'd probably do instead is: zargs -n 1 **/*.jpg | vipe > myscript
and do some quick editing in vim to modify the zargs output by hand to do whatever I need -- and then I'd run the resulting "myscript". Just fyi, "vipe" is part of the "moreutils" package [1] and lets you use your editor in the middle of a pipe.One final trick is for when you need to do in-place image manipulation. Instead of using "convert", you can use another ImageMagick command: "mogrify". It will overwrite the original file with the modified file. Of course, you should be very careful with it.
find images -name '*.jpg' | while read jpg; do
convert -geometry 200x "$jpg" "$(echo "$jpg" | sed 's/.jpg$/_thumb.jpg/')";
done
This version works even when there are spaces in a filename, whereas yours will break. $ ls -Ql
totale 4
-rw-r--r-- 1 zed users 33 set 6 07:30 " spaces "
$
$ find -name '*spaces*' | while read text; do
cat "$text";
done
cat: ./ spaces: File o directory non esistente
$
$ find -name '*spaces*' -print0 | xargs -0 cat
while read is broken with spaces
$ $ ls -Q
" spaces "
$ ls | while IFS="\n" read -r f; do ls "$f"; done
spaces
$
For lots of grim detail see David A. Wheeler's http://www.dwheeler.com/essays/fixing-unix-linux-filenames.h... find images -name '*.jpg' | while read jpg; do
convert -geometry 200x "$jpg" "${jpg%%.jpg}_thumb.jpg";
done find images -name "*.jpg" | while read -r jpg; do
convert -geometry 200x "$jpg" "${jpg%.jpg}_thumb.jpg"
done
This correctly handles spaces in file names and uses built in shell string replacement. find images -name '*.jpg' |
sed 's/\.jpg$//
s/.*/convert -geometry 200x &.jpg &_thumb.jpg/' xargs -a cmd_list.txt -I % alias %
This has been bothering me for a while. file.txt:
line1
line2
line3
line4
line5
line6
$ paste - - - < file.txt
line1 line2 line3
line4 line5 line6
Combine with the column command for pretty printing. I seem to find a use for this pretty frequently.I like the simplicity of this one but it's not very useful day to day:
$ echo *
As a replacement for "ls".Therefore every time you use it to spool one file into a pipeline, that is technically an abuse!
<file ./command --args
and even ./command <file --args
works fine in Bash and Zsh.For the downvoters: please time how long it takes to do something like `cat $file | awk '{print $1}' ` and `awk <$file '{print $1}'`
cat file | foo
foo <file
assuming foo only reads stdin so `foo file' isn't possible, is that with the latter the shell will open file for reading on file descriptor 0 (stdin) before execing foo and the only cost is the read(2)s that foo does directly from file.With the needless cat we have cat having to read the bytes and then write(2) them whereupon foo reads them as before. So the number of system calls goes from R to R+W+R assuming all reads and writes use the same block size and more byte copying may be required.
$ ls -lh foo
-rw-r--r-- 1 ori ori 954M Sep 5 22:57 foo
$ time cat foo | awk '{print $1}' > /dev/null
real 0m1.631s
user 0m1.452s
sys 0m0.540s
$ time awk <foo '{print $1}' > /dev/null
real 0m1.541s
user 0m1.376s
sys 0m0.160s
This was run from a warm cache, so that the overhead of the extra IO from a pipe would dominate.But if you add up the "user" and "sys" time in the cat example, you see that it took 1.992s of actual cpu-time... Which is actually about a 30% increase in cpu-time spent.
The perf decrease wasn't visible because you have multiple cores parallelizing the extra cpu-time, but it was there.
~/desktop$ du -h c.dat
11G c.dat
~/desktop$ time cat c.dat | awk '{ print $1 }' > /dev/null
real 0m53.997s
user 0m52.930s
sys 0m7.986s
~/desktop$ time < c.dat awk '{ print $1 }' > /dev/null
real 0m53.898s
user 0m51.074s
sys 0m2.807s
cat CPU usage didn't exceed 1.6% at any time. The biggest cost is in redundant copying, so the more actual work you're doing on the data, the less and less it matters.It is pretty much just a matter of principle.
Heh. Be careful with this, though: ^ is the caret (note spelling) according to most sources of information about these things.
Random Fun Geekery Time: Back in the Before-Before, the grapheme in ASCII at the codepoint ^ is now was an up-arrow character, which is why BASIC uses ^ for exponentiation even though FORTRAN, which came first and which early BASIC dialects greatly copied, uses .
These days, ↑ is U+2191, or ↑ in HTML.
http://www.alanwood.net/unicode/arrows.html
That is, two asterisks in a row.
Since everything's a file* in UNIX (and ilk,) it's actually not an abuse.
* Pedantic variations such as "everything's a bytestream" or "everything's a file descriptor" notwithstanding.
Do they mean the same thing?
As for cat, the utility, I'm afraid we'll never stop seeing people doing
cat file|prog1|prog2
even when it makes absolutely no sense whatsoever.If it did, I might as well do
cat file|cat|prog|cat|prog2|cat -
I mean, why not? It "looks nicer" than prog file|prog2 or
prog < file|prog2
Programmers love their cats.Besides, ken & Co. aren't daft. con would be short for concatenate. :-)
See: http://man.cat-v.org/unix-1st/1/cat http://man.cat-v.org/unix-6th/1/cat
I'm pretty sure based on the timeline of pipe (~1973, roughly V4) that the cat command (~1971 V1) predates pipes.
time read
press enter to read elapsed time. If you write your activity in the prompt and repeat it for multiple activities, you have a nice time log. You can then just copy&paste it from terminal. foo | echoin bar
This is like `foo | bar`, but I can see what's passing between them. It's a bit like `tee`, but reversed. It's what I irrationally want `foo | tee - | bar` to do. my $args = join ' ', @ARGV;
open OUT, "|$args" or die "Can't run $args: $!\n";
while (<STDIN>) {
print $_;
print OUT $_;
} function highlight() {
local args=( "$@" )
for (( i=0; i<${#args[@]}; i++ )); do
if [ "${args[$i]:0:1}" != "-" ]; then
args[$i]="(${args[$i]})|$"
break
fi
done
grep --color -E "${args[@]}"
}
This is only to be used as a filter, since it mangles filenames.I'm curious, does ack support a highlight-only mode?
#!/bin/bash
xargs -n 1 grep -l "$@"
This takes a list of files on stdin, then greps for the argument in all the files and spits out the matching files.The perk is that it can be chained:
find *.txt | narrow dog | narrow cat | narrow rabbit
This will find all the files that contain dog, cat, and rabbit. xargs -rd'\n' grep -l "$@" tar -tf tarfile.tar | xargs rm
Witless uninstall #2. Find a file that you changed just before "make install" into the wrong spot (usually config.log is a good candidate). find /bogus/installdir -newer config.log | xargs rm
Yes, this is totally unsafe. But it's an abuse... so, there you goIt's a newish feature of less (those of you with stale RHEL installs won't find it). Type '&<pattern><return>' and you'll filter down a listing to match pattern. Regexes supported. Prefixed '!' negates filter.
Wishlist: interactive filter editing (similar to mutt's mail index view filters), so you don't have to re-type full expressions.
function hl() { local R=''; while [ $# -gt 0 ]; do R="$R|$1"; shift; done; env GREP_COLORS="mt=38;5;$((RANDOM%256))" egrep --color=always $R; }
then do e.g.
whatever pipeline | hl foo bar baz | hl quux | hl '^.* frob.*$' | less -R
results in foo/bar/baz highlighted in one color, quux in another, and whole lines containing frob in another. hopefully the colors aren't indistinguishable from each other or from the terminal background :\
I use a somewhat similar setup in emacs, where a key binding adds the word under point to a syntax highlighting table, but the color is computed as the first 6 characters of the md5sum of the word.
Also, it skips blank lines. But of course that might be in the feature-not-a-bug category; and if you really want to see every line, there's always grep-quote-quote:
grep '' *.txt
In any case, fun article. :-)i don't understand his first one though. i thought both linux and bsd greps had the -H option (show filename).
here's some more:
1. sed t file instead of cat file
2. echo dir/*|tr '\040' '\012' instead of find dir
3. echo dir/*|tr '\040' '\012' |sed 's/.*/program & |program2 /' > file; . file instead of xargs or find
(of course this assumes you keep filenames with spaces or other dangerous characters off your system.) 4. same as 3. but use split to split file into parts. then execute each part on a different cpu.Normally I'd typed something like "grep -i 'something' foo | less", and wanted to just up arrow the previous line and change the grep stuff to cat. I don't know why, it doesn't really save me anything. Maybe it's the hackerish "because I can, that's why" instinct at work.
cat !$ | less
!$ is "the last argument to the previous command". less <ESC>.
it will save you even more keystrokes.The command you are looking for is yank-last-arg for bash and insert-last-word for zsh.
You can do "head -n 0" on Linux to mean "all lines".
No you can't.Honest question, btw. I'm relatively inexperienced with Linux, and I certainly haven't used BSD. I'd appreciate any critiques you may have to offer.
Yes (or perhaps tail -n +0 as that is idiomatic, which makes it clear to anyone what you are intending).
The '-' between -n and 0 means "all but the last 0 lines".