Things you (probably) didn't know about xargs
offbytwo.com
offbytwo.com
sed, awk, xargs, friends aren't used interactively like either of the above, but are rather super powerful ingredients of non-interactive environments.
Can you elaborate? I'm trying to understand how studies can show the GUI is faster than the command line, with or without xargs, grep, etc.
Sources for the studies?
Do you have other studies?
The example: "find . -name '* .py' | xargs grep 'import'" would become "find . -name '* .py' -exec grep -H 'import' {} \;".
You need to include the -H in grep to get the filename in which the match occurs.
Also note that text surrounded by asterisks in your comment become italics. Indent text by two or more spaces to reproduce text verbatim, like for code.
It's not necessarily about saving a pipe, but also, when the tool provide first-class support for the function, it's typically less prone to error. For example, the -print0 becomes unnecessary, and I've been burnt by that.
I also appreciate the writeups that don't teach poor examples. We all know how prolific copy&paste coding is. How many times have you seen "grep foo bar | wc -l" when you know it's just all-around better to "grep -c foo bar"?
I think a good intro to xargs starts with a list of things that you can't do without it. (Easy for me to say that, but of course I haven't written that piece...) It'd be great to know why to use it, not just how, you know what I mean?
Anyway, this is just off-the-cuff commentary, not criticism. Thanks for writing it up.
grep import **/*.pyThey already had a package named ack. WTF? By "getting their priorities" correct you mean catering to you?
The same problem is currently happening with node.js. Its unfortunate that node.js chose node to replace an equally ambiguous and generic name.
That being said, yes, I believe they're plain wrong on this one. They shouldn't cater to me, but to their users. Thank God they have stats.
http://qa.debian.org/popcon.php?package=ack-grep
http://qa.debian.org/popcon.php?package=ack
I count in both because I installed the ack package mistakingly. I'm probably not the only one.
Still a very large WIN for ack, the one I care about.
Plus it has --thpppt.
I usually don't bother with this kind of command line micro-optimization, but in my experience find and grep are the commands for which it's worth knowing and using every option.
So maybe a 1000 grep processes instead of one. But I've saved a pipe.
(Either that or why filenames can have spaces or LFs in them.)
I wrote a utility I named print0, which simply converts line-oriented input into null-terminated output. Very useful for building pipelines of line-oriented utilities, where each line is a filename. It's quite common to have spaces in filenames (if nothing else, user files on NAS shares), and vanishingly rare to see newlines, so I find it to be a sensible tradeoff. Things like 'find | egrep | sed | print0 | xargs -0' work as you'd expect.
(Interesting idea with print0, but I fear this is mere symptomatic relief, rather than actually fixing the problem.)
How you figure my solution to my problem is only symptomatic treatment, with technical drawbacks I already mentioned but treat as acceptable, is beyond me.
[1] As I mention in a cousin comment to this one, my utility handles all usual ASCII forms of newlines - \r, \n and \r\n (and \n\r for good measure).
When sharing servers with less UNIX savvy developers you will observe A) a notoriously polluted home-directory and B) plenty of files with funny names such as -, * , user@server.com (from failed scp attempts) all over the place.
It's indeed relatively hard to create a file called '/' by accident. But I've seen files containing the '/' along with spaces, which can be just as deadly. Oh and the popular * -file is not to be taken lightly either.
As a rule of thumb: Learn the safe way to chain these commands (find/xargs in particular) once and stick to it. And always perform a dry-run before launching the real deal.
echo {1..100} wc -l **/*.py
rm **/*~
grep 'import' **/*.py
I like zsh a lot for this small feature. rm **/*\~
rm '**/*~'One can refer to http://pubs.opengroup.org/onlinepubs/009695399/utilities/xar... for the portable flags.
It's inconsistent in comparison to most utilities, for sure.
It's consistent internally, and POSIX-compliant, at least. (iirc)
Recursively find all Python files and search them for the word ‘import’ find . -name '.py' | xargs grep 'import'*
Hmm. I don't mean to be a tweak, but you don't need xargs to do either of those things. Just:
find . -name '~' -delete
find . -name '.py' | grep 'import'
Note: I can't figure out how to get an asterisk to show up, and don't have time to look it up.
(Also, rant rant, I really don't understand why find was extended with -delete in the first place. What's next, "ls --delete" or maybe "cat --grep"?)
See section 9.1.5 http://www.gnu.org/software/findutils/manual/html_node/find_...
It walks through many of the same issues as the OP, but with more sophistication.
It also explains "+", which is used in place of the traditional ";" to essentially get xargs type argument accumulation, but within find.
That is,
find . -name '*~' -exec rm {} \+
I did not know about that. It's in Mac OS 10.6 find, for one.-print-0 | xargs -0 does not fix the race condition.
The problem is someone can swap in a symlink after the find, and before the xargs.
It does not delete the file using the entire path (which may contain a sudden symlink).
It's not possible to do this safely using xargs.
Take a look also at -execdir which does the same thing - changes to the directory first, and runs things from there. -exec is not safe and should not be used.
xargs is not safe if you are running against a directory not your own. You should use find and -execdir instead.
Yes, the original authors of posix made a mistake here.
> Also, rant rant, I really don't understand why find was extended with -delete in the first place.
I'm hoping you understand it now.
For example:
find . -name '*~' -exec 'rm {}'
The statement above executes `rm result` for every result. By contrast: find . -name '*~' | xargs rm
The example above would group the results and pass them to rm like this: rm result1 result2 result3 result4
Because I'm not an expert, I don't know how many parmeters it will pass, or if initiation of many new processes has a significant performance impact on newer systems. I would suspect that for anything involving the disk, I/O will be the bottle neck, not process turn-up time.Anyone have an opinion/insight?
-exec will be slower because for each file it has to spawn a new process (this is where most of the time goes), and then that new process has to look up the file.
xargs will be faster than -exec because it will collect a few hundred, or thousand, filenames, and pass them all to one command (the article claims 4096 is the default, on some systems it may be lower). This means that typically, only 3 processes need to spawn: `find', `xargs', and one `rm'; instead of find, and many, many `rm's.
Now, xargs is still slower than -delete because it will buffer the filenames, either waiting for the list to end, or 4000 filenames to pass to `rm'. Then, rm must look up the file from the filename.
To make my point I just set up a test situation with 1110 files ending with `~', among a total of 2221 files. I tested how fast it is to delete all the files ending with `~' using the 3 following commands:
$ find . -name '*~' -delete
$ find . -name '*~' -exec rm {} \;
$ find . -name '*~' -print0 | xargs -0 rm
| real | user | sys
-delete | 0.024s | 0.003s | 0.020s
-exec | 0.819s | 0.007s | 0.107s
xargs | 0.073s | 0.007s | 0.017s
I probably should have run those tests multiple times, and taken the average, but meh.Mind if I ask how you timed execution?
The take aways for me are:
* Use find's built-in options where possible
* Use xargs where a built-in method isn't available
* Use -exec only when nothing else will do
find . -name '.py' -exec grep -H import '{}' \;Thanks for the laugh though.
I don't think that works. That'll grep for files named '\import\',
grep -r -l import *.py
grep 'import' `find -name '* .py'`
I won't go too deep on the issues this will trigger if your filenames contain special characters. We often think of spaces, but you could have some nastier stuff. If you want to sound smart at your next geeks reunion, simply read http://www.dwheeler.com/essays/fixing-unix-linux-filenames.h...
$ getconf ARG_MAX
2097152
and if you do hit them, they're very easy to raise: $ ulimit -s 32768
$ getconf ARG_MAX
8388608
xargs is still useful for its other features, of course.linux-2.6 # git ls-files|wc -l 36747
Getting close! :)
And I would guesstimate (Linux, kernel >= 2.6.23) to still be a fairly small amount of the machines people interact with professionally through a command line.
And if in a script/snippet, you often want to cover a vast majority of the systems you _could_ end up with. Won't be System III, at least for me, but there has to be a RHEL5 system in a closet right? :)
$ bash --version
GNU bash, version 2.05b.0(1)-release (i386-pc-linux-gnu)
Copyright (C) 2002 Free Software Foundation, Inc.
$ strace bash -c '/bin/echo `seq 1 30000`' 2>&1 | grep exec
execve("/bin/bash", ["bash", "-c", "/bin/echo `seq 1 30000`"], …) = 0
execve("/bin/echo", ["/bin/echo", "1", …) = -1 E2BIG (Argument list too long)
As you can see, the argument list too long error came back from the
execve syscall, i.e., from the kernel. (Note that I shortened the strace output to make it fit the page)Thanks for the link, that's more interesting than the submission. :)
$ grep -R --include=\*.py import .but that parallelization parameter may win me back as it's cleaner than & and global vars for counters....