An Opinionated Guide to Xargs
oilshell.org
oilshell.org
> That is, grep doesn't support an analogous -0 flag.
However, the GNU grep variant does have an analogous flag:
-z, --null-data
Treat the input as a set of lines, each terminated by a zero byte (the ASCII NUL character) instead of a newline. Like the -Z or --null option, this option can be used with commands like sort -z to process arbitrary file names.
Edit: It seems that grep -0 isn't taken for something else and they should have used it for consistency? The man page says it's meant to be used with find -print0, xargs -0, perl -0, and sort -z (another inconsistency)
Takeaways: (1) There is no consistency in flag names, even --long ones (2) impressively many tools do support it! Note that some affect only input or only output. (3) All do NUL-terminated, not NUL-separated. That's fortunate — matches \n usage, and gives distinct representations for [] vs [""].
It is quite necessary, because you cannot pass an arbitrarily large command line or environment in exec system calls.
Of course, this doesn't have the problem requiring -0 because we're not reading textual lines from standard input, but working with lists of strings.
;;; This source file is part of the Meta-CVS program,
;;; which is distributed under the GNU license.
;;; Copyright 2002 Kaz Kylheku
(in-package :meta-cvs)
(defconstant *argument-limit* (* 64 1024))
(defun execute-program-xargs (fixed-args &optional extra-args fixed-trail-args)
(let* ((fixed-size (reduce #'(lambda (x y)
(+ x (length y) 1))
(append fixed-args fixed-trail-args)
:initial-value 0))
(size fixed-size))
(if extra-args
(let ((chopped-arg ())
(combined-status t))
(dolist (arg extra-args)
(push arg chopped-arg)
(when (> (incf size (1+ (length arg))) *argument-limit*)
(setf combined-status
(and combined-status
(execute-program (append fixed-args
(nreverse chopped-arg)
fixed-trail-args))))
(setf chopped-arg nil)
(setf size fixed-size)))
(when chopped-arg
(execute-program (append fixed-args (nreverse chopped-arg)
fixed-trail-args)))
combined-status)
(execute-program (append fixed-args fixed-trail-args))))) do_something | ( while read -r v; do
. . .
done )
I’ve found that it has fewer edge cases (except it creates a subshell, which can be avoided in some shells by using braces instead of parens)1. You don't need the parentheses.
2. If you use process substitution [1] instead of a pipe, you will stay in the same process and can modify variables of the enclosing scope:
i=0
while read -r v; do
...
i=$(( i + 1))
done < <(do_something)
The drawback is that this way `do_something` has to come after `done`, but that's bash for you ¯\_(ツ)_/¯[1] https://www.gnu.org/software/bash/manual/html_node/Process-S...
I think the redirection can come first, though (not at a computer to test):
< <( do_something ) while read . . . $ < <( echo foo ) while read -r f; do echo "$f"; done
-bash: syntax error near unexpected token `do'
$ < <( echo foo ) xargs echo
foo
$ bash --version
GNU bash, version 5.1.4(1)-release (x86_64-apple-darwin20.2.0)For documentation purposes, this is the exact thing I tried to run:
$ < <(echo hi) while read a; do echo "got $a"; done
-bash: syntax error near unexpected token `do'
$ while read a; do echo "got $a"; done < <(echo hi)
got hi
Maybe there is another way... $ zsh -c '< <(echo hi) while read a; do echo "got $a"; done'
got hi
My position isn't that it is a good reason to switch shells, but if you're using it anyway then it is an option.Bash uses /dev/fd when available, but also appears to have an internal implementation which silently creates named pipes and cleans them up. In Bash 5.0.18 on AIX, fake process substitution works just fine, in my testing.
One common pattern I use this for is running a bunch of checks/tests, e.g.
EXIT_CODE=0
while read -r F
do
do_check "$F" || EXIT_CODE=1
done < <(find ./tests -type f)
exit "$EXIT_CODE"
This is a more complicated alternative to the following: find ./tests -type f | while read -r F
do
do_check "$F" || exit 1
done
The simpler version will abort on the first error, whilst the first version will always run all of the checks (exiting with an error afterwards, if any of them failed)>Different shells exhibit different behaviors in this situation:
>- BourneShell creates a subshell when the input or output of anything (loops, case etc..) but a simple command is redirected, either by using a pipeline or by a redirection operator ('<', '>').
>- BASH, Yash and PDKsh-derived shells create a new process only if the loop is part of a pipeline.
>- KornShell and Zsh creates it only if the loop is part of a pipeline, but not if the loop is the last part of it. The read example above actually works in ksh88, ksh93, zsh! (but not MKsh or other PDKsh-derived shells)
>- POSIX specifies the bash behaviour, but as an extension allows any or all of the parts of the pipeline to run without a subshell (thus permitting the KornShell behaviour, as well).
paste -d \\n <(do_something1) <(do_something2) | while read -r var1 && read -r var2; do
... # var1 comes from do_something1, var2 comes from do_something2
donehttps://www.gnu.org/software/parallel/parallel_alternatives....
parallel is probably on the complex side but its also been actively developed, bugfixed and had a lot of road miles from large computing users.
What does it do that xargs and shell can't? (honest question)
Edit: To clarify, xargs usually wants to spin up a process per task. I have parallel spin up N processes and then continuously feed them.
For others that didn't know about it, see the examples here: https://www.gnu.org/software/parallel/parallel_tutorial.html...
Here's another surprising feature: https://www.gnu.org/software/parallel/parallel_tutorial.html...
Here is an example of how it works,
https://docs.computecanada.ca/mediawiki/index.php?title=GNU_...
This + restart capabilities make gnu parallel very well suited to running 1000s of compute-heavy jobs on HPC clusters.
https://github.com/tfmoraes/blender_gnu_parallel_render/blob...
If you use `xargs -P`, all processes share the same stdout and output may be mixed arbitrarily between them. (If the program being executed uses line buffering, lines usually won't be mixed together from multiple invocations, but they can be if they're long enough).
In contrast, `parallel` by default doesn't mix together output from different commands at all, instead buffering the entire output until the command exits and then printing it.
With `--line-buffer` the unit of atomicity can be weakened from an entire command output to individual lines of output, reducing latency.
Alternately, with `--keep-order`, `parallel` can ensure the outputs are printed in the same order as the corresponding inputs, which makes the output deterministic if the program is deterministic. Without that you'll get results in an arbitrary order.
These aren't technically things that xargs and shell can't do; you could reimplement the same behavior by hand with the shell. But by the same token, there isn't anything xargs can do that the shell can't do alone; you could always use the shell to manually split up the input and invoke subprocesses. It's just a question of how much you want to reimplement by hand.
For the output interleaving issue, what I do is use the $0 Dispatch Pattern and write a shell function that redirects to a file:
do_one() {
task_with_stdout > $dir/$task_id.txt
}
So if there are 10,000 tasks then I get 10,000 files, and I can check the progress with "ls", and I can also see what tasks failed and possibly restart them.You even have some notion of progress by checking the file size with ls -l.
I tend to use a pattern where each task also outputs a metadata file: the exit status, along with the data from "time" (rusage, etc.)
But I admit that this is annoying to rewrite in every script that uses xargs! It does make sense to have this functionality in a tool.
But I think that tool should be a LANGUAGE like Oil, not a weirdo interface like GNU parallel :)
But thanks for the explanation (and thanks to everyone in this subthread) -- I learned a bunch and this is why I write blog posts :)
For what it's worth, I consider oil to be closer to a unixy PowerShell rather than a more powerful bash. Note that this is not a slight, PowerShell is sweet for what it is. It (oil) really takes a hard left from the POSIX philosophy of focusing on one thing and doing it well. I'm also bitter that, if it's going to veer so far away from POSIX, that it didn't go the whole hundred and become a function language with comprehensions and such.
For what it's worth, everything you mentioned above about your approach can be done with parallel.
Functional: there are interesting shells like Elvish. But it really goes PowerShell by adding internal rich data pipelines that dont have a unixy stream-of-bytes representation. Oil does NOT go that way; it works on stuff like QSN to make pure unix interconnects more robust.
btw, oil looks very cool. I hate how many footguns are in common shells.
[1] eg writing to a tempfile and atomically renaming into place: "task_with_stdout > $dir/task_id.txt.tmp && mv $dir/task_id.txt{.tmp,}"
* does not buffer stderr
* does not check if the disk is full for a period of time during a task (thus risking incomplete output)
* does not clean up, if killed
* does not work correctly if task_with_stdout is a composed command
Given that GNU Parallel is a drop-in replacement for xargs, I am curious why you find it a 'weirdo interface'.https://docs.computecanada.ca/mediawiki/index.php?title=GNU_...
Using xargs for this kind of work is euhm... not a good idea.
If you seriously believe you can implement everything using xargs, then this (contrived) example is for you: https://unix.stackexchange.com/questions/405552/using-xargs-...
Newer versions include 'parset' which can set shell variables in parallel, which is useful if you want to 'map' values from one array to another.
For me, an essential feature of GNU parallel is that it is semantically equivalent to "sh". Imagine that you write a file that contains a long list of commands. You can pipe that file to "sh" to run the commands, or pipe it to "parallel" to do the same, but faster. If you are building the list of commands on the fly, then you can use xargs with a slightly different syntax. But somehow using "sh" or "parallel" gives a certain peace of mind due to its straightforward semantics. I never used any argument of GNU parallel apart from -j
My usage pattern: to build a list of commands explicitly then run it (possibly teeing the list into a temporary file to inspect it):
for i in one two three; do
printf "echo $i\n"
done |sh # or |parallel[0]: https://github.com/archlinux/svntogit-community/tree/package...
... which is exactly what GNU Parallel is. Your concern is even mentioned in the design documentation: https://www.gnu.org/software/parallel/parallel_design.html
https://git.savannah.gnu.org/cgit/parallel.git/tree/doc/cita...
Obviously downside to the visibility and dynamism is that it redirects stdout. You can read it back later, in order. But it’s not there for continued processing immediately.
myfunc() {
printf " %s" "I got these arguments:" "$@" $'\n'
}
export -f myfunc
seq 6 | xargs -n2 bash -c 'myfunc "$@"' "$0"It turns out that e.g. -print0 and -0 are the only safe way: line endings aren't escaped:
find . -type f -print0 | el -0 --each -x echo
GNU Parallel is a much better tool: https://en.wikipedia.org/wiki/GNU_parallelGNU xargs has --verbose which logs every command. Does that not do what you want? (Maybe I should mention its existence in the post)
xargs -P can do everything GNU parallel do, which I mention in the post. Any counterexamples? GNU parallel is a very ugly DSL IMO, and I don't see what it adds.
--
edit: Logging can also be done with by recursively invoking shell functions that log with the $0 Dispatch Pattern, explained in the post. I don't see a need for another tool; this is the Unix philosophy and compositionality of shell at work :)
Also parallels can run some of those threads on remote machines. I don't believe xargs has an equivalent job management function.
In other words:
some command | xargs -P other command | third command
This is useful if 'other command' is slow. If you buffer on disk, you need to clean up after each task: Maybe there is not enough free disk space to buffer the output of all tasks.UNIX is great in that you can pipe commands together, but due to the interleaving issue 'xargs -P' fails here. It does not live up to the UNIX philosophy. Which is probably why you unconsciously only use it at the end of a pipeline.
You can find a different counterexample on https://unix.stackexchange.com/questions/405552/using-xargs-... I will be impressed if you can implement that using xargs. Especially if you can make it more clean than the paralel version.
Parallel is the better tool but the nagware impairs its reputation.
(In any case, this surely is tangential, since the title is not "X considered harmful" for any value of X—at best it comments on a post by that title, as, indeed, you are doing.)
"A Response to Xargs Criticism"
cat input.file | ... | while read -r unit; do <cmd> ${unit}; done | ...
between 'while read -r unit' and 'while IFS= read -r unit' I can probably handle 90% of the cases. (maybe I should always use IFS since I tend to forget the proper way to use it).I suspect I'll really like your way of doing things, but an example would be very handy.
echo "foo bar" | tr ' ' '\n' | while read -r var; do echo ${var}; done
For examples in general, I guess something like "cat file.csv" could work. (the difference between using IFS= and not using it is essentially whether we want to preserve leading and trailing whitespaces or not. If we want to preserve, then we should use IFS=).I was happy to read that the author comes to the same conclusion and proposes an `each` builtin (albeit only for the Oil shell)! Like that there is no need to learn another mini language as pointed out.
I'd also perhaps argue that the reason we don't want xargs to be a built-in is precisely because of zargs and the point in your second paragraph. If it was built-in it would no doubt be obscenely different in each shell, and five decades later a standard that no one follows would eventually specify its behaviour ;)
¹ https://zsh.sourceforge.io/Doc/Release/User-Contributions.ht... - search for "zargs", it has no anchor. Sorry.
> -n instead of -L (to avoid an ad hoc data language)
Apparently GNU xargs is missing it, but BSD xargs has -J, which is a `-I` which works with `-n`: with `-I` each replstr gets replaced by one of the inputs, with `-J` the replstr gets replaced by the entire batch (as determined by `-n`).
for file in *; do
command_using_file &
done
wait
?I use variations on this all the time; pause while load is high, pause while 'x' or more things are running, sleep between invocations, etc.
It may not be as convenient for some cases, but "can't do that..." is not quite correct either.
The post is starting to feel like a hammer/nail argument, IMO.
do_something | tr \\n \\0 | xargs -0 ...Ok filenames can theoretically have newlines in them but I'd be happy to deal with that weird case. I can't recall ever having encountered it in years of using bash on various systems.
Shell pipes would then orthogonally provide the stuff like substitution that xargs does in it's own unique way (that I just can't be bothered learning) - instead you'd just pipe the find output through sed or 'grep -v' or whatever you wanted before piping into xargs.
I guess that's what aliases but I'm too lazy anymore to bother with configuring often short-lived systems all the time.
So the defaults went with principle of least surprise, pretending it's like a very long args list that you could theoretically enter at the shell, including quotes.
You could, for example, edit the args list in vi and line split / indent as you please but not impact the end result.
rm $(ls | grep foo)
will not work if you have file names that contain spaces.
Shell programming is planted thick with landmines like this.
> Besides the extra ls, the suggestion is bad because it relies on shell's word splitting. This is due to the unquoted $(). It's better to rely on the splitting algorithms in xargs, because they're simpler and more powerful.
Never can remember all the -I stuff around xargs
I wouldn't say "never use it", but I would hesitate to ever put it in a script, vs. doing a one-off at the command line.
This is far superior to futzing with xargs interactive or whatever dry run feature they have.
http://www.oilshell.org/blog/2021/08/xargs.html#more-comment...
https://lobste.rs/s/wlqveb/xargs_considered_harmful#c_kwsxtc
Fair enough, but I still favor find -exec. I find it generally less error prone, and it's never been so slow that I wished I had instead used xargs.
Also, if you're specifically using -exec rm with find, you could instead use find with -delete.
And I always fail to remember that the -exec terminator must be escaped in zsh, so using -exec always takes me multiple tries. So I only use -exec when I must (for `find` predicates).
after spending a bit of time reading the man page for find, i rarely use xargs any more. find is pretty good.
tangent:
another instance i've seen where spawning many processes can lead to bad performance is in bash scripts for git pre-recieve hooks, to scan and validate the commit message of a range of commits before accepting them. it is pretty easy to cobble together some loop in a bash script that executes multiple processes _per commit_. that's fine for typical small pushes of 1-20 commits -- but if someone needs to do serious graph surgery and push a branch of 1000 - 10,000 commits that can can cause very long running times -- and more seriously, timeouts, where the entire push gets rejected as the pre-receive script takes too long. a small program using the libgit2 API can do the same work at the cost of a single process, although then you have the fun of figuring out how to build, install and maintain binary git pre-receive hooks.
That is, find -exec is sort of "hard-coded", while find | xargs allows obvious extensions like:
find | grep | xargs # filter tasks
find | head | xargs # I use this all the time for faster testing
find | shuf | xargs
Believe it or not I actually use find | shuf | xargs mplayer to randomize music and videos :)So shell is basically a more compositional language than find (which is its own language, as I explain here: http://www.oilshell.org/blog/2021/04/find-test.html )
PS> "alice", "bob" | echo
PS> Get-ChildItem . -Include "*test.cpp","*test.py" -Recurse | foreach { Remove-Item $_.Name }
No text parsing in sight, and the object attributes can be tab-completed from the shell (e.g. I tab-completed the `$_.Name`).Also, expanding the regex into `-Include` parameters is somewhat cheating since `-Include` only takes globs, and it just so happens that that particular regex can be converted into globs.
The general equivalent is:
gci -re | ?{ $_.Name -match '.*_test\.(py|cc)' } | ri
(I used the shorter aliases because someone will probably read yours and reinforce the stereotype that PS is overly verbose.)