https://www.gnu.org/software/parallel/parallel_alternatives....
parallel is probably on the complex side but its also been actively developed, bugfixed and had a lot of road miles from large computing users.
https://www.gnu.org/software/parallel/parallel_alternatives....
parallel is probably on the complex side but its also been actively developed, bugfixed and had a lot of road miles from large computing users.
What does it do that xargs and shell can't? (honest question)
Edit: To clarify, xargs usually wants to spin up a process per task. I have parallel spin up N processes and then continuously feed them.
For others that didn't know about it, see the examples here: https://www.gnu.org/software/parallel/parallel_tutorial.html...
Here's another surprising feature: https://www.gnu.org/software/parallel/parallel_tutorial.html...
Here is an example of how it works,
https://docs.computecanada.ca/mediawiki/index.php?title=GNU_...
This + restart capabilities make gnu parallel very well suited to running 1000s of compute-heavy jobs on HPC clusters.
https://github.com/tfmoraes/blender_gnu_parallel_render/blob...
If you use `xargs -P`, all processes share the same stdout and output may be mixed arbitrarily between them. (If the program being executed uses line buffering, lines usually won't be mixed together from multiple invocations, but they can be if they're long enough).
In contrast, `parallel` by default doesn't mix together output from different commands at all, instead buffering the entire output until the command exits and then printing it.
With `--line-buffer` the unit of atomicity can be weakened from an entire command output to individual lines of output, reducing latency.
Alternately, with `--keep-order`, `parallel` can ensure the outputs are printed in the same order as the corresponding inputs, which makes the output deterministic if the program is deterministic. Without that you'll get results in an arbitrary order.
These aren't technically things that xargs and shell can't do; you could reimplement the same behavior by hand with the shell. But by the same token, there isn't anything xargs can do that the shell can't do alone; you could always use the shell to manually split up the input and invoke subprocesses. It's just a question of how much you want to reimplement by hand.
For the output interleaving issue, what I do is use the $0 Dispatch Pattern and write a shell function that redirects to a file:
do_one() {
task_with_stdout > $dir/$task_id.txt
}
So if there are 10,000 tasks then I get 10,000 files, and I can check the progress with "ls", and I can also see what tasks failed and possibly restart them.You even have some notion of progress by checking the file size with ls -l.
I tend to use a pattern where each task also outputs a metadata file: the exit status, along with the data from "time" (rusage, etc.)
But I admit that this is annoying to rewrite in every script that uses xargs! It does make sense to have this functionality in a tool.
But I think that tool should be a LANGUAGE like Oil, not a weirdo interface like GNU parallel :)
But thanks for the explanation (and thanks to everyone in this subthread) -- I learned a bunch and this is why I write blog posts :)
For what it's worth, I consider oil to be closer to a unixy PowerShell rather than a more powerful bash. Note that this is not a slight, PowerShell is sweet for what it is. It (oil) really takes a hard left from the POSIX philosophy of focusing on one thing and doing it well. I'm also bitter that, if it's going to veer so far away from POSIX, that it didn't go the whole hundred and become a function language with comprehensions and such.
For what it's worth, everything you mentioned above about your approach can be done with parallel.
Functional: there are interesting shells like Elvish. But it really goes PowerShell by adding internal rich data pipelines that dont have a unixy stream-of-bytes representation. Oil does NOT go that way; it works on stuff like QSN to make pure unix interconnects more robust.
btw, oil looks very cool. I hate how many footguns are in common shells.
[1] eg writing to a tempfile and atomically renaming into place: "task_with_stdout > $dir/task_id.txt.tmp && mv $dir/task_id.txt{.tmp,}"
* does not buffer stderr
* does not check if the disk is full for a period of time during a task (thus risking incomplete output)
* does not clean up, if killed
* does not work correctly if task_with_stdout is a composed command
Given that GNU Parallel is a drop-in replacement for xargs, I am curious why you find it a 'weirdo interface'.https://docs.computecanada.ca/mediawiki/index.php?title=GNU_...
Using xargs for this kind of work is euhm... not a good idea.
If you seriously believe you can implement everything using xargs, then this (contrived) example is for you: https://unix.stackexchange.com/questions/405552/using-xargs-...
Newer versions include 'parset' which can set shell variables in parallel, which is useful if you want to 'map' values from one array to another.
For me, an essential feature of GNU parallel is that it is semantically equivalent to "sh". Imagine that you write a file that contains a long list of commands. You can pipe that file to "sh" to run the commands, or pipe it to "parallel" to do the same, but faster. If you are building the list of commands on the fly, then you can use xargs with a slightly different syntax. But somehow using "sh" or "parallel" gives a certain peace of mind due to its straightforward semantics. I never used any argument of GNU parallel apart from -j
My usage pattern: to build a list of commands explicitly then run it (possibly teeing the list into a temporary file to inspect it):
for i in one two three; do
printf "echo $i\n"
done |sh # or |parallel[0]: https://github.com/archlinux/svntogit-community/tree/package...
... which is exactly what GNU Parallel is. Your concern is even mentioned in the design documentation: https://www.gnu.org/software/parallel/parallel_design.html
https://git.savannah.gnu.org/cgit/parallel.git/tree/doc/cita...
Obviously downside to the visibility and dynamism is that it redirects stdout. You can read it back later, in order. But it’s not there for continued processing immediately.