https://www.gnu.org/software/parallel/parallel_alternatives....
It has a lot of features that can feel excessive at first glance but if you have felt some pain in building jobs, most of it is pretty sensible and much better than rollyourown.
https://www.gnu.org/software/parallel/parallel_alternatives....
It has a lot of features that can feel excessive at first glance but if you have felt some pain in building jobs, most of it is pretty sensible and much better than rollyourown.
`nproc` is a relatively standard utility (coreutils). So, xargs -P$(nproc) gets you core (or core-proportional) parallelism.
Grouping output/Making a safe parallel grep is also easy-ish with `--process-slot-var=slot` and sending to `tmpOut.$slot`.
Jobs on remote computers can be done similarly with any kind of `arrayVar[$slot]` setup where `arrayVar` has a bunch of `ssh` targets, possibly duplicates if you want to run >1 job per host. (In pure POSIX sh you could use eval and $1, $2 positional args with shell arithmetic..)
Anyway, those three are just off the top of my head, unfairness-wise. Last I looked at the source for GNU parallel it looked like mountains upon mountains of Perl I would rather not depend upon, personally, but to each his own.
Well, there was a Rust version with zero Perl, now unfortunately archived. It wasn't 100% on a par with the original and wasn't really finished. On the other hand, built easily for Windows and helped me on a few occasions.
#!/bin/bash
if [ "${1-0}" -lt 1 ]; then # No arg / arg not a number >= 1
echo "Usage: $0 <N>"; echo "reads cmds from stdin, running up to N at once."
exit 1
fi
TMP=`mktemp -t stripen.XXXXXX`
trap 'rm -f $TMP; exit 0' HUP INT TERM EXIT
STRIPE_SEQ=1
while read cmd; do
jobs > $TMP # jobs | wc -l does not work
if [ $(wc -l < $TMP) -ge $1 ]; then
wait -n # Wait for 1/more jobs to finish
fi # Could accumulate total $? above, but would need to replace final wait.
STRIPE_SEQ=$((STRIPE_SEQ + 1))
( eval "$cmd" ) < /dev/null & # Run job in a subshell in bg
done
rm -f $TMP
wait # Wait for all to finish
and if bash ever grows some magic environment variable $NUM_BG_JOBS or you don't want auto-help or sequence numbers or etc. it can be even simpler.but still miles better than other such languages, and esp. if it would have been written in C, just as the incompatible moreutils counterpart.
parallel -k --tag --argsep -- {} echo ::: 1 -- parallel-*
Every version since 20120622 work (except for 20121022). That is code which is almost 10 years old.Compare this with PHP, whose breaking changes between releases has taken down my sites on multiple occasions.
Compare this with Python, whose breaking changes prevent me from running the overwhelming majority of Python things I've tried to use.
Subroutine signatures are an experimental feature in Perl. Or are you referring to something else?
i used parallel for years under the assumption that it was written in C and only recently learned it was written in perl when i decided to dive deeply into its documentation. if you're using a package manager to install parallel and it runs fast enough for your needs (it does) then who cares what language it was implemented in?
how bad was the perl?
I am curious how you come to that conclusion.
> `nproc` is a relatively standard utility (coreutils). So, xargs -P$(nproc) gets you core (or core-proportional) parallelism.
I follow you on this point. A bit harder on remote systems, but definitely doable.
> Grouping output/Making a safe parallel grep is also easy-ish with `--process-slot-var=slot` and sending to `tmpOut.$slot`.
I tried spending 5 minutes on coding this, but the details seem to be very hard to get right: composed commands, grouping stderr, combined with not leaving tmp files behind if killed and allowing for the total output to be bigger than the free space on /tmp. I could not do it.
Could you consider spending 5 minutes on showing in code how you would do it?
> Jobs on remote computers can be done similarly with any kind of `arrayVar[$slot]` setup where `arrayVar` has a bunch of `ssh` targets, possibly duplicates if you want to run >1 job per host. (In pure POSIX sh you could use eval and $1, $2 positional args with shell arithmetic..)
This one seemed even harder to me: It was completely unclear how you would make sure that a given number of jobs were constantly running. And how would you need to quote data, so an eval would not cause "foo space space bar" turn into "foo space bar". And how you would kill remote jobs, if the local script was killed.
If you believe this is simple, could you spend 5 minutes on showing the rest of us how you would do it in actual working code? Because it seems the devil is really in the detail.
> Last I looked at the source for GNU parallel it looked like mountains upon mountains of Perl I would rather not depend upon, personally, but to each his own.
Personally, I would take production tested code over home-made untested code any day - no matter the language in which it was written.
This is an unreasonable standard when you do not know in advance how big the output is. What do you imagine GNU parallel does? Use `df` on every host it knows about to fill every disk partition it can? That sounds like a pretty system-hostile behavior to me.
Meanwhile, putting your temp files somewhere bigger is obv. as easy as $TMPDIR or such.
Best wishes/luck. I only have 5 minutes to explain why nothing can do the impossible like read a user's mind about disk free space management or the value of partial results. All software makes some assumptions... :-)
Why is that unreasonable?
Let us say a single job outputs 10% of the free space. As long as you run fewer than 10 jobs in parallel, GNU paralel can run forever, because it spits out the output when a job is done and then frees up the space for this job, while starting the next one.
A simple example:
yes 1000000 | parallel -j10 seq | pv >/dev/null
On my laptop I get 600 MB/s which would fill /tmp in a few minutes, and it does not.When dealing with big data it is not uncommon that the total data piped between commands is way larger than the free space on /tmp (which is typically fast, where as free space on $HOME is slow - thus setting $TMPDIR to $HOME/tmp may slow down your job drastically).
If you only have 5 minutes, I hope you will use them on providing actual code to support your claim, that "The comparison is not very fair to modern day xargs."
If it takes longer than 5 minutes to code, I would say your use of "easy-ish" is unwarranted.
You leave me with the feeling that you have not thought this through and that the reason why you do not provide any code is because you are now realizing you are wrong, but you do not have the guts to admit so.
Prove me wrong by posting the code. It should be "easy-ish" :)
You can use this as the test case to implement:
yes 1000000 | parallel -kj10 "echo 'This is double spaced '{#}; seq {}" | pv >/dev/nullSpeaking of /tmp filling and questionable space management defaults:
yes 2000000000 | parallel seq | pv > /dev/null
fills my /tmp disk partition (or $TMPDIR) before emitting one byte to pv with invisible (unlinked) temp files. Not ideal. GNU sort at least shows me there are files present yet also seems to clean up on Ctrl-C.There is likely some solution to fix this in 15 kLOC of gross Perl. I did not find it in "5 minutes" (another unreasonable standard since the many 1000s of lines of GNU parallel docs take far longer to read, but you already seem to ignore my explanations of "unreasonable"). You even anticipate this in your 10% example. At least in my life, "way more" is often much more than 10x more. So, you basically contradict yourself.
As to the actual subtopic, besides being unfair/out-of-date, the comparison tableau is also incomplete - maybe willfully so, as per too common marketing dishonesty. "Proof?" People use parallelism to speed things up and need to make decisions about job granularity to not have perf killed by overhead. Some would say this matters more than 95% of the tableau evaluation points. Yet, no overhead benchmarks. Maybe they make GNU parallel look bad?
I included the example:
yes 1000000 | parallel -kj10 "echo 'This is double spaced '{#}; seq {}" | pv >/dev/null
to give you some fixed "goalposts" to aim for: Provide a solution that gives the same output byte for byte.Also you do not seem to get the point about the amount of data. I regularly have output from a single job that is bigger than RAM, but rarely have output from a single job that would fill /tmp. However, the total combined output from all the jobs will often take up more space than /tmp.
In numbers: RAM=32 GB, /tmp=400 GB, a single job=33 GB, number of jobs=1000, jobs in parallel=8.
In other words: Running all jobs and saving the outputs into files before outputting data will not be useful for me. If you want to use FIFOs I really cannot see how you can deal with output that is bigger than RAM, unless you mix output from different jobs - which again would not be useful to me. But prove me wrong by spending 5 minutes on building the solution.
As for your example:
yes 2000000000 | parallel seq | pv > /dev/null
How would you design this, if output from different jobs are not allowed to mix?If they are allowed to mix paralel gives you:
# bytes are allowed to mix
yes 2000000000 | parallel -u seq | pv > /dev/null
# only full lines are allowed to mix
yes 2000000000 | parallel --lb seq | pv > /dev/null
none of these use space in /tmp.I sit back with the feeling you are willing to spend hours complaining, but not 5 minutes on proving your assertion that it can be done "easy-ish".
Prove me wrong: Spend 5 minutes on the task you believed was "easy-ish".
If it cannot be done in 5 minutes, be brave enough to admit you were wrong.