GNU Parallel 2018
zenodo.org
zenodo.org
All the components I've wrote are optimized to the core (short of going assembly). It allowed me to scale my operations up to 10000 processes continuously for the duration of the test. All this was done on a 4G connection with roughly 40Mbps bandwidth at the time on a virtualized (VirtualBox) CentOS LAMP installation.
It blew all my expectations, without parallel it would taken more than a month. Learned a lot from the experience developing it and maybe will make a post sharing what I've learned.
In regards to being complicated, well it isn't really when you consider its use scenarios, for these sorts of tasks, it's a indispensable tool.
Think of it this way: It is far less difficult than the difficulty of the tasks where you need to use it.
One hour on that connection means a maximum of 18GB of data, which seems surprisingly small for a whole country's internet (not that I have any point of comparison), small enough that you could easily do interesting analysis of the whole lot.
True and after that I realized how small the Romanian internet really is in the big picture, surely allowed me to put things into perspective. It's small enough that it would be doable to have a near realtime update of the index since less than 10% of sites update daily and less than 1% on hourly basis such as news sites.
It correlates with traffic details of top trafficked webs in Romania where one of the top sites gets around 200k uniques a day, a small amount when you consider the big picture. It is likely that today the amount of data would be larger but not by much.
https://www.gnu.org/software/parallel/parallel_alternatives....
Every time I reach for it the magical incantations that finally get it working properly are so completely unmemorable that it’s begging to have a modern replacement (like httpie is to curl).
If someone wrote a version with a better interface it would be instantly adopted in preference to this. Might be a good weekend project.
Ah, here it is: https://gitlab.redox-os.org/redox-os/parallel
You don't see Coq users crapping on other languages for not proving their programs correct according to a formally defined specification.
Speed wise they are comparable to GNU Parallel. So the argument for Rust could be speed.
The argument against Rust could be backwards compatibility: How would you do the {= perl expr =} thing? Will it run on the current targets (CentOS 3.9 according to https://www.gnu.org/software/parallel/parallel_design.html#O...)? Will it run on systems for which you as a consultant do not have access to a Rust compiler for?
I only have quad core, but -P values higher than 4 give better performance. Maybe because the tests take different amounts of time, and it avoids them backing up. To get x7, I had -P equal to the number of tests.
real 0m1.818s
user 0m2.020s
sys 0m0.900s
top. (It completes in 1 sec, so put in a while loop - BTW top isn't behaving normally, "1" not giving multiple CPUs and not fitting on screen). Jus the top line: User 51%, System 23%, IOW 0%, IRQ 0%... | xargs -P10 -i sh -c 'command {} 1>{}.out 2>{}.err'
or maybe with tee:
... | xargs -P10 -i sh -c 'command {} | tee -a {}.out'
It is definitely not something I would dare to run in production.
xargs is part of POSIX, so it should be everywhere.
(-P/-n are not, but both GNU and BSD versions do support it.)
EDIT: Parent comment edited his post, leaving out his personal preference of language.
My first thought was, "12 pages isn't so bad." Then I loaded the page and realized you had a typo and meant _112_ page manual.
That is enough to scare me off of trying to use parallel for casual tasks.
Just because emacs is such a complex software, does that mean no one can start learning it? No. As far as GNU parallel goes simply just understand what `:::` operator does and you can start using it in your workflow today. Software is not complex for the sake of beginners, it is complex so experts have a higher variety of features.
It doesn’t require reading the whole manual, nor does the length of the manual have anything to do with parallel’s ergonomics.
Do you use bash or ssh? Have you looked at the length of those manuals? Most of parallel’s manual is advanced workflows the casual user doesn’t need, stuff like parallel commands run over a network of machines and collecting the results on the originating host.
Maybe you could show some examples of what you mean, because I think parallel is pretty easy to use. Personally I feel like it’s usually easier than Xargs, especially when dealing with tricky quoting. It does have some of it’s own syntax, like the triple-colon :::, and it’s really easy to use. Several pages of manual are about variable substitutions you can use if you want, they’re very similar to bash variable substitutions, and they’re all optional and super convenient. A bunch of parallel’s manual pages are devoted to simple examples, which is nice, and the kind of thing that’s missing from so many man pages.
> it’s begging to have a modern replacement
I’m not sure what “modern” means. You could certainly propose some changes though, or show some examples of the kind of syntax that isn’t working for you. Httpie and curl are built to do different things.
You can check it out here:
Like... the ability to run things in parallel?
Loop looks nice, don’t get me wrong. I think I’ll try it! It just seems like comparing it to Parallel is misplaced here, and a bad idea in general, since they’re not remotely related. Ole is comparing all these utilities to Parallel to show how general Parallel can be, and suggest you only need his one program and one syntax to emulate hundreds of different utilities: his argument is implicitly that people should use Parallel and not loop.
$ yes | parallel -N0 -j1 echo touch '{= $_=seq()*5 =}'.txt
vs
$ loop --count-by 5 -- touch $COUNT.txt
UNIX is about using the right tool for the right job, not the same tool for every job.
Let us remember the koan:
A Unix novice came to Master Foo and said: “I am confused. Is it not the Unix way that every program should concentrate on one thing and do it well?”
Master Foo nodded.
The novice continued: “Isn't it also the Unix way that the wheel should not be reinvented?”
Master Foo nodded again.
“Why, then, are there several tools with similar capabilities in text processing: sed, awk and Perl? With which one can I best practice the Unix way?”
Master Foo asked the novice: “If you have a text file, what tool would you use to produce a copy with a few words in it replaced by strings of your choosing?”
The novice frowned and said: “Perl's regexps would be excessive for so simple a task. I do not know awk, and I have been writing sed scripts in the last few weeks. As I have some experience with sed, at the moment I would prefer it. But if the job only needed to be done once rather than repeatedly, a text editor would suffice.”
Master Foo nodded and replied: “When you are hungry, eat; when you are thirsty, drink; when you are tired, sleep.”
Upon hearing this, the novice was enlightened.
$ seq 0 5 100000 | parallel touch {}.txt
Or you you really want the infinite loop:
$ yes | parallel touch '{= $_=seq()*5 =}'.txt
'$' must be quoted
$ loop --count-by 5 -- touch '$'COUNT.txt
This seems like a pretty weird example. All common tasks in curl are very straightforward and I've never even heard of httpie.
Let's please not change the interface to basic command line tools every few years because people want something ""modern"".
For me the trickiest thing is remembering how to deal with redirection. I find the --dry-run option to be super useful when setting up a command.
[0]: https://www.youtube.com/playlist?list=PL284C9FF2488BC6D1
I often create bash functions to avoid that shit, and then call the bash functions from GNU Parallel instead.
It used to be an issue that non-free UNIXes did not include a C-compiler by default, and then C would really not help. I am not sure if that is the case anymore.
https://manpages.debian.org/testing/moreutils/parallel.1.en....
parallel sh -c "echo hi; sleep 2; echo bye" -- 1 2 3
parallel echo hi";" sleep 2";" echo bye ::: 1 2 3
parallel -j 3 ufraw -o processed -- *.NEF
parallel -j 3 ufraw -o processed ::: *.NEF
parallel -j 3 -- ls df "echo hi"
parallel -j 3 ::: ls df "echo hi"
And if you are looking for the single page overview, why not link to the cheat sheet: https://www.gnu.org/software/parallel/parallel_cheat.pdfIf you have a hard time getting it to work and remember how, then it is likely that the fundamentals of GNU Parallel are not clear to you. Chapter 2 ("Learn GNU Parallel in 15 minutes") of the book should take care of that.
> a better interface
That is often a tough one. I, for one, like emacs' interface. Others despise it. So when you write 'better interface' I think it will be helpful if you were more specific.
Removing that requirement is a two line patch:
4723,4724d4722
< 1
< orIt's also a reasonable thing for a user of free software to modify the software and improve it by removing this kind of garbage.
Run 'parallel --citation' ONCE and have the notice silenced.
I never understood the complaint: Is it because people do not _read_ the message?
As a result of parallel's performance and invocation complexity I find myself using xargs -P more frequently than parallel.
parallel -a input_file_here -j16 --gnu 'ssh -q {} "echo 'Matrix has you'"'
parallel --slf input_file_here -j16 --nonall echo 'Matrix has you'