GNU Parallel Tutorial
gnu.org
gnu.org
I either use xargs, or this reimplementation which had all the stuff I need:
When I see it, I get irritated, but then I disable it and completely forget for another year.
It is not a legal requirement or license requirement to cite usage of parallel, and the GNU terms actually forbid making it a requirement:
https://www.gnu.org/licenses/gpl-faq.en.html#RequireCitation
The nag's wording is also annoying and wrong about academic tradition requiring citations of all tools used, whether or not the tool is integral to research. Academic tradition does require citing others' research you build your research on top of. It does not require citing MS Word because I wrote my paper in MS Word.
This is a small and stupid snag on what is otherwise a really awesome tool. My recommendation is to make the citation go away and continue using the tool, and if you use it in a way that your research actually depends on, then give the citation. If you just use gnu parallel to speed things up, and it in no way changes what you're doing or how you're thinking about it, then ignore it. Though I'm sure Ole would love to hear about projects that use parallel - what he's asking for is PR not citations, and blog posts or press or kind words that praise gnu parallel will also help him get what he needs.
Again, Ole really just wants PR here, he made that clear in the email thread. He's only asking for help, in a forceful way. If you have a reason to give him some help, and you make heavy use of parallel, it would make his day and it wouldn't hurt you.
Otherwise, there is not legal requirement nor academic tradition to cite usage of parallel in academic papers, unless the papers build directly on gnu parallel.
I really wouldn't be surprised if part of this is Ole felt like there were research papers about parallelism that built on gnu parallel's efforts directly and should have cited him, that he's been short changed by a few academics.
But simply speeding up your data processing for doing something in a field that is not related to parallel processing at all, then parallel is simply another random tool in a huge bucket of tools to get your job done, and it has no place in an academic citation.
FWIW, keep in mind a few things:
- Note that @CJefferson is afraid to use parallel because Ole made the license restrictions unclear by adding the citation request. Ole is discouraging some applications of his software by being so forceful about his request.
- Ole's citation is not from a scientific or peer-reviewed publication. It simply cannot be used in some contexts. Even the BLAS link above was in a research journal, but Parallel's is not. In many contexts, a citation of Parallel is highly inappropriate.
- The small amount of time it takes to either cite Ole or read and understand this thread is largely irrelevant to the question of whether a citation is either warranted or appropriate or allowed. Understanding the license is a one-time event, and it's very important for the license to be legally clear. Parallel's citation notice is causing confusion. The question is whether Parallel can and should be used without having to worry (forever) about the legal consequences of the contract I've agreed to by using the software.
- This approach isn't scalable. If other GNU tools, or other free software started using the same language as Ole, it would cause widespread problems.
And thus we are stuck with "please cite this so I can justify working on it, please?" instead of "You must cite this if you want to use it".
Honestly, I also don’t see how this request isn’t scalable, could you expand on that?
Anyway: there's a wonderfully specific answer to this in the FAQ:
https://www.gnu.org/licenses/gpl-faq.en.html#RequireCitation
I understand that I cannot add this to the GPL (due to the specific clauses of the GPL) and that RMS might not consider the resulting software free (though it would probably pass all of Debian’s freedom guidelines, for example), but it should be possible to have such a license in general?
I guess a good example would be a Madlib-style program, where it asks you questions and fills in the blanks in a story, often resulting in something highly amusing. The original story containing the blanks is copyrighted. The output of this program would then be a derivative work, because the original story has been modified.
But consider that the program took a story of your own (the data its processing) and output statistics, such as the word count, frequency of certain words, grammar errors, etc. This is not a derivative work. Similarly, if GNU Parallel is being used to process your input, its output isn't a derivative work.
With that said, you can have a separate EULA-type thing---which is _outside_ the domain of copyright---that imposes these terms. But that is incompatible with the terms of the GPL.
In these cases, the EULA actually may contain clauses that allow you to distribute gameplay videos. But if they do not contain exemptions, it's copyright that will restrain you, not the EULA.
http://www.develop-online.net/analysis/uploading-gameplay-co...
https://support.google.com/youtube/answer/138161?hl=en
Now in the specific case of GNU Parallel, I don't see how the output would contain any distinctive elements of the original program. As a counterexample you could not use GNU parallel to process its own source code and end up with your own copyright on the output, however.
Yes, by stating that it's a derivative work, I meant that the output would be subject to the Madlibs copyright.
GNU Parallel's output isn't subject to the GPL.
If you're thinking, well that still only is effective if your use involves making copies of the software, note that US courts have held that just using software and thereby causing it to be copied into RAM counts as making a copy for the purpose of copyright law (see MAI Systems v. Peak).
If you pay 10000 EUR you should feel free to use GNU Parallel without citing.
That sounds to me like cite or you will be taken to court for not citing.
If I cite parallel, should I cite Perl? C library? Linux kernel? Systemd?
Examples of the latter are e.g. Linux, libc, systemd etc.: All things used during production of the results, but not aimed at producing such results. The former are e.g. Matlab, the Intel MKL, ALPS, TensorFlow or similar tools and suites. A bug in the MKL may make your results wrong, while a but in the kernel will most likely just crash everything. In my experience, scientific results depend on a handful of software packets in such a way, and adding five-or-so citations to a paper is certainly no big deal and would make reproducing the results much easier.
I have no idea where GNU parallel sits there, it probably depends on how you use it – do you schedule your Monte Carlo runs with it or just rescale pictures for the website? Admittedly, it is a very grey area, but I don’t consider a request for citation unreasonable.
In other words, it effectively requires me to LIE.
No thanks.
The author just asks for citations because he's in academia, and citation backrubs is how anyone can justify getting time to do anything (including maintaining some software) in academia.
- Isn't it common courtesy to cite important tools being used for data processing? Why is this a problem?
- Wikipedia says the tool is under GPLv3, so from a bare licence perspective you should be safe to use it for whatever you want, shouldn't it?
No, no it's not. It is common to cite previous work, the research you build on top of. It is not common at all to cite tools you use, unless those tools are directly related to your research. If by "important" you mean gnu parallel counts as previous work, then citation is more of a requirement than a courtesy, otherwise citing all the tools you used can be very discourteous to your readers.
It would be extra courteous to give parallel some PR if you love it and it really helped your research get done.
> Why is this a problem?
What if all gnu tools asked the same thing parallel does? What if using emacs and find and tar all asked you to cite them when using them, with obnoxious and demanding language? I use more than 100 free tools to get any paper implemented and written, my papers would be rejected if I tried to cite them all, the vast majority are not relevant to the research.
> you should be safe to use it for whatever you want
Yes, correct! GPL does not allow requiring citations.
I understand that the copyright holder can in general specify whatever clauses in the license, even if this makes it self-contradictory, and is not bound by any statements by FSF. In this case, it becomes unclear what use of the software is actually allowed.
"When you convey a copy of a covered work, you may at your option remove any additional permissions from that copy, or from any part of it."
https://www.gnu.org/licenses/gpl.html
Note that the GPL only applies to distributors of the software, not to users, so you wouldn't be bound to such restrictions anyway.
For example, a copyright holder may say "MIT license, BUT, the Software shall be used for Good, not Evil.", and lawyers will balk. [1,2]
[1] http://www.json.org/license.html [2] https://www.cnet.com/news/dont-be-evil-google-spurns-no-evil...
But if they choose to use the GPL, they must abide by its terms. They aren't free to rewrite parts of it---their license would then be a derivative work of the GPL, which isn't permissible. Section 7 would have to be modified to say that the user can't remove such terms.
If they use the GPL and impose extra terms, state that Section 7 is invalid, then there is a contradiction in the license, and it'd be up to a court to decide. I'm sure any sensible lawyer would tell them that they should use a different license; it's senseless using the GPL if you're going to try to do such a thing.
It's of course the case that using GPL and adding such clauses does not make any sense. GNU parallel does not appear to be doing this, as the license text is unmodified GPL with no additional clauses.
> All other non-permissive additional terms are considered “further restrictions” within the meaning of section 10. If the Program as you received it, or any part of it, contains a notice stating that it is governed by this License along with a term that is a further restriction, you may remove that term.
"This License explicitly affirms your unlimited permission to run the unmodified Program. The output from running a covered work is covered by this License only if the output, given its content, constitutes a covered work."
While these additional requirements may (or may not) be possible and legal, in the case of GNU Parallel, the license is unmodified GPLv3. The project would need to identify a separate license that inherits GPLv3 and explicitly states those additional requirements, but in this case it does not. GPLv3 is used directly and verbatim on the project home page https://www.gnu.org/software/parallel/
The text the executable spits out asks for citations, but it does not identify itself as a license, it does not state that citations are a legal requirement, and it does not state a relationship to GPLv3 nor carve out exceptions or additional requirements.
I would cite parallel if it is required for the reader to understand or reproduce my research (for example, I report timings which were sped up by parallelisation). Otherwise it's just a tool, Like bash, the C standard library or the Linux kernel.
Point 2, the author is giving the code two contradictory licences. I'm not sure which applies. Do I get to choose or does he? I'm not a lawyer.
By whom? Was it the author? If so, can you cite your question on a mailing list?
# Resize all jpgs to 800x600 using 8 jobs/cores.
parallel -j 8 convert {} -resize 800x600 {.}_small.jpg ::: *.jpg
# Or get the help of a few servers (via SSH) to do the same job.
parallel -S serverA,serverB -j 8 --transfer --return {.}_small.jpg convert {} -resize 800x600 {.}_small.jpg ::: *.jpgFun tip - if your resize script prints only the output file name to stdout, you can pipe the result into another parallel command, e.g.,
parallel <resize to 800x600> ::: *.jpg | parallel <resize to thumbnail>
This way generating thumbnails runs on the small image instead of the full size image, saving more time.Also, if you're on a mac, using sips is quite a bit faster than imagemagick, and sips comes built-in.
I've never used servers with parallel for image resizing, what are the benefits? I'd have guessed it would take longer and saturate the network, versus doing it locally. Is it useful for really long running jobs when you don't want to load your local cpu? Is it actually faster sometimes, or are there other more important reasons? I could see it being useful if I didn't have a local imagemagick install, but had access to servers with it there. Maybe it'd be useful in cases where I'm running docker environments for the job processing? What other use cases and scenarios have you run into?
I often sort and process images with a macbook on my lap curled up on the sofa. Resizing locally makes the mac blazing hot with fans whining - not comfortable and kills the battery. Transfer speed is not too bad over 802.11AC. So I use it mainly as a method of moving cpu intensive work away from my lap.
Also, I sometimes use
parallel -S server,: .......
The semicolon adds the local machine to the list, it will saturate both the laptop and whatever it manages from the other computer. I have to admit I've never tested this scientifically, but it seems to be faster even with the overhead (gain of remote imagemagick seems to be more than cpu overhead of SSH file transfer).:) I have the exact same workflow: MacBook+Sofa. Totally trying out the server options today.
[1] http://www.gnu.org.ua/software/pies/
list-hosts |
xargs \
-P8 \
-n1 \
-I% \
ssh % some-commandThat on its own is super useful for what I'm working on right now. But what would make it even more useful, is: can you get GNU make to use 'sem' instead of its own jobserver? That way I could run almost everything I need to under one overall task limit, and that would be really nice to have.
(For this reason, I'm a fan of the idea that every program with its own 'parallel execution' mode should be able to interact with a common jobserver. The 'make' jobserver is, as far as I know, the simplest, and should be pretty easy to support: http://make.mad-scientist.net/papers/jobserver-implementatio... )
Are you running parallel make tasks where each task is also doing something multi-threaded or parallel? Like using make -j 8 won't work for you?
Make does have the -l load average task limiter when but I've never gotten it to work reliably, it always starts way too many jobs at first and chokes for a while before calming down. Often that won't work for me, but maybe it will help you?
I know what you mean about the load average limiter - parallel behaves like that too. I think the --delay option to parallel is supposed to solve that (I haven't tried it - will try tomorrow), but I don't know if make has anything similar.
Finally, on further reading, it definitely seems technically possible, even if it hasn't been done so far. The make documentation has a section on the jobserver protocol, which looks complete enough to write both the client and server parts: https://www.gnu.org/software/make/manual/html_node/Job-Slots...
So if nothing exists so far, it's something I might look into writing myself.
Holy moly, that's kinda crazy, but would be fun, you should totally do it! Looking forward to your blog post! ;)
Another little thing I realized a while back is that `make` (yes the crusty old make + -j flag) can be used to parallelize jobs. We do it for compiling usually, but it can be used for other jobs as well.
Make was right there under my nose I just never imagined using it for anything but compiling and building things. In that case I was forced by circumstances (was developing on a constrained ancient version of RHEL), couldn't use GNU Parallel and someone suggested `make`. The use case of obvious once a co-worker mentioned it. But it was definitely It was one of the memorable "thinking outside the box" example as they say.
"Install the newest version using your package manager or with this command:
(wget -O - pi.dk/3 || curl pi.dk/3/ || fetch -o - http://pi.dk/3) | bash
"facepalm
I agree it doesn't look ideal and may not be best practice, but what do you feel is a better realistic alternative? What is the main issue for you? Is it the lack of a hash or checksum to verify what you downloaded, and make sure you didn't get a malicious site or a compromised package?
This is far less secure then running apt-get, yum, or even pacman.