Parallel – A command-line CPU load balancer written in Rust
github.com
github.com
But this does claim to be a reimplementation of much of GNU Parallel's functionality. If it's compatible, maybe there isn't too much of an issue.
In the interest of silly UNIXism, I hereby vote for the name "serial" instead. It's just like "more" and "less", right?
There is no need for threads. Just spawn background processes:
echo 1 &
echo 2 &
wait
echo 3 &
wait
echo 4 &
...
The key is that the wait() system call will hang until any child process finishes.It's like discovering that "ls" is spawning 5 threads for internal communication.
In particular there is no sane way to async waitpid() on POSIX.
... I'm not sure I can disagree with "no sane way".
In fact, I challenge you to solve the problem "spawn process; waitpid for 15 seconds; otherwise kill hard" in Rust (or C++ if you feel like) on POSIX once with threads and once without threads by sticking to what's permitted in the standard and so that multiple processes can be waited for.
Then also measure CPU impact :)
This is false, you can call any async signal safe function. Incidentally write is one of them.
Another trick is the close-on-exit pipe.
If you change a variable, it has to be one that isn't prone to being cached in a register or the stack by the main program. POSIX's sig_atomic_t does this; in Rust you can use the normal atomic types. They are a tiny bit too careful if this is thread-local, but an ordinary thread-local variable is permitted to be cached within the same thread, and signal handlers break that.
If you take a lock, you have to do something reasonable if the lock is already held, including by the code you interrupted. So you probably shouldn't lock at all. The biggest reason for a POSIX function not to be async-signal-safe is because it wants to call malloc, which takes out a lock (at least a per-thread or per-CPU lock) on the heap. If you get signaled during a malloc, and the signal handler tries to malloc, you deadlock.
But anything that does not risk liveness or correctness problems is fair game. In particular, basically all system calls are fair game, since they're just sending a message to the kernel. C's fprintf() will want to buffer in userspace, which involves an allocation, but write() will at most buffer in the kernel, and the kernel-side code doesn't have the problem of having flow control interrupted while you're in a signal handler. Even if you were previously in a blocking write() when you received a signal, the kernel will return from its implementation of write before delivering the signal back to userspace, so there isn't a re-entrant call to the kernel-side write code. libc's fprintf() doesn't have that luxury.
(And yes, the concept of Rust on POSIX is a bit ill-defined, because POSIX is a set of C-language APIs, which can be implemented in any valid way in C, including header macros. Rust threads use pthreads, yes, but inter-thread communication doesn't involve whatever sig_atomic_t is typedef'd or #defined to.)
Worse: most signal handlers that do anything other than setting globals destroy errno in one way or another and when you go back the code that was interrupted has a good chance of malfunctioning.
Which malloc() is not. Anything that might internally allocate is out of the question. The list of functions that are safe to call in C alone is very limited and even then the question of errno arises.
https://play.rust-lang.org/?gist=ba4802a59f462cb8bce0c1bac92...
It would be significantly simpler with sigwaitinfo() if you didn't care to do a real select loop, but I assume most programs want a real select (or poll/epoll/whatever) loop.
It's not composable, but it's not impossible to work around that if you can assume non-POSIX. Platforms with kqueue just get this right because kqueue can wait for processes. Linux will let you use a different signal to alert for process completion, so pick one of the realtime signals (which is its own game of global namespaces, but hey), and library code that uses SIGCHLD won't be affected. So that's Linux, OS X, and the BSDs. I don't know of a good workaround for Solaris.
A POSIX-compliant way that makes this more composable is to make a single, dummy process to be a process group, setpgid() all your actual children into that process group, and have that process spend its time doing waitpid(0, WNOHANG) and sending you notifications over a pipe or something. That is higher overhead, but my guess is that scales better for many child processes (O(1) extra processes vs. O(n) extra threads).
In that case you are pretty much limited to polling on the thing (unless apparently you use kqueue, did not know you can wait on process events).
[0] https://www.jonathanturner.org/2015/10/lessons-from-first-12...
I added an exception as I pretty much always do but in case the owner of the site is around here, they might want to investigate this.
g1:
job1
g2:
job2
and then one goal to depend on them all and I used make -j 4> parallel dates back to around the same time. It was originally a wrapper that generated a makefile and used make -j to do the parallelization.
find . -type f -name | parallel --jobs 5 process_file
That will run "process_file foo" for ever file in a directory with 5 processes running in parallel. find . -type f -name | xargs -I {} -n 1 -P 5 process_file {} parallel --progress --wd '.' -S '4/fourworkers@blehworker.org,8/eightworkers@eightworkers.org,2/:' "./complicated_command --param1={1} --param2={2} ::: $(some_shellcommand_to_generate_param1_values.sh) ::: $(some_shellcommand_to_generate_param2_values.sh)
It's very flexible, become my go-to for speeding up parallelizable work. Next step up would be Celery or Spark or something.The code is 427 lines of C. GNU parallell is 10k+ lines of perl. Considering mmstick compares the loading times, it's easy to see where the difference comes from.
It lets you code your own replacement string (using --rpl), and lets you make composed commands with shell syntax:
myfunc() { echo joe $*; }
export -f myfunc
parallel 'if [ "{}" == "a" ] ; then myfunc {} > {}; fi' ::: a b c
It does not need a special compiler, but runs on most platforms that have Perl >=5.8. Input can be larger than memory, so this: yes `seq 10000` | parallel true
will not cause your memory to run full.You can read a lot more about the design in `man parallel_design` and see the evolution of overhead time per job compared to each release on: https://www.gnu.org/software/parallel/process-time-j2-1700MH...
In other words: Treat GNU Parallel as the reliable Volvo that has a lot of flexibility and will get the job done with no nasty corner case surprises.
It is no doubt possible to make a better specialized tool for situations where the overhead of a few ms per job is an issue and where you neither need brakes, seatbelts nor airbags. xargs is an example of such a tool, and you can have both GNU Parallel and xargs installed side by side.
A few years ago, debian made GNU parallel provide the "/usr/bin/parallel" executable, instead of moreutils. The maintainer of moreutils had some interesting things to say about that: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=597050#75
Only 5k? :D
But, yes, the criticisms are valid. I recommend moreutils.
It felt like his complaint was that that was 'bloat' since Real Men can achieve almost the same thing by just piping some output through some bash scripts they just hacked together.
But non-expert users will invariably make mistakes (e.g. get quoting wrong, not getting remote jobs to die if the controlling process is killed, or re-scheduling jobs that were killed by --timeout), and why not just have small wrapper scripts built into GNU Parallel that are well-tested, so the non-expert users can enjoy the same stability as the expert users?
Quite a few of the lines are to deal with different flavours of operating systems.
The benefit of this is that by copying the same single file you can have GNU Parallel running on FreeBSD 8, Centos 3.9, and Cygwin.
While this is technically correct, it is misleading: GNU Parallel existed before it became GNU. See details on: https://www.gnu.org/software/parallel/history.html
To add to this, parallel mentions as much in its man pages (that there is a certain startup cost, and a certain job-startup cost), and offers tips for speeding up the processing of jobs which exit fairly fast. But there's no reason you couldn't also do those things in the Rust version, so it's going to win every time. When dealing with commands which take a while to complete though, the extra overhead of the perl script would probably be negligible.
GNU Parallel has --joblog and can continue from where it left off or retry all failed jobs again.