Neat Parallel Output in Python
bernsteinbear.com
bernsteinbear.com
If you really want to share one terminal / stdout but also prevent timing-based record splitting, you could also send outputs to FIFOs / named pipes and then have a simple record boundary honoring merge program like https://github.com/c-blake/bu/blob/main/funnel.nim with doc at https://github.com/c-blake/bu/blob/main/doc/funnel.md As long as "record format" is shared (e.g. newline-terminated) this can also solve the stray print problem.
j.py:
import multiprocessing as mp, contextlib as cl, time
def f(x):
name = mp.current_process()._name
with open("o." + name, 'a+') as f:
with cl.redirect_stdout(f):
print(x * x)
time.sleep(2)
mp.Pool().map(f, range(36))
Then shell$ python3 j.py & sleep 1; multitail o.*
On Unix that should create $(nproc)-panes (i.e. screen areas) of output in your terminal, each pane logging any output each worker process did. (EDIT: For fancier progress, multitail -cT may help; See the man page.)Some Py MP expert could perhaps show how to do this with only one `open` or a nicer `name`. As a bonus - the log/status outputs are not erased but left in "o.Stuff". The cost of that is having log files to clean up.
Separate files means no contention at all and any stray prints from any libraries go into the log file. It's a separate exercise for the reader to extend this to stderr capture (either in e.* or also in o.*).
And, really, as popular as Python and MP are, all of this is probably written up somewhere.
Someone else should outline imitating `multitail o.*` in tmux, I think, as it is more involved, but a pointer to get started is here: https://github.com/tmux/tmux/wiki/Advanced-Use (That's all I have time for today. Sorry.)
Having good features to support this in standard libraries would go a long way to incentivizing devs to actually parallelize.
We do see some cool stuff under the hood from core Python devs but interest in further quality of life features seems to be lacking.
For a one off project it seems simpler to just write an html UI.
https://opentelemetry.io/docs/languages/python/getting-start... https://docs.honeycomb.io/getting-data-in/opentelemetry/pyth...
But there's so much complexity there that IMO it's best left outside of standard libraries - and it's indeed a daunting amount of new vocabulary for newcomers. I'm not aware of simpler abstractions on top of the broader telemetry ecosystem for monitoring simple parallelization, but arguably there should be one that keeps things quite simple.
[Parsl is much better, e.g., logging is built-in, but it can be a little overwhelming.]
An alternative would be to have only the main process do the updating and have the workers message it about progress, using a queue.
I had a threaded server that we were debugging which would only dump state correctly if we deleted a printf right before in a different thread. Really confused me until I figured this out.
If you’re using this in a CLI tool you’re writing in Python you might be using the library rich anyway, which provides this functionality as well including some extra features.
If anyone knows of a smaller, more focused library providing something similar to rich's Live Display functionality, I'd appreciate it.
https://docs.python.org/3/library/logging.handlers.html#logg...
You need to consume the iterator that map returns.
Please use Python as it is supposed to be used, not as Fortran.
Edit: never mind. Ignore that. There is a link on the page which I overlooked, with a more complete example. https://gist.githubusercontent.com/tekknolagi/4bee494a6e4483...
Cool stuff. Now I'm eager to find a way to make this work for multiple tqdm progress bars, running in parallel.
Also do you really use parallel from software trying to parallelise its internal workload? Note that in tfa the workers are an implementation detail of a wider program.
`parallel --eta` provides progress reporting over all tasks based upon observed completions.
> in tfa the workers are an implementation detail of a wider program.
TFA could write the workers as standalone processes then subprocess out to GNU Parallel.
At that point, one might not even want the Python outer process in this article. I often write little shell wrappers that establish non-trivial GNU Parallel or Make invocations under the hood. Generally, leaning on common workhorses for multiprocessing seems more sensible than rolling my own one-off.
parallel --latestline seq ::: {1..10}0000000
(Requires version 20220522)https://pythondialog.sourceforge.io/images/screenshots/mixed...
https://github.com/tqdm/tqdm/issues/1000#issuecomment-184208... https://github.com/tqdm/tqdm/issues/811#issuecomment-1368850...