It seems much easier to follow when you can push everything to horizontally scaled single processes for languages like Python.
It seems much easier to follow when you can push everything to horizontally scaled single processes for languages like Python.
All these are much more reliably solved with horizontal scaling.
[edit] by the way, a very useful minimal sugar on top of multiprocessing for one-off tasks is tqdm's process_map, which automatically shows a progress bar https://tqdm.github.io/docs/contrib.concurrent/
Suppose we want to know the status of the current task: how many tasks are
completed, how long before the work is ready? It's as simple as setting the
progress_bar parameter to True:
with WorkerPool(n_jobs=5) as pool:
results = pool.map(time_consuming_function,
range(10), progress_bar=True)
And it will output a nicely formatted tqdm progress bar.A multiprocessing implementation is a good prototype for a distributed implementation.
Parallelizing across machines involves networks, and well, that's why we have jepsen, and byzantine failures, and eventual consistency, and net splits, and leadership election, and discovery - so in short a stack of hard problems that in and of itself is usually much larger than what you're trying to solve with multiprocessing.
https://parsl.readthedocs.io/en/stable/faq.html
https://parsl.readthedocs.io/en/stable/userguide/monitoring....
For batch pipelines on that work many requests, having a serial workflow has a lot of the advantages you mention. Serial execution makes the load more predictable and makes scaling easier to rationalize.