Parallelising Python with Threading and Multiprocessing
quantstart.com
quantstart.com
You can write the threaded example as:
import concurrent.futures
import itertools
import random
def generate_random(count):
return [random.random() for _ in range(count)]
if __name__ == "__main__":
with concurrent.futures.ThreadPoolExecutor(max_workers=2) as executor:
executor.submit(generate_random, 10000000)
executor.submit(generate_random, 10000000)
# I guess we don't care about the results...
Changing this to use multiple processes instead of multiple threads is just a matter of s/ThreadPoolExecutor/ProcessPoolExecutor.You can also write this more idiomatically (and collect the combined results) as:
if __name__ == "__main__":
with concurrent.futures.ThreadPoolExecutor(max_workers=2) as executor:
out_list = list(
executor.map(lambda _: random.random(), range(20000000)))
In this example case, this will be quite a bit slower because the work item (in this case generating a single random number) is trivial compared to the overhead of maintaining a work queue of 200000000 items - but in a more typical case where the work takes more than a millisecond then it is better to let the executor manage the division of labour.I wasn't aware of the concurrent.futures library, thanks for pointing it out.
In [1]: import concurrent.futures
In [2]: with concurrent.futures.ProcessPoolExecutor(max_workers=8) as executor:
...: out_list = list(executor.map(lambda _: random.random(), range(1000000)))
...:
Traceback (most recent call last):
File "/usr/local/Cellar/python/2.7.6_1/Frameworks/Python.framework/Versions/2.7/lib/python2.7/multiprocessing/queues.py", line 266, in _feed
send(obj)
PicklingError: Can't pickle <type 'function'>: attribute lookup __builtin__.function failed import pip
pip.main(["install","futures"])
import random
def l(_):
return random.random()
with f.ProcessPoolExecutor(max_workers=4) as ex:
out_list = list(ex.map(l, range(1000)))
len(out_list)
#> 1000
[1] https://pypi.python.org/pypi/futures import futures as f #Include after pip.main(...Often times threads need to update shared dict/list etc... With multiprocessing this cannot be done. You can use a Queue for this but it's horribly inefficient.
Generally speaking if you need performance and Python is not meeting the requirements then you are better off using another language.
Generally I would use C++ or (gasp!) Fortran with either MPI or CUDA for these sorts of tasks if performance was the most critical factor.
I'm excited by the Julia language though!
So document that this is not just a problem with some program you wrote, but multiprocessing itself.
I'm unsure of the performance ramifications vs using concurrent.futures
Parallel Python (PP) seems to have a clunkier API, but also more functionality. I think the biggest advantage is that it can distribute jobs over a cluster instead of just different cores on the same machine. I might look into PP if I need to do things on a cluster, but I think I'll still stick with joblib when I'm on one machine.
That's just my first impression. I'd be interested to read your blog post.
www.celeryproject.org
Also, it has a very nice concept called canvas that allows you to chain/combine the data/results of different tasks together.
It also allows you to switch out different implementations of the communication infrastructure that Celery uses to communicate and dish-out tasks.
Julia addresses nearly all the problems I've found with Python over the years, including poor performance, poor threading support on multicore machines, integration with C libraries, etc. I was a big adherent of Python but as machines got more capable, the ongoing resistence to solving the GIL problem (which IronPython demonstrated can be done with reasonable impact on serial performance) I could not continue using the language except for legacy applications.
Julia looks nice but comes with its own set of problems: no inheritance, 1-based indexing, less libraries, less mature.
you have several choices for C integration in Python. SWIG, which is now generally considered a huge mess, hand-wrapping, which is a tedious pain, and dlopen/dlsym methods that talk to the C api direectly (which requires something like GCCXML to handle type recognition for complicated APIs).
I don't think PyPy's approach to transactional memory is the right direction either.
In short: multithreading on multicore machines is how you write performant software in industry. The hardware is designed for, the compilers are designed for it, and if you don't take advantage of it, you're just wasting machines.
Now people could argue that multiprocessing addresses it, but it's just message passing between different process spaces, which while a wonderful and powerful tool, is ultimately just more cumbersome (hey, I used to write big MPI/OpenMP apps that did both models at the same time).
Anyway, the ultimate existence proof is that IronPython was both faster serially and in parallel, without the GIL, than CPython. So basically we know it's possible. The Python developers have no will, inclination, or ability to make it so,.
As for interfacing with C, like I said, Cython really makes this a lot easier than the approaches you mention. You mention IronPython as not having a GIL, but then IronPython doesn't allow easy interfacing with C code, e.g., it's not compatible with numpy ...
(NB I say this as a big Julia evangelist. it has a lot of potential but is not really there yet on a number of things, this being one of them.)
One nice example using this is a shared memory, parallel sparse matrix multiplication implementation:
If they can't support true multithreading without having to pack messages or use /dev/shm, fuck em.
the right way if you want to use a thread is thread = threading.Thread(target=CALLABLE, args=ARGS)
and not
thread = threading.Thread(target=CALLABLE(ARGS))
This implements the worker pool logic already so we don't have to.
Thanks for mentioning it though.