The API was designed to be a standard that could be used by other libraries. Before if you started with thread and then realised you were GIL-limited then switching from the threading module to the multiprocessing module was a complete change. With concurrent.futures, the only thing that needs change is:
with ThreadPoolExecutor() as executor:
executor.map(...)
to with ProcessPoolExecutor() as executor:
executor.map(...)
The API has been adopted by other third-party modules too, so you can do Dask distributed computing with: with distributed.Client().get_executor() as executor:
executor.map(...)
or MPI with with MPIPoolExecutor() as executor:
executor.map(...)
and nothing else need change.This is why I chose to use it to teach my Parallel Python course (https://milliams.com/courses/parallel_python/).
Is this true?
I've been switching back and forth between multiprocessing.Pool and multiprocessing.dummy.Pool for a very long time. Super easy, barely an inconvenience.
TBH, assuming your stack allows it, gevent is my preferred mechanism for concurrency in Python. Followed by asyncio.
For places where I really need to get my hands dirty I will lean on manually controlling processes/threads.
Too bad it has a somewhat odd name, which doesn't help newbies guess what it really does.
But in almost all cases, it can replace multiprocessing.
i usually test both when i write code nowadays and concurrent.futures is useful in maybe 10% of my cases.