Parallel Python
parallelpython.com
parallelpython.com
- IPython parallel (https://ipyparallel.readthedocs.io/en/latest/ ) is pretty much the same idea but under active development, supports both local processes, small clusters, cloud computng, an HPC environments (SGE style supercomputing schedulers).
- Joblib is a great tool for common embarrassingly parallel problems - run on many cores with a one liner (https://pythonhosted.org/joblib/index.html)
- Dask provides a graph-based approach to lasily building up task graphs and finding an optimal way to run them in parallel (http://dask.pydata.org/)
EDIT: I was wrong, there is some activity in a different branch of the forum [1]. Last release almost two years ago.
[1] http://www.parallelpython.com/component/option,com_smf/Itemi...
When I ported my stuff, I discovered that some of the libraries I used were unmaintained and I replaced them by a newer library.
Q) What Python versions are supported?
A) PP was tested with Python 2.3 - 2.7 on a variety of platforms.
I also liked this book http://www.parallel-algorithms-book.com/ (free draft copy) it identifies which algorithms are best candidates for running in parallel/divide and conquer strategy, like the analysis on MergeSort vs Quicksort.
Not saying it's a bad idea, but it's a totally different approach.
We never have to leave Python 2, if we do not want to.
If you will allow me to paint with broad strokes, as I'm going to lump a lot of things together for the sake of ease of understanding, I am not speaking for every python user(s) out there. There seems to be a huge tug of war between companies/corporate users and the PSF & more independent users.
The smaller (if you will, these are going to be big in the python community) 'shops' and library developers, and users all are or have migrated to Python 3 ages ago. Django now is going to be python 3 only in the foreseeable future (I know they EOL'd their python2 compatible libraries). The Software foundation itself wants it to be essentially toast by 2020. Things like Pyramid, Flask, Bottle, and even numpy & scipy have started moving in a direction of having python3 be their 'first' update (I think in the case of Flask and Bottle, they're end of life support for python2 is happening as well, though they haven't announced anything official, they seem to be adding new features and promising features that will only make sense in python3).
Yet, big 'corporate' users like Google (Look at you, grumpy https://github.com/google/grumpy) and Apple (still shipping with 2.7...and why!) and large universities all seem to be stuck on python 2.
Because of this, this to me really hurts the community and its holding back full development of python3 and python generally going forward. Why are companies holding back the state of a language (and then complaining about it) instead of diverting their ya know, massive resources, to get their own libraries and technologies up to current stack?
Why do so many people want to stay on python2. I'll never understand.
For what its worth, the latest version of python3 (3.6, well, 3.6.1, but the release notes for this is for 3.6) haves a ton of forward thinking built in concurrency improvements and technologies that these companies could easily take advantage of!
https://docs.python.org/3/library/concurrency.html https://docs.python.org/3/library/asyncio.html https://docs.python.org/3/library/multiprocessing.html
Google's choice with Grumpy particularly astounds me. Why 2.7? Its not clear to me and in fact, its really unclear to me, why that was a 'better' choice.
Probably because Google has a lot of legact Python 2.7, and doesn't do new greenfield development in Python 3, preferring other languages (particularly Go), so...
Projects made to scratch the maker's own itch are good in that they get thoroughly dogfooded, but they also can carry a focus that reflects the creator's needs more than anyone else's.
If you look at it from that perspective and know that Google has a lot of existing python 2 code base, it makes a lot of sense.
If they wanted to stick with python, instead of grumpy, they would instead provide a way to write python modules in go.
There is still the multiprocessing module which can spawn processes, so you can effectively run code in parallel. It's pain to manage though.
I think in the end I am grateful for no parallel threading in Python as it forces me to either do things so that they can be run in naive-parallel, or to use things that have the concurrency solved out of Python.
(but it's a little hacky..)
Python's parallelism options are "share everything (but with only one thread computing at a time)", or "share nothing" in a different process on the same machine. If you're parallelizing to get extra crunch power, neither is efficient.
It's better described as having only explicit sharing (as opposed to classic threading models, which have implicit sharing.)