The multiprocessing module in Python
linuxramb.blogspot.com
linuxramb.blogspot.com
> Since a new instance of Python VM is running the code, there is no GIL and you get parallelism running on multiple cores.
Has the potential to be misinterpreted by new Python devs. As I understand it, each process gets its own interpreter with its own GIL. So its not that there is no GIL, just that there are now 3 GILs each managing separate execution spaces.
Also, at a higher level of discussion, this is just a rehash of multi-processing vs. multi-threading so perhaps it would be more illuminating to start from there and move down to the library examples.
The casual reference to multi-threading in the quote from the POSIX fork definition also adds to potential reader confusion.
No you can, just run lots of processes and make lots of GILS. Ends up still using less memory than Java in many cases. Gil will be the least of your worries.
Multi-processing breaks down when you have to communicate with high frequency between various logical runners.
I think there are vastly more problems that can be solved with multi-processing than those that require multi-threading but they do exist.
There is also another layer of simulated parallelism via coroutines...but I was just trying to point out that there was potential for confusion in this piece.
C extensions called from Python are still subject to the GIL of the calling interpreter[1].
In fact, this guaranteed thread safety has led to the creation of many non-thread safe Python C extensions which is now a critical blocker to efforts to removing the GIL[2], although there are also a number of other issues as well.
As noted in the StackOverflow, you can release the GIL but there are some tricks to doing so safely[3].
But threads of a C extensions library are not necessarily pinned to a single core by the GIL like regular Python byte code.
Although many performance oriented libraries like numpy[4] forsake multi-threading by default in many cases and rely on parallelization by the user (programmer)[5].
[1] http://stackoverflow.com/questions/651048/concurrency-are-py...
[2] https://youtu.be/P3AyI_u66Bw
[3] https://docs.python.org/3/c-api/init.html#thread-state-and-t...
[5] http://stackoverflow.com/questions/16617973/why-isnt-numpy-m...
It will run slower in any case, and if I really care about startup speed, I will just use one of the third party JDKs and AOT compile to native code.
I have coded services in Java that would never JIT before a reboot. Python can often beat un-JITed Java.
For an IO bound tasks that is using less than 100% CPU.. again Python could be identical performance or at least very similar performance to Java.. Think of a simple rest service getting 10 hits per second. Maybe the Java service will be a few micros faster, but to a client across the internet you won't care.
Now yes you are right if you are having 900 requests per second hitting your API, the Java API will crush the Python API. Not every app or service hits that though, and I would sure rather code in Python over Java.
https://engineering.instagram.com/dismissing-python-garbage-...
Maybe? You also need to assign about 2x live memory to the heap in order for GC to work properly. If you give it less, you will blow your heap and crash the JVM
> I can also give it all the memory on my server and it will translate that into performance.
No, it won't.
CMS is realistically limited to about 16 GB of ram. Past that and you are going to have routine 2, 3, 4 second pauses for full GCs. G1GC is limited a bit higher, maybe 64 GB.. but again, beyond that and you are screwed. In a high performance app, routinely you limit heap usage to around 2 GB, which means you have about 1 GB of actual memory to use.
> Java's memory usage is extremely within your control so this comparison doesn't make sense.
Honestly, it sounds like you have never had to tune the JVM or run a Java app in production, so I disagree strongly with all you have said.
from multiprocessing.dummy import Pool as ThreadPool
pool = ThreadPool(16)
res = pool.map(one_arg_fun, [arg1, arg2, ...])
A 3 line IO parallelism speedup trick. Used it fetch stuff from multiple servers recently. from threading import Thread
def blocking_function(arg):
....
pool = []
for arg in args[:16]:
f = lambda: blocking_function(arg)
thread = Thread(target=f)
thread.start()
pool.append(thread)
Edit: see below, this doesn't include job queueing but rather, hard limits your input to maximum 16 args. The point is, the code is a nice, easy way to start a number of threads working on a list of arguments.It doesn't have true CPU parallelism like multiple processes would have, but it works for IO code.
It's a little gold nuggest in the multiprocessing lib.
Python3 has asyncio as part of the stdlib but for Python2 there are a variety of coroutine libraries out there (I think the most widely recognized being gevent).
> I believe multiprocessing.dummy is just a wrapper for threading which has context switching overhead [1].
Yap. But I am fetching and waiting for data to come from servers half-way across the country. Not worried about too much thread switching overhead. The point it was still faster than a for loop with a fetch. It was replaced by 3 lines which was 10x faster.
Python releases GIL during IO so that's why thread-based pool works for IO parallelism and there is no need to fork. Was that what you mentioned? Sorry, I don't think I completely followed.
Joblib[0]: 'Embarrassingly parallel for loops'. Basically just write a generator and get multicore processing on it. Pretty straightforward.
Multiprocess[1]: This is a fork of the multiprocessing module. The biggest benefit to me (the one time I used it) is that it's easier to create shared data structures for multiprocessing (which I couldn't do with multiprocessing.Pool). Last week I had to do a really long graph calculation that only needed to return a result for a very, very small fraction of the arguments. Joblib stored all of the null results as 'None', which tore through my RAM after billions of relatively fast calculations before crashing. But using a shared dictionary and multiprocess.Pool and imap_unordered, I was able to use mutlicore processing, adding items to the dict only if the right conditions were met, and discarding the 'None's. RAM use was very minimal.
[0]: https://pythonhosted.org/joblib/index.html [1]: https://pypi.python.org/pypi/multiprocess
If you're looking for some real world examples of using multiprocessing (v3), check here: https://github.com/dpgailey/asteria/blob/master/asteria-v3/c...
Even happens if you open in a new tab.
In fact, I can't find a single browser/OS combo it doesn't happen on.
Don't have my Ubuntu machine to test it there though.
I love Python. But, it's good to remember sometimes that it runs at about 1% efficiency compared to well-optimized C. So, you paid for a whole machine and you are utilizing 0.01%. That's fine for a lot of tasks; especially tasks that are 99% I/O-bound or tasks that are just coordinating big C libraries to do 99% of the work. Just good to remember sometimes...