Python’s Multiprocessing Performance Problem
pythonspeed.com
pythonspeed.com
My default philosophy is to use python _until_ you find something that is performance sensitive, and then make a C/C++ extension for the slow bits. Pybind works great for a hybrid Python/C++ codebase (https://pybind11.readthedocs.io/en/stable/).
Then you can develop and prototype much quicker w/ Python but re-write the slow parts in C++.
Definitely more of a judgement call when threading and function call overhead enter the equation, but I've found this hybrid "99% of the time Python, 1% C++ when needed" setup works great. And it's typically easier for me to eventually port mature code to C++/Go/etc once I've fleshed it out in Python and hit all the design snags.
Pybind offers a lot of functionality, but core "good parts" I've found useful are (a) use a numpy array in Python and pass it to a C++ method to work on, (b) pass your python data structure to pybind and then do work on it in C++ (some copy overhead), and (c) Make a class/struct in C++ and expose it to Python (so no copying overhead and you can create nice cache-aware structs, etc.).
[1] https://github.com/pybind/pybind11/blob/master/tests/test_py...
[2] https://gitlab.unistra.fr/benzerara/pybind11/-/tree/master/e...
I've found cross language bindings to be pretty easy with Python. If you want something with memory management, you can use cgo
Computers are cheap, and people are expensive.
If your data is fragmented across a bunch of small containers/classes, passing it around will be expensive whichever the method (either passing to C++, or just in terms of cache efficiency).
If you just pass an array of data back and forth it's cheap.
Yes, and numpy is great, and all. Python works great as glue to marshal things to and from native code and do inexpensive (but possibly complicated) bits of control logic.
But if I'm trying to deal with large numbers of client requests, say... the lack of concurrency in python itself really hurts. Sure, I can punt almost everything to native code, but what's the point in having Python at all, then?
Not all problems have state that can be shared well across Multiprocessing or completely externalized to large lumps that travel to native code in a few calls-- I'd actually say these are special case exceptions than the rule.
At that point you'd look to Go or another language, and also carefully choose the REST API framework you're setting up.
Note that having a solution setup where the end result is "a ton of small, individual API calls" could possibly indicate a bad system architecture.
Or just a lot of clients with a fair bit of shared state which is best kept resident, which is a pretty common use case.
It's a bummer to write python code that works well, and then maxes out at 130% CPU load when you grow your usage... and not have any obvious path to scale upwards despite you having 32 threads of execution around. Then, you can rewrite some of the more expensive things in native code to squeeze a little more performance, or add indirection to store the data somewhere else so multiprocessing works.
Other languages that have more finely grained locks scale 3-4x higher with minimal thought, and much, much higher with a bit of thought about how to handle locking and data model.
> At that point you'd look to Go or another language
Well, yah... this is us complaining about Python's concurrency problems.
I don't know for certain this is the case, but I'd like to see some benchmarks about it.
I do kind of wish numpy had a stable C API that didn't require a Python interpreter though.
I hit the GIL, switched to multiprocessing, which helped. Still was about half as fast as I expected. Switched to go and using channels and got the performance I expected. Was still debating, till I got deeper into Python's crypto package. I ended up really happy with go.
Yes. I've used multiple threads in Python. It doesn't work well. Some packages, including cpickle, don't work right with multiple threads because they have static variables internally. It can work; I've had multi-threaded Python code running for years. But it's not a good approach for new work.
Python does things the CPython naive interpreter can do easily, such as letting anything modify anything else. Any code anywhere can go find something far away in another thread in another module and mess with it. Everything is a dict, so that works. This makes other things hard. Pre-compilation is hard. Optimizing is hard. JIT is hard. Threading is hard. You can't nail down stuff that probably won't change, but might.
The go was similarly very straight forward, walktree -> CheckIfNewOrchanged -> channel -> sha256/encrypt/sha256. Channels made it really easy/clear and performed quite well. I was getting near linear scaling, CPU time consumed was 8x the wall clock, and speed increased 7.9x or so. With python I was getting significantly less performance per core and worse scaling.
With 8 cores 10Gbit is 1.25GByte/8 = 160MB/sec which is ok, not great, depends on how much computation you are doing. My goal is keeping 100gbit saturated, but I am adding cores as well. I do hope to compare to Go vs Rust as well.
Don't get me wrong; I agree it's easier to build performant applications in go, and to get the performance I want, I have to set my AWS boto3 S3 settings to have massive queues.
I think Pipe isn't necessarily a drop in replacement depending on the complexity of object you want to share but I have found it significantly faster for simple things.
I'm guessing that's part of the reason the article didn't mention it (it looks like they're talking about a Pandas DataFrame which I would say is non-trivial--compared to a primitive type)
I'd think Pipe + Parquet should beat filesystem though. That really depends on storage I guess
Iirc I played a bit with msgpack and orjson to see if there was anything to gain over Pickle but I don't think it made much difference. You'd probably need to deal with structs
Looking at CPython source (3.10), on Windows, you always get a NamedPipe. On other platforms, you get a OS pipe when duplex=False otherwise a socket (socket.socketpair)
https://joblib.readthedocs.io/en/latest/
It doesn't manage to escape the python multiprocessing issue everywhere but it often does
This has the advantage of allowing work done inside a nested function, allowing large initial datasets to be shared and not have to be passed over pickle.
Threads aren't bad: threads that cause resource contention with the GIL is (possibly) bad. That is almost always done explicitly by the developer and almost never done without your knowledge.
the program still deadlocks on python3 but it works perfectly on python2, anyone know what changed between implementations that could be triggering this issue?
For some reason on OP's computer that probably causes an appearance of the program hanging, but really, it will just crash after a while exhausting some resource.
Here's the gdb stack trace:
#0 __futex_abstimed_wait_common64 (private=<optimized out>, cancel=true, abstime=0x0, op=393, expected=0, futex_word=0x17bc7c0) at ./nptl/futex-internal.c:57
#1 __futex_abstimed_wait_common (cancel=true, private=<optimized out>, abstime=0x0, clockid=0, expected=0, futex_word=0x17bc7c0) at ./nptl/futex-internal.c:87
#2 __GI___futex_abstimed_wait_cancelable64 (futex_word=futex_word@entry=0x17bc7c0, expected=expected@entry=0, clockid=clockid@entry=0, abstime=abstime@entry=0x0,
private=<optimized out>) at ./nptl/futex-internal.c:139
#3 0x00007ff326c9cc5f in do_futex_wait (sem=sem@entry=0x17bc7c0, abstime=0x0, clockid=0) at ./nptl/sem_waitcommon.c:111
#4 0x00007ff326c9ccf8 in __new_sem_wait_slow64 (sem=0x17bc7c0, abstime=0x0, clockid=0) at ./nptl/sem_waitcommon.c:183
#5 0x00007ff326c9cd71 in __new_sem_wait (sem=<optimized out>) at ./nptl/sem_wait.c:42
#6 0x000000000042765b in PyThread_acquire_lock_timed (lock=0x17bc7c0, microseconds=-1, intr_flag=0) at ../Python/thread_pthread.h:483
#7 0x0000000000625813 in _enter_buffered_busy (self=0x7ff326f330f0) at ../Modules/_io/bufferedio.c:281
#8 0x000000000045f9b3 in buffered_flush (self=0x7ff326f330f0, args=<optimized out>) at ../Modules/_io/bufferedio.c:825
#9 0x0000000000524185 in method_vectorcall_NOARGS (func=func@entry=<method_descriptor at remote 0x7ff326f611d0>, args=args@entry=0x7fffa8bc7a38, nargsf=<optimized out>,
kwnames=kwnames@entry=0x0) at ../Objects/descrobject.c:436
#10 0x000000000053bd5a in _PyObject_VectorcallTstate (kwnames=0x0, nargsf=<optimized out>, args=0x7fffa8bc7a38, callable=<method_descriptor at remote 0x7ff326f611d0>,
tstate=0x17bf410) at ../Include/cpython/abstract.h:118
#11 PyObject_VectorcallMethod (name=<optimized out>, args=0x7fffa8bc7a38, nargsf=<optimized out>, kwnames=0x0) at ../Objects/call.c:828
#12 0x000000000060283a in _PyObject_CallMethodIdNoArgs (name=0x8f6120 <PyId_flush.lto_priv.2>, self=<optimized out>) at ../Include/cpython/abstract.h:243
#13 _io_TextIOWrapper_flush_impl (self=0x7ff326f40040) at ../Modules/_io/textio.c:3038
#14 _io_TextIOWrapper_flush (self=0x7ff326f40040, _unused_ignored=<optimized out>) at ../Modules/_io/clinic/textio.c.h:685Anyways, another common pitfall in multiprocessing is attempting to serialize multithreading / multiprocessing primitives s.a. locks, variables or mutexes. My memory may fail me, but, I think, it may result in deadlock too. I think, multiprocessing code tries to guard against it, but there are some weird rules for when it's OK for serialized objects to have those primitives (I think, initialization in __init__ is fine, but not so much otherwise or something like that), but the check isn't very good / just a heuristic... But, really, I don't remember this part well.
1. Writing to stderr grabs a lock.
2. Part of the multiprocessing code (perhaps not present in Python 2) also grabs this lock.
3. If you fork at the right moment (which is quite likely with the loop) the lock is held by a thread that is now dead, and so now you're waiting for a lock to release that will never be released.
Python will be around for many years; however, now we have hit the end of Moore's law, the easiest way to speed up code is to multi-thread.
I code exclusively in Python, so I can't judge if languages like Julia are serious contenders. I also agree that building up the same ecosystem as in Python will take some time, and therefore I do not predict a rapid decline. However, many of the things I use Python for can be easily done in another scripting language.
* Django
* FastAPI/Flask
* Numpy/Pandas
* PySpark
* Jupyter Notebooks (gross I know but this is what "data analysts" and "ML/data engineers" use at many places)
Then Python will stick around for forever.
I would love if it the Go community would quit the "you don't need a framework, DUH it's GO" attitude. Make a Django/Rails for Go and there would be 10x the Go jobs.
I agree that Numpy/Pandas will introduce more migration friction, and that is why I mentioned a slow death, similar to the trajectory Java is currently on. Java is still in the top three. However, its popularity has dropped in recent years. It is worth noting that both Cobol and Fortran are still in the top 30 programming languages.
Also R is what's most often taught in school in my experience and boy does Python feel like a breath of fresh air when you've been trained in R. When you're coming out of college trained in data analysis but not software engineering per se, you've got no idea about the larger world of what other languages could offer.
By all means, if someone wants to write these frameworks, people will use them.
With more time, yeah I might choose Go or Rust and setup a couple different nodes - an API gateway, a user auth node like KeyCloak, Ory, or SuperTokens, then a Go backend.
But it would be so nice to have that all ready in one.
I do enjoy building things but I learn more and more how important it is not to focus on things that are already "solved" if they aren't part of your core business.
Distractions can scale exponentially.
There's absolutely nothing wrong with notebooks, but they are a serious PITA to take from being handed basically the musings of a data intern into a a production pipeline.
I used scare quotes because unfortunately in many places there are essentially zero qualifications to start running around doing those jobs.
There's nothing about Python's quality that's worth keeping. The reason for Python's popularity is its popularity. The reason why it won't die is inertia.
But, say, someone creates a "killer app", that runs on hardware different enough from anything we have today, and that someone hates Python (smartphones are the most recent example of such a change), then there'd be a chance to dislodge it. But I struggle to see how Python would "organically" die.
First of all, there's no such thing as a "development speed of a language", just like there isn't a development speed of ice-cream. It's just kind of a Jabberwocky: it feels like it's in English, but it doesn't really mean anything.
Development speed differs by the kind of project (eg. Web site vs filesystem), quality requirements, size of the team working on the project, expertise of the team working on the project... Needless to say that there are areas where Python is entirely not applicable, so, the development speed wouldn't even be a factor. But, even in areas where it's commonly used there are often languages that will compete for this metric, and there are certainly teams using other languages that will beat teams using Python.
But, overall, Python projects tend to be easy to start and hard to develop further and to refine. Python projects tend to fare worse in large teams. Python also doesn't attract high-caliber programmers, while also is often the first language a programmer learns, so it tends to be populated by mediocre-bad programmers (similar problem used to exist in Java before it was replaced by Python in intro to CS courses).
Finally, a huge portion of development speed rests on company's infrastructure: how quickly and reliably can developers test their code plays a tremendous role in productivity. Ironically, Python tooling is so bad that sometimes it's faster to compile a C++ program of equivalent size than to install a bunch of Python packages implementing the same thing (god forbid you are using Anaconda, because that can take hours and days in the worst case to install a handful of packages).