Solving multi-core Python
lwn.net
lwn.net
And here: https://mail.python.org/pipermail/python-ideas/2015-June/034...
I'm getting excellent scaling and performance across the board for the simple TEFB tests (https://bitbucket.org/tpn/pyparallel/src/8528b11ba51003a9821...), and have implemented something that really shows where PyParallel shines: an instantaneous wiki search REST API: https://bitbucket.org/tpn/pyparallel/src/8528b11ba51003a9821....
Quoting:
I particularly like the wiki example as it leverages a lot of benefits afforded by PyParallel's approach to parallelism, concurrency and asynchronous I/O:
- Load a digital search trie (datrie.Trie) that contains every
Wikipedia title and the byte-offset within the wiki.xml where
the title was found. (Once loaded the RSS of python.exe is about
11GB; the trie itself has about 16 million items in it.)
- Load a numpy array of sorted 64-bit integer offsets. This allows
us to do a searchsorted() (binary search) against a given offset
in order to derive the next offset.
- Once we have a way of getting two byte offsets, we can use ranged
HTTP requests (and TransmitFile behind the scenes) to efficiently
read random chunks of the file asynchronously. (Windows has a
huge advantage here -- there's simply no way to achieve similar
functionality on POSIX in a non-blocking fashion (sendfile can
block, a disk read() can block, a memory reference into a mmap'd
file that isn't in memory will page fault, which will block).)Antoine Pitrou said that Snow's proposal was similar to your work. What do you see as the similarities?
I don't agree with the subinterpreter approach. In fact, my opinion is that the best way to solve the problem is to use the approach taken by PyParallel.
Granted, I would think that wouldn't I, being the author of PyParallel and all ;-)
http://python-notes.curiousefficiency.org/en/latest/python3/...
IMHO this is the most sensible and logical approach and doesn't break any existing code. This also seems like what will likely happen anyways. Out of all the options PyPy-STM seems like the best approach because:
"pypy-stm is fully compatible with a GIL-based PyPy; you can use it as a drop-in replacement and multithreaded programs will run on multiple cores."
https://mail.python.org/pipermail/python-ideas/2015-June/thr...
However, built on top of CPython's C API are a huge pile of libraries that assume they can manipulate global or per-data-structure state without locking.
Also how is this different then problems faced when using multiprocessing style parallelism with the same library?
The readline module uses libreadline, which maintains internal global state on the C side.
The standard library has lists and dictionaries, which are not thread-safe to access from two threads simultaneously, and are not in any way locked.
Locking lists and dicts, etc would not be a reasonable thing to do. There must be some granularity in synchronization, it's clearly not a good idea to attempt to allow mutating any object from multiple threads without explicit locking.
Things like readline are issues that programmers have to deal with when writing multi threaded code in any language. There are libs which are not thread safe or which have global state or need special handling with threads (e.g. OpenGL). That's no reason to disallow threading completely.
However, this might be a big culture shock issue if Python were to suddenly start supporting proper concurrency without GIL. None of the libraries document whether they are thread safe or not, so there would be surprises ahead.
But that is not what threading means. If you want "completely independent interpreters" just fork. Threading means you share everything.
> Then the GIL can be discarded altogether
And then you start sharing variables between interpreters, and you gets lots of little baby GILs to contend with! How fun will that be? Not very.
I think the only way to do that would be to have a reentrant lock for each object individually. Isn't that potentially a big performance issue?
I think that is the main reason behind this quote from the article:
> Several attempts have been made over the years and failed to do it without sacrificing single-threaded performance.
Although PHP7 does have some workarounds to improve the performance of this mode, involving implementing thread-local storage on certain operating systems..
E.g. https://bitbucket.org/tpn/pyparallel/src/3be2954508f9938b85a...
It's slower than accessing a struct, sure:
5982: if (ctx) {
000000001E1A808B 8B 0D FF 7C 28 00 mov ecx,dword ptr [_tls_index (1E42FD90h)]
000000001E1A8091 BA 80 29 00 00 mov edx,2980h
000000001E1A8096 48 89 83 90 02 00 00 mov qword ptr [rbx+290h],rax
000000001E1A809D 65 48 8B 04 25 58 00 00 00 mov rax,qword ptr gs:[58h]
000000001E1A80A6 48 8B 04 C8 mov rax,qword ptr [rax+rcx*8]
000000001E1A80AA 48 8B 0C 02 mov rcx,qword ptr [rdx+rax]
000000001E1A80AE 48 85 C9 test rcx,rcx
000000001E1A80B1 74 70 je new_context+253h (1E1A8123h)
But think of the big picture: I use TLS everywhere for PyParallel, and PyParallel has awesome performance, so eh, net win.> One should be able to support many low-overhead independent interpreter instances on a single thread
Why? What problem does that solve?
The use of the globals in the CPython interpreter are a fairly large design mistake that has prevented Python from having as much reach as other systems. The whole embedding vs extending debate is because of those globals and that CPython has been historically difficult to embed properly. Lua on the other hand, is easy to both embed and extend.
I'd love to use `cffi` to embed Python within itself, I cannot do that.
As to supporting multiple truly independent Python interpreters in the same process, I'm not clear on what the buys you over subinterpreters. In fact, with subinterpreters you get the benefit of the main interpreter handling the C runtime global state (env vars, command line args, the standard streams, etc.). With truly independent interpreters that's more complicated.
While it's not quite what you described, if you were to suggest that subinterpreters be made even more isolated from one another than they already are then I'd agree. :) It sounds like that's really what you're after: better isolation between interpreters in a single process.
I can't speak for anybody else, but I personally haven't felt that inconvenienced by these trade offs at all.
The major one is needing to pickle objects before sending them back and forth between processes. Since I prefer to keep the thread interfaces as tight as possible anyway, this only ended up being a major problem once - when I was trying to pickle a stacktrace before sending it to the parent process.
I've had this signal handling problem in single process/single threaded code too, and in other languages.
True, it would be nice if python layered on a nicer set of APIs to deal with the nastiness.
Datapoint: I encountered the same thing on Windows, having to use the Powershell equivalent of `killall python` to deal with it.
So you have to be careful when dealing with forking and databases, which has lead us to subclass multiprocessing.Process specifically for our (common) case of wanting to continue to use the session object in the child process without having to think about recycling the Engine object. It was that code we have been trying to debug today, because in some specific cases we still run into issues (yes, we read the docs). Also, when people (usually new engineers) don't know about this they usually blindly use the multiprocessing module directly (and who can blame them) and end up spend some time debugging intermittent connectivity issues until someone says "oh, you should use the xyz module, it handles the DB stuff for you".
[1] http://docs.sqlalchemy.org/en/rel_0_9/core/connections.html#...
Yeah, provided you do this there should be no headache.
>So you have to be careful when dealing with forking and databases, which has lead us to subclass multiprocessing.Process specifically for our (common) case of wanting to continue to use the session object in the child process without having to think about recycling the Engine object.
What use case led to this being a non-negotiable requirement?
I've run into the exact issue you've described; the --lazy-apps solution may be slightly more resource intensive (since you maintain N copies of the application in memory), but ends up being much nicer to reason about.
1. http://uwsgi-docs.readthedocs.org/en/latest/articles/TheArtO...
Collecting metrics require using an external collector process (statsd) rather than just an in-proc lib (e.g., codahale metrics on the JVM).
Issues generating random numbers - I ran 8M simulations on 8 cores, but because the random seed was forked, I actually only ran 1M simulations 8 times. (Easy enough to fix, but annoying.)
Deduplication is really handy in reducing memory usage - why would you want the same (immutable) object duplicated N times?
And of course various issues with resources - you need to be careful only to connect to the DB/etc after forking otherwise stuff gets weird. And of course, this means that instead of sharing a single connection pool, you've now got one DB connection per proc, which may or may not be a big waste of resources.
IIRC, even if you do your easy fix of giving the 8 threads different seeds, your results are not trustworthy since the parallell random number streams are likely to exhibit significant correlations. There's quite a bit of algorithm research behind quality PRNGs.
This doesn't really add much that wasn't possible before, and doesn't really solve any technical or PR issue. The PR issue will never be resolved with CPython, those people who don't understand it are free to write multithreaded Java apps. But I think explicitly spinning up pthreads should be reserved for writing systems software. I'm assuming subinterpreters means they'll be allocated as pthreads because this is meant to "use all your cores". Looking forward, this proposal sounds great as long as we're on 4 or 8 cores max, at a certain point this solution starts to look like something like a gimmicky ideal created to fit the technology way back in 2015. The ultimate multicore and multinode solution at a non-systems level, if we really need every single language to solve that specific problem- would be Erlang's approach.
It's yet more technical churn rather than innovation in a language (Python3), that was born of technical churn.
To offer something constructive as well, I think an ideal solution would be to find a more implicit approach. Think gevent for pthreads or subinterpreters. That would be a lot more work to figure out, this proposal looks more like a hack. I don't think this type of "improvement" will draw people to Python3 either, I wish they'd stop throwing crap at the wall seeing what will stick. Constantly expanding the language, which is bad. The core dev team doesn't know what to do, but usually in that case it's better to do nothing. Thus Python3 is looking more and more like a playground or experimental branch as time goes on. I'm increasingly thinking my "Python3 migration" will be to Swift or Go.
The goal of the project is to achieve good multi-core support without the serialization overhead and to make that support both obvious and undeniable. While very few Python programs actually benefit from true parallelism, it's a glaring gap that I'd like to see filled.
My proposal is a means to an end. I'd be just as happy if the situation were resolved in some other way and without my involvement. However, I've found that in open source, waiting for someone else to do what you want done is a losing proposition. So I'm not going to hold my breath. The only project of which I'm aware that could make a difference is Trent Nelson's pyparallel. I hope to collaborate on that, but I'll likely continue pursuing alternatives at the same time for now. I'm certainly open to any serious recommendations on how to achieve the goal, if you're sincerely interested in making a difference. (I appreciate the the mention of gevent and Erlang, which are things I've already taken into consideration.)
As to the details of my proposal, it's very early in the project and the python-ideas post is simply a high-level exploratory discussion of the problem along with a lot of unsettled details about how I think it might be solved in the Python 3.6 timeframe. A more serious proposal would be in the form of a PEP.
Regarding your feedback, the post implies that you either misunderstood what I said or you don't understand the underlying technologies. To clarify:
* the proposal is changing/adding relatively little, instead focusing on leverage as much existing features as possible * Python's existing threading support would be leveraged * subinterpreters, which already exist, would be exposed in Python through a new module in the stdlib * subinterpreters are already highly independent and share very little global state * the key change is enabling subinterpreters to run more or less without the GIL (leaving that to the main interpreter) * the key addition is a mechanism to efficiently and safely share objects between subinterpreters * the approach is drawing inspiration in part from CSP (Hoare's Communicating Sequential Processes) * think of it as shared-nothing threads with message passing
It will most certainly improve multi-core support. It shares more in common with Erlang's approach than you think. It is neither a hack nor crap on the wall, as you put it.
Keep in mind that like just about everyone working on Python (and most open source), I'm doing this work in my spare time. I take my role very seriously and work on what I believe will benefit the community the most (relative to what I can accomplish). This is also the case for most of the other committers. Furthermore, since we all do this in our free time, the pace of Python development is super slow when compared with subsidized projects (or languages). Wanting to make a difference in a reasonable timeframe, core developers take that into consideration when deciding about working on a large Python feature. Consider both these points before rushing to judgement about what core developers are working on.
The goal fine grained parallelism python (fgpp) must be to improve the speed of code for which fat parallelism (user conconcted error prone 'multi-threading' or message passing) doesn't gain much, and for which calling out to an impl in another language or specialised library has overhead or doesn't make sense.
IMHO, the most important part of fgpp must be that it is easily reasoned about, both by compiler/runtime and programmers, so as to maximise scope for transformations and user-driven design decisions (as opposed to VOODOO). It would ideally avoid the costs of premature optimisation, or the often inefficient flattening required for vectorisation.
So I think any propsal should be addressing these questions, with appropriate benchmarks too, before it can be considered seriously.
Elixir/Erlang will probably never be fast enough, Go didn't commit to immutability and its functional programming support is minimal, Rust could have all that but it's too low-level...
I guess OCaml is as close as it comes for the moment -- it's got everything I listed except a reasonable parallel execution story, but it looks like that's being worked on at the moment.
Not very.
Not much.