Let's Remove the Global Interpreter Lock
morepypy.blogspot.com
morepypy.blogspot.com
Yes, shared memory is available in multi-processing, but it doesn't necessarily interact well with existing codes.
I've been working on adding Python support to Legion [1], a task-based runtime system for HPC. Legion wants to manage shared memory so that multiple cores don't necessarily need multiple copies of the data, when the executing tasks don't conflict (all are read-only, or access disjoint data). Legion is C++, so this mostly "just works". Some additional work is required to support GPUs, but it's still not so difficult. But with Python, if we go with multiprocessing, we have to switch to a different mechanism. Worse, Python is an optional dependency for Legion, so we can't depend on Python's multiprocessing support either.
If you have a large existing project, and a use case that can take advantage of shared memory, being forced into Python's multiprocessing scheme for parallelism is a pain.
We've been investigating using a dlmopen approach as well, based on this proof of concept [2]. Turns out that dlmopen in every available version of libc has a critical bug that prevents it from being practically useful, if you have any desire to make use of native modules. You can build a custom libc with this patch [3] but rolling a custom libc is also a massive pain.
In all likelihood we'll end up rolling our own multiprocessing to make this work. If the GIL were truly gone though, we could potentially avoid many of these issues.
[1]: http://legion.stanford.edu/
Then you can namespace your queues (and workers), and have separate queues for results handling to push data to the next stage of the pipeline, etc... With stacks of workers, configured as needed. It's all pretty high level from there. GIL has no effect here, and as a side-effect, now you can utilize a massive number of parallel processes for heavy lifting and crunching, even on different machines over the network, where-as that wouldn't be possible with a traditional threaded architecture.
Not saying this necessarily covers your use-case, but it seems strange to use dataframes as a sort of in-memory database, vs using dataframes as the framing to do the munging and heavy lifting. What are you wanting to put multiple cursors on it or something? You could do this with greenlets, for what it's worth... But as someone who has gone down that route (multiple greenlets working over shared stack) I promise doing it with multiple processes and a queue is better, and ultimately way more flexible. Especially if you use something like msgpack or protocol buffers... Then you can have any workers from multiple programming languages and development paradigms doing different work at different stages, all orchestrated and working together via Redis.
Save the dataframe in a folder that can be accessed by the gunicorn worker:
import joblib
joblib.dump(df, '/folder/shared_data.pkl')
Then in the code run by the flask / gunicorn workers themselves: import joblib
shared_df = joblib.load('/folder/shared_data.pkl', mmap_mode='r')
# use the shared_df as usual (inplace modifications are not
# authorized)
Some pandas function can have issues with read-only buffer though: https://github.com/pandas-dev/pandas/issues/17192 (caused by a currently unsolved bug / limitation of Cython) but it can work for your use case.May I ask what you consider large memory - MByte, GByte, TByte? The simplest solution is to store it as a blob on a SSD, and read it via simple file IO or put it into a DB. But I assume this was too slow, so it would be interesting to go into more details.
In the end you can do shared memory with multiprocessing in Python, which - I have to admit - requires some setup and bookkeeping work.
You often end up having to implement your own C extensions or use Numba for the core of your processing. Even with BLAS enabled, NumPy has almost zero intrinsic parallelism, np.dot() being the notable exception which releases the GIL and uses multicore by itself.
Is there any sort of list (comprehensive or otherwise) that denotes which NumPy functions are parallelism-friendly? I mean this whether it's in terms of releasing the GIL, in terms of SIMD support, or in terms of being multi-core.
np.dot() is multicore. np.load () (and family) releases the GIL. SIMD mostly depends on the build system, so if you want it you might need to build NumPy from source.
https://stackoverflow.com/questions/24022723/where-can-i-fin...
I really wish numpy/pandas/scipy wouldn't do this kind of uncontrollable parallelization.
Too bad the current project is on hold.
The closest work has been done on PyParallel https://news.ycombinator.com/item?id=7861942 but afaik it is only for windows.
Did you consider just mounting a ramdisk and storing data as files? At first glance it seems like a decent fit for sharing read-only data in memory.
Ideally, it would be a single model in memory with access from multiple threads. But that won't work right now cause GIL.
With 80% of the developers out there, they are basically assured of producing better, more stable code this way.
"Multiprocessing" is useful when you have a lot of work to do concurrently and not too much data to pass between processes. I've used Python subprocesses that way. Parallelizing your number crunching is probably not going to work very well.
> Parallelizing your number crunching is probably not going to work very well. [...]
The question is, what exactly does "number crunching" mean? We do aerial imagery analysis, so image processing in essence, which I would classify as a "number crunching" problem. A common thing e.g. is to do a time-series analysis and you can simply start multiple (2, 4, ..., N with clusters, etc.) processes for each problem. Obviously this works because most methods are computation and/or memory heavy - the additional memory requirements and "overhead" of Python itself (IMHO people overestimate the weight of starting new processes instead of threads) is completely dwarfed by the requirements (memory and CPU) of the method itself.
Interpreter state is among the most frequently accessed memory in many applications, meaning it's ideal to have it in cache. The difference between two interpreter states and one might not be big compared to the data being processed, but it's big enough to bump a lot of interpreter state out of cache, which for many programs can have drastic performance implications.
If you don't think cache locality is important, look at radix sort versus quicksort. Radix sort has a much lower O, but performs worse in most cases because of its poor cache locality.
Look, I get that there are fairly easy ways to work around these problems, but let's not just blithely pretend they aren't problems.
Sure, it's a problem for specific workloads, and Python will get there eventually - I just don't think it is a deal breaker.
https://docs.python.org/3/library/multiprocessing.html#conte...
The ``forkserver`` method eliminates most of the problems you mention: child processes are only started once, and they fork() from a totally separate process so they don't inherit all of the resources of the main process (in particular, they don't copy the whole heap). I've found this eliminates 90% of the performance-related issues I used to experience with multiprocessing.
You don't need to do setup/destroy more then once.
You're basically asking why NumPy, SciPy, Numba, etc. even exist.
They exist because Python is ridiculously fast to develop in compared to, say, C++.
To combine both your points, the best approach (if you like python) is to stick with python due to is ease of development and use libraries such as numpy as far as possible. However, if your use case is CPU bound but not served by those libraries, then you'll either need to develop your own extensions or throw away the interpreter altogether (and go with a different language).
I'm not going to run your code on my server if your code uses resources so poorly that I can't run other things I want to run on my server.
Perhaps you meant to reply to GP?
Statically and dynamically loaded binaries are resident in the kernel's page cache. Which while each process will have different locations within its process address space for each process (b/c ALSR), they _should_ be de-duplicated in RAM, ultimately all these seperate in process images will be pointing at the same physical RAM page(s).
So from a hardware cache standpoint you're mostly okay.
If you need to close that gap between the performance of multiprocessing, and multithreading, then you probably shouldn't be using Python, or any language of the same shape, in the first place.
There is one other option I'd like to see: multiprocessing style, but with multiple Python interpreter instances in the same process — one per thread. There would still be the hard delineation of data boundaries between instances, but less overhead for pushing data between them.
Unfortunately, these performance concerns often manifest well after the "rewrite it in a different language" date has expired. There are a lot of people in that boat, and they need better options.
> There is one other option I'd like to see: multiprocessing style, but with multiple Python interpreter instances in the same process — one per thread. There would still be the hard delineation of data boundaries between instances, but less overhead for pushing data between them.
If I understand correctly, the article discusses this ("subinterpreters"), but claims that there is no advantage to this approach vs multiprocessing. Presumably any overhead savings are eaten by GIL contention or some such?
Land isn't coming to you, folks, you must start rowing if you want to get there.
Rewrite bit for bit. Module for module. Package for package.
The sane way of forking a multithreaded process is to exec immediately after.
Also, you're describing multiple python processes + an extra server (redis) process - as a "simpler" solution for the limitation that Python doesn't do multi-threads well.
Of course there are a ton of use cases out there where you can scale in other ways, but threads and shared memory exist for a reason - there's no reason not to call a spade a spade and say the GIL is still a limitation.
Cache evictions can be handled by Redis natively with TTL.
For retries and failure mitigation, you can still lean on Redis via BRPOPLPUSH/RPOPLPUSH.
If you want to scale beyond one machine, you can't rely on threading to help you. So why not just do it right to begin with, and use a parallel worker queue?
It's not a matter of the GIL being a limitation, a single machine is a limitation too. Don't blame your tools because you're misusing them.
As for threading in Python... on a single machine, for one reason or another... I would still rather use multiple processes, or at the very least, would just simply use eventlet and greenthreads.
Not saying it covers all use cases, it's not a silver bullet, and it doesn't replace threading natively, but damnit, it scales better, and it's the right way to do the task at hand.
You can delete the last 6 words. Anything where multiple processes would have to read in/acquire a massive dataset to do some independent work qualifies. For instance, running some number (e.g., hundreds to hundreds of thousands) of analytical or statistical tests over a set to pick parameters, etc.
In languages which don't have a GIL, threads are almost as capable as processes, but lighter weight. Threads are almost always preferable to processes in most languages.
I understand why the GIL is still around, and don't necessarily support removing it, but it's definitely not there because it produces "better, more stable [Python] code".
But also plagued with shared state concurrency bugs, something multi-processing completely avoids so...
> Threads are almost always preferable to processes in most languages.
No, they aren't. It's too easy to write buggy code with threads, it's a flawed model. Now it's certainly true that more people choose threads than processes but that's because they vastly overestimate their ability to write bug free lock based code. Processes are better.
2. If we're taking about all languages, I'm really just not sure why you would assume threads imply locks. There are a ton of threading models out there which don't rely on explicit locking, and there are even some that don't use locking, period.
Because they remove the unsafe way of sharing state from the programmer. The issue isn't that state can be shared correct in threads, it's that it doesn't have to be done correctly and programmers are simply terrible at doing it right.
> There are a ton of threading models out there which don't rely on explicit locking, and there are even some that don't use locking, period.
It's not about locks, it's about shared mutable state. Programmers are bad at dealing with shared mutable state, regardless of how access is synchronized.
When you're jumping between C/C++ code and Python code you don't care much about the GIL... until you have a GUI which needs to be kept responsive and needs the GIL to do so.
It's a factor, sure. But, one you should weigh with other factors to determine what is best.
But sometimes you are writing a GUI app, or some "real time" code [1]. You put blocking calls onto a different thread to keep the UI responsive. But then you find that still the blocking calls freeze the UI across threads due to the GIL.
Pure Python code is not the problem in this case - the GIL gets released between statements often enough. It is long running C code. You could release the GIL manually in there, but it is not done everywhere. Also, there are often calls that are supposed to be instant (like opening a file, or starting an async operation), that take seconds under bad conditions (when the network is down).
----
[1] well, with Python probably not in the strict definition of real time, but say you are controlling some external device
There's a reason threads exist.
E.g. in the remote sensing and earth observation domain you can simply divide your problem (e.g. semantic segmentation) into (maybe over-lapping) subproblems (via e.g. tiling) and start separate processes for each image processing tool-chain.
Granted you may not utilize your resources to the full extent by only applying multiprocessing (and ignoring threading), but in my experience you can solve a lot of problems by simply applying map-reduce-like programs and optimizing for throughput.
Multi-process is just one form of concurrency, and it's not always the best one.
Relevant: "Python is Only Slow If You Use it Wrong" http://apenwarr.ca/diary/2011-10-pycodeconf-apenwarr.pdf
I think the first thing to realize is that single-threaded performance is often significantly better with the GIL than without it. I think Larry Hasting's first Gilectomy talk was extremely insightful (about the GIL in general and about performance when removing the GIL):
https://youtu.be/P3AyI_u66Bw?t=23m52s
I am not sure I would, personally, trade single-threaded performance for enabling multi-threaded applications. I view Python as a high-level rapid prototyping language that is well suited for business logic and glue code. And for that type of workload I would value single-threaded performance over support for multi-threading.
Even now, a year later, the Gilectomy project is still slightly off performance-wise (although it looks really really close :) ):
https://youtu.be/pLqv11ScGsQ?t=27m32s
As noted elsewhere, multi-processing offers adequate parallelization for this type of logic. Also, coroutines and async libraries such as gevent and asyncio offer easily approachable event loops for maximizing single-threaded resource utilization.
It's true that multi-processing is not a replacement for multi-threading. There definitely are tasks and workloads where multi-processing and its inherent overhead make it unsuitable as a solution. But for those tasks, I question whether or not Python itself (as an interpreted, dynamically typed language) is suitable.
But that's just my $0.02. If there is a way to remove the GIL without negatively impacting single-threaded performance or sacrificing reference counting for a more robust (and heavy) GC, then I am all for it. But if there is not...I would just as soon keep the GIL.
If Python would simply suck it up and eat the 20% performance hit, we could stop talking about the GIL and start optimizing code to get the 20% back.
Eliminating the GIL doesn't have to mean actually eliminating it. You could certainly have #defines and/or alternate implementations that make the fine-grained locks no-ops when compiling in GIL mode. Conversely make the GIL a no-op in multithreaded mode.
1. Big enough to need concurrency
2. Not big enough to require multiple boxes.
3. Running in a situation that can not spare the resources for multiprocessing.
4. You want to share memory instead of designing your workflow to handle messages or working off a queue.
#4 does sound appealing, but is it really worth the effort?* take out #2. if something can make use of multiple nodes it can usually make even better use of multi-core parallelization (which affects both computational and memory bandwidth performance). multi-node comes with a much higher communications overhead, so there's a relatively wide range of applications that scale well on multi-core but not multi-node.
* add that #3 comes up as soon as you have complex data structures to share. Serializing and Deserializing (by default with pickle) is a huge overhead for anything a bit more involved. If you design for this from the start you can be fine, but often these things grow and eat up bigger and bigger usecases until you run against the GIL. This basically happens with anything that has enough data and users and need - hey I heard your scheduler tool works well for the cafeteria, I'm sure it can handle our global operations right?
* about #4 - see the previous point.
Those few times, put down the hammer and use some other tool for those not naillike jobs.
#4: Even if you're just talking message passing sending a message between threads is in the 10s of nanoseconds while between processes is 10s of microseconds. That's a ~1000x slowdown on core communication. Given that CPU cores are not getting any faster, that's a pretty big hit to efficiency to take. Similarly simply moving data between processes is expensive, while moving data between threads is free.
Move means the sender no longer has a reference. As in, std::move, rust's ownership transfer, webworker's transferables, etc...
1: Yes there's a single synchronize point where the handoff happens, but this is part of sending a message at all. It's also independent to the size & complexity of the payload itself when we're talking multi-threaded instead of multi-process. You have that exact same sync point that costs the exact same regardless of whether your message consists of a single byte or a multi-gigabyte structure.
Even if you look over a large generation gap there's only a ~20% IPC improvement going from an i7-2600K to an i7-7700K ( https://www.hardocp.com/article/2017/01/13/kaby_lake_7700k_v... )
6 years & a shrink from 32nm to 14nm and all it can muster is +20%. Cores are just not getting faster by any meaningful amount.
That it is not 10x as fast I blame on AMD for not being as competitive as they could have been.
(Kaby Lake is basically a new stepping of Skylake - if intel wasn't having problems with new process nodes it likely wouldn't have been released at all, and if it was it would've been used for a one-off chip in the same generation ala the 4770K)
Python multiprocessing doesn't work well with a lot of external libraries. For example, CUDA doesn't work across forks and many system resources can be shared across threads but not processes. Python objects must be pickled to be sent to another process, but not all objects can be pickled (including some built-in objects like tracebacks).
A lot of different parallel programming models can be built on top of threads (shared memory, fork-join, message passing), and to a certain extent they can be mixed. That's not true of Python multiprocessing, which only allows a narrow form of message passing. (It's also buggy, has internal race conditions, and easily leaks resources.)
The problem for CPython is that it may not be possible to remove the GIL without breaking the C API, and a lot of the benefit of Python is the huge number of high-quality packages, many of which use the C API.
Amdahl's law bears little relevance to throughput computing (i.e. most servers).
> (It's also buggy, has internal race conditions, and easily leaks resources.)
There is also at least one memory corruption bug in multiprocessing (linked a few months back by a fellow HN reader).
And even if you're fortunate enough that Nvidia designs their GPUs to solve your problem, why should the CPU cores sit idle?
But I've never heard of someone asking for a GIL to be added to the JVM.
I've since moved to clojure, which is a language designed with concurrency from ground up. Look at clojure's `atom` - it's basically what every beginning programmer expects from globally shared variables, minus the gymnastics of handling race conditions on your own.
Also, `core.async` is such a beautiful thing to work with for writing schedulers. Compared to this, python's asyncio is an unfunny joke.
I don't think python's GIL can be removed with ad-hoc locking. Nothing sort of complete re-implementation will do.
People are just as quick to bemoan a language for not having something (generics, templates, pre-processors) because they see some perceived need, but a GIL is never one of those things.
Numba already does some of this.
Additionally, I cannot help but wonder if the answer to these problems has been the JVM all along. Especially with JVM 9 and the Truffle framework - https://github.com/securesystemslab/zippy
[1] http://www.java9countdown.xyz/ [2] https://www.infoq.com/presentations/polyglot-jvm-graal (see roughly 42:00 - 47:00)
EDIT: Wording.
from multiprocessing.dummy import Pool
pool = Pool(num_threads)
result = pool.map(your_func, your_objects)
pool.close()
pool.join()
Improve and/or complicate things from there.Often the challenge is a big amount of (hopefully read-only) data that you want to access in every 'your_func'. The naive solution is to copy the data, but this might blow your memory.
Edit: I realize I'm contradicting myself here. No shared memory is a first approximation. You can have shared memory with multiprocessing, but most objects can't be shared.
The costs of synchronizing mutable data between cores is surprisingly high. Any time your CPU thinks that the data that it has in its cache might not be what some other CPU has in its cache, the two have to coordinate what they are doing. And thanks to the fact that Python uses reference counting, data is constantly being changed even though you don't think that you're changing it.
Furthermore if you throw out the GIL for fine-grained locking, you then open up a world of potential problems such as deadlocks. Which look like "my program mysteriously froze". Life just got a lot more complicated.
It is easy to look at all of those cores and say, "I just want my program to use all of them!" But doing that and actually GETTING better performance is a lot trickier than it might seem.
the multiprocessing library allows you to launch multiple processes using your function definitions. It's almost the same as the multithreading library but does not share data.
It seems the real problem, as you pointed out, is the additional memory. I didn't consider situations where each process would need an identical large data set, instead of just a small chunk to work on.
If you're using multiple processes for CPU-bound performance, why not squeeze as much as you can out of each CPU?
Just looking at it from a financial perspective, having a great Python interpreter that doesn't have a GIL seems like a no brainer for $50,000, and it creates another reason why people should take a look at PyPy.
Side note: if you haven't looked at PyPy, check it out, along with RPython
But still I don't know anybody who uses it? It seems like the C extension API is still an issue, or am I mistaken?
Switched from CPython to PyPy, instant 3x performance boost.
This feels like a number that might in the end blow up to 10x the original estimate.
(And really, was it intended to be dangerous in this way?)
No, these libraries are already semantically broken in the same way e.g. libraries which didn't properly close their files and assumed the CPython refcounting GC would wipe there asses were broken.
They're already broken under two non-GIL'd implementations.
Expecting bad code to magically work forever is unrealistic and hinders progress.
What python doesn't have is a C api for extensions that makes sense without a GIL. So ideally a correct threadsafe C extension will continue to be correct, which probably implies that a function called "PyEval_AcquireLock" will continue to provide similar guarantees. Which means that the process for utilizing more cores with pure python code in one process will probably be a gradual upgrade process.
It might as well not be the case here, I just found it funny, 50k is our little magic number.
Neither do I think that raising $50K for Python interpreter would be an issue.
PS: I don't find Django an excellent ORM per se. On the other hand it's highly pragmatic, and their implementation of automatically-generated migrations have saved a good chunk of my time.
* High-contention parallel operations. Doing synchronization through a Manager (a separate IPC-based synchronizing broker process) is of course less preferable than, say, a futex.
* Embarrassingly parallel small tasks. This is a big one. If the operation being parallelized is short, then message-passing overhead takes up more runtime than the operation itself, like a bad Amdahl's Law scenario. Shared address space multithreading solves this problem.
* Related: parallelization without the pickling headaches! Many objects can be synchronized but not easily pickled or copied! True multithreading would really enable a large amount of use cases (map a lambda instead of a named function, anyone?) since the same Python interpreter can just pass a pointer to a single shared object.
* Related: lots of libraries (Keras, TensorFlow, for instance) make heavy use of module level globals, and aren't meant to be run on multiple cores on the same machine (TF, for instance, hogs all GPU memory). Multithreading in these deep learning environments (assuming PyPy support from those packages) is useful for parallelizing the input ingestion pipeline. But this point isn't TF/Keras dependent; I can't recall other modules but don't doubt the heavy use of module-globals that's unfriendly with fork()-ing, especially if kernel-related state is involved.
Using Python isn't always appropriate.
That's because it's been attempted over and over and over again. And each time it ends up failing due to the decrease in single-threaded performance (the bevy of necessary memory mutexes aren't free)), and the extensive amount of work required to make all of the standard libraries threadsafe.
I don't buy the $50,000 cost for a second. Sure, you might be able to safely change the interpreter for that little money, but you couldn't fix up performance and the standard library for that.
https://github.com/chrisjbillington/gil_load
In my experience, the GIL is not held for nearly as high a proportion of the time as people think it is, because properly written C extensions and blocking io always releases the GIL. So long as the proportion of time the GIL is held is not approaching 100%, then you can still get gains from threading. This is almost always the case in numerically heavy code that uses numpy or scipy, since the extensions release the GIL. Threads work almost just as well at speeding up this code as in any GIL-free interpreter.
And usually long before you consider multithreaded code, you'll want to move the bottlenecks of your code over into Cython or something, since that can give speedup factors much larger than multithreading. In which case all you need is a "with nogil:" around the the meaty bit of you Cython code, and then it too will be able to get speedups from multithreading.
Something like Pony[0] or Nim[1]? I'm not very familiar with either one, but Nim says it is inspired by Python, and on the surface Pony appears to be as well.
[0] https://bluishcoder.co.nz/2015/11/04/a-quick-look-at-pony.ht...
"Is it possible that software is not like anything else, that it is meant to be discarded: that the whole point is to see it as a soap bubble?" -- Alan Perlis
There are several things I personally _hate_ about python, but there is a cost-benefit that comes from engineering new things. What new problems are we going to be able to solve by using a new language? If the answer is clear (e.g. imperative programming vs declarative/functional programming let you solve different kind of problems) then it makes sense to do. If certain constructs enable you to completely avoid a recurring mistake (e.g. garbage collection), then it may make sense.
But this?!?!? No man, you don't need a new language to fix this.
Top comment is proposing basically Erlang or an actor model.
As for immutability... well they have to either have it or manage mutable state.
That task of engineering is not something to scoff at and I think building a new language or using an existing language with those ability would help. Erlang is not a number crunching language. But there are others such as Pony.
Much of the Python's appeal is in its huge, colossal, powerful ecosystem, with modules for everything, and things like numpy or tensorflow using it as the high-level interface language. Not breaking this is probably more important for success than efficient in-process data sharing. (Yes, process pools, queues, and a shared DB cover most of my cases.)
In fact, other than "run lots of Javascript", I'm not sure I can name a single thing Node did before Python.
How web workers aren't threads? Browsers are more widely deployed than Node, even with the same V8 engine.
Python itself is just an implementation detail of the underlying VM.
The initial implementation may need to assume single-threaded C interface support and take a global lock but it wouldn't be a stretch to have these things declare they are multithread aware and relax that restriction.
Forgive me but most of these objections seem like post-hoc rationalizations. The first step is deciding to support a GIL-less multithreaded mode. After that, solve the problems one step at a time.
It is amazing how many times accomplishing "magic" boils down to:
1. Decide we're going to solve this problem. 2. Iterate toward the solution in manageable steps.
#1 is by far the most difficult aspect :)
Also, if you use .subst(‘y’, ’n’) instead of a regex it runs in under 9 seconds locally. Thats still much slower than perl 5 (which locally takes less than half a second) but they’re making great strides at improving performance.
time yes | head -n1000000 | perl6 --profile -e 'for $*IN.readchars { .subst("y", "n").print }’ > /dev/null
says it took 86 ms. Which is pretty decent I’d say. time echo "#include <stdio.h>
main(n){char*b,*s=n=0;getdelim(&s,&n,-1,stdin);for(b=s;*s;++s)*s=='y'&&*s='n';puts(b);}">s.c|yes|head -n1000000|tcc -w -run s.c>/dev/null
real 0m0.028s
user 0m0.023s
sys 0m0.019s
28 milliseconds!I used to carry around a Perl program of my own on a printout to take to VLSI interviews. That way when I got the "Do you know Perl?" question I could bring it out and force the interviewer into MY stupid subset of Perl rather than being stuck in his stupid subset of Perl.
That's not a compliment to the language.
Why do I see Python programs that look like a weird mix of Lisp and Java? Surely it is because even with what Python enforce, there are many many ways to produce unclear code that really don't even has to do with the language used.
And why did my employer see the need for the comprehensive Python style guidelines manual... I guess Python bit just as hard as Perl.
Also now that Python is used more by newbies and non-programmers that is where more of the bad code end up (aka Perl late 90s, still pollutes the internet). The quality of Perl frameworks, libs, example code etc is actually increasing and getting easier to find.
a) Inspires much less confidence than starting with a known-correct locking model (the degenerate case being a GIL) and preserving it while improving available concurrency.
and
b) Seems at least 50/50 to end up without much in the way of tangible scalability gains once enough locking has been added to reduce the rate of crashes and data corruption to an acceptable (?!) degree. At least that was my takeaway from all the challenges Larry Hastings has documented while working on the gilectomy. Sure, they don't have to worry about locking around reference counting, but it's not like writing a (concurrent?) GC operating against concurrently executing threads isn't a significant design challenge itself with many tradeoffs to make.
Perhaps they would have done better to say "it works correctly for all programs that do not assume the built-in data structures are threadsafe". That is an accurate description, what you quoted is a reasonable approximation.
In the end I settled for C++ and QT with the native bzip2 library with a few modifications.
I don't know if the bzip2 module does this, but it probably should.
So by the time it comes to consider multiple threads, the bottlenecks that I want to paralellise are already non-GIL-holding.
I wrote a tool to measure what proportion of the time the GIL is held in a program:
https://github.com/chrisjbillington/gil_load
I encourange people to measure what fraction of the time the GIL is actually held in their multithreaded programs. Unless it's approaching 100%, go ahead and use more threads! You will get a speedup. It's my experience that this is true more often than not. The biggest exception is poorly written C extensions that do not release the GIL even though they have no need for it. But if you're writing your own in Cython it's a matter of just typing `with nogil:`.
I may be a bit naive asking this... but why would you care that much?
Looking at activity monitor on my Mac, I count 14 Google Chrome Helper Process instances each spawning upwards of 13 threads. Adobe does something similar, as do several other programs/applications on my machine. Yet, my machine is mostly idle.
I can only speak for myself here. If I want something done on my computer... I don't care if it spams my process list if that is what it takes to complete the task. Don't crash my machine, but do what you have to do to get it done quickly.
Along with that, I like it to be a single process so its easily wrappable in whatever monitoring or process-throttling application you want. I will admit I'm completely assuming that multiple processes is harder than a single process to do that with.
Also, when you get up to the 16 thread count, seeing that many processes pop up at the top of your process list is both annoying and doesn't let you know how much the application overall is using easily. It could also be scary to some users who have never seen that before and think its trying to run a whole bunch of programs.
Yes, some of those are clearly nitpicks and not good technical reasons, but this is a problem that is fixed with a good framework anyways.
Also this kind of thing should be relatively light on the GIL if done correctly. The bzip2 module releases the GIL (I assume?), as does file IO, which is most of the workload in your use case?
Ah well, at least it serves as a warning sign for budding language composers as myself. Snabel did full threading before it walked or talked:
https://github.com/andreas-gone-wild/snackis/blob/master/sna...
And to any Pythoneers with sore toes out there: pick a better language or learn to live with it, down-voting me will do nothing to solve your problems. It's a tool, we're supposed to pick the best one for the job; not decide on one for life and defend it to death. Imagine what could happen if language communities started working together rather than competing. There is no price to be won, we're all being taken for a ride.
If not for that, I'd focus on supporting some kind of pseudo-process where multiple instances of the Python interpreter could be loaded but they would only share pure-functional libs which, I assume, could be used in a threadsafe fashion... but then you run into the mutability of those libs. Well, the mutability of everything in python. Plus what happens if those libs expose anythign that you could hold a reference to - what happens to refcounting in a multithreaded Python?
Honestly, I feel like the world has passed Python by. At this point the cost of its performance limitations don't seem to be worth its payoff. Not that it's a bad language - I like Python. I just don't really feel the need to use it for anything anymore.
At this stage in Python's evolution, I view the GIL removal as a computer science project that some people will implement again, and again, just to learn or to exercise their chops. Great idea! Just don't demand that the entire community of Python developers goes down your road.
If CPython never gets rid of the GIL that suits me just fine. GIL free programming can be done on other implementations of Python like Jython and IronPython. As far as PyPy is concerned, as long as it does not disrupt the use of PyPy as a means of speeding up a CPython app from time to time, then have fun.
When your thread finishes or is ready to signal progress, you queue the event to the UI thread and forget about it.
Now I've been following this pattern for a long time and have no a experience dealing with GIL How this removal of GIL going to effect this use case if at all?
But give me some business reasons as to why removing the GIL is critical. Will is save me a ton of money? Will my stack magically just run faster?
I wonder if Google has already done so since they would benefit quite a bit from a GIL-less python.
I didn't donate to that pot but that does seem like a judicious and reasonable step to take given the assessment of STM.
I ran Larry Hastings' Gilectomy testprogram x.py: fib(40) on 8 threads HW: MacBook Pro, 2015, 8 (4+4) cores, 1Gb RAM
Jython ran the program 8 times faster, utilising all 8 cores >95%. Python ran on 1-2 cores less than 60% utilisation. (Pretty sure Jython will run 16 times faster on 16 cores)
It's 2017, why this is acceptable to GvR and the Python community is beyond me.
Jython: real 1m4.959s user 7m38.521s sys 0m2.396s
Python: real 8m19.035s user 8m16.508s sys 0m11.424s
I also pasted the fib test to pastebin: https://pastebin.com/Ryyb2K7V
Interestingly doing the same on Cpython using the multiprocessing module was ~2x slower than jython/threads. More interestingly pypy with multiprocessing was ~5x faster than jython/threads.
$ time jython fib.py 40
real 1m11.247s
user 6m14.130s
sys 0m3.012s
$ time python fib.py 40
real 2m4.067s
user 11m46.103s
sys 0m2.352s
$ time pypy fib.py 40
real 0m21.040s
user 1m51.461s
sys 0m1.892sAnyway, really good on them to finally move on killing the GIL. It's been a long-time issue - the type that only gets worse the longer you ignore it. That said, I think today Python and GIL are synonymous and the entire Python ecosystem has almost evolved around the GIL. While I'm sure there are applications that would benefit from its removal, I think in the whole, the ecosystem will not change much because of this.
At the time (a year ago) there wasn't a way to precompile using pypy, which meant shipping pypy along with gcc and a bunch of development headers for JIT-ing. Additionally a one of the extensions we used for request validation wasn't supported so we'd be forced to rewrite it. I also found that the warmup time was too much for my liking, it was several times longer than CPython's and it became a nuisance for development. I guess I could've pre-warmed it up automatically, but at that point I had better things to worry about and abandoned trying to switch.
I'm sure, given enough resources, it would be a lot better. But it's not quite as simple as switching over and realizing the performance increases without some initial investment.
It doesn't (or didn't) work when you need to rely on an extension that uses Python's C API. I haven't followed the scene in awhile so maybe that's changed. pypy's pip has so many libraries that I hardly notice, so maybe they solved that.
Unfortunately python is fundamentally slower than lua or JS, possibly due to the object model. Python traps all method calls, but even integer addition, comparisons, and so on are treated as metamethods. That's the case for Lua too, but e.g. it's absurdly easy to make a Python object have a custom length, whereas Lua didn't have a __len__ metamethod until after 5.1. I'm not sure it even works on LuaJIT either. Probably in the newer versions.
(And yeah the CPython API is still a pain point if you've got a library that uses it, although some stuff will still work using PyPy's emulation layer. It'd be great if people stopped using it though.)
You can't always inline the arithmetic ops effectively. You can recompile the method each time it's called with different types, but that's why the warmup time is an issue. This wouldn't be a problem if Python didn't make it so trivial to overload arithmetic. JS doesn't.
PS: I do realize "digital forensics" is probably not the kind of "production environment" you were thinking. Just a small datapoint about a particular branch of software that, while getting good speedups, may not benefit as much as the "X times faster" line would suggest.
You have to not have problematic libraries in your system, but honestly they're all either shitty on CPython too (literally every GUI toolkit that is not Tkinter!) or they're stuff like lxml, where the author/maintainer just has an anti-PyPy bias that they won't drop.
Initially memory tradeoff was definitely significant, somewhere around 40% or so -- it's going to vary across applications though certainly, and in a lot of cases I'm a bit happy our memory usage went up because it forces us more towards "nicer" architectures where data and logic are cleanly separated.
Not that I mean to apologize too much for it, it's something certainly to watch, but for us on our most widely deployed low-latency, high-throughput app, we traded about 40% speedup for 40% RAM on an app that does very little true CPU-bound tasks (it's an s2s webapp where per-request we essentially are doing some JSON parsing, pulling some fields out, building some data structures, maybe calling a database or two, and assembling a response to serialize ~500 times/sec/CPU core).
On more CPU-bound workflows, like one we have that essentially just computes set memberships at 100% resource usage all day long, we saw multiplicative increases, and I can't even mention how much exactly, because the speedup was so great that we couldn't run it in our data center because it started using up all our bandwidth, so I only have numbers for once it was moved into AWS and onto different machines :).
Happy to elaborate more, as you can tell, I think companies with performance-sensitive workloads need to be looking at PyPy, so always happy to talk about our experiences.
1. People like to talk about it a lot, complain about it and say their opinion of what should be done with it.
2. It's not likely to be resolved for years to come.
3. In the end, the problem has very little effect on people's lives, much much less than the amount of hype around the issue.
If you want to ask a question that warrants a response (as opposed to promoting your own effort, which is valid but does not warrant a response), please mail me, the mail is public and I'll put the responses publically on either my blog or pypy blog.
I have concerns that if such functionality will not be in the main release enabled by default(and consequently don't get as much testing), it will just bitrot and in the end, will be removed.
A fully functional PyPy that could do heavy math in multiple threads would be an amazing tool in the box, but there are plenty of risks to that (penalizing single threaded performance, for example). So this strategy makes plenty of sense to me.
They can't just do it on mainline from the outset because there are huge obstacles to overcome.. for example, that ancient foe, CPython extension interface compatibility, which assumes a single global lock covering all mutable extension data. I don't think there will ever be a way around maintaining the GIL for that, even if pure Python code can freewheel it otherwise
Could the experience gained this way (and by other projects such as the gilectemy) help with a future STM attempt ?
I wonder if we need better hardware for STM to work well too.
You can learn more about concurrency in Ruby 3 at this wonderful blog post: http://olivierlacan.com/posts/concurrency-in-ruby-3-with-gui...
In my experience this is almost never the case. Moreover, this type of synchronization is trivial to accomplish with relatively little performance sacrifice.
What is much more complicated is getting more complex logic work correctly and performantly when you are interacting with multiple different data structures from something more than a saturated loop.
Ruby does not have that much dynamism compated to "everything is dict of pointers to dicts" Python.
If done as a separate release, will that version be maintained in the future?
When you redefine a method in any language I'm aware of you just change which method the name points to. You don't modify the original method.
def fun(*args):
if not args:
return 0
return fun(*(args[1:]))
would be call-by-address after the first invocation? It could be lookup-by-name by way of code.In practice we apply speculative optimisations including inline caching and guard removal with remote dynamic deoptimisation via safe points to make it a direct call instead.
Or does it break the current support for porting cpython extensions?
Use Kickstarter or Plasso to sell a pypy pro license - its so much easier for companies to pay invoices than to donate.
If nothing else, I would pay for an official conda pypy package which works seamlessly with pandas and blas.
So I assume they're not doing a kickstarter to prevent the following from happening:
1. The internet at large will assume they're going to get a GILless PyPy that can actually run their code.
2. A separate PyPy is released that doesn't run their code.
3. People are angry that they didn't get what the thought they were gonna get, like what often happens with kickstarter backed projects.
4. With no coporate support and waning public interest due to the uselessness of a GILless PyPy, the separately released project becomes unmaintained.
Did you read the article? They said in the article they aren't asking for individual donations at the moment:
>> we would like to judge the interest of the community and the commercial partners to make it happen (we are not looking for individual donations at this point)
Plus I'm sure they will consider using Kickstarter when the time comes.
And for STM in pypy: 2nd call: $59080 of $80000 (73.9%)