Python 3 Q&A by a prolific core developer
ncoghlan_devs-python-notes.readthedocs.org
ncoghlan_devs-python-notes.readthedocs.org
If people just thought of Python 3 as the replacement to broken text model Python, like how XHTML tried to supplant HTML and failed, and then HTML5 supplanted XHTML with some quirks, the same thing happens here.
Also, I adore the comments about the GIL. I understand the benefits of using a single language in every circumstance, but (even) a JIT scripting language is not sufficient for all problems. I guess it depends on if a polyglot code base is more complex than handling multithreading and the corresponding memory complexity and gimmicks, and that is developer centric.
Still, I would much rather just stick processor bound code in openMP C functions and call those from Python when need be. It seems like the "right" answer to the performance problem, without losing much productivity.
Back in reality, though, complaining about the GIL as though its a serious barrier to adoption amongst developers that know what they’re doing often says more about the person doing the complaining than it does about CPython.
It was a solid, well-argued piece up to this point. You do yourself and the Python community a disservice by writing off your critics as ignorant. It sounds petulant and childish, and is wrong.
There are valid arguments on both sides of the GIL argument, but neither side's advocates are ignorant or bad programmers.
For writing extensions, have you considered Cython?
1) My current project at work is a GPU-accelerated keyword-matching engine. The project was started before I joined the company, so I had no say in the choice of Python. Keywords change infrequently, while we analyze a continuous stream of incoming text. There are several million keywords, ranging from small to enormous in size. Aho-Corasick (http://en.wikipedia.org/wiki/Aho%E2%80%93Corasick_string_mat...) is a pretty ideal algorithm for this scenario, which we use for the GPU matching kernel.
AC requires some preprocessing of keywords into a deterministic finite automaton (basically a suffix trie). This is very expensive for a large number of keywords with a large number of characters. The DFA grows to something like 10GB while being built.
Meanwhile, the main engine loop has to be running continuously, while updating keywords in the background. The engine is a service available to other systems on our network, so it uses multiple threads for concurrent I/O. The problem is that the GPU performance is so ridiculously high that the CPU can't keep it fed with data. I've profiled it and this is not a memory-bound problem...the CPU simply cannot keep up with the document streams that we send to it.
The concurrent I/O threads cannot reasonably be split across processes because they need a shared memory space for the data structures driving the engine. So clearly, the background keyword updating is a problem if it runs in the same process as the rest of the engine. I spent a lot of time trying to figure out how to get the keyword updating working in its own multiprocessing process. It's a complete hack to work around the failings of Python (I can go into more depth about the implementation issues if you'd like). And this is why I loathe the GIL.
We use Cython for some aspects of the code, but the keyword updating has yielded very little gain. It's difficult to rewrite parts of the keyword updater as more optimized Cython because it uses some language features that do not seem to be supported in Cython.
2) For a personal project, I need to do a lot of timeseries processing. I'm using Python to prototype, with the intention of either optimizing it eventually or possibly rewriting it in a more suitable language. I've found parsing timestamps to be particularly CPU-intensive, while working on gigabytes of data. Most data I send to a multiprocessing process will have to be returned in some form eventually, so communication costs are huge. So huge, in fact, that I only see a 10% speedup from splitting the workloads evenly across six cores. Profiling reveals that the majority of the "processing" time is actually just waiting on data getting sent back to the main process. This would not be a problem with a shared memory space.
The GIL is the single biggest target for language-advocacy FUD against Python by advocates of other languages. I would be a rich person if I got a dollar for every time I saw someone trashing Python as a toy language because of the GIL, without significant knowledge of how to use Python. It's just a much huger issue to someone without real Python experience than it is to people with Python experience.
Setting that aside, there are a vast number of use cases where someone might try to use threads when actually that's not a good solution. It actually isn't incredibly easy to come up with cases where the GIL is this huge fatal flaw. It's really sad how many times I have seen people ranting about the GIL and totally unaware of multiprocessing, unable to give a technical reason why they can't try processes, unaware of greenlets, wholly unaware that the GIL does not exist in Jython/IronPython, etc.
I'm sure that doesn't apply to you, which would put you in the minority. Thus, "often says more about the person doing the complaining".
Threads should be used judiciously, because they dramatically increase the complexity of a program and the difficulty of debugging and reasoning about it. Shared-everything is a great way to blow your foot off if you don't specifically need it. So if you throw threads at every problem indiscriminately, then that really is a weak point in your programming. (Note: I am not saying that using threads is always a bad idea)
If you are a good programmer then you should already know, and not be offended to hear, that threads are a tool of specific applicability, not a panacea to scale up everything.
Thank you for telling me what was inside my head. I was simply unaware of my emotional state and motivations until you helpfully pointed them out to me.
But seriously, "good programmers don't use threads (much)" is your counter-argument?
tldr He argues there are better ways for scaling out than threads and removing the GIL would have enormous consequence _throughout_ the whole code base so it’s removal cannot be warranted.
Reducing 10 paragraphs and 11 bullet points to “only stupid people use threads” is just poor style.
Response: The only downside? As a Python user suffering from JVM envy, I have to say that that's a SERIOUS downside!
Here's why: (1) Python is slow. Almost any real life program in pure Python will have CPU-bound components. (example: BBCode parsing on my forum).
(2) Most programs that need to scale won't need to scale beyond a single machine (the most active web forum on my continent runs on a single quad core server)
Therefore the need to be able to scale CPU-bound python programs to multiple cores on is a very real need. Even though we accept that removing the GIL is hard, let's not insult real-life Python users by suggesting that their needs are not real.
There are great process based ways to scale out – look no further than Erlang to see it’s true.
Erlang is hyped as the ideal model for concurrency, but in practice is a niche product that's primarily useful for programs that are almost pure IO - chat servers, routing components like proxy servers and packet switches.
The Erlang model does NOT apply to python, anyway, since Python processes are nothing like Erlang processes. Unlike Erlang processes, Python processes are very heavyweight and message passing between them is costly.
The big difference is that processes are much more robust and testable. The cases where threads are really needed are fringe cases and – while it’s a pity – Python doesn’t seem the right language if you don’t want to go the Jython/IronPython way.
The bigger problem is that people are used to go for threads by default although only few are able to write bug-free threaded code. It’s obsolete but prevalent performance wisdom and the fact that threads were really popular in the Java world.
Proper support of threading inherently allows more performance and flexibility than multiprocessing. On top of them, you can build powerful, Pythonic abstractions like concurrent.futures.ThreadPoolExecutor.map and STM, and on top of those, even more powerful abstractions that help the developer avoid concurrency bugs.
I'm really excited for PyPy. That is a project full of people who are not afraid to quickly iterate on powerful ideas that can make Python a high performance language that it deserves to be, instead of resorting to calling MT programming a "fringe case" and ad hominem attacks.
I understand that everyone here is acting in good faith and wants Python to be better, and the article otherwise contains lots of great information presented in a reasonable manner. You bring up lots of good points too. But other statements like the ones I mentioned are overly broad or brash.
Although it might be great stuff, the word 'Pythonic' is pretty funny next to a long Java-style name like that
What would be the first of my problems?
And this is the last time I’ve wrote this, I feel like a street-organ. >:(
And all I was saying in this thread is that the performance gap between threads and processes isn’t that big of a deal, if you run non-native code anyway. The multiprocessing module is pretty cool.
And all I was saying in this thread is that the performance gap between threads and processes isn’t that big of a deal, if you run non-native code anyway. The multiprocessing module is pretty cool.
This is a line that I hear over and over again, but I strongly disagree with. It's not always easy to predict where your performance bottlenecks will be until you actually start implementing it in some language. If I've chosen Python for a project and find I need more cores, I'm stuck with either re-implementing critical sections of code in C extensions or other languages, or using multiprocessing. And multiprocessing is not that great because it splits the memory space across processes and communication between them is extremely expensive. And there are many caveats to this which cause enormous headaches (eg., you can't fork your process while having an active CUDA context, not all Python objects are serializable, pickling is slow, marshaling doesn't work well for all data types, you must finish dequeuing large objects from a multiprocessing.Queue before joining the source process, etc.).
Yes, I could get a 10-100x speedup by re-writing everything in C. But most of the time, I would be very happy with a 6-12x performance gain from just using threads in a shared memory space.
Did you try Jython?
No, I didn't try Jython. The choice of CPython was made before I took over the project, and there are also a dozen or so dependencies which I don't think are compatible with Jython.
It is absolutely nothing like spawning OS level processes. They are micro-processes, green-processes than live inside the Erlang VM.
You are devoting multiple cores to parsing an individual user's BBcode?
This is your example of a real need to remove the GIL?
(I assume you mean truly mean bandwidth between main memory and the processor, and not to disk.)
Fine-grained caching of objects that correspond to DB rows. Most pages touch hundreds of DB rows, due to the various relationships between objects. With memcached, you have to cache at a higher granularity and contort your code quite a bit to reduce the number of gets per request.
> dicts and lists are fast, but ... your program will now be waiting on locks
In my experience, the overhead of locking is often negligible. In Java-land, you can have millions of lock operations per second. IPC involves serialization, deserialization, and context switching, in addition to actual work. Most IPC routines are built on locks, anyway.
The reason this approach is problematic is that it means the traceback for an unexpected UnicodeDecodeError or UnicodeEncodeError in a large Python 2.x code base almost never points you to the code that is broken. Instead, you have to trace the origins of the data in the failing operation, and try to figure out where the unexpected 8-bit or Unicode code string was introduced. By contrast, Python 3 is designed to fail fast in most situations[...]However, not "in general". The general problem is that bogus data leads to an irreparable situation. To fix that bug, you have to find out who broke the data. Consider for example, a circle in a linked list. You would need quite an expressive type system to prevent that statically (dependent types might suffice? Theorem provers are the ultimate hammer).
a = 1:a makeCircle rst = rst ++ makeCircle rst
The thing is that a circular list is actually an infinite list in Haskell. ;)These aren't "circles" per se, they are partially evaluated recursive structures. If you expand them you end up with evaluated list structures that are non-self referential.
"a" + b"s" is an error in Python 3.
Static types solve it earlier though, especially for little-exercised code paths which could be forgotten or missed in testing.
The programmer would probably just try to fix it via str(u"a")+"s" or u"a" + unicode("s").
Since you have brought up the general case, the origins of many bugs involve something sophisticated going wrong which would require you to encode much of your program (hopefully not too redundantly) into a Turing-complete type system. Rather than stuffing square pegs into round holes, you could just write the appropriate error handling code (or even just asserts).
alias hncomment='fold -w 77 -s | sed "s/^/ /" | pbcopy'
use:
echo "blah blah blah" | hncomment
command-v
Thus: The reason this approach is problematic is that it means the traceback for
an unexpected UnicodeDecodeError or UnicodeEncodeError in a large Python 2.x
code base almost never points you to the code that is broken. Instead, you
have to trace the origins of the data in the failing operation, and try to
figure out where the unexpected 8-bit or Unicode code string was introduced.
By contrast, Python 3 is designed to fail fast in most situations[...] print('ver: {}'.format(', '.join(str(i) for i in sys.version_info)).upper())
versus: print 'ver: {}'.format(', '.join(str(i) for i in sys.version_info)).upper()
Despite the more consistent naming (ConfigParser -> configparser), simplified api (iteritems() -> items()) and all other syntactic improvements, I somehow still find 2.x code more enjoyable. Writing small scripts in Python has kind of lost its charm for me.This, of course, is all very subjective and I'll probably grow over it in a few thousand lines of code. I hope you're all less sensitive to the little things that annoy you.
Jumping on the internet to say that “they” (specifically, the people you’re not paying a cent to and who aren’t bothered by the GIL because it only penalises a programming style many of us consider ill advised in the first place) should “just fix it” (despite the serious risk of breaking user code that currently only works due to the coarse grained locking around each bytecode) is also always an option. Generally speaking, such pieces come across as “I have no idea how much engineering effort, now and in the future, is required to make this happen”.
Fine, then don't bother. But don't insult us or call foul for our pointing out the downside impact on us of this decision. Your assumption that we are naive rubes who don't know how to code is really, really wrong.
I have a very good idea how much engineering effort would be involved in fixing the GIL, and I am well aware that Python has involved many person-millennia of gratis work, and am appreciative of both. However, I still disagree with the Python devs' obviously entrenched position that fixing the GIL isn't worth the effort, and I will continue -- even when shouted down by the likes of you -- to advocate for the GIL's removal or some equivalently good solution. (As I said earlier, I am not opposed to STM solutions, but the current one performs unacceptably without special-purpose hardware.)
Why? Because I have a single, selfish interest in this. I depend heavily on Python now, and want the language to be better. I have written many lines of Python 2 code that rely on the threading primitives in the standard library. Perhaps it was foolish of me to expect that the threading model offered by the standard library, modeled on Java's threading primitives, would some day work in the same way as Java threads do in practice. Nonetheless, I am left with a real world problem: my CPU-bound threaded Python code does not scale well to multiple cores. I need the GIL fixed, or to rewrite my Python code, or to migrate to another language that supports the standard model of threading programming that real-world programmers have been using for several decades, and which has built-in support from all major operating systems. Or, sure, wait for STM to be ready for prime time and migrate my thread-based semantics to the new STM-based semantics.
The best path right for me right now is migration to Jython or IronPython. But then we are still unsupported orphans, living in the third world of Python 2.X.
I guess it comes down to: do you want people to actually use this language to write programs they want to write, or do you want Python to be an advocacy platform for "correct" programming? Python's pragmatism has always appealed to me, so the ivory tower reaction to the practical concerns around the GIL really seem dissonant. (And this is coming from an MIT AI Lab Lisp guy who would rather write everything in Lisp. But Lisp lacks Python's wonderful, high-quality third party libraries and developer-centric pragmatism regarding unit testing, documentation, etc.)
I know you are tired of hearing people bitch about the GIL, but, really: people write multithreaded programs. They should work as advertised, using native OS threading primitives and taking advantage of the native OS thread scheduler. Why does Python offer threading primitives if the language is not meant to support, from a practical standpoint, multithreaded programs?
What standard? And whose "real world"? The need for threads has always been controversial even among OS kernel devs. UNIX/Linux/BSDs have twisted and non-trivial threading histories peppered with religious wars similar to this one. And which "several decades" are you talking about?
There is no such thing as "standard threading model". To some, a thread is just a flavor of fork() with a wrong parameter and plenty of "real world" programmers continue to believe that kernel-level threads is a hack. And please, do not make it sound like Python threads are useless. Far from it.
Python threads are not what you are used to. That's pretty much TL;DR of your comment.
...Python's pragmatism has always appealed to me, so the ivory tower reaction to the practical concerns around the GIL really seem dissonant...
I feel like they are being dragged into it though. The original motivation behind GIL support has always been a pragmatic one: removing GIL will make the entire codebase more complex, harder to hack on and will complicate and slow down the development/maintenance of the libraries. That's pretty pragmatic.
But a fairly vocal groups of users started to claim, similarly to you, that programming with threads is supposed to work like they expect it to work according to make-believe "threading standard", to which GIL supporters (correctly, IMO) replied that shared memory + locks is not the only/best approach to concurrency. It is easy to be offended by this answer but it doesn't invalidate their point.
Correct. I am used to programming language threads that work the way computer scientists and programmers have typically described them -- for example, as in this (I hope uncontroversial) Wikipedia article:
http://en.wikipedia.org/wiki/Thread_(computer_science)
When I say "standard model of threading", I am not talking about nuances of call conventions to the underlying OS thread primitives. I am talking simply about running multiple streams of instructions, bytecodes, or other units of computation in parallel, within a single OS process.
You can't talk about running multiple streams of instructions or bytecodes in parallel without talking about the nuances of how they share memory. Semantics of a multithreaded memory model are a highly "opinionated" thing -- there are lots of possible ways to define it, and the definition can have widespread effects on efficiency, ease of programming, and the guarantees that the runtime can provide. For example, an important aspect of a Python memory model would be that no Python program can SEGV the interpreter due to a race condition.
I recommend the following reading to get an appreciation for how much really goes into a memory model and how far from "simple" or "standard" it is:
http://en.wikipedia.org/wiki/Memory_model_(computing)
http://en.wikipedia.org/wiki/Java_Memory_Model
http://www.kernel.org/doc/Documentation/memory-barriers.txt
Python is a lot harder to define a good memory model for than say Java, because in Python lists and dictionaries are primitive objects. If you say: x['A'] = 1
...that is a single operation that must not corrupt the dictionary, even if multiple concurrent threads are mutating it. In practice, this means that you need to either make every such mutation wrapped in a lock (which adds a lot of locking overhead) or you need to use lock-free data structures (which are still relatively experimental and architecture-specific).I still don't think it's reasonable to conclude that typical programmers are fine with their threads not really running in parallel, or that the GIL isn't worth bothering to fix, even though fixing it would be hard. In my original post yesterday, I pointed out that as the language footprint has grown, Python's disadvantage in this respect has increased: it is much harder to remove the GIL now than it was in, say, the 1.5 era when there actually was a (problematic) GIL removal patch.
We've gotten way off track, but the original point I was trying to make was that 1) the GIL really is a problem for not-purely-theoretical programs written by competent developers, and 2) that the 2->3 transition, by complicating the language and increasing the workload for the alternative implementations, has made it less likely than ever that the GIL problem would be resolved.
And, indeed, Nick explicitly confirmed this by saying the GIL is basically a dead issue for the CPython devs. His post made many good points about the merits of the 2->3 transition, and in particular pointed out some ways that 3 has reduced work for the alternative implementations, but I remain unconvinced overall. And not out of ignorance or incompetence, as he implied.
The GIL is not a bug, it's a threading model. You wish the threading model was something else. You insist on your particular vision of an alternative threading model without acknowledging its downsides. You make no indication that you have actually considered or tried the alternative concurrency models that CPython does support, like multiprocessing, greenlets, or independent processes. You make no objective arguments for why your desired threading model is better than the ones that are currently available, except that you could avoid changing your code. You accuse Python of failing to live up to some accepted standard for what a "thread" should be, when in fact no such standard exists, especially for high-level, dynamically-typed languages like Python. If anything, newer languages are moving away from shared-state concurrency; see Erlang, Go, and Rust.
I don't think you have malicious intentions, but I urge you to reflect on what you are demanding and whether it is reasonable. What may look to you like "obvious" brokenness that demands an "obvious" fix is really a lot less clear-cut than you seem to think it is. I feel for the Python developers who have to deal with this complaining all the time.
Yes, but Python also exposes higher-level operations like table manipulation as language primitives.
> Given that Jython already allows true multithreading
That may be, but as I mentioned this has an inherent cost, both in CPU and in memory. Therefore it is not a strict improvement over CPython, just a different direction.
Just to be clear, Python does use native OS threading primitives, and it does make use of the native OS thread scheduler. Also, Python does support, from a practical standpoint, multithreaded programs, but only if the program is not CPU-bound. I think you could rewrite that paragraph to make the role of the GIL clearer.
But it's still not the case that multiple threads "normally" run in parallel, with the OS ensuring fairness, which is what (I think) most programmers would expect threads to do in a general-purpose threaded language.
"Note that potentially blocking or long-running operations, such as I/O, image processing, and NumPy number crunching, happen outside the GIL. Therefore it is only in multithreaded programs that spend a lot of time inside the GIL, interpreting CPython bytecode, that the GIL becomes a bottleneck."
The sky isn't falling, in the worst possible case you can still use Jython or whatever
The sky isn't falling, in the worst possible case you can still use Jython or whatever.
Sigh. I could port to C++ or whatever, too.
I would hardly call I/O with Python built-in functions to be a subversion of the GIL.
If Guido doesn't want to implement actor concurrency within the Python interpreter, someone could write their own Python host that allows multiple, independent instances of the Python embedded interpreter. The host would implement a C extension that allows interpreter instances to send deep copies of objects to each other.
description: http://effbot.org/zone/default-values.htm
I presume you mean that you can’t access your old Python 2 modules from a new Python 3 installation. That’s just because Python installations usually don’t share their modules (i.e. site-packages). You can try to install them using Python 3 (usually just python3 setup.py install) and see if they work or not.
Even PyPy project doesn't have a goal of running py2-compatible and py3-compatible modules inside one interpreter:
> At the end of the project, it will be possible to decide at translation time whether to build an interpreter which supports Python 2.7 or Python 3.2 and both versions will be nightly tested and available from nightly builds.
Shedding weight is sometimes a step in a good direction. But we need to draw a line at some point. As the history of Python 3 adoption shows, that point maybe was not optimal.
Edit: forgot an angle:
Also, __future__ is fixing the incompatibility of what people are using with something that they can't use. I'm not sure if it's the right problem to solve.
But look at all those current efforts at Canonical, Django, Twisted…it’s not like nothing is happening and we expect a knot to burst. Far from it! Porting started slowly and has gained a momentum by now that has surprised myself. It’s not like we changed the language completely like perl6 did.
In the result, we’ll have a better Python for it.
from __future__ import python3
(via http://www.aaronsw.com/weblog/python3)I look forward to the day PyPy is considered the real Python. Look at PyPy's homepage (http://pypy.org/), doesn't even mention Unicode as a significant feature. Instead it talks about speed, security, concurrency, and compatibility with the current real Python (2.7.2) - all the things real Python programmers care about and expect the Python developers to focus on.
PyPy may be some way off but I want to find the developers and hug them for setting the right vision and trying.
I just donated $50 and if you hate Python3 you should too.
Guess I'm not a real programmer, then, since I find proper Unicode handling to be a welcome feature, and find the implementation in Python2 is painful.