Guido van Rossum's Performance tips for Python
plus.google.com
plus.google.com
Surprisingly enough, I happen to have a talk when I discuss this precise topic, for people with a 30min to kill. http://www.youtube.com/watch?v=ZHF5Aius_Qs&feature=youtu...
I only point out that a list of optimization techniques that advocates nothing will be different than a list of optimization techniques that advocates certain abstractions at the expense of performance. And there's nothing sad about that.
It says "if you want performance don't do X because X is not a reasonable sacrifice". Then he linked to some better ones.
If you watch the linked talk, "write simple code when given the choice between simple and complex code" seems like fijal's main suggestion (and in my limited experience with PyPy has worked for me). The other main suggestion that I assume is underlying is "if you want performance try PyPy". Neither of those require sacrificing abstraction. Some of Guido's suggestions do.
He never said he doesn't advocate performance. He said performance should not be achieved by sacrificing abstractions.
>In addition, it sounds you mostly agree with regard to rewriting in C... use it as a last resort.
In CPython, not necessarily in Guido's advice, it's one of the first and most common things you hear about performance. "Just write parts of it in C".
If that was the advice of the early JVM guys ("just write parts of it in JNI"), we would never have gotten a fast JVM.
Agreed. But Guido didn't advocate performance, or anything, so fijal is trying to disagree with him about a topic he didn't state an opinion on. There's no need to make up a disagreement. Imagine if fijal instead said:
"This is a fine list of techniques if you are willing to sacrifice your abstractions. At PyPy, we're not willing to do that, and we still achieve sufficiently high performance by..."
You seem to be talking about PyPi: http://pypi.python.org/
One of the more surprising results (esp. for non-pythonists) is the fact that string formatting:
s = "%s" % some_integer
is faster than "casting" to string: s = str(some_intenger)
That's solely because looking up the 'str' symbol requires finding an element in global symbols' hashtable. This turns out to be more expensive than parsing the format string and building the result of % operator.python3.2:
>>> timeit.timeit("for i in range(100): s = '%s' % (i,)", number=100000)
2.868873119354248
>>> timeit.timeit("for i in range(100): s = '%s' % i", number=100000)
2.615748882293701
>>> timeit.timeit("for i in range(100): s = str(i)", number=100000)
2.4016571044921875
>>> timeit.timeit("for i in range(100): s = i.__str__()", number=100000)
1.8993198871612549
python2.7: >>> timeit.timeit("for i in xrange(100): s = '%s' % (i,)", number=100000)
1.9474480152130127
>>> timeit.timeit("for i in xrange(100): s = '%s' % i", number=100000)
1.6135330200195312
>>> timeit.timeit("for i in xrange(100): s = str(i)", number=100000)
2.009705066680908
>>> timeit.timeit("for i in xrange(100): s = i.__str__()", number=100000)
1.539802074432373 >>> timeit.timeit("'%s' % i", setup="i = 42", number=1000000)
0.31799793243408203
>>> timeit.timeit("str(i)", setup="i = 42", number=1000000)
0.4146881103515625
This is another argument for profiling everything.1. python3.2 str() is at least as fast as '%s' formatting
2. '%s' is slightly faster than str() with python2.7
3. theint.__str__() is faster than either alternative in all cases
for i in xrange(10000000):
s = "%s" % i
is also faster than lstr = str
for i in xrange(10000000):
s = lstr(i)
[Edit] After thinking about it a second longer, I wonder whether there is some lookup for lstr as well even though it's local. But storing the function in lstr is faster than using str so I'm not sure how this is actually implemented. I'm sure someone here will know more.Your first reply that "there's always a lookup" isn't right. Some variable accesses (specifically: local variable accesses) do simply map to array accesses.
Each "dot" incurs lookup. Your example reminds me of one idiom. Assigning nested.look.up to local var for access inside loop.
And why is that presented as something inevitable?
The interpreter/compiler could analyze that part and see that the function/name is not changed during the loop, and cache for that.
I'd guess that PyPy tries to do it that way, anyway...
You can have a lot of fun playing with the dis module to look at the bytecode that Python generates for your functions:
1. For the most part, you can just take existing Python code and have it magically transformed into not-too-horrible C. A few optional type annotations will help with the speed, but the compatibility is great right from the get-go.
2. With a little care, you can often get the inner loops of your Python code to be just as fast as hand-written C.
The tutorial is a quick read, and gets you up to speed without much effort:
If I want to make a class with a single int member, and a bunch of member functions, I can trust that will, in most cases, compile away to be just as efficient as a raw int and inline code. It is very liberating to not have to worry about the efficiency of creating another function.
1. Use built in data structures whenever possible
2. First fix your process flow, then spend more time to fix your code. Doing steps A C E F B and D might be better than A B C D E F.
3. Write direct queries with db when all else fails.
I often use python libs which functions are written in python itself and create bootlenecks. Rewrite those functions by calling "native" CPython functions, generally make an extremely large differences
We should remark a clear distinction between Python tips and CPython tips. In PyPy or other implementations some tips does just not make sense and are useless (e.g.: local scope stuff vs. outer scope stuff in loops).
I'd still like it if they could get rid of the GIL...
- IO-bound tasks: GIL is released
- CPU-bound tasks on N-core on a single host. To exploit multiple CPUs:
a) no GIL (hypothetical): N times speed up (optimistically)
(it suggests weak data dependency i.e.,
multiprocessing can be used to the same effect)
b) option a) with multiple processes (shared memory or communication-based approach)
Code complexity is the same on average (except on Windows)
c) C extensions (existing or new): speed up 10*N or more on numerical code
Cython makes it easy to write new extensions.
Currently due to dynamic nature of Python, GIL or no GIL,
C extensions might be necessary to exploit hardware fully
(though C extensions might not be an option in some projects)
- scalable to multiple hosts tasks: different processes i.e., GIL is not a problem
Benefits of GIL: - C extensions (and the interpreter) are simpler to write correctly.
Multithreaded programming is not trivial we need all the help we can get.
- no performance penalty for single-threaded code
- it encourages a synchronization through communication concurrent model (builtin in Go)
Disadvantages: - some applications have no localized performance bottlenecks.
So writing a small C extension won't help to get possible
performance benefits of running on multicore in parallel.
If performance is critical; Python might not be the right tool in this case
- there are pathological cases when performance suffers greatly due to GIL
(though other approaches would have their own pathological cases)
- a (non-informed) perception that Python can't benefit from multicoreThere are a myriad of ways around it, but it is another thing you need to take into account when writing python programs or deciding to use python for a project.
"- Don't write Java (or C++, or Javascript, ...) in Python." and then you don't need to search for workarounds.
I'm not arguing for or against the GIL here but it's something to keep in mind. "Threads run concurrently" is misleading. And it's not a matter of being pythonic, the GIL isn't even a python feature.
[1]: http://docs.python.org/c-api/init.html#thread-state-and-the-... [2]: http://www.scipy.org/ParallelProgramming#head-9e56edb190bf1e...
For heavy math workloads, which was what I was thinking of, you can use the techniques you linked to there to great effect. It is something that you need to put a little more thought into than you might in a non-locked situation like pthreads in C though.
but i suspect you were voted down because these days hn favours tribal groupthink (you criticised python) over rational thought.
Also, getters/setters are totally inane in python (and probably most other languages).
In fact, in general Guido's advice reads as a good warning to folks showing up from other languages who's first reaction might be to create a com/mycubiclefarm/exceptions/abstract/ directory and start writing SeriousBaseClassesForMyExceptions.
> the situations in which Guido is talking about being conscious of your stack frame are pretty rare in typical practice
Excessive function calls in tight loops may be expensive but hardly stack frames.
One of the reasons I'm fond of Python is that, while there is a tradeoff between flexibility and performance, it gives you the means to sacrifice flexibility to aid you in improving performance -- once you've identified what (if any) actual performance bottlenecks you face.
[0] http://c2.com/cgi/wiki?SufficientlySmartCompiler [1] http://prog21.dadgum.com/40.html
A very smart compiler could probably attempt to prove that no such modification can occur at run time throughout the whole program; but this is much harder than simply deciding whether inlining a given call is worth it or not.
In CPython land, Python is slower, but performs predictably, and if you want to speed it up you write C, which is much faster, and also performs predictably, though it takes some developer effort. In PyPy, you get some of the speed of C without the effort, but without the predictability either.
(The answer was: branch prediction.)
So, sounds like he's looking for something like v8.
if v8 can be 10-30 times faster than Python, for an equal or even more dynamic language, I don't see why Python should need to manual tune the things Guido describes in order to get some speed.
The overhead of function calls and property access in Python is part of the price for the increased flexibility of the language.
Rossums tips are for the corner case where you want to improve performance but you don't bother rewriting in C, e.g. if you have a small bottleneck in a larger program.
GVR's tips can be seen as an escalation path for optimizing. Don't jump into C before you've actually optimized your Python, because you might not need to.
In 10 years of writing Python, I've yet to hit a point where I've needed to do that. I'm kind of looking forward to it, actually.
Or does the issue simply not come up in practice, because you rarely need to redefine how assignment/accessing works?