Python 3.11 beta vs. 3.10 benchmark
github.com
github.com
Basically one can expect overall 24% increase in performance "for free" in a typical application.
Improvements across the board in all major categories. Seriously impressive.
Stack usage and multiprocessing had the largest performance increase. Even regex had 21% increase. Just wow!
And this may be the first Python3 that will actually be faster (about 5%) than Python 2.7. We've waited 12 years for this... Excited about Python's future!
-----
python-speed v1.3 using python v3.9.2
string/mem: 2400.67 ms
pi calc/math: 2996.1 ms
regex: 3201.59 ms
fibonnaci/stack: 2487.13 ms
multiprocess: 812.37 ms
total: 11897.85 ms (lower is better)
-----
python-speed v1.3 using python v3.11.0
string/mem: 2234.78 ms
pi calc/math: 2667.84 ms
regex: 2548.81 ms
fibonnaci/stack: 1149.57 ms
multiprocess: 480.25 ms
total: 9081.25 ms (lower is better)
-----
`$ time python3 -c "print('hello')"`
I have a handful of utility scripts written in python, and the overhead of starting/stopping is just massive! (Sadly, I don't have one on hand to actually print a metric...)
Numbers for "hello world" on my machine:
Java 7: 43 ms
Java 8: 29 ms
Java 11: 34 ms
Java 18: 22 ms
Python 2.7: 6 ms
Python 3.10: 12 ms
Perl 5: 1 ms
$ time python3 -c "print('hello')"
hello
real 0m0.223s
user 0m0.148s
sys 0m0.047s
and, for the record: $ python3 bench.py
python-speed v1.3 using python v3.7.3
string/mem: 46673.98 ms
pi calc/math: 98650.32 ms
regex: 38489.41 ms
fibonnaci/stack: 30723.3 ms
multiprocess: 75500.34 ms
total: 290037.35 ms (lower is better)
Might get a little better with some tweaking of performance governors / idle states enabled but, yeah... def f(x):
return x * x;
v=0.0
for i in range(100000000):
v = v + f(i)
Execution time: 15 seconds---
JavaScript:
function f(x) { return x * x; }
var v = 0.0;
for(var i = 0; i < 100000000; i++) v += f(i);
Execution time: 0.5 seconds---
So calling a function in a for loop in Python is still 30x slower than in JavaScript which is an as-dynamic language. Good to see a 24% increase, but a 3000% increase should still be possible.
I did this test in python 3.10, not python 3.11, but I assume if they did have a 30x speedup from the inlined python to python function calls this would have been mentioned, but the fastest speedup listed is 1.96x faster.
Surely PyPy would be a fairer comparison
Indeed CPython is not JIT and PyPy is, but CPython is unfortunately usually what you get, comparing with PyPy would not be fairer since the goal is to speed up "vanilla" CPython.
It's unfortunate to me that the default python interpreter, so widely used for scientific computational purposes, is so much slower than the JS interpreter you get in your web browser, and then we're being happy +24% benchmark results about this when 30x faster is known to be possible (e.g. with JIT).
I am curious but don't have the time to check would the same result be for a regex?
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Of course this is testing two native implementations of regex parsing, not the higher level scripting language handling.
Why does one have to be ok with a totally expected speed difference? This is the official interpreter of a language, they give a slow tool by default, the official tool deserves scrutiny.
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
It really is not. Python allows a huge amount of the object model to be overridden and introspected at any point. Javascript doesn't even allow operator overloading.
- Cheaper, lazy Python frames: 3-7%
- Inlined Python function calls: 1-3%
- PEP 659: Specializing Adaptive Interpreter: variable
The interpreter changes work by generating specialized code so this might have a memory impact which they are working to offset and expect to cap at 20% more.
See the in progress release notes for details: https://docs.python.org/3.11/whatsnew/3.11.html#faster-cpyth...
> During a Python function call, Python will call an evaluating C function to interpret that function’s code. This effectively limits pure Python recursion to what’s safe for the C stack.
> In 3.11, when CPython detects Python code calling another Python function, it sets up a new frame, and “jumps” to the new code inside the new frame. This avoids calling the C interpreting function altogether.
> Most Python function calls now consume no C stack space. This speeds up most of such calls. In simple recursive functions like fibonacci or factorial, a 1.7x speedup was observed. This also means recursive functions can recurse significantly deeper (if the user increases the recursion limit). We measured a 1-3% improvement in pyperformance.
https://docs.python.org/3.11/whatsnew/3.11.html#inlined-pyth...
Sadly, it isn't true parallelism, more like green threads. And they didn't manage to kill the GIL.
Python's own stackframes are heap-allocated and chained (so each function has its own stack, in essence), so its use of a single unified allocation space for stack is already an implementation detail of the interpreter.
The explanation is that Python features heavy dynamism, meaning one can easily replace any part of the system at runtime (at any moment), including builtins, code objects, and even hook into the parser, AST, import mechanism, etc.
This renders inlining a dangerous exercice: say you inline the builtin len() function, how do you know it's not the intent of a code that runs later to replace it with a different implementation?
Now, there are ways to implement inlining, but they are not as straightforward as say, with a compiler, when you know nothing is going to change afterward.
So it's something that had be delayed.
I think progress could move forward by adding a compiler flag "assume no fully free function body replacement" you know just like in every other mainstream language.
E.G: functions are objects, and any object can be called as a function. And it's just a reference to "something callable" in the namespace mapping (literally a dict in most implementations), which is mutable by definition
Also a compiler flag would not help, since most users don't compile the python VM.
Now, you could put a runtime flag, that list all stuff that are created for the first time in a namespace, and refuse to allow reassignment.
It's possible, but it would break a LOT of things and prevent many patterns.
The last attempt was to put guards, and to assume no replacement, but if a replacement occurs, then at this moment we revert locally this assumption.
The process is refined at each iteration, but there is no turn-key solution as you seem to believe.
Ah yes, a compiler flag which literally says “break the language”, that sounds like a great feature which would be used a lot.
> assume no fully free function body replacement
There is no “body replacement”, the `len` builtin is looked up as a global in the module on every access, anyone can just replace the module’s global.
You might recall that there are still holdouts who haven't migrated to Python 3. Breaking changes are generally not a good idea in mature languages.
Although Perl 6 was meant to be a Python 3000 type thing, it was spun out into its own language (Raku). Perl 7 will continue the 5.x lineage, but with saner defaults. Thus, any code that’s around now should run, but the interpreter’s name might change.
(I think…I mostly keep up with Perl for nostalgia’s sake).
The problem for them is that Perl 6 actually had an official release (and IIRC, more than one) under the Perl 6 name; the rename of the language to Raku (which IIRC was originally the name of just one VM for running Perl 6 programs) came later. So anyone who managed to keep up with the latest release of the Perl language would have two huge breaking changes (Perl 5 to Perl 6 and Perl 6 to Perl 7), the second one undoing most of the changes of the first. (AFAIK, PHP avoided all that by deciding to go back before officially releasing PHP 6, so anyone who was following the latest release just jumped directly from PHP 5 to PHP 7, avoiding the breaking changes which had been planned for PHP 6.)
Also you should think in terms of Perl 5 -> 7. Raku is the language formally known as Perl 6. Perl 7 isn't undoing anything, it will be the successor to Perl 5. Perl 7 should have been 5.32 with saner defaults, but of course it's now going to take another 20 years of bikeshedding until this is happening.
Anyone keeping up with the latest release of the Perl language has had near zero breakage for decades. (Indeed Perl has a well deserved reputation for having an outstanding track record in this regard compared to almost all other mainstream PLs.) I personally see every likelihood Perl 7 will extend that track record, though of course my crystal ball prognostications are necessarily based purely on what I see.
No one using P6 or Raku had a huge breaking change from Perl 5. No one using P6/Raku will have another one going to Perl 7.
If you presume Raku and Perl are different languages you'll get the essence of what has actually happened so far, and seems likely to be more or less true for the rest of this decade at least.
The Perl community may be attempting to "fix" that mistake with Perl 7, but the damage has been done. I think Perl 6 fractured the community so hard and scared away so many users that both Perl 7 and Raku are likely to be minor languages for the foreseeable future.
If either of them has legs, I think it's likely to be Raku. Because Raku is at least an interesting language with a lot of really interesting, powerful language features.
Perl 5/7 by virtue of its history, is mostly just another dynamic scripting language, but one with a particularly unfriendly syntax. Aside from CPAN and inertia, I think there is relatively little reason to write new Perl 5/7 code when PHP, Python, Ruby, and JavaScript are out there.
There are more interesting things happening with Ruby and Python. In fact there are Python libs for obscure stuff like CAN-bus and ISO-TP. Want to talk to a vehicle ECU? It can be done with Python.
There's also lack of momentum and quite a lot of bikeshedding in the Perl community. There were attempts to modernise Perl by rurban and others, but they were met with unecessary resistence. Without community support they all ended up as one man shows. Perl is pretty much a dead end. You are hearing this from someone who still writes Perl code every day for work. My latest proof of concept was done in Ruby and it will probably end up as production code.
I wonder when it was the last time you tried Raku. Or compared it to a Perl script with Moose.
> plus every lib has to be rewritten from scratch, including the good and mature ones from Perl.
The good and mature ones from Perl can be used from Raku with Inline::Perl5.
Inline::Perl5 is quite an ugly last resort solution. And you still need a Perl intepreter as opposed to calling on C code from D where you just need the libs.
Also, when you're talking slow: why is it that any performant Perl module, actually has most of its logic written in XS (aka C)? So I think it's shows quite a bit of hutzpah to call Inline::Perl5 "an ugly last resort solution", whereas a *lot* of upstream CPAN modules rely on the hack that is XS to make them performant. To give you an example: the pure Raku version of Text::CSV is more than 2x as fast as the pure Perl version of Text::CSV.
They didn't decide to break over print() and cosmetic changes, it runs way deeper.
It's not unreasonable. It's just what the language makes easy and useful. It's still used much less than for example method aliasing in Ruby which has about the same result.
JavaScript doesn't do this, you can replace properties of window at will. I don't think Ruby does it, or Lua. PHP probably does it under some circumstances in modern versions but only because it's remarkably un-dynamic for an interpreted language.
This is normal fare for high-level interpreted languages, for better or worse.
But that means you have to add all sorts of dependency tracking, such that you are able to deoptimise any function affected by mutation on things which were optimised in or out.
This means your complexity increases very fast, very high, before you can have anything which actually works.
JavaScript too, any function can be dynamically changed at any time, but its function calls are 30x faster (roughly, I measured it just now with a for loop doing 100 million function calls, which took 0.5s in JS, 15s in python)
This is a realistic scenario if you want to run e.g. a custom statistics function on a large array in python so it can be annoying (and yes numpy can do things but that's no reason to keep the main language encouraging less readable code where you don't define separate functions)
Javascript has the benefit of having billions of dollars allocated to it from Google/Microsoft/Apple to hire dozen of amazing engineers full time over decades to work on it.
Python has existed since 1994, yet in 2011, the Python Software Fundation budget was less than $40K: https://pyfound.blogspot.com/2012/01/psf-grants-over-37000-t...
We are talking a difference of funding of 6 order of magnitudes. Not to mention one is a language that has to stand on its own, the second has an accidental monopoly on the most popular plateform in the world.
So it had to be delayed, unlike for JS.
Also, I'm pretty sure cpython devs should ask Google/OpenAI/MicrosoftAI funding, how many millions can they waste on useless projects while not improving the core bottleneck..
Your benchmark also measures iterators and boxed integers, which were not in the JavaScript version, so it's not clear how much of the difference is due to function call overhead.
(Of course, it certainly doesn't help Python's performance that there's no simple way to do loops without iterators and integers without boxing. It would be nice if a future version could optimize the abstractions away.)
My point was that the Python version of Aardwolf's function call benchmark [1] is definitely using iterators and boxing all numbers, so it's doing a lot more than just calling a function, and the Java equivalent would look something like this:
Long f(Integer x) {
return new Long(x.longValue() * x.longValue());
}
// ...
Double v = new Double(0.0);
Iterator<Integer> range = IntStream.range(0, 100000000).iterator();
while (range.hasNext()) {
Integer i = range.next();
v = new Double(v.doubleValue() + f(i).doubleValue());
}
[1] https://news.ycombinator.com/item?id=31427506- Cheaper Python frames are enabled by making some semantic changes to the language, mostly by removing some dynamic functionality that no one uses
- "Inlined Python function calls" is a bit misleading: Python functions are not inlined, it's the "code that does the bookkeeping for calling a Python function" that is now eliminated for Python-to-Python calls.
- The most important part of the specializing interpreter is the attribute caching, which is quite significant
Source: I work on Pyston, where we do many similar things (but can't make semantic changes to the language)
What dynamic functionalities does no one use?
If a user is wanting to get or manipulate debugging information then the old frame struct will be generated when this is called.
The interpreter starts out with a generic/slow approach and as it gets more datatype info it uses more specialized and faster implementations.
In practice it is not slow at all. Plus the development iteration speed is probably second to none of all the programming languages.
If I would have my kid learn two programming languages it would be HTML and Python.
That has nothing to do with it being slow though. And as I said, all those machine learning code rely on numpy (which relies on LAPACK which is not written in python) or highly optimized cuda kernels (which again is not written in python).
Not a compelling repalcement. I spoke to the gonum developers when they first created it and told them they were wasting their time beceause the go leaders made their language intentionally be a "systems language", not a "scientific language".
In practice, you're either not using Python (sleeping, while waiting for IO operations) or not using Python (C/Fortran/whatever libraries for heavier lifting: Pandas, numpy, PyTorch, etc).
I will say that I agree with the sentiment of your comment: if your main focus is speed, you won’t be using Python, and if you’re using Python, you likely don’t care about speed.
If your metric is based on execution time, then this might be true. Many programs are faster enough where the user doesn't care, or the impact on the overall execution time is slight.
But, if your metric is compared to other languages, this is measurably false [1]. Even Python emulated in the browser is faster than CPython in many cases [2]. And, this doesn't really give a complete picture, since many CPython libraries don't actually use Python, because pure python implementations of most anything are too slow. They use python as glue to call out to compiled libraries. But, this is also the main use case of CPython: glue for not Python.
1. Python always near the bottom. See other problems too: https://programming-language-benchmarks.vercel.app/problem/b...
2. Brython benchmarks: https://brython.info/speed_results.html
A lot of numerical software is Python glue code written around compiled kernels (written in C++, Rust, Fortran, etc) but the runtime of that glue code usually does impact the total runtime in a non-negligible way. So all wins are good to take.
Django and friends might always have to be in "dynamic" mode, but possibly you could allow for some complex logic to run quickly if you use subset of the language. (e.g., like RPython but with less overhead to set up)
Also, users would not like to see performance drop permanently, only because they redefined some function (e.g. to log it’s arguments in a debugging session), so they’ll expect the system to (eventually) re-optimize code using the new state.
Languages such as JavaScript and Java do this kind of thing (Java not because programs can redefine what len means, but because the JITter makes assumptions such as “there’s only one implementation of interface Foo” or “the Object passed to this function always is an integer”), but I think both have it easier to detect the points in the code where they need to change their assumptions.
I also guess both have had at least an order of magnitude more development effort poured into them.
I would be ok with a flag that raised an exception/halted if the restricted dynamicism was encountered, if it meant appreciable performance gains.
For every project I've worked on, and every project I've really become familiar with, the "magic" that requires these crazy levels of dynamicism can relatively easily be avoided, with more "standard" interfaces, and possibly a very slight increase in complexity presented to the library user.
I used to play with python magic frequently, but eventually realized my motivation was just some meta code-flex game I was playing, and it's almost certainly never worth it, if other developers are involved.
Given that nearly everything in Python is an Object and can be modified/patched at any time, this dynamic adaptation is probably as best as one can get, without something like numba or cython, which "knows" more about specific blocks of code and can compile them down.
https://docs.python.org/3.11/whatsnew/3.11.html#pep-659-spec...
I remember reading about the GILectomy, and how actually many of the improvements, aimed at making GILectomy feasible, make single-threaded Python runs faster too. There was even the trepidation that PSF might accept these changes, but still say no to GILectomy. Is this at all related?
I'm not sure about the progress on the new form of GIL removal, but as far as I've seen it's a separate effort.
https://github.com/faster-cpython/
https://www.theregister.com/2021/05/13/guido_van_rossum_cpyt...
"A JIT, according to Shannon, will probably not arrive until 3.13 at the earliest, given the amount of lower-hanging fruit that is still to be worked on. The first step towards a JIT, he explained, would be to implement a trace interpreter, which would allow for better testing of concepts and lay the groundwork for future changes."
https://pyfound.blogspot.com/2022/05/the-2022-python-languag...
It's a fork of Python 3.9, takes out GIL and introduces optimisations to speed up both single- and multi-threaded execution (since the bar set by PSF is that no-GIL implementations must be at least as fast as GIL single threaded programs). He ends up with a net 10% speed improvement.
If he does these optimisations, and also doesn't remove the GIL, the performance boost is even larger. So, depending on how you look at it, it's either:
- A bunch of optimisations, plus a GILectomy which slows Python down, or
- A bulk change that removes GIL and speeds things up
Since these improvements were in a similar ballpark, my fear was that the improvements are taken off the branch, with GIL left in place...
Tracing GC does not run into this problem. Why Python doesn't use tracing GC is not something I am qualified to answer.
Sam Gross' work: https://docs.google.com/document/d/18CXhDb1ygxg-YXNBJNzfzZsD...
The GIL code: https://github.com/python/cpython/blob/main/Python/ceval_gil...
Py_INCREF: https://github.com/python/cpython/blob/a4460f2eb8b9db46a9bce...
The best we can hope for is having multiple interpreters[0], each with their own GIL[1].
[0] https://peps.python.org/pep-0554/
[1] https://github.com/ericsnowcurrently/multi-core-python/issue...
Are these projects related? No, except that they are being attempted in the same time frame and may complement each other.
There are tens of thousands of people who wish their python code would run a bit quicker. Many of those stand to earn/save actual money if the code was quicker, so would be happy to pay some of that towards making optimisations.
If that could be pooled together, with some kind of "$100k per 1% speed up on this set of benchmarks" metric, then developers could get properly paid for the work, and everyone would walk away happy.
As it is, everyone wishes it was faster, but realises that they alone can't pay a developer to make a dent, so nobody does it.
The CPython developers have never historically prioritised performance. Simplicity and readability of implementation have typically been higher priorities. In fact that's a key quality of Python in general that development speed is more important than runtime performance.
PHP had a similar performance binge about 6 years ago which is why it tends to run circles around similar class interpreted languages these days
Imagine what could happen if billions of dollars were invested in performance like it was done for C/C++/Java/Javascript in aggregate by various companies.
As more Python is being run in datacenters that starts to make sense, since 1% of CPU usage improvement can mean tens of millions of dollars per year in power costs.
As did IronPython, while Microsoft cared to burn money on it.
From what I remember IronPython had Microsoft hire something like 5 developers for a few years.
In terms of actual compute speed, Python is still significantly slower, and although 3.11's changes will help quite a bit, V8 is also just insane.
In terms of I/O speed, uvloop (https://magic.io/blog/uvloop-blazing-fast-python-networking/) can beat Node quite handily, so if you're more concerned with being able to handle requests than doing anything major during those requests, Python might be comparable.
python-speed v1.3 using python v3.8.5
string/mem: 3032.69 ms
pi calc/math: 6652.63 ms
regex: 2442.87 ms
fibonnaci/stack: 382.66 ms
multiprocess: 11573.28 ms
total: 24084.13 ms (lower is better)
Compare to results in the first comment in this post. Graal looks much slower except for stack manipulation.
pypy does much better:
python-speed v1.3 using python v3.9.12
string/mem: 2397.79 ms
pi calc/math: 2317.79 ms
regex: 1767.99 ms
fibonnaci/stack: 109.36 ms
multiprocess: 640.2 ms
total: 7233.14 ms (lower is better)