Faster CPython Ideas – Issue Tracker for Faster CPython Project
github.com
github.com
- without breaking anyone's code (did that for the print function etc and delayed py3 adoption for a decade but not for the most user requested feature which is speed)
- No large PRs (how a 5x speedup can take place with small patches that no one figured out yet?)
- a small team (not enough money, less risky if they tried to adopt other pypy-esque projects that are more advanced speed wise, a more comprehensive plan that would invite donations, I would certainly give. For this presentation I don't know)
To back up that fact: Python 3.10 will remove long-deprecated features, see issue tracker: https://bugs.python.org/issue41165 and What's new, the Removed section: https://docs.python.org/3.10/whatsnew/3.10.html#removed
The Python 3 debacle is a beautifully worked example of why being backwards compatible is so important, and I can 100% sympathize with them not wanting to go through that again.
Among interpreted languages python is/was one of the faster ones.
NodeJS has similar speed, and in some cases is faster. I don't know any other interpreted language though.
People are also complaining about GIL, but ironically NodeJS is single threaded.
The only production language with performance in the ballpark of Python is Ruby.
I agree that more often than not it simply isn't an issue, but CPython performance isn't much far ahead of toy languages, for a whole lot of reasons that we are all sick and tired of hearing.
So, I take this as a good sign that progress will continue.
It's often the case for end-user applications, but I find that most languages/runtimes tend to get better with time (and work of course !). There's a similar trend with the PHP runtime.
See the perf website https://perf.rust-lang.org/dashboard.html
3 packages will be installed:
expat-2.3.0_1
gdbm-1.19_1
python-2.7.18_3
Size required on disk: 18MB3 packages will be installed:
expat-2.3.0_1
gdbm-1.19_1
python3-3.9.4_1
Size required on disk: 22MBIn a way, it seems like Mark Shannon found someone to take him up on his offer on a plan for speeding up Python.
https://en.wikipedia.org/wiki/Mark_Shannon_(actor)
That mark shannon:
https://github.com/markshannon
He tried to do hotpy in 2011, and trigger the idea of FAT byte code which I believed inspired the FAT python attempt of Victor Stinner.
Now stinner is working on Hpy, dropbox on pyston, instagram on cinder, and pyjion just got a release.
Looks like Python is going to get faster one way or another.
It is always announcements, talks, conferences and if something emerges it is a bit weird like the pattern matching.
Meanwhile the Erlang people quietly produced a JIT without any advertisements.
Because almost everything you've seen as 'working code' started its life as an announcement.
And because to coordinate and discuss future work, there would need to be some announcements.
And because some announcements (based on the persons, e.g. here GvR is involved) or the funding (e.g. here MS is involved) or the specificity (e.g. here 3.10 timeframe is discussed) are more important than others.
>Meanwhile the Erlang people quietly produced a JIT without any advertisements.
Good for them. That's maybe because much fewer care for Erlang (and thus for the advertisements) related to Python (which has a much larger dev base), so the advertisements of the former are posted fewer times and discussed by fewer people.
Not only is it a large user base, but a varied one too. Flask, Django, FastAPI, Twisted...and that is just to name the web frameworks! We have scientific use, research use, cli tools. Perhaps in some cases end users (developers or not) of those tools may not be aware Python is the foundation of said tool.
Anecdotally the Erlang users I've met have been incredibly knowledgeable and in tune with the language features and development. I find that pretty cool.
In my opinion, an argument can be made that Elixir is the most prominent web framework for the language, so developers can just keep up with that and not the language if they wish. Compared to the Python ecosystem. As Erlang inevitably grows in popularity, it too may fragment.
Er, no, they didn’t, they made plenty of announcements before the release, most of which even made their way to HN.
When you like a tech, knowing people are brewing improved perfs on it is kinky. On Python, the motto has always been "it's fast enough", "if you want perfs, don't use python", "C extensions will solve this", "we don't want to make the main implementation complicated", "python dynamism and GIL make it a hard problem", etc.
So in the python world, it's particularly big news, especially given that previous attempts (gilectomy, unladen swallow, first pyston...) all died.
This is precisely the point you seem to be missing. All those things you mentioned were similar big announcements back in their day, and they all have just died by fizzling. What should be setting this one apart?
Also, most of them released working code.
So I guess we shouldn’t be excited about that, either.
I think they could get away with it for a while because none of the competing implementation would readily steal Python’s darlings, namely NumPy/SciPy and Tensorflow. Now there is competition, and it can really push CPython away to being a niche “reference implementation”, thus we are seeing these twitches.
The track record of CPython core team in the last decade makes me very wary when it comes to bringing in such features. Ship it or it didn’t happen.
P. S. Microsoft once threw a large sum of money at Kenneth Reitz to work on Requests, anyone remembers how that story ended?
However, the first release will probably suck like asyncio or type hints did.
It will take time to become usable, but it's unlikely it won't happen.
Namely, that it’s absolutely not guaranteed to work.
Mypy was Jukka Lehtosalo's work, asyncio was fixed and extended by many.
So now he needs a project at Microsoft, he changed his mind (or had it changed under pressure) and this is the project. We'll see, I agree though that the first version will be underwhelming, the following versions will have a max speedup of 50% but will be celebrated at conferences and here.
If that is cause for excitement can be seen when a product is finished.
Edit: Why is this being downvoted? Is my statement incorrect or offensive to some people?
IIRC Firefox 3 was the first browser to have what we would call a modern JS engine, with heavy emphasis on JIT compiling. Google Chrome soon followed and V8 dominated the JS perf story.
Python doesn't have such fierce competition of implementations. And besides, it has the "escape hatch" of being able to implement performance-critical code paths as C extensions, which is what heavy-lifting libraries like numpy do.
Python have many competitions, but the language lack some sort of spec and the only reference is CPython.
In the browser space their has been a lot of competition to make the fastest js runtime because it makes the browser faster.
Node.js uses the V8 engine for JavaScript which isn't just an interpreter - it includes a JIT compiler.
I think when GVR mentions "There's machine code generation in our future" that he is talking about a JIT
Performance has always been a lower-tier priority for cPython and other, faster, implementations (like PyPy) have not seen much mainstream adoption.
Because no FAANG has ever seen CPython performance as a central element of an effort central to the future of the firm, and thrown money at it accordingly.
[1] https://cffi.readthedocs.io/en/latest/index.html [2] https://doc.pypy.org/en/latest/extending.html#cffi [3] https://numba.readthedocs.io/en/stable/reference/pysupported...
This was originally a feature of ctypes but was dropped before stdlib inclusion because it's difficult to solve in the general, "public distribution" case. Nonetheless I don't think there's any need to replace ctypes to reintroduce it, it was done as a library / external tool before and still could be.
https://svn.python.org/projects/ctypes/trunk/ctypeslib/ctype...
CPython got big because of its relatively decent C-API and glue capabilities. 25% does not make a difference, decent C extensions have speedups in the order of 10-100 times (yes, times, not percent).
If I had a dollar for every time someone proposed a Python speedup ...
Microsoft is reactivating their JIT project for Python as well.
Their talk tomorrow:
"Talk: Restarting Pyjion, a general purpose JIT for Python – is it worth it?"
It looks more like Microsoft has JIT envy now that Instagram and others have open sourced (quite restricted) JIT projects and has to show presence again.
I think every CPython JIT project will be full of corner cases in both performance and behavior, so it will add to the growing database of weird Python behavior that users have to memorize.
Usually Python's high dynamism comes into the picture as main reason, however Smalltalk and SELF are just as dynamic, you can change the whole image at any given point in execution.
In Smalltalk it suffices to send a become: message to change the complete meaning of an object, and references to it, across the whole image.
Well, that's just the thing. Many have tried, many have succeeded, but their work hasn't been upstreamed so adoption is almost nonexistent. This is different. It's finally a mainline project. That's noteworthy, and long overdue given the plurality of proofs-of-concept.
So hopefully the combination of those two can make it work.
He always delivers. That's a trademark.
So yes, it's a big deal. It means something will come out of it no matter what, in opposition to everything that happened before.
As mentioned elsewhere in the thread he really has egg on his face after denying Python's performance problems for so many years. I would have more respect if he had taken a firm stance that keeping CPython's interpreter simple had pedagogical value, but instead he (and other core devs) repeatedly argued that Python did not have a performance problem, and that's why it was OK to keep it simple. It took until major pref regressions early in the Python 3 lifecycle to break out of this mindset.
The guy had no insights into language design before creating python.
He also has Mark Shannon, has the entire MS team available on the phone if needed (including the .net team, with stellar JIT), time and a good brain.
> As mentioned elsewhere in the thread he really has egg on his face after denying Python's performance problems for so many years.
He spent 20 years maintaining a project with minimal resources, so you have to make hard choices on what to work with. Perfs were not the objective of Python, and for the majority of the use cases of the time, it didn't matter much.
Now that more important things are sorted out, like unicode handling, and that Python 3 is a smooth running project, he is been given the opportunity and resources to work on the problem, so he does.
Good for us.
Honestly, for me perfs are still NOT a priority, I'd rather him work on bootstraping, but I'm glad and excited that something good will come out of this.
Well, except for working on ABC for years, right?
> He spent 20 years maintaining a project with minimal resources, so you have to make hard choices on what to work with... Now that more important things are sorted out, like unicode handling...
This is ahistorical nonsense in multiple ways.
https://docs.bazel.build/versions/master/skylark/language.ht...
If there's really a 50% general speed boost sitting around in "normal" bytecode optimization work (and having read a decent amount of CPython I believe there could be), that's a criminal amount of energy over the past decade instead wasted dealing with what the type of a codepoint sequence should be named.
https://github.com/faster-cpython/cpython
Looks like they've also got a repo with some tools for profiling bytecode:
(relevant graphic: https://ibb.co/yV6bN9H )
Thanks, Guido and team!
PyPy is already a drop-in replacement for CPython with JIT and using a fraction of the memory.
https://dev.nextthought.com/blog/2018/08/cpython-vs-pypy-mem...