Python 3.13 Gets a JIT
tonybaloney.github.io
tonybaloney.github.io
I highly recommend the blog posts if you're into learning how languages are implemented, by the way. They're incredible deep dives, but he uses the details-element to keep the metaphorical descents into Mariana Trench optional so it doesn't get too overwhelming.
I even had the privilege of congratulating him the 1000th star of the GH repo[3], where he reassured me and others that he's still working on it despite the long pause after the last blog post, and that this mainly has to do with behind-the-scenes rewrites that make no sense to publish in part.
[0] https://arxiv.org/abs/2011.13127
[1] https://sillycross.github.io/2022/11/22/2022-11-22/
[2] https://sillycross.github.io/2023/05/12/2023-05-12/
[3] https://github.com/luajit-remake/luajit-remake/issues/11
Context: I've been on a concatenative language binge recently, and his work on Forth is awesome. In my defense he doesn't seem to list this paper among his publications[0]. Will give this paper a read, thanks for linking it! :)
If they missed the boat on getting credit for their contributions then at least the approach finally starts to catch on I guess?
(I wonder if he got the idea from his work on optimizing Forth somehow?)
There's also this which seems to use the same technique:
Templates-based portable just-in-time compiler, https://dl.acm.org/doi/abs/10.1145/944579.944588
Nice to see there's still room for innovation in the VM space!
[1] https://www.usenix.org/legacy/event/usenix05/tech/freenix/fu...
[2] https://review.gerrithub.io/plugins/gitiles/spdk/qemu/+/5a24...
In fact I wouldn't be surprised if the earliest compilers were template based, as that's about the only implementation that would fit in a RAM as a compiler pass.
Regardless of the work being done in PyPy, Jython, GraalPy and IronPython, having a JIT in CPython seems to be the only way beyond "C/C++/Fortran libs are Python" mindset.
Looking forward to its evolution, from 3.13 onwards.
To me, Mojo looks like the best approach to fusing that with the Python ecosystem! (I have no doubt about it being open sourced at some point.)
https://www.intel.com/content/www/us/en/developer/articles/t...
I think people just reach for JIT more often in dynamic languages because they carry more information around and have more of a performance deficit that they want to mitigate.
I dunno. I mostly program in Fortran but JIT seems way cool. Fundamentally I don’t see why a JITer couldn’t beat my code in a very dynamic language, it would just need to find a big loop that calls a kernel a bunch of times, where the exact computation in the kernel is determined at run-time, and then jit the kernel and the loop together.
You should definitely look into Julia. It’s a beautiful language, and squarely aimed at the Fortran space. It also relies heavily on JITC and GC. That’s fine for purely scientific computing, but not so good as a general purpose language.
That’s where C, C++, Rust, and Mojo are the current major contenders, IMO.
One reason is performance. So if Python has a faster future ahead of it: Hurray!
The other reason is that the Python ecosystem moved away from stateless requests like CGI or mod_php use and now is completely set on long running processes.
Does this still mean you have to restart your local web application after any change you made to it? I heard that some developers automate that, so that everytime they save a file, the web application is restarted. That seems pretty expensive in terms of resource consumption. And complex as you would have to run some kind of watcher process which handles watching your files and restarting the application?
It's also very easy, often just adding a CLI flag to your local run command.
edit: Regarding performance, Python today can easily handle at least 1k requests per second. The vast vast vast majority of web applications today don't need anywhere near that kind of performance.
I prefer to have a local system set up just like the production server, but in a container.
Maybe using WSGI with MaxConnectionsPerChild=1 could be a solution? But that would start a new (for example) Django instance for every request. Not sure how fast Django starts.
Another option might be to send a HUP signal to Apache:
apachectl -k restart
That will only kill the worker threats. And when there are none (because another file save triggered it already), this operation might be almost free in terms of resource usage. This also would require WSGI or similar. Not sure if that is the standard approach for Django+Apache.In production, you would want to run your app through gunicorn/uvicorn/whatever on an internal-only port, and reverse-proxy to it with a public-facing apache or similar.
Set up apache to reverse proxy like you would on prod, and run gunicorn/uvicorn l/whatever like you would on prod, except you also add the autoreload flag. E.g.
uvicorn main:app --host 0.0.0.0 --port 12345 --reload
If production uses containers, you should keep the python image slim and simple, including only gunicorn/uvicorn and have the reverse proxy in another container. Etc.With those two you can just stand up an python program in a container that serves html, and put it behind whatever reverse proxy you want.
But even leaving that aside, you never know when your application will be linked somewhere or go semi-viral and not being able to serve 1000 users is all it takes for your app to go down and your one shot at a successful company to die a sad death.
The specifics of that aside, any unprepared application is going to buckle at a sudden mega-surge of users. The solution remains largely the same, regardless of technology: Make sure everything that can be cached is cached, scale the hardware vertically until it stops helping, optimize your code, scale horizontally until you run out of money. I imagine the DB will be the actual bottleneck, most of the time.
There are other reasons to not choose python for greenfield application, but performance should rarely be one IMO.
You wouldn’t want to have a long, complex call path through that code, but just parsing the body and adding it to a Celery queue before returning a 201 Created was perfectly manageable.
Once you deploy it to production, you usually run it using a WSGI/ASGI server such as Gunicorn or Uvicorn and let whatever deployment process you use handles the lifecycle. You usually don't use watcher in production.
Basically similar stuff with nodejs, rails, etc.
In prod, you don't do it. Deployment implies sending a signal like HUP to your app, so that it reloads the code gracefully.
All in all, everybody is moving to thid, even php. This allows for persitent connexion, function memoization, delegation to threadpools, etc
What? No, in reality it’s just running your app in debug mode (just a cli flag), and when you save the files the next refresh of the browser has the live version of the app. It’s neither expensive nor complex.
Definitely take a look, it's come a long way from ten years ago.
I started with cherrypy, then used twisted for years and finally aiohttp made me never look back or search for anything else.
All of the popular frameworks automatically reload. It’s not instantaneous but with e.g. Django it was less than the time I needed to switch windows a decade ago and it hadn’t gotten worse. If you’re used to things like NextJS it will likely be noticeably faster.
Think server.py and server_handlers.py, where server.py contains logic to detect a modification of server_handlers.py (like via inotify) and the base handlers which then call the "modifiable" handlers in server_handlers.py.
This is not limited to servers (anything that loops or reacts to events) and can be nested multiple levels deep and is among the top 3 reasons of why i use Python.
Reloading is instantaneous and can gracefully handle errors in the file (just print an err messge or stack trace and keep running the old code)
The long-running process is a WSGI/ASGI process that handles spawning the actual code, similar to CGI. The benefit is that it can handle how it spawns the request workers via multiple runtimes, process/threads, etc. It's similar to CGI but instead of nginx handling it, it's a special program that specializes in the different options for python specifically.
> Does this still mean you have to restart your local web application after any change you made to it? I heard that some developers automate that, so that everytime they save a file, the web application is restarted. That seems pretty expensive in terms of resource consumption. And complex as you would have to run some kind of watcher process which handles watching your files and restarting the application?
Only for development!
To update your code in production you first deploy the new code onto the machine, and then you tell the WSGI/ASGI such as Gunicorn to reload. This will cause it to use the new code for new request, without killing current requests.
It's a graceful reload, with no file watching needed. Just a "systemctl reload gunicorn"
As opposed to CGI?
x ln 0.9 = ln 0.5
x = ln 0.5 / ln 0.9
x = 6.5788
So decreasing runtime by 10% 6.5788 times results in the code running in half the original time.
Ref: https://mail.python.org/pipermail/python-dev/2016-November/1...
If you've rewritten something to better use cachelines, removed saturating memory bandwidth, etc then sure you've increased The computation rate. But that's rarely how these language specific optimizations occur.
Even in other contexts it can be ambiguous.
Yesterday I drove 60mph, today I drove 50% faster.
Yesterday I got there in 1 hour, today I got there 50% faster.
It's not really possible to tell them apart without looking at the numbers.
21%, not 19%.
It is 1.1 * 1.1 = 1.21
You are right in the opposite direction. If it got 10% slower, then it is 0.9 * 0.9 = 0.81 = 19% slower
When you say something is 10% faster, what you mean is it took 10% less time to finish. So 19% is correct.
Unless one uses a bit more esoteric definition of speed, in which a 50 percent of car speed increase, makes it go from 100 km/h to 200 km/h, such that it arrives in half the time.
Interestingly they singled out pyaes as one of the worst offenders. I've also written a pure-python AES implementation, one that deliberately takes advantage of the "long" integer representation, and it beats pyaes by about 2000%.
Where is the evidence that before python 3 was significantly slower than 2.7 before 3.6?
[1]: https://en.wikipedia.org/wiki/Benevolent_dictator_for_life
the irony :/
Test your packages on Pypy, people.
I guess that now the GIL is going away, pypy will become better at handling packages with native code like numpy
Most c extension modules should work in pypy, there's just a performance hit depending on how they're built (cffi is the most compatible).
https://doc.pypy.org/en/latest/faq.html#do-c-extension-modul...
People wanting to use Pypy usually do so because they want better performance. Having a performance hit while using pypy is disconcerting.
I was speculating that in the future, C extensions in pypy would be faster, but I now see that the GIL is actually unrelated to this performance hit. Anyway it's really a pity.
Please do not phrase that as a failure of Pypy. That is so weird.
Now, a lot of C packages work - and where they don't it's worth raising bugs: with PyPy, but also in the downstream program - occasionally they can use something else if it looks like the fix will take a while.
That said, it has happened often enough I'm cautious about where I use it. It would suck to be dependent on pypy's excellent performance, and then find I can't do something due to library incompatibility.
It's goal is to create as many 'return' instructions as it can decide to.
GIL is merely a CPython problem but synchronization can also be a compilation problem.
FWIW, the most recent changelog is at https://docs.python.org/3.13/whatsnew/3.13.html
I spent about one week implementing PyPy's storage strategies in my language's collection types. When I finished the vector type modifications, I benchmarked it and saw the ~10% speed up claimed in the paper¹. The catch is performance increased only for unusually large vectors, like thousands of elements. Small vectors were actually slowed down by about the same amount. For some reason I decided to press on and implement it on my hash table type too which is used everywhere. That slowed the entire interpreter down by nearly 20%. The branch is still sitting there, unmerged.
I can't imagine how difficult it must have been for these guys to write a compiler and succeed at speeding up the Python interpreter.
¹ https://tratt.net/laurie/research/pubs/html/bolz_diekmann_tr...
No one is disappointed by V8's 6-8% improvement with Maglev. [1]
Because V8 is (for a scripting language) insanely fast.
And Python is not, unfortunately.
In many cases, Node.js is an order of magnitude faster than CPython.
(Acknowledged: You could write your Python in C.)
(Acknowledged: PyPy exists.)
This resulted in a situation where the ecosystem is locked-in to those implementation details: CPython can't change many aspects of its own implementation without breaking the ecosystem; and other implementations are forced to introduce complex and slow emulation layers if they want to be compatible with existing CPython extension modules.
The end result is that alternative implementations are not viable in practice, as most existing libraries don't work without their CPython extension modules -- users of alternative implementations are essentially stuck in their own tiny ecosystem and cannot make use of the large existing (C)Python ecosystem.
CPython at least is in a position where they can push a breaking change to the extension API and most libraries will be forced to adapt. But there's very little incentive for library authors to add separate code paths for other Python implementations, so I don't think other implementations can become viable until CPython cleans up their API.
Jython was released 22 years ago
IronPython was released 17 years ago
To date, no Python implementation has managed to hit all three:
1. Stay compatible with any recent, modern CPython version
2. Maintain performance for general-purpose usage (it's fast enough without a warmup, and doesn't need to be heavily parallelized to see a performance benefit)
3. Stayed alive
Which, frankly, is kind of a shame. But the truth of the matter is that it was a high bar to hit in the first place, and even PyPy (which arguably had the biggest advantages: interest, mindshare, compatibility, meaningful wins) managed to barely crack a fraction of a percent of Python market share.
If you bet on other implementations being the source of performance wins, you're betting on something which essentially doesn't exist at this point.
PyPy seems pretty alive, all things considered, and for my code bases I've seen pretty dramatic speedups on the order of 2-5x. That's basically a no brainer unless I'm doing something with incompatible C extensions, which I think is the real Achilles heel of all of these alternative implementations.
It is encouraging for PyPy to see some influx of money in recent years. But I will continue to patiently wait for it to hit enough of a sweet spot of performance vs usability vs compatibility to see real adoption.
For larger programs like you sometimes it some incredibly complicated incompatibility problem. For me bitbake was one of those - could REALLY benefit from pypy but didn't work properly and I couldn't fix it.
If this works more reliably or has a faster warmup then....well it could help to fill in some gaps.
interpreted -> basic JIT -> fancy JIT
The interpreter gets you going fast. The basic JIT is extremely fast to compile but not the most performant. If the code can be JITed it quickly will be.
From there the engine can find hotspots or functions that get run a lot and use the fancy JIT on them in the background. That means the slow compile doesn’t block things but when the result can be swapped in performance can take a big jump.
At any point the engine can drop down to the interpreter if an assumption is violated (someone passes a string where they had always used numbers before) or a function is redefined.
It wouldn’t surprise me if something like that appeared as an option in Python over time to get the best of both worlds.
https://webkit.org/blog/10308/speculation-in-javascriptcore/
If the further optimizations that this change allows, as explained at the end of this post, are covered as well as this one, it promises to be a very interesting series of blog posts.
Pythonesque (sane) syntax, great AOTC language features, great performance, memory safety, excellent Python interoperability. What’s not to like?
If you think of something, contact Modular…
Maybe not absolute best practice, but it's not like PHP or Ruby or Perl, where there's rarely ever any new projects everyone blogs about using them.
All the big web frameworks are still maintained, there are still new coders learning it, it's not a language that will make people be like "Oh ew, I'm not learning that language just to work on that", etc.
> A copy-and-patch JIT only requires the LLVM JIT tools be installed on the machine where CPython is compiled from source, and for most people that means the machines of the CI that builds and packages CPython
But I doubt that's going to ever happen.
They're trying to pass data between layers of middleware, but Java has very strict typing, and the middleware doesn't know what kind of object it will get, so it has to do tons of type introspection and reflection to do anything with the data?
There's https://pypi.org/project/jsonschema-typed-v2/ but it hasn't been updated in a few years.
https://mypy.readthedocs.io/en/stable/command_line.html#cmdo...
Short of Google implementing something if Chrome hits 85-90% of all use in an attempt to dump JS it just doesn’t seem like something that would happen. I doubt any browser team would want to implement multiple languages. I doubt Google would want to switch.
Plus Python is not suited for event driven systems.
There are enough people that REALLY hate whitespace-as-syntax.
Granted, the code I see from more web and ops focused teams is miles better. But I worry the collective is not where it should be.
I once wrote an article about very simple JITs, and the first example in my article uses this style: https://blog.reverberate.org/2012/12/hello-jit-world-joy-of-...
I take some issue with this statement, made later in the article, about the pros/cons vs a "full" JIT:
> The big downside with a “full” JIT is that the process of compiling once into IL and then again into machine code is slow. Not only is it slow, but it is memory intensive.
I used to think this was true also, because my main exposure to JITs was the JVM, which is indeed memory-intensive and slow.
But then in 2013, a miraculous thing happened. LuaJIT 2.0 was released, and it was incredibly fast to JIT compile.
LuaJIT is undoubtedly a "full" JIT compiler. It uses SSA form and performs many optimizations (https://github.com/tarantool/tarantool/wiki/LuaJIT-Optimizat...). And yet feels no more heavyweight than an interpreter when you run it. It does not have any noticeable warm up time, unlike the JVM.
Ever since then, I've rejected the idea that JIT compilers have to be slow and heavyweight.
Yes, and it's practically unmaintained. Pull requests to add support for various architectures have remained largely unanswered, including RISC-V.
Rather than focussing on the raw number compare to python 3.5 or so. It's still getting significantly faster.
If they keep doing this steady pace they are slowly saving the planet!
Amdahl's Law is about expected speedup/decrease in latency. That actually isn't strongly correlated to "saving the planet" afaik (where I interpret that as reducing direct energy usage, as well as embodied energy usage by reducing the need to upgrade hardware).
If anything, increasing speed and/or decreasing latency of the whole system often involves adding some form of parallelism, which brings extra overhead and requires extra hardware. Note that prefetching/speculative execution kind of counts here as well, since that is essentially doing potentially wasted work in parallel. In the past boosting the clock rate the CPU was also a thing until thermodynamics said no.
OTOH, letting your CPU go to sleep faster should save energy, so repeated single-digit perf improvements via wasting less instructions does matter.
But then again, that could lead to Jevons Paradox (the situation where increasing the efficiency encourages more wasteful than the increase in efficiency saves - Wirth's Law but generalized and older, basically).
So I'd say there's too many interconnected dynamics at play to really simply state "optimization good" or "optimization useless". I'm erring on the side of "faster Python probably good".
Also, such a shame that it takes sooo long for crucial open source to be funded properly. Kudos to Microsoft for doing it, shame on everyone else for not pitching in sooner.
FYI Python was launched 32 years ago, Python 2 was released 24 years ago and Python 3 was released 16 years ago.
Microsoft hired Guido in late 2020 giving him freedom to choose what project he wanted. Guido decided to go back to core Python development and with approval of Microsoft created a "faster-cpython" project, at this point that project has hired several developers including some core CPython developers. This is all at the discretion of Microsoft, and is not some arms length funding arrangement.
Meta has a somewhat similar situation, they hired Sam Gross (not the cartoonist) to work on a Python non-gil project, and contribute it directly to CPython if they accept it (which they have), and they have publicly committed to support it, which if I remember right was something like funding two engineering years of an experienced CPython internals developer.
While very very very popular, Python is i think is very disliked languages, it doesnt have or it is not built around the current programming language features that programmers like, its not functional or immutable by default, its not fast, the tooling is complex, it uses indentation for code blocks (this feature was cool in the 90s, but dreaded since at least 2010)
so i guess if python become fasters, this will ensure its continued dominance, and all those hoping that one day it will be replace by a nicer , faster language are disappointed
this pessimism is the aching voice of the developers who were hoping for a big python replacement
but i also think that its true that python is not and have not been for a while considered as a modern or technically advanced language
the hype currently is for typed or gradually typed languages, functional languages, immutable data , system languages, type safe language, language with advanced parallelism and concurrency support etc ..
python is old , boring OOP, if you like it, than like millions of developers you are not picky about programming language, you use what works, what pays
but for devs passionate about programming languages, python is a relic they hope vanish
Statements like this are obviously untrue for large numbers of people, so I'm not sure of the point you're trying to make.
But certainly it's true that there are both objective and subjective reasons for using a particular tool, so I hope you are in a position to use the tools that you prefer the most. Have a great day!
So Python with mypy
If you asked me what language I would consider to be a relic that I hope would vanish, I'd go with Perl.
It is still the only beginner language that is also an industrial-strength production language. You can learn Python as your first language and also make an entire career out of it. That can't really be said about the currently "hyped" languages, even though those are very fun and cool and interesting!
Such devs are increasingly rare and, in some domains, almost nonexistent. For example, Kotlin borrowed a lot of nice features from Scala and Groovy, yet 99% of Kotlin code I've seen professionally never touched those features. Kotlin on Android seems to be overwhelmingly written by barely (or not at all) re-trained Java devs; moreover, those who never learned anything more recent than 1.8 (at least they know what lambdas are.)
In short, it's not about the language; it's about the people who use that language. Wishing a language to vanish is misguided - it wouldn't change anything. The people would just switch to the next language and would still program in the same style. You can write Fortran in every language - this is as true today as it was back in the 70s, but the percentage of people who can't be bothered to stop writing Fortran (metaphorically, in more literal meaning their first language, whatever it was) even after changing to another language got much higher. IMO due to the changes in how programming as a trade is perceived in society... but that's perhaps a rant for another time :)
LOL this is a dead giveaway you haven't been around long. There have been people kvetching about the whitespace since the beginning. Haskell went on to be the next big thing for reddit/HN/etc for years and it also uses whitespace.
https://survey.stackoverflow.co/2023/#section-admired-and-de...
If you wanted to rebut this, you'd need to argue that Julia has always been awesome and that my experience with a slow warmup was atypical. But that would be a lie, right?
And, subtext: when I wrote my first commebt in this thread, its highest sibling led with
> I think the pessimism really comes from a dislike for Python
So I weighed in as a Python lover who is pessimistic for reasons other than a bias against the language.
But your assessment of the other language you mentioned is several years out of date and made largely irrelevant by the fast pace of progress. Therefore your conclusions about the probable future of Python, which may be correct, nevertheless do not follow.
How long did it take Julia to solve its warmup issue? The language is about 12, and I last tried in earnest two years ago. So, a decade, give or take? You speak from the top of a mountain, and you say the view is nice. Sitting at the base of a similar mountain, it's the journey that I dread, because Python's recent long-term journeys have been pretty rough. And I'm just not convinced that the destination is so great.
It is an approach that traces back to original Lisp and BASIC systems, among others lesser kwown ones.
The compiler is part of the language runtime, and code gets dynamically compiled into native code.
Why is this a good approach?
It allows for experiences that are much harder to implement in languages that tradicionally compile straight to native code like C (note there are C interpreters).
So you can have an interpreter like experience, and code gets compiled to native code before execution on the REPL, either straight away, or after the execution gets beyond a specific threshold.
Additionally, since dynamic languages per definition can change all the time, a JIT can profit from code instrumentation, and generate machine code that takes into account the types actually being used, something that an AOT approach for a dynamic language cannot predit, thus optimizations are hardly an option in most cases.
If you're going to break backwards compatibility, it's not like Unicode was the only foundational problem Python 2 had.
It wasn't realistic to switch to 3.x when the libraries either weren't there or were a lot slower (due to using pure Python instead of C code).
It also wasn't realistic to rewrite the libraries when the users weren't there.
It was in many respects a perfect case study in how not to do version upgrades.
WTF has this to do with JITing the code written in Python?
It's only fairly recently that there's been critical mass of people who thought that performance trumps simplicity, and even then, it's only to a point.
This definitely wasn't true, from the user perspective. And, I'm not even convinced it's some "critical mass" of developers. These changes aren't coming from some mass of developers, there's coming from a few experts that had a clear plan, backed by the sanity of the huge disconnect that languages are actually meant for users of the language, not the developers of the language.
That made it much easier to get bigger gains than if they wanted 100% backwards compatibility with CPython.
Context: I use python for data processing and webdev. When doing data processing, Python is merely glue for libraries in compiled languages. When doing webdev, I mostly use python itself.
First, any numbers regarding benchmarks need to be treated with contempt. JITs are unbenchmarkable. No matter what you do, someone says you do it wrong. Warmed up the JIT? You did it wrong. Didn't warm up the JIT? You did it wrong. Warmed up and didn't warm up the JIT? Wrong.
You lose predictable performance characteristics due to the above. It's difficult to describe the importance of this to the people who look at Python as glue for their compiled code.
Next, on the face of it, it doesn't look like it will compose well with subinterpreters. If each subinterpreter does it's own tracing, it's going to be harder to hit the 10k watermark of jitting hot code.
This uses LLVM's JIT which is particularly slow and heavy (16MB added to the binary size) last time I tried to use it. So this limits Python attractiveness in being an embedded language.
While this remains experimental - and hence strictly optional, packaging this in distributions that use gcc would now apparently need llvm tools installed to build python. Expect feedback from distribution packagers.
Idea for improvement: when I do data processing, I'm using python to glue bits of C, Rust, and other compiled languages so this is not very useful. When I'm doing webdev, it could be useful - and I am deploying using Docker. So why not make this AOT so I can add it to a docker build step and get all the benefit without the complications of tracing jits.
If those aren't solid proving grounds for JITs, nothing is.
Print Hello, world.
That's it. Time it.
Do the same with Python.
The difference is minuscule.
People don't build CLIs with them due to cultural reasons (fashion). Developers are creatures of fashion.
java: 031s python: 013s
it's 3x slower to start up… it means that any bash script with a loop calling a java command will run 3x slower.
3x slower isn't my definition of minuscule.
I'm confused, what kind of time measurement system is this? Are those 31 seconds versus 13 seconds? 0.31s vs 0.13s? If it's the first case, something is wrong with that machine, Hello World should not take that long. If it's the second, are we talking about 0.18s? See below.
> it's 3x slower to start up… it means that any bash script with a loop calling a java command will run 3x slower.
I don't know how you write your scripts, but in practice most scripts hang waiting for IO (especially network) or waiting for a specific command to process something...
> 3x slower isn't my definition of minuscule.
Nope, but based on what you've presented so far, unless I'm misunderstanding, it's my definition of "premature optimization".
> I don't know how you write your scripts,
You know how… I put print hello world in them… that was the test you asked for?
> premature optimization
Not using something that has 3x startup time, for a short lived command is just "common sense".
I'm starting to think that the issue here is that you know java but don't know python or C, and you are unconsciously trying to get water to your own watermill.
It's just that I recognize disingenuous comments.
Java is perfectly adequate as a language for writing command line tools. You're likely using one of them weekly and don't even know it.
Just don't write your Java CLIs using Spring or DI frameworks in general :-)
doubtful, since I don't have jre installed. I had to install it just to run that benchmark.
There is a reason I didn't say anything about that aspect of JITs that you mention, it's because it's not relevant to commands which are launched from shell scripts and then exit.
JIT caches, AOT have been an option for years.
And if you are really motivated, you can write stuff in C just like in Python.
You're just sticking with what you know, calling out the others for not wanting to learn.
Not a nice attitude from where I'm at.
Or you can use PyPy of course, but that's a compatibility nightmare exactly when you're gluing together a bunch of C libraries with Python interfaces.
And in principle, JIT doesn't necessarily mean non-embeddable; LuaJIT is pretty good from an embedding perspective. Besides, I would assume they make JIT a build-time option so that people whose use case makes it problematic can use only the bytecode interpreter instead.
That said, your concerns about performance variability are warranted, and there are probably specific criticisms to be made about the choice of LLVM. LLVM is a gigantic dependency and I hope the performance wins of choosing LLVM rather than, say, crankshaft are worth it.
Why does that matter?
Presumably if you care about Python performance, you have a real world scenario that you can benchmark. Just make a benchmark that is as representative as possible and check if the JIT helps. If it doesn't help, it surely can be disabled via an option.
Because you need to have an idea of how many machines you need. And you need an idea of whether 'it will take as long as it takes' or if there's an issue in your code.
Sometimes Java decides to JIT using aes instructions. Sometimes it doesn't. Even if you run the same benchmark suite twice in a row it just does its own thing and gives wildly different results.
My learning was that one shouldn't depend on JIT performance and prefer using JNI libraries if you want consistency. Hence my comment being negative on JIT.
Hopefully this is a feature that can be disabled when you want deterministic behaviour. There's no reason to make it mandatory in a well-engineered VM.
Incorrect, LLVM is only used in the compile time. See my other comments for details.
I've also thought this. If you're pre-baking your software, why not do this step as well?
Sure when my website/app/software gets 10k concurrent users, I might regret it, but I havent regretted it yet.
Isn't this the main reason why it's only a 2-9% improvement? Not much Python code uses the while statement in my experience.
Maybe Microsoft don't know yet how to sell this thing, or maybe they are just boiling the frog. Time will tell. But I'm pretty sure your question will be repeated as soon as people will get used to the idea of Python on JIT.
Microsoft developed both JScript and Node.js. They could've continued with JScript, but obviously decided against it because JScript didn't earn the reputation they might have hoped for. Even if they invested efforts into rectifying the flaws of JScript, it would've been just too hard to undo the reputation damage.
Microsoft made multiple attempts to "befriend" Python. IronPython was one of the failures. They also tried to provide editing tools (eg. intellisense in MSVS), but kind of given up on that too (but succeeded to a large degree with VSCode).
The whole long-term Microsoft's strategy is to capture and put the developers on a leash. They won't rest until there's a popular language they don't control.
Just a guess at the pitch.
It beats Python on performance, supposedly, but compatibility has never been great.
This JIT approach is improving the performance of bits of the interpreter while maintaining 100% compatibility with the rest of the C code base, its object model, and all the extensions.
> The big downside with a “full” JIT is that the process of compiling once into IL and then again into machine code is slow. Not only is it slow, but it is memory intensive.
he talks about an IL, but what's that IL? does that mean that the future optimization will involve that IL?
shouldn't this be "python 3.13 gets a new jit compiler" because python already has a jit.
Is there an similarly accessible article about the specializing adaptive interpreter? It's mentioned in this article but not much detail is given, only that the JIT builds upon it.
I wonder if I can skip the bytecode compilation phase.
> The initial benchmarks show something of a 2-9% performance improvement.
> I think that whilst the first version of this JIT isn’t going to seriously dent any benchmarks (yet), it opens the door to some huge optimizations and not just ones that benefit the toy benchmark programs in the standard benchmark suite.
But even if you can get a 2x improvement from lots of 1% improvements (if you work really really hard), you're never going to get a 10x improvement.
Rust is never going to compile remotely as quickly as Go.
Python is never going to be remotely as fast as Rust, C++, Go, Java, C#, Dart, etc.
Trains are never going to beat jets in pure speed. But in certain scenarios, trains make a lot more sense to use than jets, and in those scenarios, it is usually preferable having a 150 mph train to a 75 mph train.
Looking at the world of railways, high-speed rail has attracted a lot more paying customers than legacy railways, even though it doesn't even try to achieve flight-like speeds.
Same with programming languages, I guess.
Two decades ago, you could (as e.g. Paul Graham did at the time) argue that dynamically typed languages can get your ideas to market faster so you become viable and figure out optimization later.
It's been a long time since that argument held. Almost every dynamic programming language still under active development is adding some form of gradual typing because the maintainability benefits alone are clearly recognized, though such languages still struggle to optimize well. Now there are several statically typed languages to choose from that get those maintainability benefits up-front and optimize very well.
Different languages can still be a better fit for different projects, e.g. Rust, Go, and Swift are all statically typed compiled languages better fit for different purposes, but in your analogy they're all jets designed for different tactical roles, none of them are "trains" of any speed.
Analogies about how different programming languages are like different vehicles or power tools or etc go way back and have their place, but they have to recognize that sometimes one design approach largely supersedes another for practical purposes. Maybe the analogy would be clearer comparing jets and trains which each have their place, to horse-drawn carriages which still exist but are virtually never chosen for their functional benefits.
In many domains, it doesn't really matter if the resulting program runs in 0.01 seconds or 0.1 seconds, because the dominant time cost will be in user input, DB connection etc. anyway. But it matters if you can crank out your basic model in a week vs. two.
I don't doubt it, but learning is only the first step to using a technology for a series of projects over years or even decades, and that step doesn't last that long.
People report being able to pick up Rust in a few weeks and being very productive. I was one of them, if you already got over the hill that was C++ then it sounds like you would be too. The point is that you and your team stay that productive as the project gets larger, because you can all enforce invariants for yourselves rather than have to carry their cognitive load and make up the extra slack with more testing that would be redundant with types.
Outside of maybe a 3 month internship, when is it worthwhile to penalize years of software maintenance to save a few weeks of once-off up-front learning? And it's not like you save it completely, writing correct Python still takes some learning too, e.g. beginners easily get confused about when mutable data structures are silently being shared and thus modified when they don't expect it. People who are already very comfortable with Python forget this part of their own learning curve, just like people very comfortable with Rust forget their first borrow check header scratcher.
I never made a performance argument in this thread so I'm not sure why 0.01 or 0.1 seconds matters here. Even the software that got you into a commercial market has to be maintained once you get there. Ask Meta how they feel about the PHP they're stuck with, for example.
What is this supposed to say? Most scripting language interpreters are written in low level languages (or assembly), but that alone doesn't say anything about the performance of the language itself.
So python programs, that already spend most of its cpu time running these libraries code, won't see much of an impact.
Granted we're still very far from that and probably won't ever reach it, but there definitely seems to be a lot of progress.
Or maybe it's just syntactic sugar around that. But sugar can be nice.
> We're getting a JIT. Now it's time to optimize the traces to pass them to the JIT.
That what makes Python flexible is what makes it slow. Restricting the flexibility were possible offers opportunities to improve performance (and allows for tools and humans to spot errors more easily).
def twice(x: int) -> int:
return x + x
print(twice("nope"))
and it should print "nopenope". Right? def twice(x: int) -> int:
if not isinstance(x, int):
raise TypeError("Expected x to be an int, got " + str(type(x)))
return x + xBesides, my point was that one of the reasons why languages with (sound-ish) static types manage to have better performance because they can omit all of those run-time type checks (and the supporting machinery) because they'd never fail. And if you have to put those explicit checks, then the type hints are actually entirely redundant: e.g. Erlang's JIT ignores type specs, it instead looks at the type guards in the code to generate specialized code for the function bodies.
That's for this benchmark:
https://pyperformance.readthedocs.io/
Note that this is with a relatively small investment as these things go, the GraalPython team is about ~3 people I guess, looking at the GH repo. It's an independent implementation so most of the work went into being compatible with Python including native extensions (the hard part).
But this speedup depends a lot on what you're doing. Some types of code can go much faster. Others will be slower even than CPython, for example if you want to sandbox the native code extensions.
- All integers are still big integers
- Use of the typing opt-out 'Any' is very common
- All functions/methods can still be overwritten at runtime
- Fields can still be added and removed from objects at runtime
The combination basically makes it mandatory to not use native arithmetic, allocate everything on the heap, and need multiple levels of indirection for looking up any variable/field/function. CPU perf nightmare. You need a real optimizing JIT to track when integers are in a narrow range and things aren't getting redefined at runtime.
With a compiler, that part is done once and, potentially, run zillions of times.
In fact, for such a fused instruction to be optimized that way on a copy-and-patch JIT it'd need to exist as a new bytecode in interpreter. A JIT that fuses instructions is no longer a copy-and-patch JIT.
A copy-and-patch JIT reduces interpretation overhead by making sure the branches in the executed machine code are the branches in the code to be interpreted, not branches in the interpreter.
This is make a huge difference in more naive interpreters, not so much in an heavily optimized threaded-code interpreter.
The 10% is great, and nothing to sneeze at for a first commit. But I'd actually like some realistic analysis of next steps for improvement, because I'm skeptical instruction fusing and other things being hand waved are it. Certainly not on a copy-and-patch JIT.
For context: I spent significant effort trying to add such instruction fusing to a simple WASM AOT compiler and got nowhere (the equivalent of constant loading was precisely one of the pairs). Only moving to a much smarter JIT (capable of looking at whole basic blocks of instructions) started making a difference.
One approach used in V8 is to have a dumb-but-very-fast JIT (ie. this), and keep counters of how often each block of code runs (perhaps actual counters, perhaps using CPU sampling features), and then any block of code running more than a few thousand times run through a far more complex yet slower optimizing jit.
That has the benefit that the 0.2% of your code which uses 95% of the runtime is the only part that has to undergo the expensive optimization passes.
V8 pre-2021 (i.e., only Ignition+TurboFan) was significantly faster than current CPython is, and the full current four-tier bundle (Ignition+Sparkplug+Maglev+TurboFan) only scores roughly twice as good on Speedometer as pure Ignition does. (Ignition+Sparkplug is about 40% faster than Ignition alone; compare that “dumbness” with CPython's 2–9%.) The relevant lesson should be that things like very carefully designed value representation and IR is a much more important piece of the puzzle than having as many tiers of compilation as possible.
That's just all JITs. Sometimes its counters for going from interpreter -> JIT rather than levels of JITs, but this idea is as old as JITs.
I recommend that people watch Brandt Bucher's "A JIT Compiler for CPython" from last year's CPython Core Developer Sprint[0]. It gives a good impression of the current implementation and its limitations, and some hints at what may or may not work out. It also indirectly gives a glimpse into the process of getting this into Python through the exchanges during the Q&A discussion.
One thing to especially highlight is that this copy-and-patch has a much, much lower implementation complexity for the maintainers, as a lot of the heavy lifting is offloaded to LLVM.
Case in point: as of the talk this was all just Brandt Bucher's work. The implementation at the time was ~700 lines of "complex" Python, ~100 lines of "complex" C, plus of course the LLVM dependency. This produces ~3000 lines of "simple" generated C, requires an additional ~300 lines of "simple" hand-written C to come together, and no further dependencies (so no LLVM necessary to run the JIT. Also "complex" and "simple" qualifiers are Bucher's terms, not mine).
Another thing to note is that these initial performance improvements are just from getting this first version of the copy-and-patch JIT to work at all, without really doing any further fine-tuning or optimization.
This may have changed a bit in the months since, but the situation is probably still comparable.
So if one person can get this up and running in a few klocs, most of which are generated, I think it's reasonable to have good hopes for its future.
In those custom SDKs, I do generate all the code at the start of the build, which takes a significant amount of time for mostly non-pertinent anymore/inappropiately done code generation.. I will really feel python3 speed improvement for those builds.
Python also has structural challenges like native extensions (don’t exist in JavaScript) where the API forces slow code or massive hacks like avoiding the C API at all costs (if I recall correctly I read that’s being worked on) and the GIL.
One advantage Python had is the ability to use multiple cores way before JS but the JS ecosystem remained single threaded longer & decided to use message passing instead to build WebWorkers which let the JIT remain fast.
The package management and build tools for python have been so atrociously bad (environments add far too much complexity to the ecosystem) that it turns many developers away from the language altogether. A system like Rust's package management, build tools, and cross compilation capability is an enormous draw, even without the memory safety. The fact that it actually works (because of the package management and build tools) is the main reason to use the language really. Python used to do that ~10 years ago. Now absolutely nothing works. It takes weeks to get simple packages working, only can do anything under extremely brittle conditions that nullify the project you're trying to use this other package for, etc.
If python could ever get it's act together and make better package management, and allow for cross-compiling, it could make a big difference. (I am aware of the very basic fact that it's interpreted rather than compiled yada yada - there are still ways to make executables, they are just awful). Since python is data science centric, it would be good to have decent data management capabilities too, but perhaps that could be after fundamental problem are dealt with.
I tried looking at mojo, but it's not open source, so I'm quite certain that kills any hope of it ever being useful at all to anyone. The fact that I couldn't even install it without making an account made me run away as fast as possible.
Can you expand on what you mean by that? I have trouble imagining a Python packaging problem that takes weeks to resolve - I'd expect them to either be resolvable in relatively short order or for them to prove effectively impossible such that people give up.
Rinse and repeat.
¯\_(ツ)_/¯ That's python!
Package consumption sucks so bad, since the sensible way of using are virtual envs where you copy all dependencies. Then for freezing venvs or dumping package versions, so you can port your project to a different system, doesn't consider only packages actually used/imported in code, but it just dumps everything in the venv. The fact you need external tools for this is frustrating.
Then there is package creation. Legacy vs modern approach, cryptic __init__ files, multiple packaging backends, endless sections in pyproject.toml, manually specifying dependencies and dev-dependencies, convoluted ways of getting package metadata actually in code without having it in two places (such as CLI programs with --version).
Cross compilation really would be a nice feature to simply distribute a single file executable. I haven' tested it, but a Linux system with Wine should in theory be capable of "cross" compiling between Linux and Windows.
Still, like you, as a beginning I would prefer a sensible package management and package creation process.
The base line should be how heavily dynamic languages like my favourite set, Smalltalk, Common Lisp, Dylan, SELF, NewtonScript, ended up gaining from JIT, versus the original interpreters, while being in the genesis of many relevant papers for JIT research.
i didn't realize they ever jitted newtonscript
Had the Newton not been canceled, probably there would be an evolution from that support.
See "Compiling Functions for Speed"
https://www.newted.org/download/manuals/NewtonToolkitUsersGu...
that seems closer to the opposite of what you were saying in the point on which we were in disagreement?
maybe i should have said that up front!
except maybe common lisp; all the implementations i know are interpreted or aot-compiled (sometimes an expression at a time, like sbcl), but maybe there's a jit-compiled one, and i bet it's great
probably with enough work python could gain a similar amount. it's possible that work might get done. but it seems likely that it'll have to give up things like reference-counting, as smalltalk did (which most of the other languages never had)
A "Lisp interpreter" runs Lisp source in the form of s-expressions. That's what the first Lisp did.
A "Lisp compiler" compiles Lisp source code to native code, either directly or with the help of a C compiler or an assembler. A Lisp compiler could also compile source code to byte code. In some implementations this byte code can be JIT compiled (ABCL, CLISP, ...).
The first Lisp provided a Lisp to assembly compiler, which compiled Lisp code to assembly code, which then gets compiled to machine code. That machine code could be loaded into Lisp and functions then could be native machine code.
The Newton Toolkit could compile type declared functions to machine code. That's something most Common Lisp compilers do, sometimes by default (SBCL, CCL, ... by default directly compile source code to machine code).
SBCL:
* (defun add (a b) (declare (fixnum a b) (optimize (speed 3))) (+ a b))
ADD
* (disassemble #'add)
; disassembly for ADD
; Size: 104 bytes. Origin: #x7006E1789C ; ADD
; 89C: 0000018B ADD NL0, NL0, NL1
; 8A0: 0A0000AB ADDS R0, NL0, NL0
; 8A4: E7010054 BVC L1
; 8A8: BD2A00B9 STR WNULL, [THREAD, #40] ; pseudo-atomic-bits
; 8AC: BC7A47A9 LDP TMP, LR, [THREAD, #112] ; mixed-tlab.{free-pointer, end-addr}
; 8B0: 8A430091 ADD R0, TMP, #16
; 8B4: 5F011EEB CMP R0, LR
; 8B8: E8010054 BHI L2
; 8BC: AA3A00F9 STR R0, [THREAD, #112] ; mixed-tlab
; 8C0: L0: 8A3F0091 ADD R0, TMP, #15
; 8C4: 3E2280D2 MOVZ LR, #273
; 8C8: 9E0300A9 STP LR, NL0, [TMP]
; 8CC: BF3A03D5 DMB ISHST
; 8D0: BF2A00B9 STR WZR, [THREAD, #40] ; pseudo-atomic-bits
; 8D4: BE2E40B9 LDR WLR, [THREAD, #44] ; pseudo-atomic-bits
; 8D8: 5E0000B4 CBZ LR, L1
; 8DC: 200120D4 BRK #9 ; Pending interrupt trap
; 8E0: L1: FB031AAA MOV CSP, CFP
; 8E4: 5A7B40A9 LDP CFP, LR, [CFP]
; 8E8: BF0300F1 CMP NULL, #0
; 8EC: C0035FD6 RET
; 8F0: E00120D4 BRK #15 ; Invalid argument count trap
; 8F4: L2: 1C0280D2 MOVZ TMP, #16
; 8F8: 0AFBFF58 LDR R0, #x7006E17858 ; SB-VM::ALLOC-TRAMP
; 8FC: 40013FD6 BLR R0
; 900: F0FFFF17 B L0
NIL
I've entered a function and it gets ahead of time compiled to non-generic machine code.Calling the function ADD with the wrong numeric arguments is an error, which will be detected both a compile and at runtime.
* (add 3.0 2.0)
debugger invoked on a TYPE-ERROR @7006E17898 in thread
#<THREAD "main thread" RUNNING {70088224A3}>:
The value
3.0
is not of type
FIXNUM
when binding A
Redefinition of + will do nothing to the code. The addition is inlined machine code.Apple's Dylan IDE and compiler was implemented in Macintosh Common Lisp (MCL). MCL then was not a part of the Dylan runtime.
I would think that Open Dylan (the Dylan implementation originally from Harlequin) can also generate LLVM bitcode, but I don't know if that one can be JIT executed. Possibly...
CLISP has a byte code machine, for which a JIT can be used.
There might be others.
how much of a performance boost does abcl get from the hotspot jit compared to, say, interpreted clisp
Not necessarily, not for dynamic languages.
With very dynamic languages you can make only very limited assumptions about e.g. function argument types, which lead you to compiled functions that have to handle any possible case.
A JIT compiler can notice that the given function is almost always (or always) used to operate on a pair of integers, and do a vastly superior specialized compilation, with guards to fallback on the generic one. With extensive inlining, you can also deduplicate a lot of the guards.
also, even mature jit compilers often only make limited improvements; jython has been stuck at near-parity with cpython's terrible performance for decades, for example, and while v8 was an enormous improvement over old spidermonkey and squirrelfish, after 15 years it's still stuck almost an order of magnitude slower than c https://benchmarksgame-team.pages.debian.net/benchmarksgame/... which is (handwaving) like maybe a factor of 2 or 3 slower than self
typically when i can get something to work using numpy it's only about a factor of 5 slower than optimized c, purely interpretively, which is competitive with v8 in many cases. luajit, by contrast, is goddam alien technology from the future
with respect to your int×int example, if an int×int specialization is actually vastly superior, for example because the operation you're applying is something like + or *, an aot compiler can also insert the guard and inline the single-instruction implementation, and it can also do extensive inlining and even specialization (though that's rare in aots and common in jits). it can insert the guards because if your monomorphic sends of + are always sending + to a rational instance or something, the performance gain from eliminating megamorphic dispatch is comparatively slight, and the performance loss from inserting a static hardcoded guess of integer math before the megamorphic dispatch is also comparatively slight, though nonzero
this can fall down, of course, when your arithmetic operations are polymorphic over integer and floating-point, or over different types of integers; but it often works far better than it has any right to. in most code, most arithmetic and ordered comparison is integers, most array indexing is arrays, most conditionals are on booleans (and smalltalk actually hardcodes that in its bytecode compiler). this depends somewhat on your language design, of course; python using the same operator for indexing dicts, lists, and even strings hurts it here
meanwhile, back in the stop-hitting-yourself-why-are-you-hitting-yourself department, fucking cpython is allocating its integers on the heap and motherfucking reference-counting them
And then there is mypyc[1] which uses mypy's static type annotations but is only slightly faster.
And various other compilers like Numba and Cython that work with specialized dialects of Python to achieve better results, but then it's not quite Python anymore.
Python to C++ translation
And here I thought that it was shocking to learn that v8 allocates doubles on the heap recently. (I mean, I'm not a compiler writer, I have no idea how hard it would be to avoid this, but it feels like mandatory boxed floats would hurt performance a lot)
Correct, to the point where at work a colleague and I actually have looked into how to force using floats even if we initiate objects with a small-integer number (the idea being that ensuring our objects having the correct hidden class the first time might help the JIT, and avoids wasting time on integer-to-float promotion in tight loops). Via trial and error in Node we figured that using -0 as a number literal works, but (say) 1.0 does not.
> i don't think local-variable or temporary floats end up on the heap in v8 the way they do in cpython
This would also make sense - v8 already uses pools to re-use common temporary object shapes in general IIRC, I see no reason why it wouldn't do at least that with heap-allocated doubles too.
I assume that integers are coerced to floats in this mode, and that there's a performance cliff if you store a non-number in such an array, but in both cases I'm just guessing.
In SpiderMonkey, as you say, we store all our values as doubles, and disguise the non-float values as NaNs.
- "C/C++/Fortran libs are Python"
- "Python is too dynamic", while disregarding Smalltalk, Common Lisp, Dylan, SELF, NewtonScript JIT capabilities, all dynamic languages where anything can change at any given moment
Also no mentality shift is expected on the "Python is too dynamic" -- which is a strange thing to say anyway -- because Python is not getting any more static due to these JIT news.
Having a Python with JIT, in many cases it will be fast enough for most cases.
Data science running CUDA workloads isn't the only use case for Python.
I don't do data science.
I know Python since version 1.6, and is my scripting language in UNIX like environments, during my time at CERN, I was one of the CMT build infrastructure build engineer on the ATLAS team.
It was never been the language I would reach for when not doing OS scripting, and usually when a GNU/Linux GUI application happens to be slow as mollasses, it has been written in Python.
Flask, gunicorn, low single digit millisecond latency. Definitely optimised for latency over throughput, but not so much that we've replatformed it onto something that's actually designed for low latency :P. Callers all cache heavily with a fairly high hit ratio for interactive callers and a relatively low hit ratio for batch callers.
But on the whole, machines are cheaper than other engineering approaches to scaling.
For us, and many others, fast enough is fast enough.
shrug. If we're talking personal experience, I've been using Python since 1.4. It's been my primary development language since the late 1990s, with of course speed critical portions in C or C++ when needed - and I know a lot of people who also primarily develop in Python.
And there's a bunch of Python development at CERN for tasks other than OS scripting. ("The ease of use and a very low learning curve makes Python a perfect programming language for many physicists and other people without the computer science background. CERN does not only produce large amounts of data. The interesting bits of data have to be stored, analyzed, shared and published. Work of many scientists across various research facilities around the world has to be synchronized. This is the area where Python flourishes" - https://cds.cern.ch/record/2274794)
I simply don't see how a Python JIT is going to make that much of a difference. We already have PyPy for those needing pure Python performance, and Numba for certain types of numeric needs.
PyPy's experience shows we'll not be expecting a 5x boost any time soon from this new JIT framework, while C/C++/Fortran/Rust are significantly faster.
Unfortunely.
> And there's a bunch of Python development at CERN for tasks other than OS scripting
Of course there is, CMT was a build tool, not OS scripting.
No need to give me CERN links to me to show me Python bindings to ROOT, or Jupyter notebooks.
> PyPy's experience shows we'll not be expecting a 5x boost any time soon from this new JIT framework, while C/C++/Fortran/Rust are significantly faster.
I really don't get the attitude that if it doesn't 100% fix all the world problems, then it isn't worth it.
> I really don't get the attitude that if it doesn't 100% fix all the world problems, then it isn't worth it.
Then it's a good thing I'm not making that argument, but rather that "Having a Python with JIT, in many cases it will be fast enough for most cases." has very little information content, because Python without a JIT already meets the consequent.
Do you agree with me that Python is already fast enough for most cases, even without a JIT?
If not, how would a 30% boost improve things enough to change the balance?
https://stackoverflow.com/questions/36526708/comparing-pytho...
But pjmlp, I use Python because it's a wrapper for C/C++/Fortran libs. - Chocolate Giddyup
In Common Lisp not anything can change at any moment. Especially not in implementations where one uses AOT compilation like SBCL, ECL, LispWorks, Allegro CL, ... and so on. They have optimizing compilers which gradually can remove dynamic runtime behavior, upto supporting almost no dynamic runtime behavior.
Stuff which is supported: type specific code, inlining, block compilation, removal of development tools, ...
JIT implementations are rare in the Common Lisp world. They are mostly only used in implementations which use a byte-code virtual machine (CLISP, ABCL, ...). Common Lisp implementations mostly compile either directly to native code or via C compilers. The effect is that native AOT compiled code is much faster.
However, last time I used it, it (1) didn’t work with many third-party libraries (e.g. SciPy was important for me), and (2) didn’t work with object-oriented code (all your @njit code had to be wrapped in functions without classes). Those two has limited for which projects I could adopt Numba in practice, despite loving it in the cases it worked.
I don’t know what limitations the built-in Python JIT has, but hopefully it might be a more general JIT that works for all Python code.
(Actually a monkey-patched version to be able to set njit arguments)
A JIT compiler is a big deal for performance improvements, especially where it matters (in large repetitive loops).
Anyone cynical about the potential a python JIT offers should take a look at pypy which has a 5x speed up over regular python, mainly though JIT operations: https://www.pypy.org/
Not pursuing JIT or efficient compilation in general was a deliberate decision way back when Python made some kind of sense. It was the simplicity of implementation valued over performance gains that motivated this decision.
The mantra Python programmers liked to repeat was that "the performance is good enough, and if you want to go fast, write in C and make a native module".
And if you didn't like that, there was always Java.
Today, Python is getting closer and closer to be "the crappy Java with worse syntax". Except we already have that: it's called Groovy.
The language is definitely getting more complex syntactically, and I'm not a huge fan of some of those changes but it's no where near Java or C++ or anything else. You can still write simple Python with all of these changes.
Read it again. It seems you were reading too fast. I'm talking about the future, not the change being discussed right now.
> It's great that the core devs are keeping up with the time.
You mistake the influence of Microsoft and their desire to sell features for progress. Python is actually regressing as a system. It's becoming worse, not better. But it's hard to see the gestalt of it if all you are looking for is the new features.
> it's no where near Java
That is true. Java is a much more simple and regular (not in the automata theory sense) language. Today, if you want a simpler language, you need to choose Java over Python (although neither is very simple, so, preferably, you need a third option).
> You can still write simple Python
I can also write simple C++ if I limit what I use from the language to a very small subset. This says nothing about the simplicity of the language...
For JS, during the time that it received its JITs, there was no cross platform native code equivalent like wasm yet. JS had to compete with plugins written in C/C++ however. There was also competition between browser vendors, which gave the period the name "browser wars". Nowadays at least, the speed improvements for the end user thanks to JIT aren't also that great, Apple provides a mode to turn off JIT entirely for security.
JavaScript JITs only emerged around 2008 with SpiderMonkey’s TraceMonkey, JavaScriptCore’s SquirrelFish Extreme, and V8’s original JIT.
Additionally a lot of the libraries/ecosystem around shared memory (https://docs.python.org/3/library/multiprocessing.shared_mem...) seems poorly conceived. If you pre-open shared memory in a ProcessPoolExecutor's initializer functions, you can't close them when the worker process exits (which might be fine, nobody knows!), but if you instead open and close a shared memory segment on every executor job, it measurably reduces performance, presumably from memory mapping overhead or TLB/page table thrashing.
But what is the counterfactual? Implementing the whole thing in Python? It seems much more work than forking/fixing matplotlib.
Well, imho the biggest problem with this approach to paralellism is that you're stepping out of the Python world with gc'ed objects etc. and into a world of ctypes and serialization. It's like you're not even programming Python anymore, but more something closer to C with the speed of an interpreted language.
That's quite surprising to learn, as I didn't think the initializer ran in a specialized context (like a pthread_atfork postfork hook in the child).
What happens when you try to close an initializer-allocated SharedMemory object on worker exit?
- Subclassing ProcessPoolExecutor such that it spawns multiprocessing.Process objects whose runloop function wraps the stdlib "_process_worker" function in a try/finally which runs your at-shutdown logic. That'll be as reliable as any try/finally (e.g. SIGKILL and certain interpreter faults can bypass it).
- Writing custom destructors of objects in your call arguments which are aware of and can do appropriate cleanup actions for associated SharedMemory objects. This is less preferred than subclassing because of the usual issues with custom destructors: no real exception handling, and objects sneaking out into long-lived/global caches can cause destructors to run late (after the interpreter has torn down things your cleanup logic needs) or not at all.
- Atexit, as you suggest. This is least-preferred because the execution context of atexit code is .... weird, to say the least. Much like a signal handler or pthread_atfork callback, it's not a place that I'd put code that does complicated I/O or depends on the rest of the interpreter being in ordinary conditions.
For machine learning speed compiling to the right CUDA / OpenCL kernel is much more crucial, so there's where the money goes.
The JavaScript VMs often break their extensions APIs for speed, but their users are more used to this.
I mean, it's great that you can write some of your code in C. But wouldn't it be great if you could just write your libraries in Python and have them still be really fast?
I don't see an indication in the article that that's the case. Am I missing something?
https://www.pypy.org/posts/2011/05/numpy-follow-up-692862769...
https://doc.pypy.org/en/latest/faq.html#what-about-numpy-num...
i'm not sure what version they gave up at
Everybody obviously wants that. The question is are you willing to lose what you have in order to hopefully, eventually, get there. If Python 3 development stopped and Python 4 came out tomorrow and was 5x faster than python 3 and a promise of being 50-100x faster in the future, but you have to rewrite all the libraries that use the C API, it would probably be DOA and kill python. People who want a faster 'almost python' already have several options to choose from, none of which are popular. Or they use Julia.
That is the problem with all the fast python implementations that have come before. Yes, they're faster than 'normal' python in many benchmarks, but they don't support the entire current ecosystem. For example Instagram's python implementation is blazing fast for doing exactly what Instagram is using python for, but is probably completely useless for what I'm using python for.
When i was a scientist, speed was getting the code written during my break, and if it took all afternoon to run that's fine because i was in the lab anyway.
Even as i moved more into the software engineer direction, and started profiling code more, most of the bottlenecks come from things like "creating objects on every incovation rather than pooling them", "blocking IO", "using a bad algorithm" or "using the wrong datasctructure for the task". problems that exist in every language, though "bad algorithm" or "using the wrong datasctructure" might matter less in a faster language you're still leaving performance on the table.
> "Python is so slow that we have to write any important code in C. And this is somehow a good thing."
The good thing is that python has a very vibrant ecosystem filled with great libraries, so we don't have to write it in C, because somebody else has. We can just benefit from that when the situation calls for it
But sure, I'm all for removing build steps and avoiding yet another layer.
That really depends.
To make the issue clear, let's think about a similar situation:
bash is nice because you can plug together inputs and outputs of different sub-executables (like grep, sed and so on) and have a big "functional" pipeline deliver the final result.
Your idea would be "wouldn't it be great if you could just write your libraries in bash and have them still be really fast?". Not if you make bash into C, tanking productivity. And definitely not if that new bash can't run the old grep anymore (which is what usually is implied by the proposal in the case of Python).
Also, I'm fine with not writing my search engines, databases and matrix multiplication algorithm implementations in bash, really. So are most other people, I suspect.
Also, many proposals would weaken Python-the-language so it's not as expressive anymore. But I want it to stay as dynamic as it is. It's nice as a scripting language about 30 levels above bash.
As always, there are tradeoffs. Also with this proposal there will be tradeoffs. Are the tradeoffs worth it or not?
For the record, rewriting BLAS in Python (or anything else), even if the result was faster (!), would be a phenomenally bad idea. It would just introduce bugs, waste everyone's time, essentially be a fork of BLAS. There's no upside I can see that justifies it.
While I’m all for making Python itself faster, it would be a shame to lose the glue language par excellence.
I think that's a pretty ignorant interpretation. Python has been built to have a giant ecosystem of useful, feature-complete, stable, well built code that has been used for decades and for which there is no need to reinvent the wheel. If that already describes the universe of libraries that you /need/ to be extremely fast and the rest of your code is IO limited and not CPU limited, why reinvent the wheel?
That makes your comment even more inaccurate because you likely don't need to write any "important" (which you are stretching to mean "fast") code in C -- you utilize existing off the shelf fast libraries that are written in Fortran, CUDA, C, Rust or any other language a pre-existing ecosystem was built in.
Try and think of a language that has mature capabilities for domains as far away as what Django solves for, what pandas solves for, what pytorch solves for, and still has fantastic tooling like jupyter and streamlit. I can't think of any other language that has the combined off the shelf breadth and depth of Python. I don't want to have to write fast code in any language unless forced to, because the vast majority of the time I can customize a great off the shelf package and only write the remaining 1% of glue. I can't see why a professional engineer would 99% of the time would need to take a remotely different approach.
There are a lot of faster python runtimes out there. Both Google and Instagram/Meta have done a lot of work on this, mostly to solve internal problems they've been having with python performance. Microsoft has also done work on parallel python. There's PyPy and Pythran and no doubt several others. However none of these attempts have managed to be 100% compatible with the current CPython (and more importantly the CPython C API), so they haven't been considered as replacements.
JavaScript had the huge advantage that there was very little mission critical legacy JavaScript code around they had to take into consideration, and no C libraries that they had to stay compatible with. Meaning that modern JavaScript runtime teams could more or less start from scratch. Also the JavaScript world at the time were a lot more OK with different JavaScript runtimes not being 100% compatible with each other. If you 'just' want a faster python runtime that supports most of python and many existing libraries, but are OK with having to rewrite some your existing python code or third party libraries to make it work on that runtime, then there are several to choose from.
Python with it's C API basically gives you the keys to the kingdom on a machine code level. Modifying something that has an API to connect to essentially anything is not an easy proposition. Of course, it has the advantage that you can make Python faster by performance analysis and moving the expensive parts to optimized C code, if you have the resources.
They test it against the top 500 modules on PyPI and it's currently compatible with about half:
https://www.graalvm.org/python/compatibility/
But investment continues. It has some other neat features too like sandboxing and the ability to make single-binary programs.
The GraalPython guys are working on the HPy effort as well, which is an attempt to give Python a properly specified and engine-neutral extension API.
What I would've preferred is they leave all that stuff alone, add nice features like async/await that don't break existing things, and make important changes to the runtime and package manager. Python's packaging is so broken that it's almost mandatory to have a Dockerfile nowadays, while in JS that's not an issue
And if you have a complicate computation graph, there are already JITs on this level, based on Python code, e.g. see torch.compile, or TF XLA (done by default via tf.function), JAX, etc.
It's also important to do JIT on this level, to really be able to fuse CUDA ops, etc. A generic Python JIT probably cannot really do this, as this is CUDA specific, or TPU specific, etc.
There has been a lot of competition to make browsers fast. Nowadays there are 3 main JS engines, V8 backed by google, JavaScriptCore backed by apple, and spidermonkey backed by mozilla.
If python had been the language embedded into web browsers, then maybe we would see 3 competing python engines with crazy performance.
The alternative interpreters for python have always been a bit more niche than Cpython, but now that Guido works at microsoft there has been a bit more of a push to make it faster
Good to see progress anyways.
SHA256 in pure Python would be unusably slow. In Javascript it would be at least usably slow.
Javascript is fast. Browsers are fast.
I did, and doing it in the browser was so bad that it was unusable. I suspect that it's not the crypto that's slow but the file reading. But anyway...
> SHA256 in pure Python would be unusably slow
None would do that because:
> Python's SHA256 is written in C
Hence why comparing "pure python" to "pure javascript" is mostly irrelevant for most day to day tasks, like most benchmarks.
> Javascript is fast. Browsers are fast.
Well, no they were not for my use case. Browsers are really slow at generating file checksums.
And as noted upthread that's a significant part of the uptake of Python in scientific fields, and why pypy despite the heroic work that's gone into it is often a non-entity.
This is a major problem in scientific fields. Currently there are sort of "two tiers" of scientific programmers: ones who write the fast binary libraries and ones that use these from Python (until they encounter e.g. having to loop and they are SOL).
This is known as the two language problem. It arises from Python being slow to run and compiled languages being bad to write. Julia tries to solve this (but fails due to implementation details). Numba etc try to hack around it.
Pypy is sadly vaporware. The failure from the beginning was not supporting most popular (scientific) Python libraries. It nowadays kind of does, but is brittle and often hard to set up. And anyway Pypy is not very fast compared to e.g. V8 or SpiderMonkey.
Reee.
And even these issues are part of the greater problem of late stage capitalism that in general produces god awful stuff with questionable value. E.g. vast majority of industry code is such.
Care to list some of those details ? (I have zero knowledge in Julia)
It's my assesment that the problems listed in there are a cause why Julia will not take off and we're largely stuck with Python for the foreseeable future.
There is something big going on in caching the binaries, so there's a chance the TTFX will get workable.
The slowest part in the JavaScript version seems to be reading the file, accounting for 70–80% of the runtime in both Firefox and Chromium.
Have you tried to do this in Python?
A Node comparison would be more appropriate.
https://docs.modular.com/mojo/manual/
edit: The main point I forgot to mention - it aims to compete with "low-level" languages like C and Rust in performance
1. Javascript is a less dynamic language than Python and numbers are all float64 which makes it a lot easier to make fast.
2. If you want to run fast code on the web you only have one option: make Javascript faster. (Ok we have WASM now but that didn't exist at the time of the Javascript Speed wars.) If you want to run fast code on your desktop you have a MUCH easier option: don't use Python.
I have seen this mentioned multiple times, someone as a good reference explaining what makes python more dynamic than JS ?
You could probably optimistically optimise some code, assuming it doesn't use any of the dynamic features of Python. You're going to get crazy performance cliffs though.
Python's users can always swap out performance critical components to another language. So Python development delivered more when it focussed on improving strengths rather than mitigating weaknesses.
In a way, Python being slow is just a sign of a healthy platform ecosystem allowing comparative advantages to shine.
Python's syntax is still nicer for mathy stuff, to the point where I'd go into job coding interviews using Python despite having used more JS lately. And I'm comparing to JS because it's the closest thing, while others like Java are/were far more cumbersome for these uses.
Indeed, reading the blog post build much higher expectations
For example, if the JIT compiler realizes the program is adding two integers it could potentially replace the code with two MOVs and a single ADD. However, what about the error handling in the case of an overflow? Python switches to its internal BigInt representation in this case and cannot rely on architecture specific instructions alone once the result gets too large to fit into a register.
Modern programming languages are all about trading performance for convenience and that is what makes them slow — not because they are running an interpreter and not compiling to machine code.
> The initial benchmarks show something of a 2-9% performance improvement.
Which is underwhelming (as mentioned in the article), especially if we look at PyPy[0]. But it's a step forward nonetheless.
It's a bit less underwhelming if you consider that only function objects with loops are being JITed. nb: for loops in Python also use the JUMP_BACKWARD op.
'Twas the night before Christmas, when all through the code
Not a core dev was merging, not even Guido;
The CI was spun on the PRs with care
In hopes that green check-markings soon would be there;
...
...
...
--enable-experimental-jit, then made it,
And away the JIT flew as their "+1"s okay'ed it.
But they heard it exclaim, as it traced out of sight,
"Happy JIT-mas to all, and to all a good night!"
https://github.com/python/cpython/pull/1134652024, right?