Python interpreter written in rust reaches 10000 commits
github.com
github.com
>please use the original title, unless it is misleading or linkbait; don't editorialize.
This makes it much easier.
Not game changing for the Python world, but I think it's pretty neat.
Have they got a big new compiler idea?
Pretty sparse. Uses cranelift tho, which is what Firefox is using for wasm
For comparison, here's a complete JIT-for-befunge using cranelift: https://github.com/serprex/Befunge/blob/master/barfs/src/jit...
Rust is seeing ~5x improvement in compile time using cranelift for debug builds instead of llvm. Cranelift is much better suited as a jit library than llvm is. Bit of a pain postgres went the llvm route
Unladen Swallow is seared in my mind. It was a failed project, back in 2009, to improve the speed of CPython by using LLVM as a JIT.
First, the Unladen Swallow team (IIRC) spent a lot of time fixing bugs in LLVM.
Second, LLVM isn't fast at compiling code, at least not for a JIT. This is legitimately surprising, because the official LLVM Tutorial implements a JIT.
Third, LLVM used to stand for Low Level Virtual Machine. I don't know when it stopped standing for that; clearly it hasn't for a long time. But with “Virtual Machine” in the title, you can see why people might have thought it would be suitable for implementing a dynamic language. cf. GraalVM these days.
I have only developed in Python professionally, and when I play around in Go and Rust, I really miss the ability to sketch things out in an IPython session.
CPython doesn't have a JIT, it has an interpreter. So it spends a lot of time in this loop: https://github.com/python/cpython/blob/main/Python/ceval.c
But you still have a repl with pypy. The downsides of JIT is that compiling bytecode to assembly can take up time (hurting startup performance, but that can be mitigated by not applying jit aggressively) & some programs have very dynamic behavior which the JIT has to eventually give up on (& go back to interpreting) or run off some pathological performance cliff where it takes up a bunch of memory & runs 10x slower than interpreter
Ruby 3 introduced a jit to their reference implementation
There are also Jython and IronPython, although I don't know how successful/popular or actively-maintained those implementations are.
Someone recently pointed out on IRC that Python is transitioning away from a language with "one implementation, and the implementation is the spec" to "one preferred implementation with an informal spec, but many alternative implementations".
Now that PyPy v7.3.7 seems to support 3.8 without any issues that I've encountered, I'd strongly consider evaluating it for use in production code. It might be interesting to run some async webserver + data processing benchmarks or something across the various implementations. Maybe a good benchmark would be measuring total throughput on something like "process/sanitize some user-provided text file, then serve a prediction from a PyTorch model".
A significant performance improvement in python would benefit many ds related tasks.
Doesn't matter how fast a GPU you have; Python and the GIL is the bottleneck.
Fortunately, with a tool like DVC or even Make, you usually don't have to (or want to) put that code in the same script as the actual machine learning part. So you can theoretically run the former with PyPy and the latter with CPython, if you really need to maximize both.
But my post was more oriented towards non-data-science uses of Python, like writing an API server or a web crawler or a TUI application. I think the "serve a prediction from a PyTorch model" part threw off the conversation a bit!
It's been "transitioning" for as long as it existed. The alternative forks all eventually die.
I don't think this is a new thing at all. When I first heard of Python a decade and a half ago, the "pitch" in the official docs back then was that "Python" was a programming language while "CPython" was the reference implementation of it. And the docs were (and still are) quite careful to point out which things are implementation details of CPython and are not to be relied upon if portability to other implementations or future versions of CPython are desirable features.
I don't know what you mean about an "informal spec", the the full formal specification of Python (the language) is here: https://docs.python.org/3/reference/index.html
I think GP meant "informal" in the sense that python wasn't originally meant to be a multi-implementation language. Though the oldest alternative implementation goes back to 1997, alternatives were always, until relativel recently, second class. I don't have the tools or the time to quantify what I mean by "recently" or "second class", but I hope you know what I mean.
The single biggest marker of being "second class" that I do know of is how the new features always gets introduced first by Cpython, then every other implementation plays catch up. This is unheard of in true multi-implementation languages, what I have in mind is C, C++, Java and Javascript. I'm sure there are plenty more, but those just, off the top of my head, are the most prominent languages whose features are introduced first in completely implementation-agnostic way, then all implementations start racing to get it complete.
I don't mean to say that Cpython maintainers just wake up in the morning and decide to add new syntax or semantics to the language, PEP documents are quite formal and implementation-agnostic, but I always had the impression it's something by Cpython devs for Cpython devs, and supporting that is how a Cpython implementation always appear first and other implementations lag behind by a varying amounts.
Heck, I probably should have mentioned Stackless too, which apparently is still being actively developed.
Nuitka probably is worth mentioning along these lines too.
You can't really "fake" a global lock. Either you're globally locking everywhere with the global lock or you aren't.
But if you are interested on removing the GIL, you gonna love 2022: https://docs.google.com/document/u/0/d/18CXhDb1ygxg-YXNBJNzf...
I understand that there are currently a large number (or at least a vocal number?) of HN readers with an interest in Rust. I have been here long enough to witness Lisp, Haskell, Node, Golang, Julia, and even .NET Core all go through similar cycles.
However, Rust is the ONLY one on that list for which "___ written in Rust" is a nearly automatic trip to the top of the front page. It's weird. It feels like astroturf at times, and even if it's legitimate good faith then it's still overbearing.
Just...... is there literally ANYTHING noteworthy about "_____" other than the programming language it was written in? I'm not sure how noteworthy that is, by itself, even if you have a strong interest in that language.
- it makes compiling your entire project into a standalone executable trivial
- the interpreter can be compiled to WASM way more easily than with something like pyodide
- you can provision python with cargo, which mean no more fiddling with pyenv, deadsnake, epel, etc. and yet getting a consistent, to the minor version, python distribution
- rewritting you hot path in rust becomes first class citizen. Since the python story is, start with python, and when you need to scale, you can always create an extension later, this is really attractive.
All that has only value, of course, if the project reaches a good compat and is supported.
But still, the possibilities it offer are not negligible.
Interpreted != not compiled.
But I find it noteworthy (inasmuch as I'm writing this comment) because the phenomenon is real. I've heard first-hand horror stories of shops that use commits per day as a serious metric, and the result is horrible.
It’s getting rid of the GIL and having existing Python code continue to work, with C extensions and all.
(It's entirely normal for words to have both specific and general meanings.)