LPython: Novel, Fast, Retargetable Python Compiler
lpython.org
lpython.org
At least as fast as C++ is a bold claim, but this is an interestingly documented process. I'm keen.
I'm quite partial to Nuitka at this point, but I'm open to other Python compilers.
Yes, offloading to GPU we want to support naturally via NumPy syntax. We will look at this very soon, most likely via annotating that a given array lives on a GPU or host, and then array copy will copy it from host to device, etc.
However, it really is different from projects like this one, in that it doesn't attempt to obtain C-like speed (but does hope to do some optimisations). For example, x+=1 will still dynamically dispatch depending on the runtime type of x, and (if it's an int) do the normal Python arbitrary precision operation. But those will be called from machine code rather than interpreted byte code.
(Essentially, it unrolls the main loop of the CPython interpreter, which is written in C, for every byte code operation, and eliminates every case of the switch statement inside except the one that corresponds to this operation. That's what gets compiled.)
> Nuitka is the Python compiler. ... It then executes uncompiled code and compiled code together in an extremely compatible manner. Nuitka translates the Python modules into a C level program that then uses libpython and static C files of its own to execute in the same way as CPython does.
I'm happy to stick by the tool's own description of itself. But if you don't accept that, and Nuitka still isn't a compiler by you, sure. Go ahead with your pedantic and strictest definition of the meaning of a 'compiler'.
Learn to read English, before insisting on your pedantry as some sort of truth, in disregard of the tool's own intent/description. (And I do hope you take issue with everyone's description of Typescript as a compiled programming language too.)
I think my comment being devoid of content might have caused some frustration. So to make the best of everyone's time:
You can look into cython and pythran to see what I'm talking about. cython lets you optimize code step by step via generating an html page with your code and highlighting lines that still require the use of the python runtime. It lets you add types and cdef function definitions in order to reduce your dependency on the python runtime.
Another good example is pythran, which takes your python code and turns it into c++ code to be compiled by a c++ compiler. I understand that this isn't a direct compilation to machine code, but a middle step which lets you compile the output to machine code.
Then there is numba and taichi which have just-in-time compilation decorators. Taichi also provides a sophisticated runtime which lets you run parts of the code on a GPU.
Surprisingly, the best performance I've experienced among these examples was numba + numpy, even though numba alone can sometimes have optimizations that surpasses all compilation efforts, because it turns your loops into mathematical formulas and runs them at O(1) complexity when it can.
Two questions come to my mind:
- presumably, since it is compiled, it does static checks on the code? How many statically-detectable bugs that are now purely triggered at runtime can be eliminated with LPython?
- does it deal with the unholy amounts of dynamism of Python? Can you call getattr, setattr on random objects? Does eval work? Etc. Quite a few Python packages use these at least once somewhere...
The examples in the article appear to mainly revolve around numerical calculations, so I suspect the target audience is people doing scientific computing who need to "break out" into a compiled mode for heavy CPU calculations (similar to numba or even Julia) from time to time when their calculations aren't vectorisable.
I've noticed a split between the needs of software engineers on the one hand, who need expressive abstractions to manage systems of extensive rather than intensive complexity, and scientific programmers or model-builders on the other hand, who are much more likely to just use the primitives offered by their language or library as their needs revolve around implementing complicated algorithms.
Yes, it does static checks at compile time. The only thing that we do at runtime (imperfectly right now, eventually perfectly) are array/list bounds checks, integer overflow during arithmetic or casting and such. Those checks would only run in Debug and ReleaseSafe modes, but not in Release mode. So for full performance, you choose the Release mode. For 100% safe mode, you would need to use ReleaseFast.
- does it deal with the unholy amounts of dynamism of Python? Can you call getattr, setattr on random objects? Does eval work? Etc. Quite a few Python packages use these at least once somewhere...
It deals with it by not allowing it. We will support as much as we can, as long as it can be robustly ahead of time compiled to high performance code. The rest you can always use just via our CPython "escape hatch". The idea is that either you want performance (then restrict to the subset that can be compiled to high performance code) or you don't (then just use CPython).
If you have some ideas how to best communicate this, let me know.
The best approach that I know right now is to say that we support a strict subset of Python, and if it compiles, it will be fast. The rest of Python you have to call explicitly and it will be slow and you get a CPython dependency in the binary that we generate, but you can do it. That way there is a clear distinction what gets compiled via our compiler into high performance machine code and what gets dispatched via CPython (it will eventually get "compiled", but it will just call CPython).
They have some benchmarks vs regular python here :
https://github.com/mypyc/mypyc-benchmark-results/blob/master...
One difference is that MyPyC compiles your code to a C extension, so your are still dependent on python. On the other hand you can call regular python libraries with the normal syntax while, in LPython, the "break-out" syntax to regular libraries isn't straightforward
In any case super exiting to see work going into AOT python
https://github.com/lcompilers/lpython.org-deploy/pull/37
So now we have 25 compilers there.
Yes, the current syntax to call CPython is low level, you have to create an explicit interface. We can later make it more straightforward, such as using a `@python` decorator to a function, where inside you just do CPython. We always want to make it explicit, since it will be slow, and by default we want the LPython code to always be fast.
With that said, I love the fact that this can generate executables, and I'm looking forward to trying it out in the future! Python compilers are really cool. I recently used Nuitka to build a standalone version of an app I wrote for my own use, so that I can run it on hosts that don't have Python installed (or don't have the right Python packages installed yet). This seems to be focused much more on speed.
One thing which I didn't understand from the homepage: can I take vanilla Python code and AOT compile it with this?
Yes, I noticed too, thanks. We currently don't have a dedicated documentation for LPython and a lot of the LFortran documentation applies. We will eventually have a dedicated LPython documentation.
> With that said, I love the fact that this can generate executables, and I'm looking forward to trying it out in the future! Python compilers are really cool. I recently used Nuitka to build a standalone version of an app I wrote for my own use, so that I can run it on hosts that don't have Python installed (or don't have the right Python packages installed yet). This seems to be focused much more on speed.
We focus on speed, but we definitely create executables (no Python dependencies), just like any other C or Fortran compiler would. You only get a CPython dependency if you call into CPython (explicitly). Indeed creating such standalone executables simplifies the deployment and packaging issues: all the packages must be resolved at compile time, once you compile your application, there are no more dependencies on your Python packages (unless any of them calls into CPython of course).
> One thing which I didn't understand from the homepage: can I take vanilla Python code and AOT compile it with this?
Only if that vanilla Python compiles with LPython (since any LPython code is just Python code, a subset). So in general no. If you call CPython from your LPython main program (let's say), then you will get a small binary that depends on CPython to call into your Python code. I thought about somehow packaging the Python sources into the executable, similar to how PyOxidizer does it, but that's a project on its own almost. We will see what the community wants. If we can make LPython support a large enough subset of CPython, I think I would like to use LPython as is, since it's nice to not have any Python dependency and everything being high performance, essentially equivalent to writing C++ or Fortran.
Will LPython have the ability to generate AoT compiled libraries in addition to executables?
Indeed, I think for large projects you want:
* generate a binary
* fast compilation in Debug mode
* fast runtime in Release mode
Sometimes you want JIT, so we support it too, but I personally don't use it, I write the main program in LPython and just compile it to a binary.
> Will LPython have the ability to generate AoT compiled libraries in addition to executables?
Yes, you can do it today. Just compile to `.o` file and create a library. We use this library feature in production at my company today. If you run into issues, just ask us, we'll help.
And good point about JIT--it does have its place. What I should have said is that I wish there were better options with high quality support for _both_ AoT and JIT. Most of the options I'm aware of have good support for one, but poor or no support for the other. I'm curious to see how well LPython bridges the gap.
"LPython is built from the ground up to translate numerical, array-oriented code into simple, readable, and fast code."
Both are valid and consistent approaches with their pros and cons, I listed some of them here: https://fortran-lang.discourse.group/t/fast-ai-mojo-may-be-t....
Is it superset if it doesn't implement all of Python? It seems more like a set that has a non-empty intersection with Python.
edit: after trying quickly, it seems that lpython really requires type annotations everywhere, while codon is more permissive (or does type inference)
Internally they each parse the syntax to AST, then have some kind of an intermediate representation (IR), do some optimizations and generate code. The differences are in the details of the IR and how the compiler is internally structured.
Regarding the type inference, this is for a blog post on its own. See this issue for now: https://github.com/lcompilers/lpython/issues/2168, roughly speaking, there is implicit typing (inference), implicit declarations and implicit casting. Rust disallows implicit declarations and casting, but allows implicit typing. As shown in that issue they only meant to do single line implicit typing, but (by a mistake?) allowed multi-statements implicit typing (action at a distance). LPython currently does not allow any implicit typing (type inference). As documented at the issue, the main problem with implicit typing is that there is no good syntax in CPython that would allow explicit type declaration but implicit typing. Typically you get both implicit declaration and implicit typing, say in `x = 5`, this both declares `x` as a new variable as well as types it as integer. C++ and Rust does not allow implicit declarations (you have to use `auto` or `let` keywords) and I think we should not do either. We could do something like `x: var = 5`, but at that point you might as well just do `x: i32 = 5`, use the actual type instead of `var`.
From "Show HN: Python Tests That Write Themselves" (2019) https://news.ycombinator.com/item?id=21012133:
> pytype (Google) [1], PyAnnotate (Dropbox) [2], and MonkeyType (Instagram) [3] all do dynamic / runtime PEP-484 type annotation type inference [4]
Hypothesis (@given decorator tests) also does type inference IIUC? https://hypothesis.readthedocs.io/en/latest/
icontract and pycontracts do runtime Preconditions and Postconditions with Design-by-Contract patterns similar to Eiffel DbC; they check the types and values of arguments passed while the program in running and not just at coding or compile time.
I use Numba quite a bit to does up slow Pandas operations. Cool to have another alternative.
That obviously has to be the goal, but is it really feasible to be faster than good C/C++ or Fortran? I did some research into the Python Compiler landscape and came to the conclusion that it almost always boils to LLVM. So, if you want to have fast code, just help the compiler make the most of your code and you'll be 99% of the way there and as fast as possible without significantly more effort.
Would you agree with my layman's understanding of this topic?
Regarding LLVM: my experience so far is that LLVM is indeed amazing what it can do in terms of optimizations. It's very very good. However, it is not all LLVM. As our benchmarks in the blog post show, we compare Numba, Clang and LPython, all three of which use LLVM. But we get vastly different performance for what seems to look like identical initial code. To know exactly why, we would have to meet with the Numba and Clang developers and study this, I suspect Clang lowers to LLVM too soon, and uses C++ to do abstractions (like `std::vector` or `std::unordered_map`) and perhaps it can't quite get the top performance this way. Numba perhaps doesn't get all the types as tight as LPython, or perhaps implements some things not as efficiently, or perhaps doesn't apply as good optimizations before lowering to LLVM. I suspect LLVM gets the best performance if the compiler generates as straightforward LLVM IR code as possible, without layers and layers of abstractions that might not end up being "zero cost" in practice.
Does typing have to be extensive or does the majority of it get inferred with perhaps just function and class boundaries needing annotations?
And if the latter, does the typing get inferred _after_ the initial ssa pass, so a name has a context-specific type?
I think - the performance gains are coming from the Python syntax being transliterated to an LFortran intermediate representation (much like Numba converts code to LLVM IR).
Any calls to CPython libraries are made using a special decorator, which might be doing interop using the CPython API. My guess is, this will come with a performance penalty. More so if you’re using Numpy or Scipy, as you’ll be going through several layers of abstractions and hand-offs.
This is because Numba, Pythran and JAX (in a way) get around this by reimplementing a subset of Numba/Scipy/other core libraries. Any call to a supported function is dynamically rerouted to the native reimplementation during JIT/AOT compilation.
I’d be interested in seeing how far LPython can tolerate regular Python code, with a ton of CPython interop and class use.
In any case, glad to see more competition. Not to take anything from the authors - this is a massive effort on their part and achieves some impressive results. Mojo has VP money behind it - AFAIK, this is a pure volunteer driven effort, and I’m grateful to the authors for doing it!
Granted, I'm optimistic for Mojo's potential but I do wish I could run it locally. Modular's pricing model also remains to be seen.
> This is way too similar to Mojo (language) by Modular. At least at some level conceptually. All in a good way.
Yes, the main difference is that Mojo is (or will be) a strict superset of Python, while we are a strict subset (but you can call the rest of Python via a decorator).
> I think - the performance gains are coming from the Python syntax being transliterated to an LFortran intermediate representation (much like Numba converts code to LLVM IR).
Correct. We use the same IR as LFortran and then we lower to LLVM IR.
> Any calls to CPython libraries are made using a special decorator, which might be doing interop using the CPython API. My guess is, this will come with a performance penalty. More so if you’re using Numpy or Scipy, as you’ll be going through several layers of abstractions and hand-offs.
Yes, it calls CPython, so it's slow.
> This is because Numba, Pythran and JAX (in a way) get around this by reimplementing a subset of Numba/Scipy/other core libraries. Any call to a supported function is dynamically rerouted to the native reimplementation during JIT/AOT compilation.
We do as well: we support a subset of NumPy directly (eventually most of NumPy). We also support a very small subset of SymPy. Over time we add more support to more basic libraries. The rest you can call via CPython, but slow. For SymPy we'll experiment building it on top (at least some modules, like limits) and compile using LPython. Given that any LPython code is just Python, this might be a viable way, as long as we support enough of Python directly.
> I’d be interested in seeing how far LPython can tolerate regular Python code, with a ton of CPython interop and class use.
We support structs via `@dataclass`, but not classes yet (although LFortran does to some extent, so we'll add support soon to LPython as well). For regular Python call LPython will give nice error messages suggesting to type things. Once you do and it compiles, it will run fast.
> In any case, glad to see more competition. Not to take anything from the authors - this is a massive effort on their part and achieves some impressive results. Mojo has VP money behind it - AFAIK, this is a pure volunteer driven effort, and I’m grateful to the authors for doing it!
We are supported by my current company (GSI Technology) as well as by NumFOCUS (LFortran), GSoC and other places; we have a very strong team (5 to 10 people). In the past I was supported by Los Alamos National Laboratory to develop LFortran. I have delivered SymPy as a physics student with no institutional support. So I have experience doing a lot with very little. :)
Haven't heard from pypy for a while now.
Please don't take this as "why u no faster?!" as much as I really would enjoy understanding what it takes to bring pypy up to cpython parity. I do see https://doc.pypy.org/en/latest/contributing.html#your-first-... says that pypy has a lot of layers and (I'd suppose) a fraction of the number of people compared to cpython's contributor base, but like I said it's hard to understand from the outside whether it's just a lot of rocks to break, or there are genuine novel optimization or computer science tricks that have to be solved when rolling out a newer version
Above all, thanks so much for your work on pypy - it has really helped me a lot and not from its speed but from my ability to use in in places where getting cpython to run is harder than "curl pypy.tar.bz2 | tar -xjf"
We could target MLIR later, right now we are just targeting LLVM.
What's novel about it?
If this could be used in mainstream web dev with the level of speed detailed, python might eat Javascript's lunch.
*bias - python obsessed
Auto swagger generation, validation pipes and so on.
On the risk of starting a holywar: why.
My approach is that it is still unclear to me what exactly will be possible in the future, while I know exactly how to deliver these compilers today. I suspect a traditional compiler will be more robust and also a lot faster than an LLM for tasks like translation to another language or compilation to binary. And speed of compilation is very important for development from the user perspective.
Conclusion: I don't know what the future will bring, but I suspect these compilers will still be very useful.
Local variables are on stack. Lists, dicts, sets and arrays use heap, unless the length is known at compile time, where we might use a stack, but I think we need to make it configurable, since one can run out of stack quite easily this way.