Nuitka: An extremely compatible Python compiler
nuitka.net
nuitka.net
I first found it back in 2015 when I worked at a company where we built a python based desktop application as a part of our industrial control system. Nuitka provided better performance then pyinstaller/cx_freeze while it was still simple to work with. Back then there were still a few incompatibilities but I've followed the project throughout the years and it's mature like a fine wine since.
* https://news.ycombinator.com/item?id=8771925 (2014, 135 comments)
* https://news.ycombinator.com/item?id=10994267 (2016, 52 comments)
* https://news.ycombinator.com/item?id=15354613 (2017, 60 comments)
[0]: https://nuitka.net/doc/developer-manual.html#ssa-form-for-nu...
"Python Interpreters Benchmarks"
A small recommendation for these release pages would be to link to the tag in whatever source host you're using! I just want to see the code.
https://github.com/bellard/quickjs/blob/master/quickjs-opcod...
The point isn't just not having a JIT. You can run v8 without a JIT. The point is good performance without a JIT.
> Only output bytecode in a C file. The default is to output an executable file.
> @item -e
> Output @code{main()} and bytecode in a C file. The default is to output an executable file.
https://github.com/bellard/quickjs/blob/master/doc/quickjs.t...
Looks like the only compilation strategy is directly to ARM Thumb assembly or did I miss something?
"Microsoft MakeCode: from C++ to TypeScript and Blockly (and Back)"
https://www.youtube.com/watch?v=tGhhV2kfJ-w
and the respective C++ FFI,
Theoretically you could "precompile" JavaScript by running your code on some known input and caching the VM's state. Then you can "hydrate" your VM with that state. But this would only speed up the VM warmup (which is not that long) and requires significant browser/VM buy-in, so it is probably infeasible.
The only way to produce better performance in JavaScript is to restrict it to a less-dynamic subset (e.g. objects must have static keys) and add at least some static typing so that a compiler can resolve some references ahead of time. That's why asm.js was a thing. But asm.js has been superseded by WASM, which better in most regards.
One problem is that this GC doesn’t interact with the browser’s GC in any way. So you have painful memory management interactions when (for example) a DOM event handler references a WebAssembly object which in turn holds a reference to a browser object, possibly with cycles (so simple reference counting isn’t enough).
You can fall back to manual memory management here, but that’s painful to use. To make this work seamlessly, you need a way to trace references across both heaps in one swoop. Last I checked, there was standardization work underway to enable that.
See :
https://github.com/pyodide/pyodide
And an example featuring a pure client side jupyter instance :
For example, the homepage of Nuitka says they only just added support for constant folding and propagation. That's such low hanging fruit it's crazy they have to do that for themselves.
You can imagine generic libraries for Python/JavaScript/Ruby that will turn the AST alone into its most optimal form and then let some other backend worry about code generation or VM implementation.
Then there is the question of optimizations that are (typically) easier/possible only at the code-generation phase.
All of this is solvable of course, but there needs to be some will to do so (or in the case of companies, enough commercial benefit).
The 2-3x improvements are on specific micro benchmarks.
The main goal is getting a standalone executable that works with c extensions.
At my company, for a web app we run in production, we strive to get every response out of our infrastructure in under a millisecond, everything above that except for a few select endpoints is considered as a bug. By using sensible technology choices, it's not even that hard to do. A RDMS like postgres can answer to queries in microseconds.
Nice bonus, we can provide real time computation features our competitors could only dream of, just by not using dog-slow technology.
Our customers are not techies and you know what? When they use our product, the first comment is usually "wow, it's so fast".
Just to expand on your point:
Sure you can let a chunk of code take a few seconds longer than it could with little overall effect.
But stack a hundred of these up and all those seconds compound and suddenly become immensely important.
A suit of armour with a million chinks isn't very effective armour.
(As an aside, depending on your use case, it can be worth optimizing the big bottlenecks, high use, computationally expensive, etc. components and not worry too much about the rest)
Counterpoint: https://en.wikipedia.org/wiki/Chain_mail
I am agreeing with you, I don’t understand your response.
I remember reading an article somebody was describing that instead of even adopting project to work with Cython he/she outright started their project in Cython and found it beneficial.
Also is this actually faster than just python? I tried it out with some bad expensive looping using lists with python3.9 and it was 100ms slower (1.3sec vs 1.4sec).
Just in general why should I be using Nuitka?
> In the future Nuitka will be able to use type inferencing based on whole program analysis. It will apply that information in order to perform as many calculations as possible in C, using C native types, without accessing libpython.
I believe the difference with PyPy is that PyPy tries to do this using just-in-time compilation and nuitka uses ahead-of-time compilation.
I did try this. With the "stock" python3 (3.8.something) I couldn't get nuitka to compile my code (or the example from the docs). With python3.9 I could start the compilation, but I didn't have libpython-dev shared object and finally with python3.9-dev I was able to make the binary. However after transfering it to another Ubuntu 20.04 machine I couldn't run it since it was linked against libpython3.9-dev.
Just as side note. The same python code (i.e. the .py file) would run on both machines out of the box.
The only dep would be libc, and if you compile against an old version, it should be forward compatible.
Mind sharing that code? I'm sure someone would like to be able to look into this.
python -m nuitka --clang --follow-imports main.py
repo: https://github.com/MRCIEU/gwas2vcf
If someone can make the program run faster by whatever means, it will make a bunch of people quite happy.
import time
start = time.time()
list = []
for i in range(10000000):
list.append(i)
sum = 0
for item in list:
sum = sum + i
print(time.time() - start)
I am not superduper python expert, but I know that arrays/lists are the stupidest thing you can use so that's why I did stupid things with them. So before someone posts more optimized version the code was suppose to be shit since I assumed that the compiler would optimize it.I expected the binary to essentially be
1. start = time.time()
2. pre-populated list since it should be constant e.g. list = [1,2,3,...] or completely remove it since it is not used anywhere
3. pre-filled sum since it too is constant or same as above, removed since it is not used anywhere
4. getting second time.time() and printting the difference
But there was no compilation time optimization. Granted this wasn't promised, but I just assumed it was in there since it is compiling.Have you tried xrange?
And I am fairly sure xrange() is not in Python3
Come to think of it, you don't particularly know what sum = sum + i is doing either. There could be an __radd__ implementation in whatever your custom range spits out that indexes the WWW, renders a Mandelbrot or whatever.
And yes, of course overloading range and __radd__ that way would also be considered stupid, but if the loops were optimized away, it wouldn't be a drop in replacement for Python.
Since the `sum` is never accessed by anyone it can be optimized away and now since the only accesser of `list` is gone it too can be optimized away.
Thus both loops most certainly CAN and SHOULD be removed from the final compiled code. Only reason Python can't do it is because it is not compiled, so it has to actually run through the code at runtime and see what happens.
This is basic static analysis and any other compiler for any other language would do this. This is literally why we have keywords like `volatile` in compiled languages.
I don't understand your first argument at all. Why wouldn't the compiler know what the standard `range()` function does? The second one only makes sense since we aren't type hinting, however that too could be reversed by analyzing the code.
If we can't trust the functions in standard library then what is the benefit of compiling the code?
I'm sure it knows what the standard range() function does. It just doesn't know that that's the standard range function you use when the code is being run. Any Python program is also a module. Any module can be imported. Any imported module's global namespace is unknown at the time you write the program.
Imagine I have your code, above, in a module called nextlevelwizard, and I write my own little progam as this:
import builtins
def myrange(n):
cur = 0
while cur < n:
print(cur)
yield cur
cur += 1
builtins.range = myrange
import nextlevelwizard
Now, suddenly, your loop has a side effect. Side effects cannot be "optimized away" - the whole point of programs are generally to produce some sort of side effect!This is a silly example, of course, but the dynamic nature of Python means that it really is very hard to skip steps you assume are irrelevant.
Even in the example you gave it could replace the loop with a bunch of print calls since the result (yields) do nothing.
Compilers are hard.
Is this a good idea? If I understand correctly, if an HTTPS request fails, it falls back to HTTP. So an attacker could just block the HTTPS request, then MITM the HTTP request.
https://news.ycombinator.com/item?id=28377545
Very nice for deployment and scripting.
If you're willing to give up some compatibility, you could get faster code that can be checked statically by transpiling to one of the many great choices from a set of statically typed languages.
This is the approach taken by py14, pyrs and their successor py2many, on which I've been working for some months.
Anyway what are your incompatibilities?
https://github.com/adsharma/py2many/blob/main/doc/langspec.m...
But, given LuaJIT and V8, it's obviously possible to generate high-performance code without such restrictions.
As for PyPy adoption, it's hindered more by the lack of API compatibility with CPython native extensions.
If you're finding it easy to write a quick command line tool in python, but end up rewriting it in another language for size/performance/self contained deployment, this may be a good fit.
There are several challenges to solve - translating stdlib of one language to another (see plugins.py) and bridge semantic gaps (rust doesn't like mutable global state, python has them).
Without metaclasses and dynamic types I think it's better to call the language something else, and say it's Python-inspired (like Elm is Haskell-inspired)
https://github.com/adsharma/py2many/tree/main/tests/cases
are run with cpython interpreter and verified for compatibility.
Even when there is a desire to innovate (design by contract or pattern matching as an expression), hope to do so without breaking cpython (as long as you stick to the subset).
"compatibility comes at a cost."
And use "python -m nuitka", not "nuitka" alone.
"-m" is a very important option in python that solves a lot of path problems.