Nuitka: a Python compiler written in Python
github.com
github.com
* FBS (FMan's Build System) - worked well for a while, but we couldn't upgrade Qt past a certain version and he version we used was only happy with Python 3.6.
* Beeware - our current solution. Very, very happy with this; works smoothly, lets us use the latest Python and Qt versions and lets us make professional-looking installers that we can distribute.
Buying the Pro version still leaves one to deal with Qt licensing does it not?
I was suggesting that, for you to make more $$$ from it, its better to provide a free trial period for the said version.
It's documented very well, extremely easy to use and the pro version is priced very well.
Question - do you have any comparisons on the size of the resulting distributed executable compared to Nuitka, pyinstaller?
This is the issue I've faced. Nontrivial programs (that, say, import pandas, numpy, selenium) are huge, implying a ~250MB+ download for the user - and then maybe up to 1GB in disk when its all unpacked.
I sent him an email years back, just thanking him for his work and asking how to donate. I was surprised when he wrote a very thoughtful and kind reply, acting rather shocked why someone would give him money. He's a real gem of a person, and I'm glad his work is getting (rightly) recognized.
For individual functions you can see 2-4 times speed up over pure python.
Most importantly, I've never seen it result in slower than cpython performance.
Yes, but so much of the Python ecosystem uses the C API that the "last resort" and the "common case" are one in the same most of the time. `psycopg2cffi` IIRC that's not very well supported. It looks like it was updated last in January of 2021 but before that the last update was from 2018. This was what prevented us from using it in 2020. Moreover, there are other packages besides Postgres drivers; that's just the one that I recall running into problems with. To be clear, I want Pypy to be successful, and the project is nothing short of amazing.
> I'm pretty sure I saw other PEP 249 implementations for PostgreSQL the last time I checked (which admittedly was a few years ago).
There was a pure Python version, but it didn't seem battle-tested and there was no indication of its quality or performance. I don't want to pull a package like that into production for something as important as a database driver. It's been a couple years since I looked into it as well, so perhaps things have changed for the better in the interim.
Yeah, that's shame, because 1) that API is mildly awful, and 2) it really restricts implementation choices to the extent that making another implementation of the language is a problem just because of the need to fake to numerous existing C extensions that your implementation choices are the same as CPython's (when really your implementation is likely to work very differently on the inside). I guess Python people really programmed themselves into a corner here.
It's not too different to programming the CPython API, and will give optimum code in Pypy and CPython.
I didn't notice any particular speedup for this app - manipulating lots of JSON Python data structures is not really PyPy's sweet spot. Whereas for some processing of large genomic data files I saw a substantial speedup.
It can be a much broader term, for example you will run into people here arguing (with some justification!) that a "transpiler" doesn't exist since it's a mere subset of the broad definition of a compiler.
I'm just reporting on the mental image that calling such-and-such a compiler forms in my mind's eye. I expect I'm not alone in that.
When I last tried Nuitka I couldn't get it to work with pycryptodome. Sounds like I should maybe give it another go.
I’ve chosen Go rather than Python for a few small projects recently. Not for performance, parallelism or a particular fondness for the language, but just because Go can build truly standalone executables.
I’m still wary of that experience and will avoid Python where I have such deployment needs unless the language natively comes up with such a build solution.
FWIW, I think borg(backup solution written in Python and C) uses pyinstaller to get a single binary executable. It may be of interest to you.
> The created binaries can be made executable independent of the Python installation, with --standalone and --onefile options.
The proposed solution by Pyinstaller is to build your program in the "oldest" version you support, which is silly compared to build systems like Go's, which are true standalone.
Having said that, I'm curious about this solution, since it seems to claim a true standalone build...
Unfortunately this might include libreadline.so, which is licensed under GPL, making your resulting executable unable to be under a proprietary license.
There are ways to solve this issue, but one has to search and read documentation (and code, in my case -- when I was researching it the docs were not clear).
Yes - and last time I used it, it created either a large folder or a compressed archive containing all of those libraries. Only the latter gives a truly standalone executable - but it's very slow to start up because it has to extract the archive to disk every time it runs.
It sounds like Nuitka has a solution for this problem, at least on Linux: "[the binary] will not even unpack itself, but instead loop back mount its contents as a filesystem".
Still doesn't statically link C libraries (or at least I didn't find the setting for it), or other libraries for that matter.
Pyinstaller binary build depends only on: libdl.so.2, libz.so.1 and libc.so.6.
Nuitka binary build depends on: libdl.so.2, libz.so.1 and libc.so.6 AND libpthread.so.0 (for the loopback mounts I suppose).
The one that always creates problems is libc.so.6, which usually is not present 4 year old systems...
As an example, here's a little project which I also release as a single binary file that works everywhere I've tested, with no deps:
Metrics are fun!
It's really not that hard to build standalone binaries for Mac, Linux and Windows using GitHub Actions. Especially if you use a reasonably modern language like Go, Dart or Rust.
> so all you've accomplished is going from O(n*2) files to O(n)
You go from N*M*2 where `M` is the number of target platforms to M (no need to use big-O notation as far as I can tell).
Nuitka 0.6.0 released - https://news.ycombinator.com/item?id=18092837 - Sept 2018 (14 comments)
Nuitka: A Python compiler - https://news.ycombinator.com/item?id=17683932 - Aug 2018 (6 comments)
Nuitka: A Python compiler - https://news.ycombinator.com/item?id=16980704 - May 2018 (4 comments)
Nuitka: a Python compiler - https://news.ycombinator.com/item?id=15354613 - Sept 2017 (60 comments)
Nuitka Progress in 2015 – Python Compiler - https://news.ycombinator.com/item?id=10994267 - Jan 2016 (52 comments)
Nuitka: a Python compiler - https://news.ycombinator.com/item?id=8771925 - Dec 2014 (135 comments)
Nuitka — A Python Compiler - https://news.ycombinator.com/item?id=1746340 - Oct 2010 (33 comments)
> Future
> In the future Nuitka will be able to use type inferencing based on whole program analysis. It will apply that information in order to perform as many calculations as possible in C, using C native types, without accessing libpython.
They already claim a 300% speed boost, but with those type inferencing things, it will probably go a lot of further... Can't wait for this project to mature.
I guess compiling to WebAssembly should be trivial then?
> The mypy project has been using mypyc to compile mypy since 2019, giving it a 4x performance boost over regular Python.
Does anyone know how mypyc and Nuitka compare in practice?
Disclaimer: just a happy user.
I'd be looking more into this, but anyone else knows any other dynamically typed language which compiles to native code, not JITed or bytecode (I remember there were few Lisps out there).
If I could find the message I sent to the Nuitka mailing list (or did I suggest it via email? Can't remember) I'd publish that here.
;)
What did I say which was wrong? Honestly thought it was a mildly interesting anecdote!
Sheesh! You people! :)
Julia supports both a REPL and AOT compilation.
There are some Forth compilers for various systems. https://www.thefreecountry.com/compilers/forth.shtml
I wrote a small app, which is simply:
autopep8tool.py:
import autopep8; autopep8.main()
Then had this .bat file to do the .exe for me: autopep8tool.make_exe.cmd:
@echo off
pushd %~dp0
call "%VS140COMNTOOLS%\..\..\VC\bin\amd64\vcvars64.bat"
p4 edit ..\tools\autopep8tool.\*
p4 edit ..\tools\python38.dll
for /d %%i in (*.dist *.build) do rd %%i /s /q
echo import autopep8; autopep8.main() > autopep8tool.py
call nuitka --standalone --assume-yes-for-downloads --unstripped --windows-dependency-tool=pefile autopep8tool.py
copy autopep8tool.dist\autopep8tool.\* ..\tools
copy autopep8tool.dist\python38.dll ..\tools
for /d %%i in (*.dist *.build) do rd %%i /s /q
popd
works great so far, and even that the produced .exe is 10mb + 4mb python38.dll, it's great logistical saver (IMHO), as I can rotate the tool to other teams, depots, without requiring proper python (which is even trickier on windows). It's probably not completely sandboxed (hermetic), but works well so far.calling autopep8.exe --help even gives the right help.
So great tool (though I've been told about others, like google's subpar, etc.)
Granted, my case is really not general one, but nuitka fit well there (like I really wanted to avoid another 'freeze' tool that unpacks python to a temp folder, and calls it from there just to run autopep8.py).
That to be said, I was not able to get nuitka working with some of the other popular format lib (I'll need to look again, it takes a bit of setup)
You can also prepend the python shebang before the ZIP and mark it as executable in the filesystem. youtube-dl gets distributed like that.
I need to, because otherwise I wouldn't be able to package it for Windows users.
The app is also closed source.
I cannot easily package the application for Linux users, because of so many different flavours of Linux, so many different combinations of Python and Qt versions.
Which is why I've given up on a pure Linux version and I'm in the middle of adapting it to work under WINE alongside Elite: Dangerous also running under WINE, instead.
Using Pyinstaller it works out in the end, but after everything, it would have been better to use a language meant for user-facing applications.
edit: including virtualenvs and conda environments have 41 python.exe files on my workstation right now.
https://news.ycombinator.com/item?id=8772124
GvR and the old boys currently present themselves as CoC proponents, and strategically use the CoC against anyone who criticizes them.
Compare with his current Microsoft sponsored activities. If anyone would use that language against "his" (i.e. Mark Shannon's) project, they'd be out and suffer public defamation in no time.
I haven't followed Python dev, so can't comment about how the CoC is being used in practice.
In tough situations, CoC powers are inversely proportional to the height of your position on the totem pole.
I myself pronounce it so it rhymes with and sounds identically to "cock", for several reasons.
The quote was about a presentation in 2013, not the project as it is now. Guido's complaint is mostly that they were equating "compiled" with "performance" and were not rigorous in backing that up. The presenters may have been young. But I'd have made the same complaint. The blog thing was just snarky.
That was a period of time when he has admitted he wasn't doing so well as a human being and this is pretty low on the snark-meter compared to other BDFLs.
The project page as it is today says very little about performance 'wins' as I read today. In fact, it points out losses, in the same vein as the RedHat issue with Python linked to libpython posted the other day here. Their point doesn't seem to be "make your code go faster", but as an easier mechanism of distribution for some cases.
I'm not getting the "his" shot. I haven't seem him take credit for the work, just that headlines do because journalists are lazy. I may have missed something.
But the MS team's activities today are things he would have dismissed years ago. He's learned and changed his mind. If he was completely consistent with himself from 2013, I'd think he wasn't that bright. (I'll never understand those kinds of arguments about politicians)
You're drawing a line from criticism of a 2013 paper, through his current career, to a malicious intent or conspiracy. That's a lot more defamatory than a valid (but a little rude) critique. You're attacking a person's character, not their work output.
That is pretty unacceptable.
> In fact, it points out losses,
Because they have cleared away enough chaff to start encountering these problems with the design of the cpython code itself.
There isn't much point in saying X is faster if significant harassers in the python community will make a point of walking out of your talks and later belittling your work online, right?
I don't think so. Call me cynical, but I suspect that Guido went where the money was (one does not simply go out of retirement). Otherwise there would be an apology to the projects he's been consistently berating for years.
Good thing Kay Hayen has the perseverance and quite mad planning skills. Today we have a compiler that works better with each release and is capable of delivering results today, while van Rossum's work at Microsoft as of now is largely vaporware I don't have very high hopes for.
$ python -m pip install nuitka
$ python3.8 -m nuitka hello.py
install Nuitka with no problems and compiled "hello world", also with no problems.
It created a 4,060,104 byte hello.bin with no symbol table, so no need to run strip to make it smaller, which again worked like a breeze.
(The developers say on GutHub they will work on smarter ways to keep the dependency list small, which can at present escalate for modules like pandas that themselves have >1000 dependencies.)
Def worth a look for anybody that wants to distribute python developed system without installing python or the inevitable many dependencies of projects. It just makes distribution a whole lot more predictable.
And to the reader who asked about GraalPython its on my todo list - have it installed but life, children, and work keep getting in the way.
I've also used it in a production estate, and it works wonderfully well.
Would recommend!
From my experience some libraries are hard to compile (e.g. BeautifulSoup) but not impossible they need some tweaking.
Sometimes I struggle with embedding data for some reason (especially since I'm spoiled with Go 1.16 `// embed`).
Presumably this just bundles a python runtime and a copy of the dependencies and hides it all inside some sort of huge file?
Basically Graal Python and Nodejs each provide a custom interpreter for the target lang, the main goal of which is to provide interoperability with the Graal polyglot ecosystem. So you can run your python code under the GraalPython interpreter and it will run JITed fast and can import libs from other Graal-supported langs.
But as far as outputting executable binaries, Graal only provides that for JVM projects and LLVM languages like C/C++/Rust.
So it's not impossible, but you have to build your own Java wrapper project that loads the Graal Python interpreter class in code and then runs your python lib inside that.
I expect eventually that boilerplate step can be automated as part of the Graal build tools.
If I've got this wrong I welcome any corrections!
clearer, less-garbled version as an answer here: https://stackoverflow.com/a/67331258/202168
I don't even use the dynamic nature of Python - monkey patching etc.
I wish more people wrote algorithms in Python for understandability then compiled them into a compiled language for performance. I find that Python resembles pseudocode because it is syntactically sparse, unlike C++. In Python I don't need to worry about ownership or memory. It probably needs Pointers to be a low level system language.
I just wish Python had parallelism.
I've tried multiprocessing but I found it buggy for my use case.
I tried to create a worker queue and submit items to worker threads with JoinableQueue but eventually the system deadlocks and no progress is made. I'm not sure what the problem is. It never finishes.
It's a sentence correlator using multiprocessing - it generates correlations of words in your sentences.
https://github.com/samsquire/notebook/blob/master/sentence-c... https://github.com/samsquire/notebook/blob/master/workers.py
Maybe somebody can catch what I'm doing wrong.
edit: it may be because i've forgotten to run task_done on joinable queue. edit2: yes it was that! it works!
Python does support parallelism, so long as you're calling into a C library to do the work, because they (usually) release the GIL while doing the work. That sounds like a cop out but using C libraries is very common to do CPU intensive work anyway. Examples include numpy (e.g. you can call np.dot from multiple threads in Python and they will genuinely use multiple cores) and all the rest of the scipy stack, and even modules in the standard library (e.g. when using zipfile, even to process an in-memory buffer, the GIL is released).
You probably do not want to do that in the first place. The reason that the GIL still exists is that the overhead of fine-grained locking for Python datastructures would negate the speedup from multithreading.
Other language (like GoLang) do support parallelism but they have a more advanced runtime and don't use a GIL.
In the case of Python, the single-threaded implementation was already there, the implementation that removed the GIL in favor of fine-grained locking failed to deliver the goods.
If you want to actually take advantage of parallelism, you want data structures that are amenable to it, which is basically sections of arrays without a lot of potential read/write conflicts over multiple threads. If you're at that point, you probably want to use C/C++/Rust anyway, you don't want speed up something that's already 100x slower (the interpreter) by parallelizing it. Python offers that with C-extensions like NumPy.
Not if you're reading most of the time and are using a modern concurrent GC.
I have definitely worked on Python projects in the past that were easily parallisable and worked absoluately fine. I simply created threads with threading.Thread, passed messages around with queue.Queue and used some standard Python modules to process those messages. The messages were independent enough from each other that I didn't have any issues with marshalling access to them, while being large enough that the overhead of mulithreading was still outweighed by its benefits. Rewriting the whole thing in C++ (etc.) would have made negligable difference to performance.
I can certainly imagine a program similar to what you describe, where there are lots of messages that operate on a single monolythic data structure, which couldn't easily be made to work with true parallelisation in Python.
I don't know what they are but I think most people want a single invocation of np.dot in a single thread to use multiple worker threads (as described in https://scipy.github.io/old-wiki/pages/ParallelProgramming) by pushing the work to a thread-aware matrix library.
Usually, this will be safe because np.dot only reads from its arguments, and produces a fresh new array to hold the result.
I just checked the documentation and found there is an out=argument to pass a destination array. (In any case there are other numpy functions that do modify their argument.) If you pass the same array to that parameter in two different threads then I don't know what happens. Maybe numpy locks the GIL to write its output. Maybe it produces C-style undefined behaviour (this would be my guess ... and my preference).
> I think most people want a single invocation of np.dot in a single thread to use multiple worker threads
Maybe? I'm not convinced that's what I want! But it's pretty irrelevant either way to my point; np.dot was just an example. My real point is that lots of CPU-heavy Python functions are implemented in C and release the GIL, so in practice parallelism with Python is often possible.
99% of people want this: np.dot and others use all the cores on your system by default. This accelerates the common use case, which is individual user on a multicore machine who wants single result ASAP. Then, there's an env var that lets you configure the underlying matrix library to use only a single (or maybe another number) of cores, and you work with the multithreaded/multiprocess libraries to run many independent calculations (possibly on shared read-only source data); then that scheduler has a tunable set to saturate the cores or IO on the system.
This is what most BLASses use and it's how data analysts and supercomputer folks can get along.
> I've worked in this field awhile.
This statement undermines your point more than it supports it. There is no "this field" for numpy. Just working in some field that uses it is no qualification for knowing whether users as a whole would like it multithreaded automatically. "Data analysts" and "supercomputer folks" no doubt include many users but certainly not all and possibly not even most (for all we know).
> 99% of people want this: ...
I can believe this might be true. I don't think I ever claimed otherwise. In fact I said "maybe"! I only claimed that I don't want this.
I will make one more claim though: I don't think there's any way that you could know this, even if it's really true.
But yes, I agree and already knew that it's how BLAS libraries work by default, including OpenBLAS which is what the prebuilt numpy wheels use. Even then, as you're probably already aware by the sounds of it, that's only for sufficiently large matrices. If you're just using small matrices and vectors (e.g. in 3D to represent real-world coordinates) then processing will be done directly in the calling thread (because the overhead of using multiple threads would be greater than the benefit).
Actually, you can use Numba to get the same effect while still writing Python. Numba allows you to simply apply a decorator to critical Python code and it is JITted into C-like performance. I do this in my code, and it has worked great for my uses.
Would support for parallelism have benefits other than increasing speed? If not, I'd rather see an increase in Python's single-threaded performance - which is what Guido and his colleagues at Microsoft are currently working on.
I wonder if Python will ever take the JVM approach of having a TemplateInterpreter and mapping each bytecode to assembly directly. Of course you would have to do it for each platform.
http://openjdk.java.net/groups/hotspot/docs/RuntimeOverview....
I think you object to the name python3, because there is a duck typed interpreter that's popular.
There is demand for iterative creation of software: first get the logic out with the least friction and then think about types, resource leaks, good software engineering practices etc.
pip3 install py2many
py2many --nim=1 foo.pyStandard ML was where I started. As a language it's great, as an ecosystem it's limited. OCaml seems to be where the action is (and even then its packaging/dependency management isn't great - but then it can't be worse than Python's).
F# has a very good reputation but I haven't used it a lot myself.
A pure python implementation of python stdlib ready be transpiled could be that bridge.
Also it's possible to support multiple source syntaxes in such an ecosystem. ML based, python based or even a hybrid. In the end all I need is a well documented AST.
In this case, using Numba for the performance-critical aspect was trivially simple, just the addition of a decorator. And I've found that is generally pretty close to true, at least for the kinds of things where I care about performance.
That's why I'm still in the Python world rather than the OCaml world for my algorithmic work. I really like OCaml as a language and would actually prefer to use it, but Python/Numba seems to give me substantially better performance, as well as all of Python's standard libraries.
And Numba lets you turn off the GIL, so you can get multiple cores going at once.
YMMV, of course. I'm not making a general claim that would be true for everyone. But I think it may be worth mentioning that it could be true for more uses than one might assume.
\s? :D
I understand that you (and I for that matter) can't stop using python altogether rn. Just wanted to point out that these two desirable qualities where build into Julia from the get go.
My gestalt sensation is the community is still too small, the tooling too unpolished, to be ready for anything production-worthy. If there were a bulletproof 0rpc library, I think it would go a huge way towards growing the community and mindshare through successive approximation. But the other problem is many of the folks on the forum seem totally disinterested in this problem.
There's a "plugin" option for numpy, in Nuitka, that makes it work. There are similar plugins for a handful of packages.
https://github.com/Nuitka/Nuitka/blob/develop/Standard-Plugi...
The list of plugins is:
dill-compat Required by the dill module
eventlet Required by the eventlet package
gevent Required by the gevent package
multiprocessing Required by Python's multiprocessing module
numpy Required for numpy, scipy, pandas, matplotlib, etc.
pmw-freezer Required by the Pmw package
pylint-warnings Support PyLint / PyDev linting source markers
qt-plugins Required by the PyQt and PySide packages
tensorflow Required by the tensorflow package
tk-inter Required by Python's Tk modules
torch Required by the torch / torchvision packages