Python-based compiler achieves orders-of-magnitude speedups
news.mit.edu
news.mit.edu
"Currently, there are several Python features that Codon does not support. They mainly consist of runtime polymorphism, runtime reflection and type manipulation (e.g., dynamic method table modification, dynamic addition of class members, metaclasses, and class decorators). There are also gaps in the standard Python library coverage. While Codon ships with Python interoperability as a workaround to some of these limitations, future work is planned to expand the amount of Pythonic code immediately compatible with the framework by adding features such as runtime polymorphism and by implementing better interoperability with the existing Python libraries. Finally, we plan to increase the standard library coverage, as well as extend syntax configurability for custom DSLs."
Yes there are already subsets like this, but its not as helpful if it isnt standard.
If you code a website, fast api and django, the two most popular framework to do so, heavily rely on them to make you productive.
Most databases I’ve dealt with will happily outstrip Python for a good chunk of the common queries.
And write SQL directly for medium complexity queries.
Assume your python program is fully static and well behaved from a compiler perspective. JIT compile it, and observe what it does, so that you can invalidate the compiled code and re-JIT it if it does something overly dynamic.
Incidentally, what JVM did. (I’m sure now it’s been tweaked beyond recognition)
It's a compiler for python code that can create stand alone executables, and up to 4 times the speed of the initial code.
Best of all, it's extremely reliable, with a high level of support of event the tricky things like the scientic and gui stacks.
Could it compile an app that uses Pillow and AggDraw and ReportLab and OpenPyXL with a TKInter GUI into a standalone app I can give to a coworker? That would be extremely useful!
And I have to ask, does a powersnail live in a powershell? =) (also a language I like that needs better GUI and EXE packaging)
You can use a superset language of python called cython that generate C code. It can be used to generate C bindings or fast python (for cpython) modules implemented in a python like code.
You can use a really fast language like C, Rust, C++,... create python wrappers with cython, swig, Boost.python, cffi,... and use python like glue code.
Python is not a fast languages as others, but there are tricks to make fast programs.
Converting a few thousand rows to python/django objects takes _time_. I can't quantify anything, because it's been too long, but I remember it being fairly significant. When I profiled it, the majority of the time was spent calling __setattr__ a few million times.
Like you said, it depends on your use case. If your queries are slow, then optimize your database queries. But if your queries are fast and your responses are still slow, then investigating pypy is definitely worth it. You can also play around with .values_list or something in Django, so that you get 'raw values' instead of objects (but there's still a cost to building them up).
Overall top performing frameworks (JS, Java, and Rust) at 650K rps. That's 7x over the top Python based framework at: 86K rps.
And another very popular Python framework (flask) gets just to 2K rps. That's 325 times worse to the best.
And that's the "single DB query" benchmark: https://www.techempower.com/benchmarks/#section=data-r21&tes...
https://github.com/TechEmpower/FrameworkBenchmarks/wiki/Proj...
That was a couple of years ago at this point, and I've not been in the python ecosystem since then, but I can only imagine things are getting better in that regard rather than worse.
If you start with an incompatible, highly performant interpreter, the compatibility "distance" is difficult to measure and could create unknown performance cost. For example, PyPy doesn't support C modules due to the differing memory layout.
* C extensions
* Instagram's forking server model
I gave a talk that touched on some of this last year: https://2022.ecoop.org/details/ICOOOLPS-2022-papers/5/Cinder...
I wonder why the recording is not up...
Quick, transactional HTTP exchanges (GET, POST, etc.) aren't really its thing-- there's no time for the compiler to get warmed up; the request is complete before pypy has gotten out of bed.
But if you have to do really complex view rendering (graphs or something) where it would take cpython ~10s or more to process, then pypy will leave cpython in the dust.
Considering at least 2 people have gone to look at the source and then come here to comment, it would have been a net benefit for all involved. Plus, what does it say about the potential quality of your compiler if you can't even make correct English statements? This seems easier to get right than if( x = *p++ )
2. There are lots of good developers who aren't capable of making any statements in English.
For people with native or fluent English, for sure. For the others, probably not.
I thought it would have higher memory usage? (based only on reading)
code here: https://github.com/avinassh/fast-sqlite3-inserts
my blog post: https://avi.im/blag/2021/fast-sqlite-inserts/
Python does not need to be rescued. Its fast enough. For 90% of applications, you are kidding yourself if you need more speed.
These are enhancements on Python, where you want to run stuff even faster on par with other languages.
Is it "fast enough"? fast enough 90% of the time? Or just fast enough to leave you uncertain that it's even a good choice?
Sorry, we have productive choices now that don't leave me worrying about this situation. I still like it though, and if they could solve those issues I'd probably use it a lot more!
If you are trying to run a lot of executions in parallel where you need the full instruction set of the CPU over a GPU (namely, the compare/jump operations, otherwise known as if statements, for loops, e.t.c), then you are most likely either extremely resource constrained (like writing code for a microprocessor or embedded system), at which point you will still use C/C++, or writing something like a video game, which again is C/C++/C# (or Swift for Apple).
For most every other use case, Python is simply applicable with its vast array of libraries. For example, at my work, we use FastUI with orjson for a backend that needs to handle some significant TPS. Its fast enough. Could we write the entire thing in another language and use less ECS containers/EC2 instances? Sure, and we will save on cloud costs but lose massively on developer costs.
Moreover, people have huge python apps that they can't just rewrite and python just isn't fast enough. This has happened so many times. So many man hours have been spent optimizing python code, that we have over a dozen different implementations in just this thread alone and it doesn't include 3 that I know of.
Python's current Achilles heel is actually it's performance. It's slow as fuck. Those 10% of applications matter. And faster performance won't hinder anything for people writing 100 line scripts.
Execution speed is more than execution speed - you need to be correct, and being fast enough is quality; faster may be wasteful.
Which part don't you like:
That python is relatively slow? https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Or that people are trying to fix its slowness? See the thread you're currently in.
Python is not "slow" in the sense that it not applicable to use when you wanna run real world applications.
Yours is the equivalent argument for buying a BMW M3 over a Toyota Corolla because it can do a faster lap time around a track and thus is better for your commute, except in the real world with traffic and traffic lights, on a 30 minute commute home, you will probably arrive at your destination 1 minute quicker in the BMW over a Corolla.
My concern is: there have been a few projects already that are, from the outside, more or less the same approach and set of tradeoffs as this. And they haven't been that successful. Given that this is treading familiar ground I would expect some words about how this is different, and the lack thereof makes me a bit skeptical to say this will become successful when others did not.
Success of projects like this is not usually based on merit, but on how many people you can convince to go along with it until it eventually becomes a thing of its own sustainability. So, the recipe here would be to:
1. Be "good enough" and easy enough to get started such that early adopters have a great first experience speeding up something important to them. (hook them) 2. Be open and friendly to potential incoming contributors, letting them land changes, have a say in the discussion, and generally be part of it all. (community build) 3. Encourage people to share their successes and hopes / dreams for how great $X is on their blogs, HN, social media, etc. (propaganda) 4. Goto 1.
In this case, step 3 will work best by highlighting that "You actually don't need most of the dynamic features of Python" as the central narrative.
One big caveat is that Codon choose to not use Python semantics for `%` so the basic test of `print(-2 % 5)` fails unless you run it with `-numerics=py`... which should just be the default behavior -- and a great first community patch / discussion!
I wonder how Kim Kardashian programming language looks like. I guess low level but with garbage collector. :D
Hey dolls, let me introduce you to the Kimmie programming language, it's like totally fab and easy to use!
To declare a variable, just use the hashtag symbol and the variable name, like this:
#my_var
To assign a value to the variable, use the word "like" followed by the value, like this:
#my_var like 10
To print out a message, use the word "OMG" followed by the message in double quotes, like this:
OMG "Hello, dolls!"
To add two variables together, use the word "add" followed by the two variables, like this:
#var1 like 5 #var2 like 7 #sum like add #var1 #var2
import re
def trashtalk_interpreter(code): variables = {} code_lines = code.split("\n")
for line in code_lines:
if line.startswith("#"):
var_name, _, value = line.partition(" like ")
if "add " in value:
_, var1, var2 = value.partition("add ")
var1 = var1.strip()
var2 = var2.strip()
variables[var_name.strip()] = variables[var1] + variables[var2]
else:
variables[var_name.strip()] = int(value)
elif line.startswith("OMG"):
message = re.findall(r'"(.*?)"', line)
if message:
print(message[0].format(**variables))
# Sample code
code = '''
#my_var like 10
OMG "Value of my_var: {my_var}"
#var1 like 5
#var2 like 7
#sum like add #var1 #var2
OMG "Sum of {var1} and {var2} is {sum}"
'''trashtalk_interpreter(code)
class KimmieInterpreter:
def __init__(self):
self.variables = {}
def interpret(self, code):
for line in code.split("\n"):
tokens = line.split()
if len(tokens) == 0:
continue
if tokens[0] == "#":
self.variables[tokens[1]] = None
elif tokens[0] == "#my_var":
self.variables[tokens[1]] = int(tokens[3])
elif tokens[0] == "OMG":
print(tokens[1][1:-1])
elif tokens[0] == "add":
var1 = self.variables[tokens[2]]
var2 = self.variables[tokens[3]]
result = var1 + var2
self.variables[tokens[1]] = resultThis language isn't here to make friends.
And exposed naked primitives ... :-)
https://docs.exaloop.io/codon/general/differences
So more limited types (integers) and more type checking and collections have to have one kind of thing in them.
There are other python compilers though, like https://github.com/Nuitka/Nuitka
I wonder really what the advantages/disadvantages of these are?
So aside from that tiny issue at the center of the decade-long Python 2 to 3 migration debacle, it's virtually identical!
Sounds like an issue that could easily bite someone in the behind and cause quite nasty bugs.
The fine print is strong with this one. It makes me wonder why they didn't just start with RPython.
But it still shocks me just how much money and manpower is thrown at trying to bikeshed and optimize and compile Python and its libraries, while the Nim compiler is essentially a community hobby project that has made the concept of a "compiled Python" a reality already. The orders of magnitude in scale difference, and the qualities of the output products, are staggering.
I'm kind of starting to see what Guido is talking about when he says Python is a legacy language that's probably on its way out. Even in the interpreted world, languages like Janet and other newcomers are performing fascinating experiments, often doing more with less.
I checked out Nim a few years ago because I have a large Python project that I'd like to move to a compiled language. In my experience, if you get on the forums and do any kind of comparison between Python and Nim, you will quickly get responses of "Nim isn't Python, so quit trying to make it like Python".
I think Nim would have been much more successful if it was more like Python, and if Python compatibility was one of its goals. Just as an example, a subrange a..b in Nim is closed on the right, unlike Python, and a..<b is the open Python version. They could just as easily have made a..b open and used a..=b for the closed interval, for Python compatibility.
I'm not saying Nim had to be the perfect Python compiler and compile all Python code unmodified. Based on what I've read, Python is too dynamic for this. But in cases where there was the choice to either be compatible with Python or "do something unique", Nim often takes the unique path, and not always for any good reason IMO.
Migrating to a new language is not easy when you have millions of lines of code.
You can adopt it incrementally, but then you could just as well switch to a language with higher default performance, more language features that just work, unified tooling etc. and adopt that incrementally?
Wow, what a way to mischaracterize what Guido said.
His point was about languages evolving to be more abstract than Python or any of the ones you mentioned. Programming is going to become more and more abstract to the point where you will be able to program in natural language through speech. In the mean time, we still have to write code manually.
And look, there are plenty of valid criticisms of Python, but you are kidding yourself if you don't think its going to be one of the primary languages of the future. There is a reason why it has the 2nd most gihub repos (behind JS, because of hard dependency on it for web stuff).
And the simple reason is this: the vast, vast majority of applications don't need the fastest possible speed, its much more important to be able to develop fast, and have it be right. Its easier and cheaper to throw another EC2 instance in your stack rather than pay a developer to write stuff from scratch whereas in Python you can just import the relevant library for your needs and be up and running much faster, not only the short concise syntax used, but also the introspection into the running language because of its interpreted nature. And this allowed the snowball effect to happen, where developers could quickly write relevant libraries, which in turn allowed other developers to quickly import those libraries and write their libraries, slingshotting Python into a language that is used not only for bleeding edge ML stuff but to run backend web stacks with no issues.
And in the cases where you do need speed, this is where these compilers come in, and its a 100% valid use of manpower and money. Think of it as another library.
Every other language that focuses on things like static typing, whatever type of inheritance the designers think is best, memory safety, and all the other theoretical CS stuff completely misses the above point, and for that reason alone, it will never become mainstream. Rust is not going to happen, Nim is not going to happen, Julia is not going to happen, Scala is not going to happen, Elixir is not going to happen. Sure, there will be a significant amount of code written in those, but the popularity will never come close to Pythons. You may not like it, but you know this is true.
We have already seen this cycle happen with Haskell where functional programming was the next best thing. you would constantly see posts about it at the same frequency you now see posts about Rust, and look where Haskell is now.
The problem with Rust is that they have the unsafe operator. When using a 3d party library, I have no idea if someone put a bunch of unsafe code in there, so all memory safety guarantees go out the window. Sure, you can grab the raw source and compile it yourself, but then that introduces a whole bunch of friction into the dev process.
And the reason unsafe is in Rust is because you can't write standard library stuff, especially with performance in mind, using traditional Rust constructs.
In the end, Rust doesn't give you anything over a compiled C extension to Python, that can be written as memory safe in the sense that it just receives a buffer of data to process with preallocated memory, runs said processing, and returns the data. This is pretty much the standard way that ML works except the compiled extensions just get put on the GPU rather than CPU, and the overhead of the translation layer is extremely small in comparison.
It's not a Python replacement in any sense of the word.
https://github.com/facebookincubator/cinder/
Disclaimer: I used to work on it.
The BigInteger "issue" pretty much makes something like Fibonacci a worst case scenario for it.
Not to nit-pick...this has been characterized by a team who tested and compared a large set of languages against a wide range of application code. The number is, if I remember correctly, about 78x slower. I don't think "orders" of magnitude is entirely fair. Yes, Python is slow. I have made the mistake of trying to use it for time-critical embedded applications. Never again.
Aside from this admittedly pedantic observation, the first thing that crossed my mind with regards to this tool --which sounds fantastic-- is that you would have to trust the correctness and reliability of your code to this translation layer. Not sure how to think about this other than to keep a mental note of it if using this tool.
In other words, the acceleration isn't measured against raw C implementations (where the 78x factor I quoted would be relevant). It is measured against Python or PyPy.
How much faster does Codon make your Python code. The answer seems to be somewhere around the 5x to 10x range.
In that context, and in the context of actual applications rather than hand-picked tests (how much can we optimize a loop), "orders of magnitude" seems to be an exaggeration.
BTW, MIT does this kind of thing all the time with their press releases. They have a brand to support with outlandish claims about everything that comes out of there. Those with frequent exposure to this kind of press release are wise to this. I've seen it for decades. It's marketing.
For me, when someone says "orders of magnitude" it means "massive". I tend to say "10 times faster", "50 times faster" even "100 times faster". I probably start using "orders of magnitude" faster at 1000x or when I am trying to explicitly make an impression on a mathematically-challenged audience. "Orders of magnitude" sounds great to that crowd.
I have never, in 40 years in CS/Engineering, heard anyone use powers-of-two when they say "orders of magnitude". Doing so would open you to serious misinterpretation. Engineers might say something like "a factor of 2 to the n" or something like that.
Likewise, modern low level languages should have syntactical conveniences and optional whitespace.
It’s 2023. We don’t need to keep having this war.
I don't like it for my use cases, but whenever I read about it on HN it's supposedly the best-tooled, finest artifact of performance engineering ever built.
I mean AFAIK the hard part of Python is that the language allows dynamic overwriting of attributes (or something like that). Is that feature actually needed for projects like Django, FastAPI, numpy, etc?
Maybe I'm wrong, but the main idea I'd like to ask is, can we make a compiler for a subset of that language with C-API compatibility?
However, maintaining C-API compatibility means you need to set up lots of data structures exactly how the C API requires, and maintaining and updating those ends up losing you lots of your benefits of JITing.
You could, hypothetically, introduce an entirely new API, which allowed for faster dynamic recompiling, but then you'd need to get every package anyone cares about to switch to that.
So you can interpret (and later AOT compile as well) LLVM bitcode and python, and this approach will allow cross-language optimizations as well, which were not available at all before. But feel free to add a bit of JS/Java, etc to your code as well!
Sure. Crossing that FFI boundary is going to be expensive. But there’s lots of techniques to mitigate it or even in the limit eliminate it. if I recall correctly you can JIT a fast call that knows how to invoke the FFI directly without the extra indirection layer. Basically a fancy runtime LTO.
I think a huge part of it is CPython’s interest in keeping the core codebase as simple as possible which seems to be the overriding reason for why the global lock still hasn’t been removed (which iirc even Ruby pulled off at some point). Also the reason there’s no JIT afaict and why Pypy got started to prove it is possible to JIT (and frequently sees substantial gains vs cpython). The problem they’ve had is that CPython is a moving target and it’s hard to keep a parallel runtime up to date on a shoestring amount of funding. That’s why you see alternate approaches like numba (JIT’ed Python) which are less of a departure and Cinder (better budget). To me this seems like a CPython project actively hostile to JIT than C data structures meaning you lose some benefit to FFI overhead. Performance is a virtuous cycle too - when there’s enthusiasm about a language you get more and more people paid to make your language fast. For a while companies tried. Google gave up. Facebook only has it as a fork with a public plea for the maintainers of CPython to mainline literally anything.
The CPython maintainers feel like the biggest obstacle. No?
For example, JNI only exposes handles and you need to convert an handle to a pointer, so the runtime knows for the time being that handle is special and being used by native code.
When it is only an opaque handle, lots of optimizations can happen and the native code won't see them.
Don’t get me wrong. I’m not passing a value judgement on the maintainers. But the reasons don’t feel technical to me.
Normally I'd say you mean the interface you use when you call native machine code from Python, but I don't see how this would slow things down.
Yes.
FastAPI depends on really slow pydantic (disclaimer: I'm the author of the faster typedload).
All those dynamic typechecking modules rely on the dynamic nature of the language. The alternative would be to having to generate code at compile time instead.
pydantic is also in the process of being rewritten in rust to be not so slow any longer, and in the process it will become incompatible with anything else than cpython (the normal python runtime). Which in turns means fastapi won't be able to run on anything else (unless they decouple from pydantic… which probably won't be easy).
Check out Pythran, that is exactly what they've done.
That's exactly how PyPy works.
PyPy is a Python interpreter, a drop-in replacement for CPython 2.7, 3.8 and 3.9"Oh... a native compiler. That cheats by not really honoring Python. Got it."
I hope this time we will see better results.
To share objects between threads, some synchronisation is needed, for example to update reference counts. There are a few ways to do this:
- make the user add locks; the problem with this if it goes wrong it can crash the interpreter and make it impossible to debug the problem from within python, which is not user-friendly, and lots of existing code will break. Competent users are already doing this though, so it's nearly free. - add fine-grained locking/synchronisation for object internals within the interpreter. This slows everything down, even if you're not using threads. - Lock the whole interpreter state whenever a thread is running. This makes threads less useful (no speed-up from threading pure-python code that isn't doing IO; you have to use multiprocessing for that), but it's cheap as you only need to lock/unlock when you're doing something slow anyway (IO, thread switching, native code).
I think this explains why GIL removal has not been successful yet despite much work: the alternatives slow down single-threaded code, which is not worth it when nearly all sensible uses of threading don't benefit either.
main.py:15:1: error: syntax error, unexpected 'async'
One must note that this is impossible, unless you have chosen to handicap the C-implementations while benchmarking. Borderline unethical IMO to put forth such a claim.
So not impossible and therefore not unethical.
I mentioned JIT because it seems to be based on a similar principle at least, that of optimizing things on the programmer's behalf by looking at the program's usage and not just by looking at how to speed up the code generally.
Ok, but that's not what they are claiming - their claim (at least based on what the article is saying) is more about one toolchain vs another, i.e. "if you use our compiler (that takes python code as input) then the resulting executable will run as fast (or possibly faster than) programs created by all the popular compilers (that take C/C++ code as input)." The sales pitch is that they've got magic sauce in their compiler, and you get to use Python as well.
Humanity wasted close to 50 years optimizing compilers for one garbage language. Wasted unimaginable efforts, money and developer hours... and all could've been avoided if the same people dedicated a fraction of those resources to language design.
Same thing happened with Java. And now the existence of a well-developed compiler became an argument in its own right in favor of choosing a bad language.
There's no need or reason to try to make Python run faster. It's a trash language. At best, it deserves a credit for being funny 30 years ago... but that had worn out pretty fast. Now it's just dumb. Improving its compiler will be again a resource sink for the programming community that, in the best case, may hope to produce something of value by accident, independent of its main goal...
It works, it doesn't confuse me, it's easy to find libraries and examples, and when it's too slow -- which is surprisingly rare -- I have other options to turn to. If that's trash then so be it; call me a raccoon because I'm there for the trash.
Name one. I'll tell you why it's trash.
> Clearly it's at least good enough
Trash (literal, from the dumpster) is often good enough to eat. What's your point exactly?
> I find it much more pleasant to use
It's a skill issue.
> it's easy to find libraries and examples,
This has nothing to do with the language. Give Julia 20 years of popularity, and it will have just as many useful libraries and examples.
What do you mean by language design here? Is it the user-facing bits and the ergonomics? Because it seems to me (as a non python dev) that that's the bit that python devs really like.
Fast food is universally more popular than any healthy food that requires time to cook.
People choose to buy low-quality goods in general, trading quality for immediate effect all the time. Take any industry, any product kind, you will see that consumption is skewed towards paying extra for immediate gain rather than paying for quality to minimize waste over time.
The fact that you chose to rely on the opinion of non-experts in the field to assess the value / quality of a particular technology only means that you don't understand what quality is about. You are confused between wants and needs.
Python has several disjoint domains where it's used. So, for example, when it comes to statistics, then J, R or Julia would all be better than Python. Not ideal, but still a lot better.
When it comes to infrastructure and ops, then Erlang would've been a lot better. Still not ideal because of how existing implementation deals with deployment (too complicated), but that's not a feature of the language, and could be worked on in the same way how OP wants to work on Python compiler.
When it comes to Web... well, I'm not a specialist... and I find everything about Web revolting, so it's hard for me to think about alternatives. On the other hand, there's rarely a language that doesn't come with a Web framework / some tools that allow it to be used to make Web applications. So, just, basically, throw a dart, and wherever it lands it's going to be better than Python with a very high probability.
Python is also used to teach intro to computer science. And there's a lot of problems with this idea. Firstly, I don't believe that intro to CS should be taught by way of learning to program. It should give an overview of what CS is about, give some foundation, basic concepts from important fields... just like intro to math does, for example. But if we still have to have intro to CS the way we do today, then Scheme would be a lot better for this. Assembly is also a good pick, but for a different reason.
Now, to address the "in general" part: I don't believe that languages like Python should be used universally in different domains. What I believe we need is a language like OMeta, which we specialize for the domains we want to program in, so that we can keep language and compiler mechanics separately from syntax of specific domains. Ironically, this was even obvious at the time of ALGOL design, but nobody waited for it to be implemented and went with the quick-and-dirty solution instead.