Python 3.11 vs 3.10 performance
github.com
github.com
So I remember a guy recently came up with a patch that both removed GIL and also to make it easier for the core team to accept it he added also an equivalent number of optimizations.
I hope this release was is not we got the optimizations but ignored the GIL part.
If anyone more knowledgable can review this and give some feedback I think will be here in HN
Personally, cosmo is one of those projects that inspires me to crack out C again, even though I was never understood the CPU's inner workings very well, and your work in general speaks to the pure joy that programming can be as an act of creation.
Thanks for all your contributions to the community, and thanks for being you!
And most people do care for single-threaded speed, because the vast majority of Python software is written as single-threaded.
This is a self-fulfilling prophecy, as the GIL makes Python's (and Ruby's) concurrency story pretty rough compared to nearly all other widely used languages: C, C++, Java, Go, Rust, and even Javascript (as of late).
As a python dev, pythons multiprocess/multithreading story is one the largest pain points in the language.
Single threaded performance is not that useful while processors have been growing sideways for 10 years.
I often look at elixir with jealousy.
I'd really appreciate if python included concurrency or parallelism capabilities that didn't disappoint and frustrate me.
If you've tried using the thread module, multiprocessing module, or async function coloring feature, you probably can relate. They can sort of work but are about as appealing as being probed internally in a medical setting.
The ergonomics are such that it's not difficult to use.
Why can't or shouldn't we have a mechanism comparably fantastic and easy to use in Python?
It is architecturally comparable to async Javascript programming, which imho is a shoehorn solution.
https://github.com/colesbury/nogil/
Interesting article about it here:
https://lukasz.langa.pl/5d044f91-49c1-4170-aed1-62b6763e6ad0...
Since c-extension wheels are basically built for single python versions anyways, this is potentially manageable.
If only people had asked for it before the Python 3 migration so it could have been done with all the other breaking and performance harming changes. But no, people only started to ask for it literally yesterday so it just could not ever be done, my bad. /sarcasm
If anything the whole Python 3 migration makes any argument against the removal of the GIL appear dishonest to me. It should have been gone decades ago and still wasn't included in the biggest pile of breaking changes the language went through .
I would not be surprised. It is highly likely that the optimizations will be taken, credit will go to the usual people and the GIL part will be extinguished.
When you get an error there's a verbose explanation you can ask for, which describes the problem, gives example code and suggests how you can fix it. The language has a longer ramp-up period because it contains new paradigms, so little touches like this help a lot in onboarding new devs.
To be fair, if decorating the error with information about a filename is what you needed, since Rust's Error types are just types nothing stops you making your function's Error type be the tuple (std::io::Error, &str) if you have a string reference, or (though this seems terribly inappropriate in production code) leaking a Box to make a static reference from a string whose lifetime isn't sufficient.
Traceback (most recent call last):
File "calculation.py", line 54, in <module>
result = (x / y / z) * (a / b / c)
~~~~~~^~~
ZeroDivisionError: division by zero
In this new version it's now obvious which variable is causing the 'division by zero' error.https://docs.python.org/3.11/whatsnew/3.11.html#enhanced-err...
Stack trace is useful, especially in understanding how the code is working in the system.
But if the goal is solve the problem with the code you've been working on, existing traces are way too verbose and if anything add noise or distract from getting to productive again.
I could see tracebacks get swifty terminal UI that shows only the pinpointed error point that can be accordion'd out to show the rest.
But: it's kind of a virtue that the most important error message comes last. Far too many (ahead of time) compilers output too many errors. The most likely error to matter is the first one, but it is often scrolled way off screen; the following errors might even just be side effects that mislead or confuse the new users.
The old Python stack trace, circa 2015, was the best in the world.
The new ones are even better.
The gap between Python and Rust, and the JVM languages/C/C++, is increasingly widening.
Stack trace is one of the areas Python and Rust are making the case for why they're the languages of the future.
Clang and GCC error messages have come a loooong way in the last 10 years or so and their quality is impressive.
To be fair, I'm currently stuck at python 3.6 ATM and I hear that python has also improved a lot since.
Although, personally, I enjoy python list comprehensions.
[x for x in y if x is not z]
Sometimes though, facilitated by Jupyter's notebooks making executing as you build the code very easy, I create a beast like this: [os.path.join([x for x in y if x is not z]) + '/bin' for a in b if 'temp' not in a]
Yes this is fictional and probably contains an error but you get the point, you can filter lists of lists like this, but it's really unfriendly to anyone trying to understand it later.Anything more will come back to bite you later.
That's one expression because it used to be part of a giant comprehension, but I moved it into a function for a bit more readability. I'm considering moving it back just for kicks though.
My philosophy is: if you're only barely smart enough to code it, you aren't smart enough to debug it. Therefore, you should code at your limit, to force yourself to get smarter while debugging
[x for stop in range(5) for x in range(stop)]The "What's New" page has an item-by-item breakdown of each performance tweak and its measured effect[1]. In particular, PEP 659 and the call/frame optimizations stand out to me as very clever and a strong foundation for future improvements.
The more o used Python, and the more Python I saw written, the more I am convinced this is not a good or reasonable argument. A good chunk of the Python devs I’ve interacted with are mystified by the concept of virtual environments, regularly abuse global variables, struggle to read the docs, and structure their code poorly.
Telling these people “oh just write this section in C” will go nowhere. It’s asking them to figure out: installation, compilation, packaging, FFI, tool chains, and foot-guns of language(s) that are arguably more dangerous and whose operation varies from “not super straightforward” to “arcane” based on your dev and deployment environment and application needs.
Unfortunately sometimes you have to provide a library for consumption to programmers that only speak python so you don't have many options.
Certainly that's something you can do, but unfortunately for Python it opens the door for languages like Julia, which are trying to say that you can have your cake and eat it too.
In what context, under what operation? Depending on context, the difference makes sense and is what one would expect - tuples are immutable, with fixed size known at compile time, and stack-allocated; arrays are mutable, dynamic in size, and usually heap-allocated. That's why StaticArrays.jl [1] exists, for when you need something in between/the best of both worlds.
> I like Julia but its easy to write slow julia unless you keep the performance tips in mind.
I very much agree with this statement though. But following a very few basic ones like "avoid non-constant global variables", "look out for type instabilities", "use @views and @inbounds where it makes sense" gets you most of the way there, for eg. about 3-5x of the time a C program would take. Most of the rest of the tips on the Performance Tips [2] page are to squeeze out the last bits of performance, to go from 5x of C to near-C performance.
[1] https://github.com/JuliaArrays/StaticArrays.jl [2] https://docs.julialang.org/en/v1/manual/performance-tips/
But if you compare concretely typed heap allocated arrays in Julia and C, there's no real performance difference (Steven Johnson has a nice notebook displaying this https://scpo-compecon.github.io/CoursePack/Html/languages-be..., and if you see the work on LoopVectorization.jl it makes it really clear that the only major difference is that C++ tends to know how to prove non-aliasing (ivdep) in a bit more places (right now)). So the real question is, did you actually want to use a heap allocated object? I think this really catches people new to performance optimization off guard since in other dynamic languages you don't have such control to "know" things will be placed on the stack, so you generally heap allocate everything (Python even heap allocates numbers), which is one of the main reasons for the performance difference against C. Julia gives you the tools to write code that doesn't heap allocate objects, but also makes it easy to heap allocate objects (because if you had to `malloc` and `free` everywhere... well you'd definitely lose the "like Python" and and be much closer to C++ or Rust in terms of "ease of use", which would defeat the purpose for many cases). But if you come from a higher level language, there's this magical bizarre land of "things that don't allocate" (on the heap) and so you learn "oh I got 30x faster from Python, but then someone on a forum showed me how to do 100x better by not making arrays?", which is somewhat obvious from a C++ mindset but less obvious from a higher level language mindset.
And FWIW, this is probably the biggest performance issue newcomers run into, and I think one of the things to which a solution is required to make it mainstream. Thankfully, there's already prototype PRs that are well underway, for example https://github.com/JuliaLang/julia/pull/43573 automatically stack-allocates small arrays which can prove certain escape properties, https://github.com/JuliaLang/julia/pull/43777 is looking to hoist allocations out of loops so even if they are required they are minimized in the loop contexts automatically, etc. The biggest impediment is that EscapeAnalysis.jl is not in Base, and it requires JET.jl which is not in Base, and so both of those need to be made "Base friendly" to join the standard library and then the compiler can start to rely on its analysis (which will be nice because JET.jl can do things like throw more statically-deducible compile time errors, another thing people ask for with Julia). There's a few people who are working really hard on doing this, so it is a current issue but it's a known one with prototyped solutions and a direction to get it into the shipped compiler. When that's all said and done, of course "good programming" will still help the compiler in some cases, but in most people shouldn't have to worry about stack vs heap (it's purposefully not part of the user API in Julia and considered a compiler detail for exactly this reason, so it's non-breaking for the compiler to change where objects live and improve performance over time).
Anything where performance matters should be written in languages where JIT and AOT compilers come as standard options on the reference implementations.
CPython is already executed in pure C, the only difference it being a very slow, unoptimized C.
The code for simple adding of two numbers is insane. It jumps through enough hoops to make you wonder how is it even running everything else.
And with how dependency happy everything is these days, I avoid trendy projects like the plague.
Having read about some of the changes [1], it seems like the python core committers preferred clean over fast implementations and have deviated from this mantra with 3.11.
Now let's get a sane concurrency story (no multiprocessing / queue / pickle hacks) and suddenly it's a completely different language!
[1] Here are the python docs on what precisely gave the speedups: https://docs.python.org/3.11/whatsnew/3.11.html#faster-cpyth...
[edit] A bit of explanation what I meant by low-hanging fruit: One of the changes is "Subscripting container types such as list, tuple and dict directly index the underlying data structures." Surely that seems like a straight-forward idea in retrospect. In fact, many python (/c) libraries try to do zero-copy work with data structures already, such as numpy.
When you have so many large companies with a vested interest in optimization I believe that Python can become faster by doing realistic and targeted optimizations . The other strategies to optimize didn’t work at all or just served internal problems at large companies .
What exactly do you do with Python that slows you down ?
So many successful projects/technologies started out that way. The Web, JavaScript, e-mail. DNS started out as a HOSTS.TXT file that people copied around. Linus Torvalds announced Linux as "just a hobby, won't be big and professional like gnu". Minecraft rendered huge worlds with unoptimized Java and fixed-function OpenGL.
I took a graduate level data structures class and the professor used Python among other things "because it's about 80 times slower than C, so you have to think hard about your algorithms". At scale, that matters.
It prevents you from taking advantage of multiple cores. Doesn't really impact straight-line execution speed.
A data structures course is primarily not going to be concerned with multithreading.
Being fast isn't contradictory with this goal. If anything, this is a lesson that so many developers forget. Things should be fast by default.
It absolutely is contradictory. If you look at the development of programming languages interpreters/VMs, after a certain point, improvements in speed become a matter of more complex algorithms and data structures.
Check out garbage collectors - it's true that Golang keeps a simple one, but other languages progressively increase its sophistication - think about Java or Ruby.
Or JITs, for example, which are the latest and greatest in terms of programming languages optimization; they are complicated beasts.
In over 90% of my work in the SW industry, being fast(er) was of no benefit to anyone.
So no, it should not be fast by default.
Maybe better to elaborate on what it should be, if not fast? Surely you aren’t advocating things should be intentionally slow by default, or carelessly inefficient?
There’s a valid tradeoff between perf and developer time, and it’s fair to want to prioritize developer time. There’s a valid reason to not care about fast if the process is fast enough that a human doesn’t notice.
That said, depends on what your work is, but defaulting to writing faster, more efficient code might benefit a lot of people indirectly. Lower power is valuable for server code and for electricity bills and at some level for air quality in places where power isn’t renewable. Faster benefits parallel processing, it leaves more room for other processes than yours. Faster means companies and users can buy cheaper hardware.
Languages that tried to be all things to all people really havent done so well.
-- Donald Knuth
https://softwareengineering.stackexchange.com/a/80092
> Yet we should not pass up our opportunities in that critical 3%. A good programmer will not be lulled into complacency by such reasoning, he will be wise to look carefully at the critical code; but only after that code has been identified.
Note that all of the given benchmarks are microbenchmarks; the gains in 3.11 are _much_ less pronounced on larger systems like web frameworks.
Because the core team just hasn't prioritized performance, and have actively resisted performance work, at least until now. The big reason has been about maintainership cost of such work, but often times plenty of VM engineers show up to assist the core team and they have always been pushed away.
> Now let's get a sane concurrency story
You really can't easily add a threading model like that and make everything go faster. The hype of "GIL-removal" branches is that you can take your existing threading.Thread Python code, and run it on a GIL-less Python, and you'll instantly get a 5x speedup. In practice, that's not going to happen, you're going to have to modify your code substantially to support that level of work.
The difficulty with Python's concurrency is that the language doesn't have a cohesive threading model, and many programs are simply held alive and working by the GIL.
From the very page you've linked to:
"Faster CPython explores optimizations for CPython. The main team is funded by Microsoft to work on this full-time. Pablo Galindo Salgado is also funded by Bloomberg LP to work on the project part-time."
This is in very active development[1]! And seems like the Core Team is not totally against the idea[2].
[1] https://github.com/colesbury/nogil
[2] https://pyfound.blogspot.com/2022/05/the-2022-python-languag...
* Statically allocated ("frozen") core modules for fast imports
* Avoid memory allocation for frames / faster frame creation
* Inlined python functions are called in pure python without needing to jump through C
* Optimizations that take advantage of speculative typing (Reminds me of Javascript JIT compilers -- though according to the FAQ Python isn't JIT yet)
* Smaller memory usage for frames, objects, and exceptions
Dang that certainly does sound like low hanging fruit. There's probably a lot more opportunities left if they want Python to go even faster.
Given that Python is interpreted, it's quite unclear what this could mean.
Also, what does it mean to "call" an inlined function?? Isn't the point of inline functions that they don't get called at all?
> During a Python function call, Python will call an evaluating C function to interpret that function’s code. This effectively limits pure Python recursion to what’s safe for the C stack.
> In 3.11, when CPython detects Python code calling another Python function, it sets up a new frame, and “jumps” to the new code inside the new frame. This avoids calling the C interpreting function altogether.
> Most Python function calls now consume no C stack space. This speeds up most of such calls. In simple recursive functions like fibonacci or factorial, a 1.7x speedup was observed. This also means recursive functions can recurse significantly deeper (if the user increases the recursion limit). We measured a 1-3% improvement in pyperformance.
switch (op) {
case "call_function":
interpret(op.function)
...
}
Now it does: switch (op) {
case "call_function":
... setup frame objects etc ...
pc = op.function
continue
....
}
Not sure if they just "inlined" it or use tail-call elimination trick.Python is compiled. CPython runs bytecode.
(If Python is interpreted, then so is Java without the JIT).
Because it's developed and maintained by volunteers, and there aren't enough folks who want to spend their volunteer time messing around with assembly language. Nor are there enough volunteers that it's practical to require very advanced knowledge of programming language design theory and compiler design theory as a prerequisite for contributing. People will do that stuff if they're being paid 500k a year, like the folks who work on v8 for Google, but there aren't enough people interested in doing it for free to guarantee that CPpython will be maintained in the future if it goes too far down that path.
Don't get me wrong, the fact that Python is a community-led language that's stewarded by a non-profit foundation is imho its single greatest asset. But that also comes with some tradeoffs.
Who is working on python voluntarily? I would assume that, like the Linux kernel, the main contributors are highly paid. Certainly, having worked at Dropbox, I can attest to at least some of them being highly paid.
second, CPython is just a C interpeter, there isn't much assembly if any.
third, contributing to CPython is sufficiently high-profile you could easily land a 500k-job just by putting it on your CV.
No, those are not the real reasons why it hasn't happened before.
The developer who worked at Microsoft to make iron python I think that was his full-time project as well and it was definitely faster than cpython at the time
Not because they deliberately only wanted to build an alternative interpreter.
Given the popularity, you’d think Python would have had several generations of JITs by now and yet it still runs interpreted, AFAIK.
JavaScript has proven that any language can be made fast given enough money and brains, no matter how dynamic.
Maybe Python's C escape hatch is so good that it’s not worth the trouble. It’s still puzzling to me though.
Yeah but commercial Smalltalk proved that a long time before JS did. (Heck, back when it was maintained, the fastest Ruby implementation was built on top of a commercial Smalltalk system, which makes sense given they have a reasonably similar model.)
The hard part is that “enough money“ is...not a given, especially for noncommercial projects. JS got it because Google decided JavaScript speed was integral to it's business of getting the web to replace local apps. Microsoft recently developed sufficient interest in Python to throw some money at it.
Google wasn’t alone in optimizing JS, it actually came late, Safari and Firefox were already competing and improving their runtime speeds, though V8 did doubled down on the bet of a fast JS machine.
The question is why there isn’t enough money, given that there obviously is a lot of interest from big players.
I'd argue that there wasn't actually much interest until recently, and that's because it is only recently that interest in the CPython ecosystem has intersected with interest in speed that has money behind it, because of the sudden broad relevance to commercial business of the Python scientific stack for data science.
Both Unladen Swallow and IronPython were driven by interest in Python as a scripting language in contexts detached from, or at least not necessarily attached to, the existing CPython ecosystem.
Even if it wasn't good, the presence of it reduces the necessity of optimizing the runtime. You couldn't call C code in the browser at all until WASM; the only way to make JS faster was to improve the runtime.
> JavaScript has proven that any language can be made fast given enough money and brains, no matter how dynamic.
JavaScript also lacks parallelism. Python has to contend with how dynamic the language is, as well as that dynamism happening in another thread.
There are some Python variants that have a JIT, though they aren't 100% compatible. Pypy has a JIT, and iirc, IronPython and JPython use the JIT from their runtimes (.Net and Java, respectively).
PyPy did implement a JIT for Python, and it worked really well, until you tried to use C extensions.
But still, it’s kind of surprising that PHP has a JIT and Python doesn’t (official implementation I mean, not PyPy).
Think about it: if you have some Python application that's having performance issues you can either dig into a foreign codebase to see if you can find something to optimize (with no guarantee of result) and if you do get something done you'll have to get the patch upstream. And all that "only" for a 25% speedup.
Or you could rewrite your application in part or in full in Go, Rust, C++ or some other faster language to get a (probably) vastly bigger speedup without having to deal with third parties.
Programmers love to save an hour in the library by spending a week in the lab
Or you can just throw more hardware at it, or use existing native libraries like NumPy. I don't think there are a ton of real-world use cases where Python's mediocre performance is a genuine deal-breaker.
If Python is even on the table, it's probably good enough.
Instead, there are big companies, who are running let's say the majority of their workloads in python. It's working well, it doesn't need to be very performant, but together all of the workloads are representing a considerable portion of your compute spend.
At a certain scale it makes sense to employ experts who can for example optimize Python itself, or the Linux kernel, or your DBMS. Not because you need the performance improvement for any specific workload, but to shave off 2% of your total compute spend.
This isn't applicable to small or medium companies usually, but it can work out for bigger ones.
edit: ray https://github.com/ray-project/ray is also pretty easy to use and powerful for actual parallelism
Do you compare it to threads and pools, or judge it on its merits as an async framework (with you having experience of those that you think are done better elsewhere, e.g. in Javascript, C#, etc)?
Because both things you mention "demands on how you build your code" and "limited scope" are part of the course with async in most languages that aren't async-first.
I don't see how "asyncio is annoying and can only be used for a fraction of scenarios everywhere else too, not just here" is anything other than reinforcement of what I said. OS threads and processes already exist, can already be applied universally for everything, and the pool executors can work with existing serial code without needing the underlying code to contort itself in very fundamental ways.
Python's version of asyncio being no worse than someone else's version of asyncio does not sound like a strong case for using Python's asyncio vs fixing the better-in-basically-every-way concurrent futures interface that already existed.
You are confusing concurrency and parallelism.
> Threading is "implicit" context switching all in the same process/thread
No, threading is separate native threads but with a lock that prevents execution of Python code in separate threads simultaneously (native code in separate threads, with at most on running Python, can still work.)
But when you try to do things that aren't a map-reduce or Pool.map() pattern, it suddenly becomes pretty warty. E.g. scheduling work out to a processpool executor is ugly under the hood and IMO ugly syntactically as well.
Are you talking about this example? https://docs.python.org/3/library/asyncio-eventloop.html#asy...
However, batteries are not included. For example, it provides no HTTP client/server. It doesn't interop with any synchronous IO tools in the standard library either, making asyncio a very insular environment.
For the majority of problems, Go or Node.js may be better options. They have much more mature environments for managing asynchrony.
[1] https://docs.python.org/3/library/asyncio-eventloop.html#asy...
I believe the reason is that python does not need any low-hanging fruits to have people use it, which is why they're a priority for so many other projects out there. Low-hanging fruits attract people who can't reach higher than that.
When talking about low-hanging fruits, it's important to consider who they're for. The intended target audience. It's important to ask ones self who grabs for low-hanging fruits and why they need to be prioritized.
And with that in mind, I think the answer is actually obvious: Python never required the speed, because it's just so good.
The language is so popular, people search for and find ways around its limitations, which most likely actually even increases its popularity, because it gives people a lot of space to tinker in.
Do we have completely different definitions of low-hanging fruit?
Python not "requiring" speed is a fair enough point if you want to argue against large complex performance-focused initiatives that consume too much of the team's time, but the whole point of calling something "low-hanging fruit" is precisely that they're easy wins — get the performance without a large effort commitment. Unless those easy wins hinder the language's core goals, there's no reason to portray it as good to actively avoid chasing those wins.
Oh, that's not how I interpret low-hanging fruits. From my perspective a "low-hanging fruit" is like cheap pops in wrestling. Things you say of which you know that it will cause a positive reaction, like saying the name of the town you're in.
As far as I know, the low-hanging fruit isn't named like that because of the fruit, but because of those who reach for it.
My reason for this is the fact that the low-hanging fruit is "being used" specifically because there's lots of people who can reach it. The video gaming industry as a whole, but specifically the mobile space, pretty much serves as perfect evidence of that.
Edit:
It's done for a certain target audience, because it increases exposure and interest. In a way, one might even argue that the target audience itself is a low-hanging fruit, because the creators of the product didn't care much about quality and instead went for that which simply impresses.
I don't think python would have gotten anywhere if they had aimed for that kind of low-hanging fruit.
What I'm describing, which is the sense I've always seen that expression used as in engineering, and what GP was describing, is: this is an easy low-risk project that have a good chance of producing results.
E.g. If you tell me that your CRUD application suffers from slow reads, the low-hanging fruit is stuff like making sure your queries are hitting appropriate indices instead of doing full table scans, or checking that you're pooling connections instead of creating/dropping connections for every individual query. Those are easy problems to check for and act on that don't require you to try to grab the fruit hard-to-reach fruit at the top of the tree like completely redesigning or DB schema or moving to a new DB engine altogether.
Easily obtained gains; what can be obtained by readily available means
https://www.merriam-webster.com/dictionary/low-hanging%20fru...
the obvious or easy things that can be most readily done or dealt with in achieving success or making progress toward an objective
"Maria and Victor have about three months' living expenses set aside. That's actually pretty good …. But I urged them to do better …. Looking at their monthly expenses, we found a few pieces of low-hanging fruit: Two hundred dollars a month on clothes? I don't think so. Another $155 for hair and manicures? Denied."
"As the writers and producers sat down in spring 2007 to draw the outlines of Season 7, they knew, Mr. Gordon said, that most of the low-hanging fruit in the action genre had already been picked."
"When business types talk about picking low-hanging fruit, they don't mean, heaven forbid, doing actual physical labor. They mean finding easy solutions."
https://dictionary.cambridge.org/dictionary/english/low-hang...
something that is easy to obtain, achieve, or take advantage of
"The easy changes have all been made. All the low-hanging fruit has been picked."
"When cutting costs, many companies start with the low-hanging fruit: their ad budgets."
"For the beauty-care industry, the teen demographic is a new category for them - low-hanging fruit."
"I'm a great believer in picking low-hanging fruit. Start with what's easy, and go higher later."
"This legislation is some of the low-hanging fruit - the issues that we can agree upon across parties and regions."
If everyone knows it's never going to reach the performance needed for high performance work, and there's already an excellent escape hatch in the form of C extensions, then why would people be spending time on the middle ground of performance? It'll still be too slow to do the things required, so people will still be going out to C for them.
Personally though, I'm glad for any performance increases. Python runs in so much critical infrastructure, that even a few percent would likely be a considerable energy savings when spread out over all users. Of course that assumes people upgrade their versions...but the community tends to be slow to do so in my experience.
Ah! So they are so tall that picking the low-hanging fruit would be too inconvenient for them.
Talk about stretching an analogy too far.
Yes, and I would argue it already exists and is called Rust :)
Semi-jokes aside, this is difficult and is not just about removing the GIL and enabling multithreading ; we would need to get better memory and garbage collection controls. Parts of what's make python slow and dangerous in concurrent settings are the ballooning memory allocations on large and fragmented workloads. A -Xmx would help.
In addition to the other point here about speed not being a target in the first place.
And why they were not addressed: because starting a certain points, the optimization are e.g. processor specific or non intuitive to understand. Making it hard to maintain vs a simple straight forward solution.
I know I’m a bit more focused on the business side of things than a lot of techies here on HN, it’s a curse and a gift, but what I read the 10-60% speed increase as is money. The less resources you consume, the less money you burn. Which is really good news for Python in general, because it makes it more competitive to languages like C# for many implementations.
This comes from the perspective of someone who actually things running Typescript on your backend is a good idea because it lets you share developer resources easier so that the people who are really good at react can cover for the people who are really good at the backend and the other way around in smaller teams in non-software-engineering organisations.
if you get non descriptive errors, it means you haven't followed proper exception handling/management
its not the fault of Python but the developer. go ahead and downvote me but you if you mentioned parent's comment in an interview, you would not receive a call back or at least I hope the interviewer is realizing the skill gap.
https://rpython.readthedocs.io/en/latest/
There is also the Dynamic Language Runtime, which is used to implement IronPython and IronRuby on top of .NET:
There are bunch of transpilers which might provide a path from statically typed python3 to native binaries. py2many is one of them.
The downside is that all of the C extensions that python3 uses become unusable and need to be rewritten in the subset of python3 that the transpiler accepts.
(Though JS had some potential to do that by the power of it's browser role, before compiling interpreters to WASM became a better way of getting browser support than compiling another scripting language source to JS.)
It's just it also has sooo much historical baggage and normative lifestyle assumptions along with it, and now it's passed out of favour.
And historically support in the browser was crap and done at the wrong level ("applets") and then that became a dead end and was dropped.
> Mypyc compiles Python modules to C extensions. It uses standard Python type hints to generate fast code. Mypyc uses mypy to perform type checking and type inference.
Could you expand on this?
[0] http://www.parrot.org/ states:
The Parrot VM is no longer being actively developed.
Last commit: 2017-10-02
The role of Parrot as VM for Perl 6 (now "Raku") has been filled by MoarVM, supporting the Rakudo compiler.
[...]
Parrot, as potential VM for other dynamic languages, never supplanted the existing VMs of those languages.
All approaches to "make Python fast again" tend to lock down some of the flexibility so aren't realy general purpose.
It was designed for "native-like" execution in the browser (or now other runtimes.) It's more like a VM for something like Forth than something like Python.
It's a VM that runs at a much lower level than the ones you'd find inside the interpreter in e.g. Python:
No runtime data model beyond primitive types - nothing beyond simple scalar types; no strings, no collections, no maps. No garbage collection, and in fact just a simple "machine-like" linear memory model
And in fact this is the point of WASM. You can run compiled C programs in the browser or, increasingly, in other places.
Now, are people building stuff like this overtop of WASM? Yes. Are there moves afoot to add GC, etc.? Yes. Personally I'm still getting my feet wet with WebAssembly, so I'm not clear on where these things are at, but I worry that trying to make something like that generic enough to be useful for more than just 1 or 2 target languages could get tricky.
Anyways it feels like we've been here before. It's called the JVM or the .NET CLR.
I like WASM because it's fundamentally more flexible than the above, and there's good momentum behind it. But I'm wary about positioning it as a "generic high level language VM." You can build a high-level language VM on top of it but right now it's more of a target for portable C/C++/Rust programs.
https://mail.python.org/archives/list/python-dev@python.org/...
Folks not desperate for the improvements might want before jumping in.
In the last 10 years, I looked into pypy I think three or four times to speed things up, it didn't work even once. (I don't remember exactly what it was that didn't play nice... but I'm thinking it must have been pillow|numpy|scipy|opencv|plotting libraries).
"1x faster" to me says twice the speed of the original, but I believe they use it to mean "the same speed".
But you're right, I think "1.96x as fast" sounds more correct.
edit: but now that I think about it "X as fast" or "X faster" in terms of a duration will always sound a bit weird
Great news.
As a result, it will be quite some time before we see a version 4.
Semver often means that major is the language version, and minor is the runtime version
There was a heap of features added between 3.5 and 3.11, for example. Enough to make it a completely different language.
(It actually changes a whole lot, because there's a whole lot of code already out there written in Python. 25% faster is still 25% faster, even if the code would have been 100x faster to begin with in another language.)
And that's even assuming that the code would have existed at all in another language. The thing about interpreted GC languages is that the iterative loop of creation is much more agile, easier to start with and easier to prototype in than a compiled, strictly typed language.
It's about on par with Ruby, only Lua and JIT compiled runtimes beat it (for very understandable reasons).
Common Lisp and Java begs to disagree.
Python is used in so many different areas, in some very critical roles.
Even a small speedup is going to be a significant improvement in metrics like energy use.
The pragmatic reality is that very few people are going to rewrite their Python code that works but is slow into a compiled language when speed of execution is trumped by ease of development and ecosystem.
[0] https://github.com/markshannon/faster-cpython/blob/master/pl...
Does not sounds like Rust.
(But other than that I agree with your point.)
Doing side by side comparisons between golang and python on Lambda last year, we halved the total execution time one a relatively simple script. Factor of 100, I assume, is an absolute best case.
Also, because of FFI, it's not like the two languages are opposed to each other and you have to choose. They can happily work together.