Python 3.11 is faster than 3.8
jott.live
jott.live
zig: Emulated 600 frames in 0.24s (2521fps)
rs: Emulated 600 frames in 0.37s (1626fps)
cpp: Emulated 600 frames in 0.40s (1508fps)
nim: Emulated 600 frames in 0.44s (1367fps)
go: Emulated 600 frames in 1.75s (342fps)
php: Emulated 600 frames in 23.74s (25fps)
py: Emulated 600 frames in 26.16s (23fps) # PyPy
py: Emulated 600 frames in 33.10s (18fps) # 3.11
py: Emulated 600 frames in 61.43s (9fps) # 3.10
Doubling the speed is pretty nice :D Still the slowest out of all implementations though :P[0] https://github.com/shish/rosettaboy
EDIT> updated the nim compiler flags to build in release mode like most other languages, thanks @plainOldText!
nim: Emulated 600 frames in 0.44s (1367fps)
Do you happen to know the right flags for release-mode Go? Last time I checked (admittedly years ago) I thought they just had the one “reasonably fast and reasonably debuggable” build modePerhaps benchmarks should be compared on equal footing, say, with the default release flag or all optimizations turned on, otherwise they're improper.
But really the original author/poster should do some set on his box. I cannot even compile/run all his things. The point of my comment was just to give a reference for how far off impressions can be from build flags. PGO (available to Nim, c++, but maybe not to Rust yet?) is a whole other set of maybe nothing burgers or maybe big improvements.
(But, btw, I could not agree more that all experiments in this entire general space should have various big, bold disclaimers. Over-concluding from these things is rampant.)
For something super jumpy like a simulator, I would find it unsurprising for PGO to make up (or surpass) the difference to Zig in both Nim and C++. 20 years ago there was this ACOVEA [1] project to try to discover great sets of gcc flags that could often find 2X improvements in object code speed for me.
The range from build flags/procedures is often much greater than the supposedly interesting cross-language variation. These things often more measure developer experience/persistence than something intrinsic (and build flags/procedures are only part of that experience/persistence).
There are various debug things you can turn on, such as the race detector, the memory sanitizer, or coverage tracking.
I wonder what optimizations Zig does that lets it generate machine code / LLVM bitcode faster than both C++ and Rust (which certainly have larger teams backing them), at least in the case of this Gameboy emulator project.
[1] half the time I’d make the compiler crash; the other half it would generate a binary which crashes at runtime, with weird heisenbug behaviour like “adding a print statement to log how far down a function I am causes the code to stop crashing at all” — like right now there is a load-bearing print statement which shows the address of the SDL Window object, because otherwise the compiler seems to optimise the Window out of existence and then a few lines later it segfaults on the null pointer…
Is the most glorious thing I ever heard.
that sounds like you have undefined behavior in your program
https://stefan-marr.de/2022/10/cost-of-safety-in-java/
Unsafety buys you a little more performance, but the baseline should the safe version, not the unsafe version. It is like having to explain on a case by case basis why you aren't using lead pipes for this application.
Always default to safety. The difference between safe and unsafe native code is usually single digit percentage points. Or weeks on a Moore scale.
Heck, I see 1.33x differences in run times between the first & second run of the rust branch of his benchmark, but only 1.05x diffs in run times between 1st & 2nd Nim branch.
In my experience, reasoning about things like this is rarely as simple as "bounds checks" which are (often) the most highly predictable branches.
zig: Emulated 600 frames in 0.23s (2880fps)
rs: Emulated 600 frames in 0.35s (1691fps)
cpp: Emulated 600 frames in 0.40s (1519fps)
nim: Emulated 600 frames in 0.44s (1367fps)
go: Emulated 600 frames in 1.78s (338fps)
php: Emulated 600 frames in 9.28s (65fps)
py: Emulated 600 frames in 33.10s (18fps)
That's after Python cheated with native implementation of bitblt while all other implementations are doing it pixel by pixel, which is more correct.Applying BS scaling to the above table, Nim would score 1736, only slightly faster than the Rust. Who knows what little flag tweaks in cpp/rs/zig could similarly re-arrange the numbers?
Now, why is -d:lto (or -d:release) not the default build mode? Well, because it takes time and (some) people already complain about compile times and workflow (sometimes). Trade offs abound almost as much as over-conclusion. ;-)
PHP is traditionally used solely for websites. Some of those have grown rather large, to the point that having engineers optimize the language is cheaper than buying more servers.
Python, on the other hand, is first and foremost a scripting language. When performance does matter, you often end up using a wrapper around a C library, like NumPy. This means there is relatively little money in optimizing the Python interpreter.
It would seem sensible that since Facebook poured a lot of resources in optimizing PHP, Google would have done the same for Python.
Also, Python is the first or second most popular language.
Well they sort of half assed tried years ago - they employed GvR at one point. Then they gave up and they continue to lean heavily on C++, Java and there is thing they developed in the interim: Go.
> Also, Python is the first or second most popular language
Hogwash. For all the open source and other development that occurs and is “indexed” on internet discussion forums there is countless boring ass shit behind the scenes in sweatshops around the world and corporate back rooms. PHP, Java, and C# are still probably more popular to start.
Much of which is also in Python.
I say this as someone that has made their career mostly on the “top 5” TIOBE languages - but that ranking is a bunch of crap for the overall commercial software world.
C? Give me a fucking break. Never in 20 years have I known there to be more C than Java jobs.
And now everyone and their dog including doctors and other non tech-savvy professionals are doing Intro to Python courses (and then never touching it again) so they too can feel like they know something about ML or stats. Lot of noise.
By what metrics? Redmond quarterly from June shows Python at #2, and PHP at #4.
https://redmonk.com/sogrady/2022/10/20/language-rankings-6-2...
I've looked at various sources since the last 5 years and not much changed. I hoped some languages like Nim, Crystal or Zig would pick up a bit of steam, but no.
Go and Rust moved up a bit and now seem safer bets to invest some time into.
> By what metrics?
Hmm, why don’t you answer your own question?
> Redmond quarterly from June shows Python at #2,
That clearly is one that puts it “first or second”, yes.
But still wanted to know by what metrics the GP was making their claim.
For Python, the C integration probably makes things harder.
In PHP, doing `$a->b` accesses field `b` of object `$a`. The implementation is essentially type check (making sure `$a` is an object) and hash table lookup. If this fails, then it calls `__get` method if it exists.
In Python on the other hand, `a.b` involves `__getattr__`, `__getattribute__`, `__mro__`, `__get__`, `__set__`, `__delete__` and `__dict__`. Here is a description of how the attribute access works: https://docs.python.org/3/howto/descriptor.html#overview-of-....
(Measures Python 3.10.4 against PHP 8.1.5 - so expecting these to change a bit).
It's generally faster than Python in doing the same things as Python does, unrelated to web too.
Meta/FB is also pretty invested in Hack (was once a php dialect, again, out of the loop), maybe they contributed a thing or two?
At Wikimedia we adopted it which has cut our CPU usage by half and has saved a few hundred of servers. We had some Facebook engineers helping which involved patching the Linux kernel while at it. Those were good times.
Eventually PHP 7 followed up with a similar approach and had more or less the same performance as HHVM. Facebook went then to focus on the Hack language (a dialect of PHP with strong typing) and eventually phased out back compat with Zend.
From what I remember, Sara Golemon at Facebook has done a lot of outreaching to Open Source project and gave us a lot of assistance (as well as others at Facebook).
I remember it being slower than PHP4 for a while
TIL MySQL 6
I recall a couple of times over the past 15-20 years where a current python significantly outperformed a current php version, but my recollection is that php was usually a bit faster, or sometimes a lot faster.
PHP 7 was released on dec 2015, and gained significant speed bumps and better memory usage, with the average php execution time being cut in half, or more, generally without any code change whatsoever. It was quite remarkable.
The path from 7.x-8.1 so far has generally seen incremental speed bumps again - usually somewhere between 3-8% improvements per release. Obviously this is going to be dependent on use cases, but overall it's been a fairly steady set of speed improvements over the last 7 years.
Of course, people then go and build up sculptures of objects that will be thrown away after every request, and that stuff makes everything slow (that style of code is why PHP wasn't good enough for Facebook, IMHO), but you can build trash sculpture in any language.
I'm also wondering if all your #[inline(always)] is slowing things down.
For the benchmark 600 frames, the Rust emulator calls it 2500049 times, but the Zig emulator only calls it 1808121 times. This is very roughly the ratio between the reported times.
At least one source of this discrepancy is that the Zig emulator doesn't think any sprites are active, but that doesn't account for all of it, and I'm hitting some crashes trying to run it with the display, so I'm probably not going to investigate it further.
(I also hit the same sort of strange Zig compiler unreliability as some other comments have mentioned, where I had to add some logging in seemingly random places to get the Zig emulator to run at all.
And the argument parser just blatantly returns a dead pointer to the ROM path; when I first ran it it just gave "error: InvalidUtf8" until I tracked that down.)
IIRC each of the #[inline] statements was tested and each made a noticable performance improvement. That’s especially true for things like RAM::get() - since the gameboy does I/O by having different chunks of RAM act differently (some address ranges are just RAM, some are hardware controls, some read data from the cartridge, etc) you can replace the hundred-line generic get() with a single instruction if you happen to know that you are looking up one hard-coded address, and that address has no special behaviour.
[1] https://mypyc.readthedocs.io/en/latest/introduction.html
--- a/py/src/cpu.py
+++ b/py/src/cpu.py
@@ -260,7 +260,6 @@ class CPU:
param = None
cmd_str = cmd.name
self.PC += 1
- self._debug_str = f"[{self.PC:04X}({ins:02X})]: {cmd_str}"says
# opcache in 8.1 gives a nice speedup (25s to 10s)
Is the 23.7 seconds above using the 8.1 opcache?
Unfortunately, "just write the performance-sensitive bits in C" is pretty impractical because it only works when you're handing the C routine a relatively small amount of data relative to the amount of work to be done on that data (otherwise the costs of marshaling to C data structures will quickly eat up any gains from processing in C). And unfortunately, it turns out that a whole bunch of code works this way, so the whole bargain of "slow interpreter + easy C extensions" breaks down for a lot of real-world applications, but now we're locked into it.
Processing loops adds a lot of overhead because the interpreter has to make the case jumps after each loop. Figuring out a way to minimize that overhead by using built-in data structures and stdlib library will speed up your code by an order of magnitude.
Don't forget, the built-in types are already running in C.
FFI can pervade an ecosystem making changes (including performance optimizations) to the host language more difficult. It also tends to complicate build and deployment stories. For example, from my Mac I can trivially build a native binary that will Just Work on any Linux system, even one without a libc. Contrast that with Python where I can’t even install popular packages onto most non-Ubuntu Linux distros. And there’s a lot of other things like that which crop up with pervasive FFI.
I’m convinced that the FFI sweet spot is “difficult but possible” and native performance should be good enough 99% of the time.
[profile.release]
lto = true
- With Zig I kept running into compiler bugs, plus no package manager (I’ve vendored SDL and Clap into the source tree)
- C++ I’d occasionally shoot myself in the foot in ways that other languages would have caught, plus no package manager (OS-level package management does an OK job, so long as you don’t mind using old versions, and faffing about with different operating systems acting very differently)
- The pain from Rust was one time where the compiler wanted me to specify a lifetime, and I didn’t understand, so I just spammed lifetime specifiers in various places until it compiled. I’ve been using Rust for a couple of years now and I still don’t really understand lifetimes, but thankfully 99% of the time I can avoid them.
- Nim was a relatively nice language but massively lacking in available libraries (like even parsing command line arguments took me a day just trying to find a library which worked)
- Go is pretty nice, my main pain is the tolerable but constantly-annoying verboseness of error handling (`err := foo(); if err != nil {return err}` compared to rust’s `foo()?`)
- PHP I just hate on a deep and personal level thanks to years of being a PHP4/5 developer. The language is actually mostly-ok-ish these days, but the standard library is still full of frustration like inconsistent parameter orders within a family of functions.
- Python is all-round really nice to write, but the test suite takes like 20 minutes to run, which really messes with my flow-state
" I’ve been using Rust for a couple of years now and I still don’t really understand lifetimes"
Seems like a major pain point.
They do look intimidating to start with, admittedly, and I'll concede that's a negative point for Rust. But it does get better if you practice for a bit.
I ran a quick pprof and indeed it's spending a lot of time in cgo:
Showing nodes accounting for 28720ms, 67.67% of 42440ms total
Dropped 145 nodes (cum <= 212.20ms)
Showing top 10 nodes out of 53
flat flat% sum% cum cum%
13080ms 30.82% 30.82% 16600ms 39.11% runtime.cgocall
4720ms 11.12% 41.94% 4750ms 11.19% main.(*RAM).get
2840ms 6.69% 48.63% 33070ms 77.92% main.(*GPU).tick
1970ms 4.64% 53.28% 3720ms 8.77% runtime.mallocgc
1450ms 3.42% 56.69% 1470ms 3.46% main.(*RAM).set
1160ms 2.73% 59.43% 41350ms 97.43% main.(*GameBoy).tick
1000ms 2.36% 61.78% 3160ms 7.45% runtime.exitsyscall
890ms 2.10% 63.88% 1610ms 3.79% main.(*CPU).tick_interrupts
820ms 1.93% 65.81% 850ms 2.00% runtime.casgstatus
790ms 1.86% 67.67% 5530ms 13.03% main.(*CPU).tickThe other languages are also using SDL via their respective interacting-with-C interfaces - what makes Go special here?
https://stackoverflow.com/questions/28272285/why-cgos-perfor...
Even if it is SDL slowing it down, Go FFI being slow is still a real disadvantage compared to the other languages, and you can't just pretend like it doesn't exist in this case.
By these standards a 10 year old cpu with a beefy GPU will beat any new cpu as well.
It absolutely does. A simple for loop in the standard Python interpreter will literally take 100x longer than the same thing in a language like C/C++, try for yourself if you don't believe me. CPython is unbelievably slow.
I couldn't reproduce 100x (no optimization flags, otherwise it won't do anything)
Apple clang version 14.0.0 (clang-1400.0.29.102) -> 0m0.347s
ruby 3.1.2p20 -> 0m11.314s
Python 3.8.12 -> 0m19.662s
So ruby is almost 60% faster, but the C version is "only" 32x faster than ruby. 55x python.I thought the differences would be smaller these days
$ brew install pypy3
pypy3: The x86_64 architecture is required for this software.$ pyenv install pypy3.9-7.3.9
I like using PyEnv for managing my Python versions. It will natively compile Python builds and should be doing the same on M1 (which is what I'm using). `pyenv install --list` shows you what is available.
EDIT: Not sure why they don't have newer versions of PyPy there (I don't use PyPy) but all it takes is a PR to here: https://github.com/pyenv/pyenv
ED> Looks like CPython prefers using a dict as a lookup table for opcodes (which is what this implementation does), while PyPy prefers having a long series of if-statements. Hmm.
Or maybe https://docs.python.org/3/library/array.html rather than list.
Best wishes
On my machine (very similar, a macbook pro m1 max):
python 3.10: 52s
python 3.11: 35s
pypy 3.9.12: 10s
(This test is basically a perfect test for JITs: one loop repeated many times)https://gist.github.com/llimllib/7af8144a92d3c2e1fc58be62988...
python 3.10: 60s
python 3.11: 46s
pypy 3.9.12: 6s
Looks like pypy performs comparatively better on x86_64(I could be wrong)
edit: oh, huh I think you're right that it's running here on rosetta:
$ file installs/python/pypy3.9-7.3.9/bin/pypy
installs/python/pypy3.9-7.3.9/bin/pypy: Mach-O 64-bit executable x86_64
Wonder if there's a way to run a native version?edit 2: there are nightlies here: https://buildbot.pypy.org/nightly/py3.9/
Running the latest, a native binary gives more than 2x speedup:
# first you have to allow all the unsigned binaries to run
$ xattr -dr com.apple.quarantine pypy-c-jit-106295-5dd3b18303e2-macos_arm64/bin/*
# then we get 3.5s:
$ time pypy-c-jit-106295-5dd3b18303e2-macos_arm64/bin/pypy nbody.py 10000000
-0.169075164
-0.169077842
real 0m3.522s
user 0m3.468s
sys 0m0.045s
pypy continues to impress! $ pip install numpy
<snip building wheel>
$ python --version && python -c "import numpy; print(numpy.identity(5))"
Python 3.9.12 (05fbe3aa5b0845e6c37239768aa455451aa5faba, Mar 29 2022, 09:54:47)
[PyPy 7.3.9 with GCC Apple LLVM 13.0.0 (clang-1300.0.29.30)]
[[1. 0. 0. 0. 0.]
[0. 1. 0. 0. 0.]
[0. 0. 1. 0. 0.]
[0. 0. 0. 1. 0.]
[0. 0. 0. 0. 1.]]Should make essentially no difference then, since I rather doubt the Python implementation of n-body can leverage the GPU, or strains the RAM so much that the 200GB/s of the Pro (IIRC) would be an issue.
What’s wrong with strict typing and ahead of time compilation as long as it compiles fast? Doesn’t this prevent many runtime errors that can occur in Python?
Python is a higher level language than C++ so it requires less effort but newer compiled/typed languages offer more of what Python is good at
Nothing, as long as the compilation is fast. C++ compilation is not fast. The large C++ projects I've worked on over the last few years, compilation is bordering on 2 hours for a full build, and 10-60s for incremental builds. At one point, incremental changes were taking 15 minutes to link at one point (resolved by [0]). Go is a great example of fast compilation and strict typing IMO.
[0] https://devblogs.microsoft.com/cppblog/improved-linker-funda...
> In 2007, build engineers at Google instrumented the compilation of a major Google binary. The file contained about two thousand files that, if simply concatenated together, totaled 4.2 megabytes. By the time the #includes had been expanded, over 8 gigabytes were being delivered to the input of the compiler, a blow-up of 2000 bytes for every C++ source byte.
https://go.dev/talks/2012/splash.article#TOC_5.
I also hate how just adding some definition to a .h file that's only referenced in one .c (or .cpp) file will recompile loads just because the header file changed. Maybe there's ways to improve on that (ccache?), but I mostly write C to contribute to open source projects (rather than my own) and it can be an annoying wait.
A 128 cores threadripper is a great workstation to compile C++ code fast :D
(Unless a hypothetical Node.js rewrite would be comparable on speed. I wonder how it would do on memory.)
My system has a 3900x, 32GB ram, and a fast NVME SSD for reference.
The last chromium build took 1h:45m to complete. This build had LTO enabled and -j18 passed to ninja.
Without LTO it would probably be closer to an hour I think but I don't seem to have a non-LTO build in the logs.
Your estimation seems pretty accurate!
I think an even better example might be OCaml. Ocaml's compilation speed (last time I checked) was on par with Go, but it provides a much nicer (IMO) type system.
That’s why I was suggesting newer compiled languages incorporating some of what we’ve learned over the past 4 decades. eg type inference
Similarly, Rust compilation speed is much slower than GoLang's, whose type system does relatively very little for you. I haven't seen exact numbers, but I've seen people say Rust and C++ have similar compilation speeds, and both are, in general, slow to compile.
I see Nim as clearly better if you write a larger software. It's a shame it's usage is low and thus there aren't many Nim resources.
Nothing, it's all good things: It's easier to write, easier to debug, and the compiler (not the user) catches bugs.
I ran a very basic benchmark against a local web application: I got 413.56 requests/second on 3.10 and the exact same code gave me 533.89 requests/second on 3.11.
That's a big enough increase that I think it's worth actively upgrading projects. Usually I wait for a few months for things to settle in first.
If you are doing heavy matrix math, numpy runs at FORTRAN speed, tools like scikit-learn and Tensorflow also get high performance by doing the heavy lifting outside Python.
But remember that the exercise wasn't "to run the n-body simulation", or even "to run maths", but just to see how fast "plain Python" is with the release of 3.11 - using numpy and scipy, which rely on compiled code that they can hand work off to, would make any runtime value you get completely meaningless for the purposes of benchmarking pure Python =)
Because "you wouldn't write this code" applies to those examples as well. In JS you'd tap into a C-compiled native library the exact same way numpy/scipy does in Python. And in C, if you absolutely needed the fastest performance, the code would be full of micro optimizations and a sprinkling of assembler.
I'm tired of people comparing languages but then leaving out the major winning points for Python. Numpy & co. are integral part of Python, no serious ddv would use just "pure" Python for numerical methods, ever. So let's compare real'world Python, shall we? I doubt Go has a chance then. /rant
If you want to test baseline performance, you want a little program that's small enough to be easily understood, but just elaborate enough to hit enough of the standard programming patterns. So it's basically irrelevant what this code actually does, we're just looking at how much faster Python has gotten. Because that's something we care about. We don't care that "go is faster" or "C is faster", we use Python and we like to know that the newer versions actually are substantially better than the older versions we used to (or still have to) work with.
Not everyone will be aware that this meant as praise. ;-)
FORTRAN codes persist today because (1) the old school memory model of FORTRAN is fast, and (2) it is so easy to write numeric codes that do the wrong thing with rounding and numerical instability. There's a reason why Foreman Acton wrote a book titled Numerical methods that (usually) work.
https://www.amazon.com/Numerical-Methods-that-Work-Spectrum/...
Code something up in C, Haskell, oCAML or CUDA and you miss out on the 40+ years of experience people have had with a FORTRAN code from the 1970s.
Are you saying FORTRAN avoids these problems, or that it is prone to them? If the former, how does it do it?
Today somebody who doesn't know numerics frequently codes something up for the wrong reasons (e.g. to learn a new language, because they think the 1970s FORTRAN code is obsolete, ...) and never did the testing to know that the code they wrote is numerically stable or not.
That is, you might think it is pretty easy to code something numerical up, and sometimes it is, but frequently you write something that's a little bit wrong and sometimes you write something that's terribly wrong, sometimes it isn't even wrong.
It's not that FORTRAN is necessarily more accurate than another language, but that you can trust a code that has been around for 40 years and codes that have been around 40 years have been written in FORTRAN.
---
As for the memory model I think about it the most when I write embedded programs for my Arduino.
It drives me nuts that C diddles the stack pointer around meaninglessly when for most of the programs I write there are a small number of parameters that decide the size of all the arrays (like an old FORTRAN program) and local variables, recursion and all that are a source of problems and not solutions.
The only reason I write C for that thing at all is that some of the programs I write are performance sensitive and I could get a bigger boost running the C code on an ARM or ESP32 than I could get writing AVR8 assembly and eliminating meaningless loads, stores and other activity that C does "just because".
It's sequential code with fiddly side effects. I know I've written one.
But I'm generally curious if this is in-fact possible in someway.
> That benchmark shows that Python is slow. If you avoid writing your code in Python it can be really fast!
Whilst I'm here: shameless self-promotion for Equinox and Diffrax:
https://github.com/patrick-kidger/equinox https://github.com/patrick-kidger/diffrax
Which are neural network and differential equation libraries for JAX.
[Obligatory I-am-a-googler-my-opinions-do-not-represent-your-employer...]
I think the delta between JAX and PT on commodity hardware is just a little bit too small to really create much splash on HN.
But, the language before has a bigger shift to making it faster due to the amount of resources it has now that are dedicated to making CPython faster similar to what happened with JavaScript. I hope CPython eventually is as fast as JavaScript.
Please think about your actual workload and take them with a grain of salt.
For instance, for most web apps, you spend a large amount of time waiting on database responses. It looks nothing like the modeling the tests in the article do. Benchmarks are not typical workloads.
Don't just assume because some random person on hn says zig is faster you should rewrite your business apps.
1) Network speeds have increased dramatically compared to CPU speeds.
2) People don't optimize code very much.
3) Web apps tend to do more work per request than they did in the 90s.
Regardless, I've never seen an app saturate it's network pipe, but I've seen plenty saturate all their CPU cores, including relatively well-tuned ones. For instance, I wrote a Netty-based reverse proxy app once, and while I got it to run far faster than the typical app in that company, it was still CPU-bound in all my tests.
We picked Python Asyncio based Tornado async for our services. As soon as we hit scale we were getting CPU bound.
Deep diving into profiling came up with JSON parsing being the culprit.
It's a painful problem to crack once you hit those limits. In some cases you could genuinely skip parsing the JSON (by returning a raw JSON containing string to client) such as when you are simply getting data from cache or db.
In other cases you simply can't skip it. For eg when you are interfacing with a 3rd party library that will only speak JSON. At that stage you are stuck.
You could try and use a wrapper around faster native JSON parser (say uJson) but it will be a trade-off between the parsing time and the time taken to copy the string to the FFI parser and copy back the results. And deal with all the complexity that that entails.
Or you could hand it off as a job to an async queue (this might be the canonical architectural approach to prevent blocking the event loop) but then you have just shifted the problem to a different place where you'll still need to throw more instances at the problem. And this adds extra latency.
I too was in the "don't optimize prematurely" camp but picking Python today for new services IMO would be taking that principle a bit too far.
Especially considering the ergonomics that modern languages like Golang or Rust offer.
If you hit that doing something novel and obscure, sure. If it is literally parsing JSON, you spend a little effort researching non-stdlib JSON parsing libraries, pick one of the several stable much-faster-than-stdlib ones, and move on.
Once the team has the know-how to write a service in Rust/Golang or even the modern pleasant to write Java, it becomes hard to justify why we'd pick Python for a new service at all.
He very much did not explicitly mention that its already done, with established results, and that the described “pain” isn’t something you have to take on at all.
Once again, I said to look at what your load actually is. If you're CPU bound, then yes, buy a bigger instance, optimize, or switch languages.
It's like how if you are putting up a blog, you probably shouldn't be looking at running kubernetes clusters for just that blog.
And that's not even taking into account frameworks and ORMs like Hibernate, which itself can multiply the CPU used several-fold on top of JDBC, or whatever lower-level interface it wraps. I've never measured frameworks in dynamic languages, but I have no reason to believe they aren't similarly inefficient.
And, yes, one of the optimizations I did for my reverse proxy app was to upgrade the JSON library, which brought a significant performance boost. But it's not the only source of CPU usage for apps, nor was it the only major optimization I successfully applied.
True, but that has nothing to do with what OP said:
> you spend a large amount of time waiting on database responses.
I mean I could buy it if Python was maybe 5-10x slower than "fast" languages, but the benchmarks people are throwing around show it is 50-100x slower. Every Python codebase I have used (apart from one-off scripts I guess) has eventually run into the "ok it's slow, how can we make it faster?" barrier.
Now does it matter? Not all codebases are the same. If I'm serving api requests, does it being 55ms vs 15ms make a difference? 110ms vs 75ms? If you're writing analysis on a huge data set and an iteration takes 100ms vs 50ms? yeah that could be the difference between weeks vs days.
Almost like someone should look at what they're doing before saying 'IshKebab said python is slow, so we should rewrite all our code'
But usually 90% of the response time for the api request is the time it took the SQL server to execute the relevant queries.
I make software that requires searching over large datasets (image recognition for construction drawings). It's a web app running python on the server and it feels instantaneous for users as they're searching.
Most people aren't even doing that - they're just pushing and pulling data from a db. Building maintainable software is what really matters.
That's why I'd never start a new project in Python (among other reasons). It backs you into a poor performance corner.
Yes, it's the difference between being able to serve 66 r/s or 18 r/s.
Still, for web apps there are still huge gaps based on language and framework used. [0]
[0] https://www.techempower.com/benchmarks/#section=data-r21
To me ahead of time compilation is a convenience that in many cases catches bugs before they present all their glory to a customer.
As for "generally ugly syntax" - beauty is in the eyes of the beholder. I am for example multilingual and do not get hung up on syntax unless it resembles brainfuck. However having fn instead of function, not having brackets when supplying parameter list, having variables to use some special characters does not qualify for "beauty". It is just a different way of doing the same thing and often feels that it is done for the whole purpose of being different or half arsed attempt to make parsing simpler.
It's not entirely subjective. For example a language that is inconsistent in it's use of syntax could fairly be described as "objectively ugly".
My point is - I don't buy that "language aesthetics are entirely subjective".
In my experience, having done quite a lot of dynamic language benchmarks, Javascript is about as fast as PHP and both are about 6x faster than Python.
Has someone looked at the Python code used here, if it has any obvious gotchas?
~/p/not_that_slow python3.11 modified.py 5000000
-0.169075164
0.183753791
Time: 6.335 s
~/p/not_that_slow python3.11 main.py 5000000
-0.169075164
-0.169083134
Time: 48.515 s
Other than this, the only modification I made is to include a `print` statement at the end to show the time taken.EDIT: That is not actually equivalent, my bad. The iterator is consumed and the execution ends early, I didn't realize the algorith iterated over it multiple times.
And itertools is a real gem, full of lots of goodies to handle list manipulation tasks.
I’m in no way criticizing the benchmark but that’s the reason for the starker difference.
For a little more information, this is using Bun 0.2.1, which uses JavaScriptCore (JSC), a part of WebKit which powers Safari. Since I'm running on an M1 Pro (Apple's ARM chip), there is probably somewhat of a benefit in using JSC.
There are many flaws in this benchmark but the order of magnitude looks right to me: https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Improvements look amazing relative to old Python. But compare it to PHP, Javascript, Lua. Ok, it might be better than Ruby sometimes.
Imagine the alternate timeline where numpy is removing some of its C function calls because Python is fast enough on its own. I'm not sure we can make it there from here.
Who has a C++ repo that isn’t in a low level language because it has to be?
There is no production repo where python gets fast enough to replace Rust/C++.
Still great to see improvements because there are repos where python replaces other languages.
That missing word is important.
_sigh_
Why do people still say that?
1. There are no interpreted "languages", only interpreters for said language. One can compile or interpret anything.
2. When Java didn't have a Jit, it was still called "compiled language", even though it was running bytecode, same as Python.
I think this "interpreted" vs "compiled" distinction is an anachronism. Pure interpreters are almost extinct.
> (Javascript) It's a JIT compiled language with far more investment
Oh, so it's compiled now?
It's indeed due to investment, it wasn't always the case.
> C++ is a compiled language
Or is it?
Would be curious to see the performance of Java Rust Ruby W/e
I can estimate by using the energy consumption of each language as compared to C (https://storage.googleapis.com/cdn.thenewstack.io/media/2018...) but real numbers are appreciated.
Except that Pythonistas often delegate such computations to libraries like numpy that achieve C-level speed. Python is supposed to be a glue language, JS is supposed to be a language that can run on browsers, so of course the two have different goals and performances.
The only request I'd have is more granular garbage collection (marking FFI memory as refcounted for more immediate destruction).
If one would compile both, the Python version and the Javascript version, to WebAssembly and execute that, how would the performance differ then?
Python 3.11.0 native: 0m43.755s
Python 3.11.0 WebAssembly (on node.js): 1m31.305s (about 2.1x as long).
For what it is worth, I also maintain a Python --> Javascript transpiler (https://www.npmjs.com/package/pylang) and I tried it on exactly this benchmark (with some very minor changes):
PyLang: 0m19.386s (on node.js) about half as long as Python 3.11.0 native
The article uses bun, which uses Safari's JS runtime, which sometimes has significantly better performance than node.js for these sorts of benchmarks (especially for Webassembly), in my experience.
These tests were all on an M1 Macbook Pro.
JavaScript has a runtime that updates as it executes. It notices a loop is being run frequently and then compiles a new version of the loop that runs faster.
Python doesn't have this functionality.
WebAssembly is for running things in the browser and would slow everything down (Python, JS, C++). The tests in the post are run natively (on my command line).
Javascript is, for the most part, completely sandboxed and separate. Which means the JIT for Javascript is totally free to do neat things like move an object from one region of memory to another.
CPython, on the other hand, has to care about things like "was this object created in C? Is it pinned to that memory location? Does this method end up delegating to C?"
And you see that in the complexity it takes to introduce native methods in both. Moving and using data from JS to Webassembly is a PITA. Accessing data in a Python object from c ( and vice versa) is trivial.
That triviality is what gets in the way of JIT optimizations.
Bun (the runtime I used for the post), has a native C interface that's far simpler to use than Python (even with pybind/nanobind): https://github.com/oven-sh/bun#bunffi-foreign-functions-inte...
Deno (a v8-based runtime) has the same thing: https://deno.land/manual@v1.25.4/runtime/ffi_api
These use the standard C-ABI directly without requiring headers (unlike Python).
Node.js standardized a more Python-like C API called Node-API: https://nodejs.org/api/n-api.html
I believe Python can continue to add performance and is not hamstrung by the implementation.
Take the deno example, in order to invoke a C function you have to tell deno "load this lib, this dynamic function is here, and the method signature looks like this".
Look at what types are permitted to be passed between the two as well. deno isn't exposing deno objects to C, it's exposing only simple primitive types. The cpython FFI, on the other hand grants C full access to python objects.
That one layer of abstraction and distance makes all the difference and is hard to backport in.
There's a reason alternative pythons (pypy, graalpy) don't really support it and struggle to support libs that rely heavily on it (like tensorflow).
We've been using this feature heavily in Shumai.
I think you are vastly overestimating the complexity associated with this (user exposed ref-counting/garbage collection) and may not be totally up to date on what's implemented.
No, I think you are misunderstanding the problem.
I'm not saying that, with a good 3rd party lib, generating FFI APIs can't be easy and slick. I'm saying that non-python languages have more complicated FFI APIs that practically necessitate these sorts of generation libraries (like bun).
When you grab a bit of memory out of an object in C with python, you are reaching directly into python's internal representation of the object and tickling the bits there.
When you do that with deno/javascript/others, there's a layer of abstraction introduced to keep our native method from directly tickling the bits that the VM is aware of.
That's the problem.
From C, you can create a new python object, store it off in a global variable, send it back to the python vm, and later in a thread go tickle some of the bits and see that tickling in the Python VM. Because that object is the same one used by Python and C.
The complication that arises with CPython is many very popular librarys (numpy, tensorflow, pandas) rely HEAVILY on the fact that the objects they are working with in C are the same ones python uses. That's why they've been so slow to port to pypy if at all.
And that ability for C libraries to very deeply interact with the VM is exactly the problem that makes it hard to improve CPython's JIT. That's the reason other python JITs, like pypy, either don't or have very limited support for CPython's FFI capabilities.
The reason FFI works so well with other languages is they drew very tight and clear boundaries around how interactions work and who owns what when.
So please, stop spamming "bun". It's a non sequitur.
Can't it be automated? Accessing data is often done in patterns that are very similar.
Perhaps we should access Python from Rust, would that help?
"pypy" is an alternative runtime for Python which can do JIT optimizations like Javascript interpreters do. Another commenter in a different subthread already posted a quick speed comparison involving pypy, it's much faster than the standard python runtime (called CPython).
But pypy isn't necessarily a synonym for Python (which almost always means CPython). It doesn't have 100% compatibility with 3rd party libraries (it's probably 99.9% but just 1 unsupported dependency is enough to prevent your whole project from going pypy).
And also with most Python projects, the python code isn't the slow part of the stack. If you're making a web app, your cache and database is probably where 99% of your request time is spent. For scientific computing, data analysis, or ai training, you're usually using 3rd party libraries written in a "fast" language like C/Fortran/Rust with python hooks, and your python just acts as glue code to pass data between these library calls.
It's almost always the runtime. Unless the language spec makes it particularly difficult to implement a feature in a way that's performant while still respecting the spec.
Which is why the discussion on which languages are "faster" is seldom productive. Languages are not fast or slow by themselves. And, even when they are on average, implementing stuff on Assembly is no guarantee of speed. Quite often, higher level abstractions lead to optimizations that are difficult to replicate if one is operating at a lower level.
As a very trivial example, consider short circuit evaluation. It's something that you explicitly have to code in assembly, but most languages have implemented out of the box. Or garbage collection - allocating and deallocating memory on demand is memory efficient, but it's not necessarily what you want to do for performance. In many cases, deferring garbage collection allows for often-accessed memory to be retained, where as a C implementation would be constantly allocating and freeing. Fixing that requires programmer effort.
What I think matters the most is the ability of a language (+runtime) to access lower layers when needed (ASM blocks in C, FFI in Python).
When browsing Code Golf, most often than not, the fastest implementation is in assembler. And sometimes it can be 2x - 4x faster than next fastest implementation which usually is C++.
There was no equivalent competition between major organizations to make Python fast. Folks have tried, but not on the same scale. Yes, factors like language design and existing ecosystems entanglements played a role here, but if Python had the investment in performance Javascript had (financial and expert attention), it'd be dramatically faster today (though it may have required a fork or similar, depending on how tied to simplicity Guido was feeling).
To be clear, I'm not saying nobody tried to make Python faster, or that Python devs don't know what they're doing. I know very well both of those things aren't true. But, the kind of sophistication required to make Python fast in absolute terms is very hard to build and maintain, and for various reasons nobody who has found Python to be too slow has found "invest in a bunch of person-years of VM engineer attention dedicated to performance work" to be their best path.
Python is dynamic class-based OO with multiple inheritance. JS is dynamic prototypical OO, though it has recently added convenience syntax implementing single-inheritance class-based OO on top of it.
They are not the same.
but no. while both languages share some similarities (dynamically typed, some form of "objects" and "classes") there are a lot of differences, both at "language level" and at "implementation level"
Relative to python? I ran the benchmark in the post with both node.js and bun, and running it with bun consistently took ~60% more time. That was with bun v0.1.10 though. I tried with bun v0.2.1 but it just crashed.
mitata 'node sim.ts 10000000' 'bun sim.ts 10000000' 'deno run --allow-read sim_deno.ts 10000000'
cpu: 11th Gen Intel(R) Core(TM) i7-1185G7 @ 3.00GHz
runtime: shell (x86_64-unknown-linux-gnu)
benchmark time (avg) (min … max)
----------------------------------------------------------------------------------
node sim.ts 10000000 803.35 ms/iter (773.08 ms … 848.1 ms)
bun sim.ts 10000000 1.23 s/iter (1.19 s … 1.4 s)
deno run --allow-read sim_deno.ts 10000000 1.37 s/iter (1.27 s … 1.69 s)
summary
node sim.ts 10000000
1.53x faster than bun sim.ts 10000000
1.71x faster than deno run --allow-read sim_deno.ts 10000000I had to laugh a few months back when someone suggested we switch a newish project to python from PHP 8.1 because "python is built in C" and "it's better for multithreaded math functions". A lot of misinformation out there.
Also yeah, that guy wasn’t great. As i said i had to laugh.
If you want to build real software, you should be looking at Go or Rust, not Python. They are already fast, they already have concurrency, and now with Go's generics it's safe to say that their type systems are exactly what are needed to build real software in 2022.
Basically you write a ton of 3k line scripts, have them talk to each other a message queue like RabbitMQ and have a SQL Server e.g. PostgreSQL server do all the heavy lifting for you.
You can't do any program like that but 80% of programs that are written in businesses can be done like that. Typical CRUD stuff.
And you can develop it at 3x times the speed then if it was done in a more traditional language like Java/C# etc..
Python is a powerful language if you use it for the right use cases.