A JIT compiler is a big deal for performance improvements, especially where it matters (in large repetitive loops).
Anyone cynical about the potential a python JIT offers should take a look at pypy which has a 5x speed up over regular python, mainly though JIT operations: https://www.pypy.org/
- "C/C++/Fortran libs are Python"
- "Python is too dynamic", while disregarding Smalltalk, Common Lisp, Dylan, SELF, NewtonScript JIT capabilities, all dynamic languages where anything can change at any given moment
Also no mentality shift is expected on the "Python is too dynamic" -- which is a strange thing to say anyway -- because Python is not getting any more static due to these JIT news.
Having a Python with JIT, in many cases it will be fast enough for most cases.
Data science running CUDA workloads isn't the only use case for Python.
I don't do data science.
I know Python since version 1.6, and is my scripting language in UNIX like environments, during my time at CERN, I was one of the CMT build infrastructure build engineer on the ATLAS team.
It was never been the language I would reach for when not doing OS scripting, and usually when a GNU/Linux GUI application happens to be slow as mollasses, it has been written in Python.
Flask, gunicorn, low single digit millisecond latency. Definitely optimised for latency over throughput, but not so much that we've replatformed it onto something that's actually designed for low latency :P. Callers all cache heavily with a fairly high hit ratio for interactive callers and a relatively low hit ratio for batch callers.
But on the whole, machines are cheaper than other engineering approaches to scaling.
For us, and many others, fast enough is fast enough.
shrug. If we're talking personal experience, I've been using Python since 1.4. It's been my primary development language since the late 1990s, with of course speed critical portions in C or C++ when needed - and I know a lot of people who also primarily develop in Python.
And there's a bunch of Python development at CERN for tasks other than OS scripting. ("The ease of use and a very low learning curve makes Python a perfect programming language for many physicists and other people without the computer science background. CERN does not only produce large amounts of data. The interesting bits of data have to be stored, analyzed, shared and published. Work of many scientists across various research facilities around the world has to be synchronized. This is the area where Python flourishes" - https://cds.cern.ch/record/2274794)
I simply don't see how a Python JIT is going to make that much of a difference. We already have PyPy for those needing pure Python performance, and Numba for certain types of numeric needs.
PyPy's experience shows we'll not be expecting a 5x boost any time soon from this new JIT framework, while C/C++/Fortran/Rust are significantly faster.
Unfortunely.
> And there's a bunch of Python development at CERN for tasks other than OS scripting
Of course there is, CMT was a build tool, not OS scripting.
No need to give me CERN links to me to show me Python bindings to ROOT, or Jupyter notebooks.
> PyPy's experience shows we'll not be expecting a 5x boost any time soon from this new JIT framework, while C/C++/Fortran/Rust are significantly faster.
I really don't get the attitude that if it doesn't 100% fix all the world problems, then it isn't worth it.
> I really don't get the attitude that if it doesn't 100% fix all the world problems, then it isn't worth it.
Then it's a good thing I'm not making that argument, but rather that "Having a Python with JIT, in many cases it will be fast enough for most cases." has very little information content, because Python without a JIT already meets the consequent.
Do you agree with me that Python is already fast enough for most cases, even without a JIT?
If not, how would a 30% boost improve things enough to change the balance?
https://stackoverflow.com/questions/36526708/comparing-pytho...
But pjmlp, I use Python because it's a wrapper for C/C++/Fortran libs. - Chocolate Giddyup
In Common Lisp not anything can change at any moment. Especially not in implementations where one uses AOT compilation like SBCL, ECL, LispWorks, Allegro CL, ... and so on. They have optimizing compilers which gradually can remove dynamic runtime behavior, upto supporting almost no dynamic runtime behavior.
Stuff which is supported: type specific code, inlining, block compilation, removal of development tools, ...
JIT implementations are rare in the Common Lisp world. They are mostly only used in implementations which use a byte-code virtual machine (CLISP, ABCL, ...). Common Lisp implementations mostly compile either directly to native code or via C compilers. The effect is that native AOT compiled code is much faster.
However, last time I used it, it (1) didn’t work with many third-party libraries (e.g. SciPy was important for me), and (2) didn’t work with object-oriented code (all your @njit code had to be wrapped in functions without classes). Those two has limited for which projects I could adopt Numba in practice, despite loving it in the cases it worked.
I don’t know what limitations the built-in Python JIT has, but hopefully it might be a more general JIT that works for all Python code.
(Actually a monkey-patched version to be able to set njit arguments)
The base line should be how heavily dynamic languages like my favourite set, Smalltalk, Common Lisp, Dylan, SELF, NewtonScript, ended up gaining from JIT, versus the original interpreters, while being in the genesis of many relevant papers for JIT research.
i didn't realize they ever jitted newtonscript
Had the Newton not been canceled, probably there would be an evolution from that support.
See "Compiling Functions for Speed"
https://www.newted.org/download/manuals/NewtonToolkitUsersGu...
that seems closer to the opposite of what you were saying in the point on which we were in disagreement?
maybe i should have said that up front!
except maybe common lisp; all the implementations i know are interpreted or aot-compiled (sometimes an expression at a time, like sbcl), but maybe there's a jit-compiled one, and i bet it's great
probably with enough work python could gain a similar amount. it's possible that work might get done. but it seems likely that it'll have to give up things like reference-counting, as smalltalk did (which most of the other languages never had)
A "Lisp interpreter" runs Lisp source in the form of s-expressions. That's what the first Lisp did.
A "Lisp compiler" compiles Lisp source code to native code, either directly or with the help of a C compiler or an assembler. A Lisp compiler could also compile source code to byte code. In some implementations this byte code can be JIT compiled (ABCL, CLISP, ...).
The first Lisp provided a Lisp to assembly compiler, which compiled Lisp code to assembly code, which then gets compiled to machine code. That machine code could be loaded into Lisp and functions then could be native machine code.
The Newton Toolkit could compile type declared functions to machine code. That's something most Common Lisp compilers do, sometimes by default (SBCL, CCL, ... by default directly compile source code to machine code).
SBCL:
* (defun add (a b) (declare (fixnum a b) (optimize (speed 3))) (+ a b))
ADD
* (disassemble #'add)
; disassembly for ADD
; Size: 104 bytes. Origin: #x7006E1789C ; ADD
; 89C: 0000018B ADD NL0, NL0, NL1
; 8A0: 0A0000AB ADDS R0, NL0, NL0
; 8A4: E7010054 BVC L1
; 8A8: BD2A00B9 STR WNULL, [THREAD, #40] ; pseudo-atomic-bits
; 8AC: BC7A47A9 LDP TMP, LR, [THREAD, #112] ; mixed-tlab.{free-pointer, end-addr}
; 8B0: 8A430091 ADD R0, TMP, #16
; 8B4: 5F011EEB CMP R0, LR
; 8B8: E8010054 BHI L2
; 8BC: AA3A00F9 STR R0, [THREAD, #112] ; mixed-tlab
; 8C0: L0: 8A3F0091 ADD R0, TMP, #15
; 8C4: 3E2280D2 MOVZ LR, #273
; 8C8: 9E0300A9 STP LR, NL0, [TMP]
; 8CC: BF3A03D5 DMB ISHST
; 8D0: BF2A00B9 STR WZR, [THREAD, #40] ; pseudo-atomic-bits
; 8D4: BE2E40B9 LDR WLR, [THREAD, #44] ; pseudo-atomic-bits
; 8D8: 5E0000B4 CBZ LR, L1
; 8DC: 200120D4 BRK #9 ; Pending interrupt trap
; 8E0: L1: FB031AAA MOV CSP, CFP
; 8E4: 5A7B40A9 LDP CFP, LR, [CFP]
; 8E8: BF0300F1 CMP NULL, #0
; 8EC: C0035FD6 RET
; 8F0: E00120D4 BRK #15 ; Invalid argument count trap
; 8F4: L2: 1C0280D2 MOVZ TMP, #16
; 8F8: 0AFBFF58 LDR R0, #x7006E17858 ; SB-VM::ALLOC-TRAMP
; 8FC: 40013FD6 BLR R0
; 900: F0FFFF17 B L0
NIL
I've entered a function and it gets ahead of time compiled to non-generic machine code.Calling the function ADD with the wrong numeric arguments is an error, which will be detected both a compile and at runtime.
* (add 3.0 2.0)
debugger invoked on a TYPE-ERROR @7006E17898 in thread
#<THREAD "main thread" RUNNING {70088224A3}>:
The value
3.0
is not of type
FIXNUM
when binding A
Redefinition of + will do nothing to the code. The addition is inlined machine code.Apple's Dylan IDE and compiler was implemented in Macintosh Common Lisp (MCL). MCL then was not a part of the Dylan runtime.
I would think that Open Dylan (the Dylan implementation originally from Harlequin) can also generate LLVM bitcode, but I don't know if that one can be JIT executed. Possibly...
CLISP has a byte code machine, for which a JIT can be used.
There might be others.
how much of a performance boost does abcl get from the hotspot jit compared to, say, interpreted clisp
Not necessarily, not for dynamic languages.
With very dynamic languages you can make only very limited assumptions about e.g. function argument types, which lead you to compiled functions that have to handle any possible case.
A JIT compiler can notice that the given function is almost always (or always) used to operate on a pair of integers, and do a vastly superior specialized compilation, with guards to fallback on the generic one. With extensive inlining, you can also deduplicate a lot of the guards.
also, even mature jit compilers often only make limited improvements; jython has been stuck at near-parity with cpython's terrible performance for decades, for example, and while v8 was an enormous improvement over old spidermonkey and squirrelfish, after 15 years it's still stuck almost an order of magnitude slower than c https://benchmarksgame-team.pages.debian.net/benchmarksgame/... which is (handwaving) like maybe a factor of 2 or 3 slower than self
typically when i can get something to work using numpy it's only about a factor of 5 slower than optimized c, purely interpretively, which is competitive with v8 in many cases. luajit, by contrast, is goddam alien technology from the future
with respect to your int×int example, if an int×int specialization is actually vastly superior, for example because the operation you're applying is something like + or *, an aot compiler can also insert the guard and inline the single-instruction implementation, and it can also do extensive inlining and even specialization (though that's rare in aots and common in jits). it can insert the guards because if your monomorphic sends of + are always sending + to a rational instance or something, the performance gain from eliminating megamorphic dispatch is comparatively slight, and the performance loss from inserting a static hardcoded guess of integer math before the megamorphic dispatch is also comparatively slight, though nonzero
this can fall down, of course, when your arithmetic operations are polymorphic over integer and floating-point, or over different types of integers; but it often works far better than it has any right to. in most code, most arithmetic and ordered comparison is integers, most array indexing is arrays, most conditionals are on booleans (and smalltalk actually hardcodes that in its bytecode compiler). this depends somewhat on your language design, of course; python using the same operator for indexing dicts, lists, and even strings hurts it here
meanwhile, back in the stop-hitting-yourself-why-are-you-hitting-yourself department, fucking cpython is allocating its integers on the heap and motherfucking reference-counting them
And then there is mypyc[1] which uses mypy's static type annotations but is only slightly faster.
And various other compilers like Numba and Cython that work with specialized dialects of Python to achieve better results, but then it's not quite Python anymore.
Python to C++ translation
And here I thought that it was shocking to learn that v8 allocates doubles on the heap recently. (I mean, I'm not a compiler writer, I have no idea how hard it would be to avoid this, but it feels like mandatory boxed floats would hurt performance a lot)
Correct, to the point where at work a colleague and I actually have looked into how to force using floats even if we initiate objects with a small-integer number (the idea being that ensuring our objects having the correct hidden class the first time might help the JIT, and avoids wasting time on integer-to-float promotion in tight loops). Via trial and error in Node we figured that using -0 as a number literal works, but (say) 1.0 does not.
> i don't think local-variable or temporary floats end up on the heap in v8 the way they do in cpython
This would also make sense - v8 already uses pools to re-use common temporary object shapes in general IIRC, I see no reason why it wouldn't do at least that with heap-allocated doubles too.
I assume that integers are coerced to floats in this mode, and that there's a performance cliff if you store a non-number in such an array, but in both cases I'm just guessing.
In SpiderMonkey, as you say, we store all our values as doubles, and disguise the non-float values as NaNs.
Not pursuing JIT or efficient compilation in general was a deliberate decision way back when Python made some kind of sense. It was the simplicity of implementation valued over performance gains that motivated this decision.
The mantra Python programmers liked to repeat was that "the performance is good enough, and if you want to go fast, write in C and make a native module".
And if you didn't like that, there was always Java.
Today, Python is getting closer and closer to be "the crappy Java with worse syntax". Except we already have that: it's called Groovy.
The language is definitely getting more complex syntactically, and I'm not a huge fan of some of those changes but it's no where near Java or C++ or anything else. You can still write simple Python with all of these changes.
Read it again. It seems you were reading too fast. I'm talking about the future, not the change being discussed right now.
> It's great that the core devs are keeping up with the time.
You mistake the influence of Microsoft and their desire to sell features for progress. Python is actually regressing as a system. It's becoming worse, not better. But it's hard to see the gestalt of it if all you are looking for is the new features.
> it's no where near Java
That is true. Java is a much more simple and regular (not in the automata theory sense) language. Today, if you want a simpler language, you need to choose Java over Python (although neither is very simple, so, preferably, you need a third option).
> You can still write simple Python
I can also write simple C++ if I limit what I use from the language to a very small subset. This says nothing about the simplicity of the language...