Not necessarily, not for dynamic languages.
With very dynamic languages you can make only very limited assumptions about e.g. function argument types, which lead you to compiled functions that have to handle any possible case.
A JIT compiler can notice that the given function is almost always (or always) used to operate on a pair of integers, and do a vastly superior specialized compilation, with guards to fallback on the generic one. With extensive inlining, you can also deduplicate a lot of the guards.
also, even mature jit compilers often only make limited improvements; jython has been stuck at near-parity with cpython's terrible performance for decades, for example, and while v8 was an enormous improvement over old spidermonkey and squirrelfish, after 15 years it's still stuck almost an order of magnitude slower than c https://benchmarksgame-team.pages.debian.net/benchmarksgame/... which is (handwaving) like maybe a factor of 2 or 3 slower than self
typically when i can get something to work using numpy it's only about a factor of 5 slower than optimized c, purely interpretively, which is competitive with v8 in many cases. luajit, by contrast, is goddam alien technology from the future
with respect to your int×int example, if an int×int specialization is actually vastly superior, for example because the operation you're applying is something like + or *, an aot compiler can also insert the guard and inline the single-instruction implementation, and it can also do extensive inlining and even specialization (though that's rare in aots and common in jits). it can insert the guards because if your monomorphic sends of + are always sending + to a rational instance or something, the performance gain from eliminating megamorphic dispatch is comparatively slight, and the performance loss from inserting a static hardcoded guess of integer math before the megamorphic dispatch is also comparatively slight, though nonzero
this can fall down, of course, when your arithmetic operations are polymorphic over integer and floating-point, or over different types of integers; but it often works far better than it has any right to. in most code, most arithmetic and ordered comparison is integers, most array indexing is arrays, most conditionals are on booleans (and smalltalk actually hardcodes that in its bytecode compiler). this depends somewhat on your language design, of course; python using the same operator for indexing dicts, lists, and even strings hurts it here
meanwhile, back in the stop-hitting-yourself-why-are-you-hitting-yourself department, fucking cpython is allocating its integers on the heap and motherfucking reference-counting them
And then there is mypyc[1] which uses mypy's static type annotations but is only slightly faster.
And various other compilers like Numba and Cython that work with specialized dialects of Python to achieve better results, but then it's not quite Python anymore.
Python to C++ translation
And here I thought that it was shocking to learn that v8 allocates doubles on the heap recently. (I mean, I'm not a compiler writer, I have no idea how hard it would be to avoid this, but it feels like mandatory boxed floats would hurt performance a lot)
Correct, to the point where at work a colleague and I actually have looked into how to force using floats even if we initiate objects with a small-integer number (the idea being that ensuring our objects having the correct hidden class the first time might help the JIT, and avoids wasting time on integer-to-float promotion in tight loops). Via trial and error in Node we figured that using -0 as a number literal works, but (say) 1.0 does not.
> i don't think local-variable or temporary floats end up on the heap in v8 the way they do in cpython
This would also make sense - v8 already uses pools to re-use common temporary object shapes in general IIRC, I see no reason why it wouldn't do at least that with heap-allocated doubles too.
I assume that integers are coerced to floats in this mode, and that there's a performance cliff if you store a non-number in such an array, but in both cases I'm just guessing.
In SpiderMonkey, as you say, we store all our values as doubles, and disguise the non-float values as NaNs.
The base line should be how heavily dynamic languages like my favourite set, Smalltalk, Common Lisp, Dylan, SELF, NewtonScript, ended up gaining from JIT, versus the original interpreters, while being in the genesis of many relevant papers for JIT research.
i didn't realize they ever jitted newtonscript
Had the Newton not been canceled, probably there would be an evolution from that support.
See "Compiling Functions for Speed"
https://www.newted.org/download/manuals/NewtonToolkitUsersGu...
that seems closer to the opposite of what you were saying in the point on which we were in disagreement?
maybe i should have said that up front!
except maybe common lisp; all the implementations i know are interpreted or aot-compiled (sometimes an expression at a time, like sbcl), but maybe there's a jit-compiled one, and i bet it's great
probably with enough work python could gain a similar amount. it's possible that work might get done. but it seems likely that it'll have to give up things like reference-counting, as smalltalk did (which most of the other languages never had)
A "Lisp interpreter" runs Lisp source in the form of s-expressions. That's what the first Lisp did.
A "Lisp compiler" compiles Lisp source code to native code, either directly or with the help of a C compiler or an assembler. A Lisp compiler could also compile source code to byte code. In some implementations this byte code can be JIT compiled (ABCL, CLISP, ...).
The first Lisp provided a Lisp to assembly compiler, which compiled Lisp code to assembly code, which then gets compiled to machine code. That machine code could be loaded into Lisp and functions then could be native machine code.
The Newton Toolkit could compile type declared functions to machine code. That's something most Common Lisp compilers do, sometimes by default (SBCL, CCL, ... by default directly compile source code to machine code).
SBCL:
* (defun add (a b) (declare (fixnum a b) (optimize (speed 3))) (+ a b))
ADD
* (disassemble #'add)
; disassembly for ADD
; Size: 104 bytes. Origin: #x7006E1789C ; ADD
; 89C: 0000018B ADD NL0, NL0, NL1
; 8A0: 0A0000AB ADDS R0, NL0, NL0
; 8A4: E7010054 BVC L1
; 8A8: BD2A00B9 STR WNULL, [THREAD, #40] ; pseudo-atomic-bits
; 8AC: BC7A47A9 LDP TMP, LR, [THREAD, #112] ; mixed-tlab.{free-pointer, end-addr}
; 8B0: 8A430091 ADD R0, TMP, #16
; 8B4: 5F011EEB CMP R0, LR
; 8B8: E8010054 BHI L2
; 8BC: AA3A00F9 STR R0, [THREAD, #112] ; mixed-tlab
; 8C0: L0: 8A3F0091 ADD R0, TMP, #15
; 8C4: 3E2280D2 MOVZ LR, #273
; 8C8: 9E0300A9 STP LR, NL0, [TMP]
; 8CC: BF3A03D5 DMB ISHST
; 8D0: BF2A00B9 STR WZR, [THREAD, #40] ; pseudo-atomic-bits
; 8D4: BE2E40B9 LDR WLR, [THREAD, #44] ; pseudo-atomic-bits
; 8D8: 5E0000B4 CBZ LR, L1
; 8DC: 200120D4 BRK #9 ; Pending interrupt trap
; 8E0: L1: FB031AAA MOV CSP, CFP
; 8E4: 5A7B40A9 LDP CFP, LR, [CFP]
; 8E8: BF0300F1 CMP NULL, #0
; 8EC: C0035FD6 RET
; 8F0: E00120D4 BRK #15 ; Invalid argument count trap
; 8F4: L2: 1C0280D2 MOVZ TMP, #16
; 8F8: 0AFBFF58 LDR R0, #x7006E17858 ; SB-VM::ALLOC-TRAMP
; 8FC: 40013FD6 BLR R0
; 900: F0FFFF17 B L0
NIL
I've entered a function and it gets ahead of time compiled to non-generic machine code.Calling the function ADD with the wrong numeric arguments is an error, which will be detected both a compile and at runtime.
* (add 3.0 2.0)
debugger invoked on a TYPE-ERROR @7006E17898 in thread
#<THREAD "main thread" RUNNING {70088224A3}>:
The value
3.0
is not of type
FIXNUM
when binding A
Redefinition of + will do nothing to the code. The addition is inlined machine code.Apple's Dylan IDE and compiler was implemented in Macintosh Common Lisp (MCL). MCL then was not a part of the Dylan runtime.
I would think that Open Dylan (the Dylan implementation originally from Harlequin) can also generate LLVM bitcode, but I don't know if that one can be JIT executed. Possibly...
CLISP has a byte code machine, for which a JIT can be used.
There might be others.
how much of a performance boost does abcl get from the hotspot jit compared to, say, interpreted clisp