Can JIT be Faster?
tirania.org
tirania.org
But this disregards one of the major selling points of profiling + optimizing JIT: you can apply optimizations that cannot be proven safe and you can get away with it. This means you can not only be "as fast as C", you can implement a JIT that will mop the floor with C.
Good JITs are already taking advantage of this, and the very first comment on Miguel's article points this out. This (four-year-old!) article by Charles Nutter explains it a bit more in the context of the JVM:
Anyway, as I said below, there is nothing that stops you from implementing a JIT for C, it's just not cost effective from a "where do i put money to get better performance" perspective. If you did this, these JIT's would do just as well as the ones you think will "mop the floor" with C.
Something that seems to be missed here is that the whole point of compilers (and JITs) having intermediate IR is to abstract away the language to a point that you can optimize and codegen stuff.
So talking about a "statically compiled" or "dynamically compiled" language is well, silly. Most popular languages could be compiled either way[1], it just happens that we don't because it's not "better enough' in the real world.
If you want a real world example, LLVM will happily JIT C/C++ code for you, for example. It doesn't care that it's C/C++. It will even run it interpreted if you like.
[1] Some languages would still require a support library with an interpreter/jit/compiler if you statically compiled them, in order to handle dynamic class loading.
Edit: the old C2 wiki actually has something useful to say about this.
For example performance critical code that has a lot of conditional branches, but only a few are taken at runtime, will often work faster in a JIT.
(Of course, the typical solution is not to write performance critical code with a large number of branches, i'm just giving you examples)
I remember a study posted here a few months back that supported the notion that most web apps have a certain "steady state" that they reach after they start up, where almost all the types of variables are known, and most of the funkier "dynamic" changes don't happen. If one can get a JIT to remember the tracing information from session to session, it should be able to run as fast as or faster than C.
If you bundle JIT with garbage collection, bounds checking, and all the usual things that come with the JVM/CLR and execute on current hardware, probably not.
If you were to JIT a language like C it would probably be faster as you could optimize to specific hardware. OpenCL uses JIT for example and it's many times faster than static compilation (because it can take advantage of extra hardware at runtime).
The issue to contend with is the reality that about 30 years of the industry have gone into making C fast, everything from compilers to the way CPUs are designed. There's likely no inherent reason any style of compilation / execution couldn't beat static compilation, it's just we need the same level of resources and hardware support for those execution models.
For instance with garbage collection and bounds checking you no longer need an MMU.
That said, although bounds checking eliminates some of the reasons for MMUs, it does not really solve many other features that we take for granted today that depend on MMUs, like on-demand-loading.
Garbage collectors and VMs typically use the MMU to improve the performance of their own operations. You use the MMU to cause execution to stop, you use MMUs to setup redzones on the stack to avoid checking on every function entry point for how much space is left on the stack and VMs use the MMU to map their metadata into their address space without having to load them from file and provide the services on demand.
Microsoft had a research OS called Singularity that was written in C# mostly and could be run in real memory mode because they didn't need the barriers.
Also, I do not accept that OpenCL argument. IMO, OpenCL is more similar to static compilation than to a JIT. It is as if you recompiled your kernels whenever you changed your video card (perhaps also when you change its configuration); it does not do things that a JIT would do, such as "hm, most of the pixels are black; let's optimize for that case".
Even if you didn't want to do that, on demand CFL reachability formulations of pointer analysis can calculate reasonable pointer results for individual pointers fast enough if it became important.
Realistically however, no JIT is going to do advanced pointer analysis, unless you have CPU to burn. As for "conclusively prove that aliasing does not happen", you don't actually have to, because if it is truly going to improve performance, you can insert runtime checks.
if (&a == &b) <do super fast thing> else <slower fallback code>
Nobody does it because the cases in which it would help are small, and most of those folks are willing to do profiling/training runs, and get about as good results, even if it requires more pain.
As Herb said, the languages were made for different tradeoffs. Speaking as a compiler guy, yes, you can make almost any of them as fast as each other given enough time and effort.
But putting time and effort into a JIT may not be as effective as choosing a different language.
But the people who advocate C++ wouldn't recognize that language as a C++, because all the overhead of a JIT runtime violates the "only pay for what you need" dictum.
(E.g. consider the uptake of Managed C++, C++/CLI, etc.; which effectively are what you describe.)
The features Java redefined were mainly for security reasons, and pointer safety, not because it made it possible to run with a JIT
As a proof, you can JIT C/C++ with LLVM right now if you wanted to.
It's not a great JIT because it has no adaptive compilation heuristics, etc, but it works just fine.
C/C++ very specifically leave most of the execution environment undefined. They do not require anything that makes JIT'ing require language level changes. Your JIT may be unsafe (IE crash as easily as a regular C/C++ program), or otherwise require basically native access (IE not really viable for sandboxing in a browser), but it would work fine
Things like Managed C++/etc required language changes because they wanted it to interoperate with a common intermediate language and runtime.
There are things that you can do in C++ that you can't do in CLR (again, because the CLR chose safety/etc over other considerations), and vice versa (IE garbage collection at a language level), so managed C++ required language changes to support this.
You can compile standard C++ with C++/CLI just fine, and have it run on a fine managed platform. "IJW" was the slogan: It Just Works. Still didn't go anywhere.
(The CLR did not choose safety over other considerations. It has a safe subset, but it is not like the JVM; it has unsafe operations too. When you write IL for it, you are not even required to use GC memory; you can manipulate things at the level of untyped blobs of data if needs be.)
So, the most "bang for your buck" may be using a JIT for a language like Python. Sure, it won't be as fast (at least, not until a huge amount of work has been done on it), but it can get faster. Trying to beat C + a good C compiler is ... ambitious.
Doing it for Python will be impossibly hard, yes.
There's no way you can make broad statements (like "nothing will beat C" or "fortran is actually faster than C for numerical stuff") without having a few exceptions. Actually, Pypy is faster than C, in extremely specific circumstances (i.e. ones which were set up by the Pypy people to show off, see "PyPy faster than C on a carefully crafted example" ... thought they point out there are things you could do in C to make it win).
The most bang for you buck in JITs is for dynamic languages, because while stuff like type inference can exponentially explode, you'll only really have 1 type in many cases (i.e. a float), and the JIT can pick that up. And getting Python within an order of magnitude of C would be a massive win.
This helps solve the problem that compiler writers often have multiple alternatives they can generate, but have to make hard decisions as to which to pick.
I'm not saying that implementing this will be easy, but then again all of the easy things have already been done.
I also think software people still have a innate fear of "wasting" CPU/memory/disk which is hard to overcome.
Compare with other databases that put an inordinate amount of effort and tuning into coming up with the one true query plan per query.
Wait, what? Hotspot does a "quick compile" on the first pass? When did this happen?
Note that it does just interpret by default.
Could you point to somewhere on the interweb which discusses this in detail?
Do they want people to have a hard time reading what they write?