Someone is missing the boat: custom ISA is the future.
Someone is missing the boat: custom ISA is the future.
Custom ISA are 99% of the case useless marketing porn for hardware provider.
The reality is that most Linux distributions in 2020 are still binary compiled with SSE41 instructions (maybe AVX at best). Instructions that shipped decades ago.
AVX2 and AVX512 being barely supported in some HPC parts and the libc itself.
SIMD is good, but fragmentation of instruction set is much more of a problem in most case.
Nobody will dare to support your fancy new instruction if it covers 2% of the CPU used world wide and make their software a pain in the butt to distribute.
On that Linus is perfectly right.
The nice thing about JITCs is:
1. They run parallel to the application so can use those spare cores your program isn't using.
2. Especially true as they run a lot during startup when your app is probably single threaded anyway, even for something like a web server where it's going to be multi-threaded soon.
3. Upgrading the VM upgrades the compiler, so new instructions can be used almost immediately, as long as your software is tested/runs on the newest VMs.
JVM has historically tried to auto-vectorise everything and not been all that good at it, but now it's getting vector extensions that generalise to newer instruction sets fairly well, so we might start to see more HPC done in higher level languages.
That's theory. Practice is that JIT-compiled language generally performs even worst than compiled language with an outdated instructions set.
And there is reasons to that, compiler passes are expensive, including vectorization, and you specially do not want anything expensive running in your critical JIT passes.
Some language like Julia still do it but to the price of a very long compiling and starting time.
Something that has never been acceptable over JVM or in Nodejs.
Generally JIT outperforms AOT for a language like Java by about 20%. For Scala, it's even more. For a more dynamic language like JavaScript or Ruby you don't even try to AOT it at all the difference is so huge.
Even in languages that are normally AOT compiled, profile guided optimisations make a big difference. For C++ I've seen figures in the 15%+ range. That's a lot! Most projects don't use PGO though because it's a pain to deploy.
C2 does auto-vectorisation and it's not a slow compiler. It runs in parallel with the app, it's not a problem. The difficulty with auto-vectorisation isn't how much time you have to compile it, it's more that matching all the different loop constructs to the instructions is a very complicated problem that ends up needing tons of special cases in the compiler. It's a problem of code complexity rather than runtime performance. And it's opaque to the user: if they change their code structure a bit and it's no longer recognised as a vectorizable template, it'll stop being vectorised and performance drops off a cliff. You can't really see that though as a developer ahead of time because it's all just "best effort" optimisation. Game devs in particular hate that and would rather have a compiler error than silently bailed out optimisation passes.
That's why Java is switching to explicit vectorisation. It's not about JITs not having enough CPU time, it's about giving the developer reliable and predictable performance even in the face of arbitrary refactorings.
I personally have a problem when compromises are made for the general case to satisfy a special case, especially when the special case motivations appear to be bound mostly in marketing reasons. Even more so if the special case is well-known to adversely impact performance in the general case.
Before I saw this post, I didn't even know ARM had some complex javascript instruction. Contrast that with - You can't watch or read a review for any current gen Intel server product without being presented with at least one AVX benchmark. Is this marketing or is it practical engineering? For how many applications is AVX availability actually a hard constraint?
Name two
Did you miss the part about custom instructions?
Custom extensions aren't being made to beat benchmarks.