TornadoVM: Running Java on GPUs and FPGAs
infoq.com
infoq.com
So basically Java syntax for some kind of restricted C/CUDA dialect. How can you even say you're running Java if you don't have objects or dynamic allocation? Everytime the promise of a general purpose programming language running on GPUs is made, this is actually what is delivered, eg. marketing fluff, not compilers actually getting smarter in any fashion
And when you think about how a GPU works, it completely makes sense. A high level language for a GPU will not look like Java
Recursion is now supported in CUDA via dynamic parallelism (or faking the stack) but it is not performant at all.
The ideal language for GPU programming is closer to Julia or a MATLAB-like language where arrays and matrices are first-class.
That pulled up "Chapel" in my head and a couple of clicks later the net pulled this up:
"PGAS (Partitioned Global Address Space) programming models were originally designed to facilitate productive parallel programming at both the intra-node and inter-node levels in homogeneous parallel machines. However, there is a growing need to support accelerators, especially GPU accelerators, in heterogeneous nodes in a cluster. Among high-level PGAS programming languages, Chapel is well suited for this task due to its use of locales and domains to help abstract away low-level details of data and compute mappings for different compute nodes, as well as for different processing units (CPU vs. GPU) within a node. In this paper, we address some of the key limitations of past approaches on mapping Chapel on to GPUs ..."
https://pldi19.sigplan.org/details/CHIUW-2019-papers/4/GPUIt...
Can someone explain this further? I am genuinely interested in learning what makes Java etc not suitable for GPUs.
Java is a dynamically-allocated, garbage-collected language. So far as I know, GPU code doesn't have the ability to dynamically allocate memory. (I mean, I guess you could, but you'd probably need one memory pool per parallel computational unit. Otherwise, you'd need a global lock on the global memory pool.)
Java is a general programming language. GPUs are not general-purpose processors.
That's off the top of my head. There may be other reasons as well.
It is also possible to do native memory allocation in Java.
Imposing restrictions on a Java subset is hardly any different of imposing restrictions on C and C++, which apparently is ok to do.
C/C++ structs(classes) on the other hand represents a memory layout so an array of "objects" would be laid out as a single contigious block of memory. (and this works well for a GPU)
There is/was a propsal for Java for value types but it hasn't yet been approved. https://openjdk.java.net/jeps/169
(As a side note, C# classes are used as refernces like in Java but they also have "struct" in C# that behaves like C structs, and this is what the Unity boost compiler leverages and could be used for a GPU variant with less restrictions)
Grated it isn't perfect, hence the value types proposal.
Here is one of the CUDA commercial offerings for Java and .NET,
Only if escape analysis says it didn't go anywhere. Which works for iterators but it's not going to work for input & output arrays to this system, especially since it's asynchronous meaning it did escape the stack frame.
A pointer can point to object of any arbitrary sub type in any memory location. Suddenly you have both branching and random memory access. If your program is branching heavily then you may end up with 1 piece of data being computed by a core that can handle 64 pieces of data at once and random memory access kills you because GPUs don't have huge caches like CPUs. The way GPUs deal with high latency is that they simply batch a large amount of work and switch to a different "thread" during memory access. If all your threads are busy loading from memory then you won't see any speedups.
Finally GPUs do not actually have 4096 cores. The RX Vega 64 only has 64 cores and those cores only run at around 1.3GHz. A Ryzen with 8 cores running at 4GHz with out of order execution can trivially outperform a GPU that is only using 1.6% (1/64) of its theoretical performance.
The simplest is "drop me off anywhere on the route, but only pick up on demand at scheduled stops"
The more complex is "wear one diversion every now and then"
Some small buses actually self-optimize to a locally efficient route. Post-Buses in rural locations.
It is surely possible GPU can segment, into small disjoint sets of "like" and so continue to offer parallelism, but at reduced intensity? Or, have point inefficiencies in the computation, toward the overall goal of "all alike"
This fits really nicely with scaling up on the cloud. If you have a parallel workload that you need to auto-scale very quickly, you could migrate from small CPU to big CPU, to AVX512, to GPUs to FPGAs, all automatically. Maybe it won't produce code that can beat hand-tuned Verilog but it'll do it a lot faster and cheaper.
Re: no objects. That doesn't actually mean no objects. Recall this is based on Graal. Graal is really, really good at scalar replacing objects automatically, i.e. converting allocations into local variables, even when they're being e.g. allocated in a method and then returned. This is how Truffle gets such good performance: it just recursively inlines everything into one giant method and then wires up all the allocation sites to use sites. It turns out you can really use a lot of objects this way, yet at runtime they all disappear.
Given it's based on the same compiler infrastructure I'd expect that to also be true of TornadoVM; perhaps a TornadoVM developer can chime in. Probably you can do method calls and working with wrapper types without hitting the limitation, or at least, it can be added without too much effort (I guess the hard part is flattening the types into arrays as value types, which may require a more sophisticated optimisation pass).
It's pretty hard to buy a computer these days that doesn't have some sort of GPU. Either it has on board graphics or it has a board connecting it to a screen.
That's OpenCL's big thing. Which TornadoVM is sitting on top of. And OpenCL never really took off, it's in a bit of a weird spot currently being partially deprecated.
It's a fine idea on paper, but there doesn't seem to be much of a market demand for it. Probably because it's not a thing anyone really needs. Step debugging is about the only reason to bother at all, but you really don't need any sort of heterogeneous compute migration system framework thing to pull that off. It's just a local flag or even a re-compile option.
> This fits really nicely with scaling up on the cloud. If you have a parallel workload that you need to auto-scale very quickly, you could migrate from small CPU to big CPU, to AVX512, to GPUs to FPGAs, all automatically.
Scaling up in the cloud is a perfect example of why TornadoVM itself isn't that interesting. Just like you can't write a normal multithreaded single process application and have it magically scale out over a network, nor can you here write a normal CPU process and have it automatically migrate between CPU & GPU. You have to design your program from the ground up around concepts like SIMT, memory coalescing, DMA bandwidth/latency, etc... These aren't just "here's some tricks to go _even faster_" things, these are "you must do this or don't even bother using the GPU at all" things.
To make this work you don't need to port a different language or VM to the platform, you need the compute equivalent of Kubernetes or Hadoop or whatever other cloud scaling infrastructure system you want. Otherwise this largely boils down to a Java to OpenCL transpiler. Cool, but it's not advancing the state of art here, either.
Also, TornadoVM supports method calls and invitations to native code (e.g., Math library). The examples in the presentations are for simplicity but we have some use cases, such as KFusion Kinet with ~7k lines of Java code for computer vision (https://github.com/beehive-lab/kfusion-tornadovm).
They basically do. There's a minimal "core" that does things like branching and that core manages a bunch of threads that all execute the same instruction in parallel. But there's a lot of those minimal cores still, so transistors spent on it need to be worth it. Spending transistors to make a nicer programming model means fewer transistors spent going faster.
Alternatively the design you're looking for is Intel's Xeon Phi, which died. Intel's Xe looks maybe more traditional GPU in architecture but TBD, it might be less restrictive in this regard. Maybe.
Using dynamic memory liberally, having lots of indirect branches and non-local memory access - that's pretty bad for performance on the CPU as well, even though it is optimized for that. It's just that people don't notice because they're used to programming that way.
It's not any different from, say, JavaCard, which doesn't even have java.lang.String or garbage collection.
> And when you think about how a GPU works, it completely makes sense. A high level language for a GPU will not look like Java
It can still have Java syntax, which can also leverage existing IDEs and other libraries (eg. for unit testing).
And getting Java to even run on these is going to remove so many features that there's no real reason to pick it anyway. It's sorta like how JavaCard that removes keywords like "new", "long" and "throw". Like at that point it's a stretch to even call it Java.
Write the performance critical parts in this restricted subset to leverage the hardware capability.
Write the rest of the application in ordinary JVM languages.
This is restricted enough to essentially be a whole other language.
At the moment, I'm integrating some avatar dressing code I'd written. That'll be a pretty good test of its utility. But so far, pretty impressed.
It was a big effort to overcome that to have basic trig functions good enough to make an "asteroids" type game.
I read about it not being available. The textbooks said so. I checked. It was a compile error to use float.
Here are a couple of links. Remember this was a very long time ago. Like 20 years.
https://books.google.com/books?id=O4m0TrliwscC&pg=PA297&lpg=...
https://www.javaworld.com/article/2076023/go-wireless-with-j...
With this you can write the critical method in subset of java, test it using standard java tooling and then run it on gpu.
All your tools will work because it's a subset of java - so you can use the same code formatter, static analysis tools, build scripts, continuous integration, IDE highlighting and completion.
You can use IDE refactoring over the whole codebase instead of doing it in 2 phases on 2 different languages (or more often doing the performance-sensitive parts by hand).
That's a big difference, much better than maintaining codebase in 2 languages one of which usually isn't supported by your tools.
When Khronos realised they should have had their own PTX (SPIR) and higher level programming models, the race was already lost.
https://arcb.csc.ncsu.edu/~mueller/cluster/nvidia/2.0/Progra...
When did C started supporting templates?
And that's what their docs say too.
C++ class was added in 2010 as answered in another thread, in CUDA 3.0.
I thought you were pretty sure about CUDA 5.0.
Which is easily proven when reading the documentation for the first set of CUDA releases.
As for the rebooted hardware design,
https://developer.nvidia.com/cuda-toolkit-21-january-2009
> C++ templates are now supported in CUDA kernels
Or maybe 2010 with CUDA 3.0, to make up for your C++ classes.
https://developer.nvidia.com/cuda-toolkit-30-downloads
> C++ Class Inheritance and Template Inheritance support for increased programmer productivity
Let me check the date, I think we are in 2020, so if my math doesn't betray me, it looks like about 10 years to me.
From where I am standing, I am pretty sure you never used CUDA.
OpenCL advertised "write once, run everywhere" approach, it was completely bonkers when it comes to HPC. Optimizations that one applies on different hardware make code look completely different. It has its niche, but not in HPC. I am no longer in this industry, so maybe things changed. But at the time, some teams picked up OpenCL and quickly dropped, as it didn't give enough bang for a buck, and tooling didn't compare with rather polished CUDA stuff.
"Optimizations that one applies on different hardware make code look completely different."
Unless you got C program where this statement is true.
I'm still having hard time to believe that kernels optimized for SIMD architecture will be useful on CPU and vice versa. And OpenCL people advertised this, if memory serves me well.
Given Objects are not supported, I am having a hard time seeing any valid use case for this over C,C++,Rust etc...
Why would you need to avoid that?
Even Rust wouldn't be that much, if it wasn't being built on top of LLVM.
Also, just as a separate point can someone point me to a design that acheives 5TFlops on Stratix 10, let alone the claimed 10TFlops - because my understanding is to get that performance you would need to run the fabic at 1GHz and use 100% of the DSPs - which is frankly hilariously impossible.
With those compiler specializations, we aim to close the performance gap between hand-tuned code and generated code.
Obviously he's still hit an extremely useful sweet spot right there. I don't know many non-masochists who would choose to use VHDL over Java, that's for sure. Who out there would really rather work out the intricacies of VHDL instead of just compiling a Java class and calling it a day?
Ditto for the GPU. You could learn Vulkan, but if what you're doing is just GPGPU type stuff, why?
Just throw it in a Java class and call it be done with it.
You don't need to be a genius, but you need to know a lot about low-level stuff. "Simple" matrix-vector multiplication is a task where quirks of hardware already make quirks of the language fade in comparison. You need to find out how to split your task into blocks to minimize access to global memory, you need to manage your shared memory, etc, etc. _Somewhat_ performant matrix-vector multiplication algorithm looks nowhere close to textbook definition because of this. So sticking Java or whatever popular language on the problem is not going to make it much more accessible, as you still need very specific knowledge to not waste electricity by writing an algorithm that is 5-10 times slower than it should be, because all it does is waiting for global memory.
Cool project, even if there are a lot of features missing atm. It's always nice to see alternatives to raw VHDL or Verilog for FPGA programming.
We did some preliminary work on executing some parts of the interpreter on a GPU this year: https://github.com/jjfumero/jjfumero.github.io/blob/master/f...
And just after I write it, an article on AMD's work on ROCm pops up.
Exciting times.
You can find more information here: https://dl.acm.org/doi/10.1145/3313808.3313819
The runtime can be added via a jar file, but lambda-based operations must be converted to OpenCL/Cuda/PTX/LLVM or other low-level GPGPU language.
Aparapi did the latter using runtime byte code instrumentation. TornadoVM also does the same thing as a JDK compiler plugin. AeminiumGPU did the same using a transpiler [0] before the actual Java compilation step.
TornadoVM compiles from Java bytecode to OpenCL as well. But additionally, it optimizes and specializes the code by interleaving Graal compiler optimizations, such as partial escape analysis, canonicalization, loop unrolling, constant propagation, etc) with GPU/CPU/FPGA specific optimizations (e.g., parallel loop exploration, automatic use of local memory, parallel skeletons exploration such as reductions). TornadoVM generates different OpenCL code depending on the target device, which means that the code generated for GPUs is different for FPGAs and multi-cores. This is because of OpenCL code is portable across devices, but performance is not portable. TornadoVM addresses this challenge by applying compiler specialization depending on the device.
Additionally, TornadoVM performs live task migration between devices, which means that TornadoVM decides where to execute the code to increase performance (if possible). In other words, TornadoVM switches devices if it knows the new device offers better performance. As far as we know, this is not available in Aparapi (in which device selection is static). With the task-migration, the TornadoVM's approach is to only switch device if it detects application can be executed faster than the CPU execution using the code compiled by C2 or Graal-JIT, otherwise it will stay on CPU. So TornadoVM can be seen as a complement to C2 and Graal. This is because there is no single hardware to best execute all workloads efficiently. GPUs are very good at exploiting SIMD applications, and FPGAs are very good at exploiting pipeline applications. If your applications follow those models, TornadoVM will likely select heterogeneous hardware. Otherwise, it will stay on CPU using the default compilers (C2 or Graal).
Some references:
* Compiler specializations: https://dl.acm.org/doi/10.1145/3237009.3237016
* Parallel skeletons: https://dl.acm.org/doi/10.1145/3281287.3281292
* Live task-migration: https://dl.acm.org/doi/10.1145/3313808.3313819