LLVM Backend for the VideoCore4, Raspberry Pi 2 VPU
github.com
github.com
LLVM backends can be in a plugin, but you have to tell the frontend which target to use (so it can use the correct data-layout etc). Its the frontends that don't support plug-in targets. Clang, for example, hardcodes the supported targets.
A combination of using enums as a 'target triple' and lack of dynamic target registration conspire against easy-to-use out-of-tree targets :(
I wish we had a more stable API (maybe with a transformation layer) for the backends that may not probably need much of the new features. If LLVM had a maintained transformation layer from the "current" SelectionDAG to a stable one, it would have helped a lot.
Have you considered this is a deliberate decision, with a goal of having most targets in-tree? :)
(I'm someone who was stuck on 3.6 for ages because of the pain of keeping up with the trunk)
Hardware companies are very touchy about the patents, and not publishing an ISA-related work may often be the only way to protect from the trolls.
Such companies would often encourage (privately or publicly) clean room 3rd party backends, and I'm pretty sure that the one we're discussing here is exactly of this kind. But they would be scared to death to allow any of their employees to even send a tiny patch to a compiler backend, if it could expose some microarchitecture knowledge that was not published.
This constant churn kind of turns the MIT license into the GPL but only for small companies and individuals who can't keep up with the churn. I guess that is why Apple likes LLVM.
You cannot load targets as plugins in clang :(
Not publishing any of the firmware work yet since it can't boot ARM yet (but SDRAM init reliably works across all boards). Once I get ARM working to some extent, I'll probably clean the code up and publish it too.
Regarding the boot loader, have you seen my piface boot loader? http://cowlark.com/piface/ It sounds like what you have is way more sophisticated, as I never got SDRAM init working properly, but you never know --- there might be something useful there.
I was actually considering having another go at a VC4 compiler, this time with pcc; there's an OS I want to port. I'm really happy to know that I no longer have to!
- Setting up exception vectors and enabling exceptions
- Reclocking VPU from PLLC
- UART initialization
- SDRAM initialization
- Copying an ARM blinker stub to 0x0
- ARM power domain initialization
- PLLB initialization
- Enabling passthrough mapping for ARM
- ARM AXI interface initilaization
I'm still trying to figure out what I'm missing in order to get it to work but I suspect it's related to not properly setting up the ARM PLL. I was going to port the RPi clock management driver from Linux but I don't have the time at the moment.As far as the compiler goes, I think it works reasonably well, though I only tested it by compiling my own code with it, there may be things that cause it to error out (anything that involves the frame pointer like VLAs/some C++ features). Also code quality is not ideal since it still doesn't make use of conditional instructions (aside from conditional branches) and doesn't implement AnalyzeBranch to eliminate redundant branches.
I started cleaning up TableGen to turn multi-instruction asm prints into glue DAG in SelDAGtoDAG but I still have to do it for like 4 instructions, which is pretty much a requirement to have MC code emission.
Random question: how did you do 64 bit arithmetic? I couldn't find any kind of add-with-carry or subtract-with-carry instruction. I was semi-resigning myself to have to do a compare-and-test as well as the add, which would have more than doubled the amount of work.
Also, nice work reverse engineering the ARM controller --- how did you get the info?
I think I'm on the right track but I don't have ARM working yet, most likely due to clock misconfiguration. Can probably fix it when I have more time.
Sidenote, I wish #raspberrypi-internals was more active :(
Although the VideoCore in the Pi is particularly interesting because it's responsible for the early stages of the Pi's boot process (the ARM cores are actually turned off when you initially apply power). Right now, that's all a big Broadcom-proprietary binary chunk; good compilers would be the first step towards freeing that code.
QPU is a different beast, in VC4 it cannot even run arbitraty C code. In VC5 it can, but inefficiently.
To fully use the computational power of the GPU, you have to make use of its parallelism. That means dealing with the fact that you can have hundreds or thousands of "waves" (things with a register file and a program counter) in flight simultaneously, and each "wave" corresponds to many (in AMD's case, 64) threads in the conventional sense.
It is the last part that makes the biggest difference compared to regular CPUs, because it changes how you have to think about control flow. If/else-statements must be compiled in such a way that the wave goes through both branches if the threads in the wave branch differently (if all threads branch the same way, you can of course skip the other branch).
The first part makes a big difference as well, of course. GPUs care far less about single-threaded performance, so there is no out-of-order or speculative execution, and the memory latency is high. When a wave has to wait, the latency is made up for by scheduling another wave instead. That is, there is a high level of what is called "hyper-threading" on the CPU.
http://www.broadcom.com/docs/support/videocore/VideoCoreIV-A...
Note that in this case, the author is writing the bootloader firmware so performance isn't a major concern, though.
https://github.com/hermanhermitage/videocoreiv
The VPU is basically a general purpose RISC processor with some fancy vector instructions on top. In fact, most of the firmware that runs on it is written in C.
The assembler/linker I'm using is not ideal, I want to get MC code emission working eventually. I saw that you mentioned limitations on ld/st, why not use lea for data?
For example:
BB1_12: # %sdram_clkman_update_end.exit2
mov r0, 2114982312 # long
ld r2, (r0)
lea r0, .str8(pc) # PCrel load
lea r1, __FUNCTION__.sdram_init_late(pc) # PCrel load
bl xprintfWould appreciate any pointers, no pun intended initially.
Is it possible to use the LLVMLinux patches to run the kernel directly on the VC4?