QPU is a different beast, in VC4 it cannot even run arbitraty C code. In VC5 it can, but inefficiently.
To fully use the computational power of the GPU, you have to make use of its parallelism. That means dealing with the fact that you can have hundreds or thousands of "waves" (things with a register file and a program counter) in flight simultaneously, and each "wave" corresponds to many (in AMD's case, 64) threads in the conventional sense.
It is the last part that makes the biggest difference compared to regular CPUs, because it changes how you have to think about control flow. If/else-statements must be compiled in such a way that the wave goes through both branches if the threads in the wave branch differently (if all threads branch the same way, you can of course skip the other branch).
The first part makes a big difference as well, of course. GPUs care far less about single-threaded performance, so there is no out-of-order or speculative execution, and the memory latency is high. When a wave has to wait, the latency is made up for by scheduling another wave instead. That is, there is a high level of what is called "hyper-threading" on the CPU.
http://www.broadcom.com/docs/support/videocore/VideoCoreIV-A...
Note that in this case, the author is writing the bootloader firmware so performance isn't a major concern, though.
https://github.com/hermanhermitage/videocoreiv
The VPU is basically a general purpose RISC processor with some fancy vector instructions on top. In fact, most of the firmware that runs on it is written in C.