The instruction set's here:
https://github.com/hermanhermitage/videocoreiv/wiki/VideoCor...
...but to summarise:
- 32 registers which can hold either integers or 32-bit floats;
- ARM-like 32-bit 3op instructions, with a limited set of 16-bit 2op instructions;
- integrated 32-bit floating point instructions, using the same registers as everything else;
- some 48 bit instruction forms (allowing a full 32-bit payload! No ghastly ARM constant pools or PowerPC-style split payloads!);
- two cores, with integrated interrupt handler;
- integrated DSP with 80-bit vector instructions working on a 64x64x8bit vector working area, which TBH looks like it was stolen from another processor completely and which I frankly don't understand;
- ARM-style conditional execution (for some instructions);
- ARM-style multiregister loads and saves (and pushes and pops);
- DSP-style saturated arithmetic and fixed-point support;
- no ALU-setting-condition-flags instructions, or add-with-carry or subtract-with-carry operations, which makes 64-bit arithmetic really painful (if you need cmp to set the condition flags, what's the carry flag even for?)
It's actually a really nice thing to write assembly for. It's orthogonal enough to be understandable, while quirky enough to allow some really satisfying optimisations. e.g. there's the addcmpb instruction, which while add a value to a register, compare with another register, and branch based on a condition code, all in a single 32-bit operation --- it's basically a loop in a box.
It's also the only time I've ever seen 6-bit floating point constants...
Initially the instructions did all set the status flags but it caused a tight feedback loop in the processor. The choice was between a higher clock frequency for all instructions or better 64-bit arithmetic.
None of the initial video applications needed 64-bit support so it lost out, although I did get to put in the divide instruction just so my Doom port could run faster :)
Are you allowed to tell us what the C compiler used internally was based on? I know there are some very easy to port proprietary compilers which commonly don't see the outside world, and I'm wondering whether it was one of those, or whether some poor sucker had to port gcc.
Well, it's 3D, so you pretty much need perspective divide at the very least…
As it happens, while we were waiting for this compiler to be made for us, I ported GCC to the architecture for my own use. I don't remember it being all that painful, just a few pages of machine description and everything seemed to work fine.
This only supported the scalar instruction set. However, when we needed an MP3 decoder I found that it really needed 32bit precision to meet the audio accuracy, so I also made a different port of gcc that targeted the vector processor. I changed the float data type so any mention of float actually represented 16 lanes of a 16.16 fixed point data type implemented on the vector processor. From what I recall, mp3 decode required 2MHz of the processor for stereo 44.1kHz.
There isn't a fundamental difference in how Intel CPUs start up and how the Raspberry Pi starts up, as I sometimes see people implying.