ld r1, @r2 ; r1 = *r2
add r3,r1,r4 ; r3 = r1+r4
wouldn't set r3=*r2+r4 because the memory access hasn't finished by the time the add runs. ld r1, @r2 ; r1 = *r2
add r3,r1,r4 ; r3 = r1+r4
wouldn't set r3=*r2+r4 because the memory access hasn't finished by the time the add runs.Both load delay slots and branch delay slots are allowing the microarchitecture (a simple 3-stage pipeline) to dictate architecture, which is a classic way to store up pain for the future.
And to give you an example of how you had to calculate things:
There were, if my memory serves me correctly 6 pipeline stages on the 5k numbered 0-5, plus instruction prefetch which was numbered -1. You subtracted the stage in which the instruction took effect from the stage in which a subsequent instruction needed to see that effect, and the result was the number of intermediate stages that all ALUs would need to go through.
Worst case scenario would be if you were modifying RAM that would be read as an instruction; it wouldn't take effect until stage 5, and instruction prefetch was stage -1 so you needed to make sure all ALUs were busy for 6 clock cycles. In theory you could do the math to figure out the scheduling for each ALU, but I just dropped 6 SSNOPs in there, since it was a code path that was only hit during loading of a new process, 6 wasted clock cycles was not a concern.
Note that this is unrelated to interlocks, as any Modified Harvard Architecture will require some sort off synchronization when changing the instruction stream. However, most ISAs have a single instruction that stalls the pipeline and discards any prefetched instructions (e.g. isync on Power). They added one in later revisions of the MIPS ISA as well.
Another fun thing was that there was no interrupt-safe way to disable interrupts, as the interrupt-enabled bit was in a word-sized register along with other values that could legitimately be changed by an ISR. This was also fixed by later revisions of the ISA.