Having to optimise for three very different microarchitectures with the G3 (short narrow pipeline, no AltiVec, small caches), G4 (short wide pipeline, focus on in-order AltiVec, useful OoO otherwise, large L2 yet memory bandwidth starved), and G5 (long and wide pipeline, focus on super-scalar OoO, AltiVec just glued on, more memory bandwidth) must have been "fun".
How many variants of the codecs did you have to implement to make it work?