I experimented with cmov for a lot of cpu-bound graphics routines years ago and found that it was rarely faster than branching and never faster than a branchless version w/o cmov. The folks on c.l.a.x said this was due to cmov being microcoded on modern CPUs and so under the hood your still just branching and the only benefit really was instruction cache.