With a conditional move the processor executes both sides of the branch, but only "commits" the side that actually should be taken. Mis-predicting a branch on a modern OOO superscalar processor can be much more expensive than executing both sides.
But the writer claims that his CMOV version takes more time than the one with the branch for big arrays. I'd expect that the access to the non-cached RAM dominates, and we see that for short arrays CMOV is faster, which is what is to be expected.