Let's say you want to measure if code is faster with branch or cmove, so you make a microbenchmark. In that case the branch predictor has little other branches to keep track of, so it definitely has enough resources to predict that branch well.
In a real world program that function may only be called once in a while and the branch predictor may consider other branches more important.
I wonder if cmove has advantages even if your microbenchmark tells you it doesn't. One side effect would probably be that it takes pressure off of the branch predictor, which we have no good way of measuring in microbenchmarks.