He literally presents x86 and ARM assembly dumps where shift right generates one instruction, and divide generates that same instruction plus several others.
Then, he feels the need to run an unnecessary benchmark (most likely screwing it up somehow) and concludes there is no difference!
But how can there possibly be no performance difference, in general, between the CPU running an ALU instruction and running that same instruction plus several other ALU instructions?!?
It's almost unbelievable.
As to how he screwed up the benchmark, my guesses are that either he failed to inline the function (and the CPU is really bad), or failed to prevent the optimizer from optimizing the whole loop, or didn't run enough iterations, or perhaps he ran the benchmark on a different VM than what produced the assembly (or maybe somehow the CPU can extract instruction level parallelism in this microbenchmark, but obviously that doesn't generalize to arbitrary code).