Instruction count is not everything but when something is 10x larger, unless the 1x is extremely complicated, I think it is safe to expect that the 10x version will run slower. This 1x version is extremely trivial despite that one branch.
.L4:
mov edx, edi
and edx, 1
cmp edx, 1
sbb eax, -1
shr edi
jne .L4
ret
In the first run, jne might be mispredicted (equating to a pipeline flush, ~15 cycles) but in all the other consecutive runs it will be predicted correctly. This means that the larger the amount of iterations is, cost of the first misprediction becomes more and more negligible.
Vectorized execution OTOH is not cheap and does not necessarily result in faster code so purely seeing it in assembly would not mean much without actually measuring it. Scalar versions of the same code may be faster than the compiler-generated auto-vectorization code.
Each SIMD instruction has a latency attached to it (see https://www.intel.com/content/www/us/en/docs/intrinsics-guid...) and the majority of those SIMD latencies are in between 1 and 7 CPU cycles so definitely not free as in a free beer.
That said, I am surprised by the results - in my mental model of how things work this does not add up. Rust assembly is ~50 instructions, roughly ~30 of them being vector (SIMD) instructions and the rest of ~20 instructions being the superset of (only) 6 instructions used in C++ assembly. I cannot see how those 6 instructions, even with the branch mispredict cost of ~15 cycles, can run ~2.5x slower than the pretty much convoluted version of the SIMD popcnt.