I thought it would be interesting to compare the behaviour of (very) different AArch64 processors on this code.
I ran your code on an Oracle Cloud Ampere Altra A1:
sum_slice time: [677.45 ns 684.25 ns 695.67 ns]
sum_ptr time: [689.11 ns 689.42 ns 689.81 ns]
sum_ptr_asm_matched time: [1.3773 µs 1.3787 µs 1.3806 µs]
sum_ptr_asm_mismatched time: [1.0405 µs 1.0421 µs 1.0441 µs]
sum_ptr_asm_mismatched_br time: [699.79 ns 700.38 ns 701.02 ns]
sum_ptr_asm_branch time: [695.80 ns 696.61 ns 697.56 ns]
sum_ptr_asm_simd time: [131.28 ns 131.42 ns 131.59 ns]
It looks like there's no penalty on this processor, though I would be surprised if it does not have a branch predictor / return stack tracking at all. In general there's less variance here than the M1. The SIMD version is indeed much faster, but by a smaller factor.And on the relatively (very) slow Rockchip RK3399 on OrangePi 4 LTS (1.8GHz Cortex-A72):
sum_slice time: [1.7149 µs 1.7149 µs 1.7149 µs]
sum_ptr time: [1.7165 µs 1.7165 µs 1.7166 µs]
sum_ptr_asm_matched time: [3.4290 µs 3.4291 µs 3.4292 µs]
sum_ptr_asm_mismatched time: [1.7284 µs 1.7294 µs 1.7304 µs]
sum_ptr_asm_mismatched_br time: [1.7384 µs 1.7441 µs 1.7519 µs]
sum_ptr_asm_branch time: [1.7777 µs 1.7980 µs 1.8202 µs]
sum_ptr_asm_simd time: [421.93 ns 422.63 ns 423.30 ns]
Similar to the Ampere processor, but here we pay much more for the extra instructions to create matching pairs. Interesting here that the mismatched branching is faster than the single branch.I guess absolute numbers are not too meaningful here, but a bit interesting that Ampere Altra is also the fastest of the 3 except in SIMD where M1 wins. I would have expected that with 80 of these cores on die they'd be more power constrained than M1, but I guess not.
Edit: I took the liberty of allowing LLVM to do the SIMD vectorization rather than OP's hand-built code (using the fadd_fast intrinsic and fold() instead of sum()). It is considerably faster still:
Ampere Altra:
sum_slice time: [86.382 ns 86.515 ns 86.715 ns]
RK3399: sum_slice time: [306.94 ns 306.94 ns 306.95 ns]