It depends! Consider these three functions, in Rust: https://gist.github.com/steveklabnik/da82b51f9dbedd532192492...
One uses a loop, one uses fold with a closure, one uses point-free style. On my machine, summing a vector of a million fives gives
running 3 tests
test sum_fold ... bench: 147,129 ns/iter (+/- 9,101)
test sum_loop ... bench: 148,016 ns/iter (+/- 10,294)
test sum_pointfree ... bench: 146,230 ns/iter (+/- 9,114)
(there is some variance here between runs, sometimes the +/- for each changes a bit.)Cargo's bench profile doesn't use debug assertions, and hence, no overflow checks by default. Let's turn them on:
running 3 tests
test sum_fold ... bench: 147,638 ns/iter (+/- 12,654)
test sum_loop ... bench: 143,485 ns/iter (+/- 7,255)
test sum_pointfree ... bench: 143,050 ns/iter (+/- 7,549)
Looks like they're being eliminated by LLVM anyway.Using godbolt, both the loop and the fold versions (I didn't bother to look at pointfree) are both getting vectorized, and while they're not _exactly_ identical, they're quite similar.
(disclaimer: YMMV, I might have made a mistake, always inspect assembly for your situation, etc)