There is usually a bunch of weird specifics when it comes to calculating performance, but this one happens to be quite simple as we bottleneck on 1 write per clock, and the add instruction latency of 1, so none of the hairy stuff should come into play.
Not knowing what cpu compiler and settings are used I can't do much to replicate the test.