I think they're misusing "forward propagation" and "backward propagation" to be basically mean "post training inference" and "training".
they seem to be assuming n iterations of the backward pass, which is why it's larger...
n iterations would be a constant factor, which is omitted from asymptotic complexity